"remote photoplethysmography"
Patent Information
- Application Number
- PCT/AU2026/050284
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure AU2026050284_01102026_PF_FP_ABST
Abstract
Description
" Remote Photoplethysmography"Cross-Reference to Related Applications
[0001] The present application claims priority from Australian Provisional Patent Application No 2025901016 filed on 28 March 2025, the contents of which are incorporated herein by reference in their entirety.Technical Field
[0002] This disclosure relates to estimating a physiological signal of a subject based on remote photoplethysmography.Background
[0003] Remote photoplethysmography (rPPG) is a non -invasive technology used to measure physiological signals from image data, captured using a standard camera, for example. Such physiological signals may be used to determine heart rate, blood pressure and other measurements from image data. This technology leverages the principles of photoplethysmography, which detects blood volume changes in the microvascular bed of tissue. A benefit of rPPG lies in its ability to provide contactless physiological measurements, which can be particularly useful in various scenarios. For instance, it allows for unobtrusive sensing of vital signs in health monitoring systems, making it suitable for use with young children, patients with burnt or sensitive skin, cognitive impairments, and those undergoing certain medical procedures or severe illnesses. Furthermore, it provides distinct advantages for those in remote communities without access to the otherwise required specialised medical devices. Additionally, rPPG can be used in remote health monitoring, enabling continuous observation of vital signs without the need for physical contact.
[0004] Some approaches to remote photoplethysmography utilise machine learning. However, these approaches suffer from degraded performance in the presence of complex real-world scenarios such as subject motion and dynamic illumination. In particular, these approaches cannot explicitly and generally employ the 3D surface of a subject as the surface which encodes the relevant features for rPPG estimation. Moreover, methods which operate on video incur significant VRAM usage and computational cost as the machine learning models are applieddirectly to the video. These approaches are less amenable to deployment in cloud-based or edge-based environments.
[0005] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.
[0006] Throughout this specification the word “comprise”, or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.Summary
[0007] Disclosed herein are a method and system for estimating a physiological signal of a subject based on remote photoplethysmography. Some embodiments of the disclosed method and system leverage the spatial structure of a physiological surface of a subject over time to determine the changes in the spatial structure over time.
[0008] According to an aspect of the present disclosure, there is provided a method for estimating a physiological signal of a subject based on remote photoplethysmography, the method comprising:receiving image data capturing a physiological surface of the subject across multiple frames having a temporal sequence;creating a spatio-temporal graph from the image data, the spatio-temporal graph representing a spatial structure of the physiological surface across the multiple frames according to the temporal sequence;evaluating a trained graph model, trained to generate an output corresponding to a prediction of the physiological signal of the subject, on the spatio-temporal graph to generate an output; andestimating the physiological signal of the subject based on the output.
[0009] It is an advantage to create the spatio-temporal graph as this enables an extraction of spatio-temporal features of the physiological surface of the subject. As some physiological signals manifest in a spatio-temporal pattern on the physiological surface of the subject, extracting such spatio-temporal features enables the physiological signal to be accurately determined from the image data.
[0010] In some embodiments, the spatio-temporal graph comprises edges representing the spatial structure of the physiological surface and nodes associated with a visual feature of the physiological surface at a location of the physiological surface related to the respective node.
[0011] In some embodiments, the visual feature comprises colour of the physiological surface at the location related to the respective node.
[0012] In some embodiments, the method further comprises creating multiple spatial graphs from the image data, each of the multiple spatial graphs representing one of the multiple frames; and creating the spatio-temporal graph comprises connecting the multiple spatial graphs according to the temporal sequence.
[0013] In some embodiments, the method further comprises, for each of the multiple frames, determining a mesh projected onto the respective frame based on anatomical landmarks of the subject captured within the respective frame, wherein creating multiple spatial graphs is based on the mesh.
[0014] In some embodiments, the mesh for at least one of the multiple frames comprises multiple regions and each of the multiple regions is associated with a visual feature corresponding to a spatial average of pixels of the respective frame within the respective region.
[0015] In some embodiments, edges of the at least one of the multiple spatial graphs are based on spatially adjacent regions, the spatially adjacent regions comprising a subset of the multiple regions whose boundary polygon shares one or more vertices.
[0016] In some embodiments, the edges of the at least one of the multiple spatial graphs are weighted based on the spatially adjacent regions.
[0017] In some embodiments, the edges of the at least one of the multiple spatial graphs are weighted using a Gaussian weight based on a distance between the spatially adjacent regions.
[0018] In some embodiments, the mesh comprises a fixed topology for each of the multiple frames, wherein each of the multiple spatial graphs comprises the edges of the at least one of the multiple spatial graphs based on the fixed topology.
[0019] In some embodiments, the edges of the spatio-temporal graph correspond to the edges of each of the multiple spatial graphs.
[0020] In some embodiments, the mesh is a two-dimensional representation of a three-dimensional surface mesh of the physiological surface of the subject.
[0021] In some embodiments, the trained graph model comprises one or more pooling operations configured to reduce an input graph by reducing a number of nodes of the input graph within a spatial dimension or a temporal dimension.
[0022] In some embodiments, reducing the number of nodes of the input graph within the spatial dimension comprises determining a subset of nodes of the input graph by clustering the nodes of the input graph based on their corresponding positions within a canonical coordinate space.
[0023] In some embodiments, reducing the number of nodes of the input graph within the spatial dimension comprises reducing the number of nodes by a factor of between 3 and 5.
[0024] In some embodiments, reducing the number of nodes of the input graph within the spatial dimension comprises computing an average or a maximum of visual features associated with the nodes of the input graph.
[0025] In some embodiments, the one or more pooling operations are configured to reduce a number of edges of the input graph by determining one or more connections between the nodes of the input graph within the spatial dimension.
[0026] In some embodiments, the trained graph model comprises one or more graph convolution operations configured to perform a convolution of node features across a spatial dimension at an associated time of an input graph.
[0027] In some embodiments, the trained graph model comprises one or more convolution operations configured to perform a convolution of node features across a temporal dimension for an associated node of the input graph.
[0028] In some embodiments, the trained graph model comprises a prediction head comprising one or more of an up-sampling block, a max pooling layer and a fully connected layer.
[0029] In some embodiments, the method further comprises determining a physiological measurement based on the physiological signal, the physiological measurement being one of rate of change in blood volume;heart rate;heart rate variability;respiration rate;depth of anaesthesia;blood pressure;blood flow; andblood oxygenation.
[0030] In some embodiments, the image data captures a face of the subject.
[0031] In some embodiments, the method further comprises, for each of the multiple frames, determining a mesh projected onto the respective frame based on anatomical landmarks of the subject captured within the respective frame; and creating the spatio-temporal graph from the image data comprises applying a sampling operator to the image data and the mesh, the sampling operator being configured to calculate a visual feature of the physiological surface for each region of the mesh.
[0032] In some embodiments, the sampling operator is configured to calculate a distribution for each region of the mesh and the sampling operator is configured to calculate the visual feature for each region based on the respective distribution.
[0033] In some embodiments, each of the distributions is parameterised by a mean parameter and a covariance parameter, the mean parameter and covariance parameter being trainable parameters.
[0034] In some embodiments, the mean parameter is a weighted sum of vertices of the mesh, the vertices corresponding to the anatomical landmarks of the subject.
[0035] In some embodiments, weights of the weighted sum are based on a score for each region, the score being indicative of an overlap between the region and the respective distribution.
[0036] In some embodiments, the sampling operator is configured to calculate the covariance by calculating a Jacobian based on weights of the weighted sum.
[0037] In some embodiments, the sampling operator is configured to calculate the visual feature by combining pixel values with the associated distribution.
[0038] In some embodiments, the sampling operator and the trained graph model have been trained simultaneously using a shared loss function.
[0039] In some embodiments, the sampling operator is trained using a top-K gate configured to select top-K regions of the mesh during forward pass of the training.
[0040] According to an aspect of the present disclosure, there is provided software that, when executed by a computer, causes the computer to perform the method of any one of the previously described embodiments.
[0041] According to an aspect of the present disclosure, there is provided a system for estimating a physiological signal of a subject based on remote photoplethysmography, the system comprising one or more processors configured to perform the method of any one of the previously described embodiments.
[0042] Optional features provided in relation to the method, equally apply as optional features to the software and the system.Brief Description of Drawings
[0043] An example will be described with reference to the following drawings:
[0044] Fig. 1 shows an example of a spatio-temporal graph (STGraph) creation, which shows a facial STGraph that disentangles the colour features and geometric structure of the facial surface from video through modelling the facial surface within video as a 3D mesh.
[0045] Fig. 2 illustrates a system for estimating a physiological signal of a subject based on remote photoplethysmography.
[0046] Fig. 3 illustrates a method for estimating a physiological signal of a subject based on remote photoplethysmography.
[0047] Fig. 4a shows an example of a spatio-temporal graph.
[0048] Fig. 4b shows another example of a spatio-temporal graph, which is an alternative arrangement of the spatio-temporal graph shown in Fig. 4a.
[0049] Fig. 5 shows an example of a frame for which a mesh is projected onto the frame.
[0050] Fig. 6 shows the node features Xtand adjacency matrix A(0)of the spatial graph Q at time step t, which are derived from the attributes and relationships of the triangular 3D mesh faces.
[0051] Fig. 7 shows an example diagram of spatio-temporal graph construction.
[0052] Fig. 8 shows the architecture of the trained graph model in an example embodiment.
[0053] Fig. 9 shows how Hierarchical Spatial Graph Pooling (HSGP) block maps a fixed input graph G characterized by the node features h{k}and adjacency matrix A(k)to a fixed coarsened graph G characterized by the pooled node features h'(k)and adjacency matrix A'(k)based on pre-defined node clusters.
[0054] Fig. 10 shows a cross-dataset evaluation from UBFC-rPPGMMPD across folds.
[0055] Fig. 11 shows a visualization of STGraphs constructed from 3D landmarks, 2D landmarks, and bounding boxes.
[0056] Fig. 12 shows a visualization of the activations of the disclosed method in layers I = 3 with respect to the 3D facial mesh for a video from PURE (left) and across the MMPD dataset (right).
[0057] Fig. 13 shows linear equations describing the triangle edges, which are derived from the triangle’s vertex positions.
[0058] Fig. 14 shows a depiction of the image-space mask for a given edge initially assigning pixels below the line as True.
[0059] Fig. 15 shows the intersection of the three interior-facing edge masks used to compute the interior image-space mask for the triangle.
[0060] Fig. 16 shows a visualization of the clusters of node position which maps the initial | Vw|= 852 nodes to | Vw|= 213 coarsened nodes.
[0061] Fig. 17 shows a visualization of the coarsening of the adjacency matrix from| N(k}|= 852 nodes to | N(k}|= 213 coarsened nodes, and then from | N(k}|= 213 nodes to| Vw|= 53 coarsened nodes.Description of Embodiments
[0062] Disclosed herein are a method and system for estimating a physiological signal of a subject based on remote photoplethysmography, which may be referred to as rPPG estimation. The disclosed method and system create a spatio-temporal graph (STGraph) from image data which directly encodes features of a physiological surface, such as its spatial structure, over time. Moreover, the spatio-temporal graph also captures features as they change in time, providing relevant information that can be leveraged to estimate a physiological signal. A trained machine learning model, such as a trained graph model, can then process the spatio-temporal features encoded within the STGraph to provide an accurate physiological signal.
[0063] Some physiological signals manifest in a spatio-temporal pattern on the physiological surface of the subject. For example, subtle changes in skin colour occur as blood volume pulse may perfuse across the 3D facial surface, manifesting a spatio-temporal pattern. Hence, the disclosed method and system aim to isolate and leverage these features encoded within imagedata (such as video) to enable contactless physiological measurements such as pulse rate.Constraining a machine learning model to leverage these surface features may have robust and performant physiological signal estimation.
[0064] An example is present in Fig. 1, which shows a facial STGraph that disentangles the colour features and geometric structure of the facial surface from video, thereby providing a motion-robust representation upon which to model the relevant spatio-temporal features. The disclosed method and system aim to employ the 3D physiological surface as the machine learning model’s primary operand, which can improve accuracy and reduce computational costs.
[0065] The disclosed method achieves high performance with fewer parameters and less computational cost. In particular, the disclosed method provides an optimized GPU-accelerated algorithm that enables real-time node feature computation with minimal VRAM usage. The lightweight nature of the disclosed method enables efficient and performant estimation of a physiological signal from the disclosed STGraph, potentially enabling deployment on edge and mobile devices. Furthermore, due to the low dimensionality of the input as compared to video, the disclosed method can be scaled to large batch sizes, enhancing throughput during inference.
[0066] The disclosed method also provides strong interpretability and fine-grained visualizations, which indicate that it effectively leverages important spatial structure of a physiological surface during rPPG estimation. The experiments described herein also demonstrate the generalization capability and the robustness of the disclosed method. Given the importance of privacy preservation in the design of a machine learning model, the disclosed STGraph sparsely encodes features of a physiological surface, providing improved privacy benefits over video-based methods. The trained graph model provides strong and explicit interpretability characteristics, which contribute to trustworthiness that may be useful for deployment of this technology.System
[0067] Fig. 2 illustrates an example embodiment of a system (denoted as system 200) for estimating a physiological signal of subject 210 based on remote photoplethysmography. Subject 210 comprises physiological surface 215. In this example, subject 210 is a human and physiological surface 215 is the face of subject 210. Fig. 2 is one example of a configuration of system 200. However, system 200 is not strictly limited to this configuration and this may be onepossible embodiment of system 200. It is noted that system 200 of Fig. 2 is only meant to illustrate an example and a preferred system which is capable of performing the disclosed method.
[0068] System 200 comprises device 220, which may be a smartphone, computer, tablet, server device, or any other similar device. Device 220 comprises processor 221. Device 220 comprises memory 222, which comprises non-volatile memory 223 and / or volatile memory 224. Processor 221 may communicate with memory 222 by communicating with non-volatile memory 223 and / or volatile memory 224. Non-volatile memory 223 is a non-transitory computer readable medium and may be an optical disk drive, hard disk drive, solid-state drive, flash memory, storage server, cloud storage or another equivalent type of memory. Volatile memory 224 may be cache, RAM or another equivalent type of memory.
[0069] Memory 222 may store data to be retrieved for later use. For example, memory 222 may store image data, such as individual images and / or video data. Memory 222 may also store the spatio-temporal graph and other graphs, such as spatial graphs. Memory 222 may further store the output of the trained graph model when applied to the spatio-temporal graph and the estimated physiological signal that is based on the output. In essence, memory 222 may store any data (such as parameters or the like) that is necessary to perform the disclosed method. The data thereof may be stored in memory 222 in the form of a JSON format file, XML format file or another equivalent data format file. The image data may be indicative of a two-dimensional image, which may be stored on memory 222 as Joint Photographic Experts Group (JPEG) format, RAW image format or a similar / equivalent image format, for example.
[0070] The methods described herein may comprise applying one or more trained machine learning models. These machine learning models may be stored on memory 222 by storing the weights that form the respective models, for example. Memory 222 may also store any output values calculated by processor 221 applying the one or more trained machine learning models, or any other variable or data necessary to perform such methods described herein.
[0071] Software, that is, an executable program stored on non-volatile memory 223 causes processor 221 to perform methods for estimating a physiological signal of a subject based on remote photoplethysmography. While the singular of “processor” is used herein, it is meant to also encompass multiple processors that are individually or together configured (e.g., programmed) to perform the methods disclosed herein. As such, processor 221 may refer tomultiple central processing units (CPUs) and / or graphical processing units (GPUs) that are configured to collectively perform the methods disclosed herein.
[0072] Once executed, the software may cause processor 221 to (and hence, processor 221 may be configured to) receive image data capturing a physiological surface of the subject across multiple frames having a temporal sequence; create a spatio-temporal graph from the image data, the spatio-temporal graph representing a spatial structure of the physiological surface across the multiple frames according to the temporal sequence; evaluate a trained graph model, trained to generate an output corresponding to a prediction of the physiological signal of the subject, on the spatio-temporal graph to generate an output; and estimate the physiological signal of the subject based on the output.
[0073] Device 220 further comprises camera 225 that is in communication with the processor 221 via the input / output (I / O) port 226. Camera 225 may be any type of image sensor, such as, but not limited to, a webcam, digital camera and smart device camera. In some embodiments, camera 225 is integrally contained in device 220, despite being drawn as disjointed entities in Fig. 2. Camera 225 captures image data of subject 210, which is then communicated to the processor 221 via the I / O port 226. For example, camera 225 may be configured to capture images or video of subject 210. In particular, camera 225 may be configured to capture physiological surface 215 of subject 210 across multiple frames having a temporal sequence. In some embodiments, device 220 may be remotely located from camera 225. For example, camera 225 may be a webcam which is in the same location as subject 210. Hence, camera 225 may receive image data captured by camera 225 via a wired communication (such as Ethernet) or wireless communication via the Internet or via Bluetooth. However, in other embodiments, device 220 may not comprise a camera.
[0074] Software may provide a user interface (such as a graphical user interface) presented to the user on device 220. The user interface may be configured to accept input (via buttons or text fields etc) from the user, via a touch screen or a device attached to device 220 such as a keyboard or computer mouse. These devices may also include a touchpad, an externally connected touchscreen, a joystick, a button, and a dial. In an example, the user interface may display multiple subjects or different image data of one subject, and the user may select image data or subject by interacting with the user interface. The user interaction may then cause processor 221 to perform a method for estimating a physiological signal of a subject based on remote photoplethysmography based on the selected image data or subject.Method
[0075] Fig. 3 illustrates an example embodiment of a method (denoted as method 300) for estimating a physiological signal of subject 210 based on remote photoplethysmography. Fig. 3 is to be understood as a blueprint for a software program and may be implemented step-by-step, such that each step in Fig. 3 may be represented by a function in a programming language, such as, but not limited to, Python, C++ or Java. The resulting source code is then compiled and stored as computer-executable instructions on non-volatile memory 223, which causes processor 221 (or multiple processors or a distributed computing architecture) to perform method 300.
[0076] At 301, method 300 includes receiving image data capturing a physiological surface of the subject across multiple frames having a temporal sequence.
[0077] At 302, method 300 includes creating a spatio-temporal graph from the image data, the spatio-temporal graph representing a spatial structure of the physiological surface across the multiple frames according to the temporal sequence.
[0078] At 303, method 300 includes evaluating a trained graph model, trained to generate an output corresponding to a prediction of the physiological signal of the subject, on the spatiotemporal graph to generate an output.
[0079] At 304, method 300 includes estimating the physiological signal of the subject based on the output.
[0080] In the context of the present disclosure, “physiological signal” may refer to a measurable signal within subject 210 that provides information about their biological processes or health and may be given as a time series. Such physiological signals can be obtained using EEG (electroencephalogram), EKG (electrocardiogram), or EMG (electromyogram), but in the context of this disclosure, the physiological signal is estimated based on remotephotoplethysmography.
[0081] As discussed above, in the example of Fig. 2, subject 210 is a human and physiological surface 215 is the face of subject 210. It is noted that the disclosed method may be particularly applicable to estimating a physiological signal of subject 210 by capturing the face of subject 210, as will be discussed in this disclosure. However, the disclosed method is not limited tosubject 210 being a human and physiological surface 215 is the face of subject 210. For example, physiological surface 215 may be any skin area of subject 210, such as the wrist, neck or leg. Moreover, physiological surface 215 may be an eye or the tongue of subject 210. Preferably, physiological surface exhibits a spatio-temporal pattern, such as a pulse or more specifically, a change in colour due to blood flow. With regards to subject 210, subject 210 may be an animal. In particular, subject 210 may be an animal with exposed skin that exhibits a spatio-temporal pattern, such as pulse or a change in colour due to blood flow.Receiving image data
[0082] Processor 221 receives 301 image data capturing physiological surface 215 of subject 210 across multiple frames having a temporal sequence. Processor 221 may receive 301 image data from camera 225. However, processor 221 may also receive 301 image data that was stored on memory 222. Processor 221 may also receive 301 image data from an external source, such as a server or an external storage device. The image data may be communicated to processor 221 via the Internet or Bluetooth, for example.
[0083] The image data may be indicative of a video, in other words, video data (e.g., a recording, reproducing, or broadcasting of moving visual images). For example, the image data may be a video recording of subject 210 over a period of time. As such, the image data may comprise multiple individual images, which may be referred to as frames. The image recording device (e.g., camera 225) may be configured to capture the image at a frame rate. For example, the image recording device may capture image data at 24 frames per second (fps), 30 fps, 60 fps, 120 fps, 1-15 fps, or any other frame rate. This means that the number of individual images received by processor 221 (i.e., the image data) may depend on the period of time and the frame rate (for example, a period of time of 1 minute at 24 fps would produce 1,440 individual images (frames)).
[0084] The image data comprises multiple frames, which has a temporal sequence. In other words, the multiple frames are temporally related, in the sense that each of the multiple frames may be captured at a different point in time. More explicitly, each of the multiple frames may capture physiological surface 215 of subject 210 and each of the multiple frames may correspond to a different point in time. In other words, the multiple frames may have an order based on the temporal sequence. It is noted that processor 221 may receive 301 image data comprising multiple frames having a temporal sequence, but the multiple frames may be received byprocessor 221 in no particular order. Moreover, in some situations, processor 221 may receive a subset of image data. For example, image data may comprise 5 frames and processor 221 may receive frame 2, frame 5 and frame 3, in this order. In this example, processor 221 still receives 301 image data comprises multiple frames having a temporal sequence (e.g., a temporal order). In other words, the temporal sequence or temporal relationship between the frames is preserved. As such, processor 221 may arrange the multiple frames according to the temporal sequence (e.g., processor 221 may arrange the frames as frame 2, frame 3 and frame 5).Creating a spatio-temporal graph
[0085] Processor 221 creates 302 a spatio-temporal graph from the image data. Processor 221 may create 302 the spatio-temporal graph by applying an algorithm, mathematical operation or machine learning model to the image data. A spatio-temporal graph may be a type of data structure used to represent and analyse data that varies across both space and time. The spatiotemporal graph may combine spatial and temporal information to capture the dynamic relationships and dependencies between entities over time. The spatio-temporal graph represents a spatial structure of physiological surface 215 across the multiple frames according to the temporal sequence. In other words, the spatio-temporal graph may be a graph representation of the image data. Further, the spatio-temporal graph may represent the spatial structure of physiological surface 215 over time and hence, the spatio-temporal graph may represent changes in the physiological surface 215 over time. The spatial structure may refer to the spatial arrangement or inherent geometry of physiological surface 215. For example, if physiological surface 215 is the face of subject 210, then the spatial structure may represent the relative geometry of the facial features (e.g., nose, eyes, mouth etc).
[0086] The spatio-temporal graph may be a data structure comprising data elements indicative of nodes and edges. For example, the nodes may represent spatio-temporal features, while the edges may represent the relationship between these spatio-temporal features. In some embodiments, the spatio-temporal graph comprises edges representing the spatial structure of physiological surface 215 and nodes associated with a visual feature of physiological surface 215 at a location of physiological surface 215 related to the respective node. Processor 221 may extract these visual features from the image and associate them with a time (such as an ordered place in the temporal sequence) to create 302 the spatio-temporal graph.
[0087] It is noted that some edges may represent the spatial structure of physiological surface 215 and other edges may represent a temporal relationship between the spatial structure of physiological surface 215. In other words, some edges may represent how parts of the spatial structure are connected in time. The nodes of the spatio-temporal graph may represent a location of physiological surface 215 at one time. As such, each of the nodes of the spatio-temporal graph may represent a visual feature of physiological surface at one time. In some embodiments, the visual feature comprises colour (such as an RGB value) of physiological surface 215 at the location related to the respective node e.g., skin pigment. However, other visual features are possible.
[0088] In some embodiments, processor 221 may create multiple spatial graphs from the image data. For example, each of the multiple spatial graphs may represent one of the multiple frames. A spatial graph is a type of graph that represents spatial relationships between entities. In the disclosed method, a spatial graph may represent the spatial structure of physiological surface 215 at one time (e.g., the time when the respective frame was captured). Any of the multiple spatial graphs may comprise edges representing the spatial structure of physiological surface 215 and nodes associated with a visual feature of physiological surface 215 at a location of physiological surface 215 related to the respective node. In these embodiments, processor 221 may create a spatio-temporal graph by connecting the multiple spatial graphs according to the temporal sequence. As the multiple frames have a temporal sequence, the multiple spatial graphs may also the same temporal sequence. “Connecting” in this sense may refer to establishing edges between “time-adjacent” nodes.
[0089] For explanatory purposes, consider the example shown in Fig. 4a, which shows an example of a spatio-temporal graph (denoted as spatio-temporal graph 400). In this example, the image data comprises three frames and each of the three frames is represented by a spatial graph (e.g., spatial graphs 410, 420, 430). Each of spatial graphs 410, 420, 430 comprise nodes (such as node 411 of spatial graph 410, and node 421 of spatial graph 420) and edges (such as edge 441 of spatial graph 410). It is noted that the nodes of each of spatial graphs 410, 420, 430 correspond to nodes of spatio-temporal graph 400. Similarly, some edges of spatial graphs 410, 420, 430 correspond to edges of spatio-temporal graph 400. However, spatio-temporal graph 400 further comprises temporal edges that connected temporally adjacent nodes (represented by short-dashed lines). For example, edge 417 connects node 411 and node 421. These temporal edges also form the edges of spatio-temporal graph 400. Spatio-temporal graph 400 in Fig. 4a may be considered to be a three-dimensional graph.
[0090] Fig. 4b illustrates an alternative representation of spatio-temporal graph 400. Spatiotemporal graph 400 in Fig. 4b may be considered to be a two-dimensional graph. Additionally, or alternatively, spatio-temporal graph 400 in Fig. 4b may be considered to be a “flattened” representation of spatio-temporal graph of Fig. 4a. It is noted that both representations of Fig. 4a and Fig. 4b contain the same data (e.g., nodes and edges), but that arrangement of the data may be different in each representation.
[0091] In some embodiments, processor 221 may create the multiple spatial graphs based on anatomical landmarks or other markers of physiological surface 215. In some examples, the anatomical landmarks may correspond to nodes of the multiple spatial graphs. For example, if physiological surface 215 is the face of subject 210, the nose, eyes and / or mouth may be represented as nodes in the multiple spatial graphs.
[0092] In some embodiments, processor 221 determines a mesh projected (or superimposed) onto at least one of the multiple frames based on anatomical landmarks of subject 210 captured within the respective frame. Processor 221 may repeat this for each of the multiple frames. In these embodiments, processor 221 may create the spatio-temporal graph and / or the multiple spatial graphs based on the mesh. Creating the spatio-temporal graph and / or the multiple spatial graphs based on the mesh may improve the accuracy of estimating the physiological signal, as it reduces subject motion in the video, for example.
[0093] For explanatory purposes, consider the example shown in Fig. 5, which shows an example of a frame (denoted as frame 500) for which a mesh (denoted as mesh 550) is projected onto frame 500. In this example, some of the vertices of mesh 550 (represented by black filled squares) are placed at some of the anatomical landmarks of subject 210 (such as the peak of the nose, middle of the mouth and vertices of the eyes). In some examples, the vertices of mesh 550 may correspond to the nodes of the corresponding spatial graph.
[0094] In some embodiments, the mesh (such as mesh 550 in Fig. 5) is a two-dimensional representation of a three-dimensional surface mesh of the physiological surface of the subject. For example, the mesh may be a projection of a three-dimensional mesh of a two-dimensional surface. This may enable a physiological signal to be estimated on physiological surfaces that are non-planar e.g., a face.
[0095] In some embodiments, the mesh comprises multiple regions. For example, referring to Fig. 5, the edges of mesh 550 define the multiple regions (such as region 560). In this example, each of the multiple regions are triangular regions. However, in other examples, the multiple regions may be another shape. Rather than the nodes of the spatial graph corresponding to the vertices of mesh 550, the nodes of the spatial graph may correspond to the multiple regions. In other words, each node of the spatial graph may correspond to one of the multiple regions of mesh 550.
[0096] It is noted that each of the nodes of the spatial graph may be associated with a visual feature of physiological surface 215 at a location of physiological surface 215 related to the respective node. For example, with reference to Fig. 5, if the nodes of the spatial graph correspond to the vertices of mesh 550, then the location may correspond to the vertices of mesh 550. Hence, in this example, the associated visual feature may be determined at the vertex.
[0097] However, in other embodiments, where the nodes of the spatial graph correspond to the multiple regions of mesh 550, the location may be a point within the corresponding region (such as the centre point of the region, for example). In some embodiments, each of the multiple regions may be associated with a visual feature. In some examples, the visual feature corresponds to a spatial average of pixels of the respective frame within the respective region. Performing a spatial average of the pixels may reduce the impact of pixel-level noise and therefore, may provide an advantage. For example, if the visual feature is colour, then processor 221 may calculate the mean RGB value of the pixels within the region (e.g., convex hull triangular mesh face projected onto the image-plane).
[0098] In some embodiments, the edges of the at least one of the multiple spatial graphs are based on spatially adjacent regions. The spatially adjacent regions being a subset of the multiple regions. For example, the spatial adjacent regions may be a subset of the multiple regions whose boundary polygon shares one or more vertices. In other words, the spatial adjacent regions are those which share an edge with the respective region. Similarly, the spatial adjacent regions are those which share a vertex with the respective region.
[0099] In some embodiments, the edges of the at least one of the multiple spatial graphs correspond to binary values based on the spatially adjacent regions. For example, for one of the multiple regions (such as region 560 in Fig. 5), the nodes associated with the spatial adjacent regions and the node associated with the region may be connected by edges with a value of “1”.Similarly, regions that are not spatial adjacent may be connected by edges with a value of “0”. In essence, there may be no edges between nodes that correspond with non-spatially adjacent regions.
[0100] In some embodiments, the edges of the at least one of the multiple spatial graphs are weighted based on the spatially adjacent regions. For example, the edges may be based on an inverse weighted UV distance (e.g., the distance in a pre-defined canonical UV-coordinate space) between the spatially adjacent regions. Such a UV-coordinate space may simply be Euclidean space, for example. Similar to the binary values embodiment, regions that are not spatial adjacent may be connected by edges with a value of “0”. However, the edges corresponding to the spatial adjacent regions may be weighted based on the relative geometry between the regions. The value of the edges may be between “0” and “1”, or another scale. For example, the edges may be based on distances between the midpoints of the regions. As such, regions that are closer together (e.g., the corresponding midpoints are closer) may correspond to an edge correspond to a value close to “1”. Similarly, regions that are further apart may be associated with an edge with a value close to “0”. Using weighted edges may incorporate more spatial information into the spatial graph and hence, the spatio-graph temporal graphs, thereby leading to more accurate estimates of the physiological signal. In some examples, spatial relationships and similarity metrics between nodes may be used to improve the model’s ability to contextually leverage the surface.
[0101] In some embodiments, the edges of the at least one of the multiple spatial graphs are weighted using a Gaussian weight based on a distance between the spatially adjacent regions. For spatial edge features, nearby facial regions may often exhibit spatially correlated colour changes, so it may be advantageous to exploit this spatial coherence to improve the prediction of the trained machine learning model. “Gaussian weight” in the context of the present disclosure may refer to being based on a Gaussian function or Gaussian distribution. For example, a decaying Gaussian weight may be used, which provides an inverse relationship between the spatially adjacent regions (i.e., regions that are closer together are weighted higher than regions that are further apart). In other embodiments, the edges may be weighted based on non-Gaussian functions.
[0102] It is noted that the edges of the multiple spatial graphs and the spatio-temporal graph may be represented by a data structure, such as an array or matrix. In some examples, the edges are represented by an array, which may be referred to as an adjacency matrix. Each element inthe adjacency matrix may represent the spatial adjacency of two regions of mesh 550. The elements of the adjacency matrix may be binary values, if the edges correspond to binary values, or weighted values, as described above. For the spatio-temporal graph, there may be a corresponding adjacency matrix whose elements represent spatial and temporal adjacency between nodes. In some examples, there may be a data structure representing the nodes of the spatio-temporal graph and a data structure representing the edges (e.g., an adjacency matrix), and both data structures may be used as input into a trained machine learning model (e.g., a trained graph model).
[0103] In some embodiments, the mesh comprises a fixed topology for each of the multiple frames. In this context, topology may refer to the arrangement or layout of the elements (nodes and edges) of the mesh (such as mesh 550 in Fig. 5). In general, the topology of physiological surface 215 does not change significantly between frames. For example, if physiological surface 215 is the face of subject 210, then the spatial structure of the face should not vary between frames. As such, the mesh (such as mesh 550 of Fig. 5) may not vary between the multiple frames. Hence, the spatial adjacency (or spatial relationships) between the multiple regions of the mesh may stay the same across the multiple frames. Therefore, in some embodiments, each of the multiple spatial graphs may comprise the edges of the at least one of the multiple spatial graphs based on the fixed topology. In other words, each of the multiple spatial graphs may have the same edges as a result of the fixed topology. If the edges are represented by an adjacency matrix, then the adjacency matrix for each of the multiple spatial graphs may be the same. This may reduce computational resources as only one adjacency matrix may be used to represent the edges of each of the multiple spatial graphs.
[0104] In general, nodes are temporally connected to themselves alone. In other words, the same node exists across the multiple frames. As such, the edges representing temporal connections in the spatio-temporal graph may be easily predicted. For example, processor 221 may not determine the temporal relationship in the same way that it may determine the spatial relations of the multiple regions. This may be referred to as temporal symmetry. As such, the adjacency matrix of the spatio-temporal matrix may be significantly reduced as a result of this temporal symmetry. In some embodiments, the edges of the spatio-temporal graph correspond to the edges of each of the multiple spatial graphs. In other words, in some embodiments, the adjacency matrix of the spatio-temporal graph may be equivalent to the adjacency matrix of any one of the multiple spatial graphs as a result of the temporal symmetry (noting that the adjacency matrix of the multiple spatial graphs may be equivalent due to the fixed topology). This maysignificantly reduce computational resources, as the adjacency matrix of the spatio-temporal graph contains far fewer elements.Evaluating a trained graph model
[0105] Processor 221 evaluates 303 a trained graph model, trained to generate an output corresponding to a prediction of the physiological signal of the subject, on the spatio-temporal graph to generate an output. It is noted that applying a trained machine learning model (such as the trained graph model) to the spatio-temporal graph, as opposed to applying a trained machine learning model directly to image data (such as a video), reduces the computational resources as the spatio-temporal graph contains less information than a video. That being said, the spatiotemporal graph may capture spatio-temporal features of subject 210 which have a strong correlation to the physiological signal of subject 210. As such, it may be said that the spatiotemporal graph contains a higher density of relevant information for predicting the physiological signal as compared to a video.
[0106] In some embodiments, the spatio-temporal graph may be represented by different data structures. For example, the edges of the spatio-temporal graph may be represented by an adjacency matrix, as described above. Further, it is noted that the visual features of the nodes of the spatio-temporal graph may be the input of the trained graph model. As such, processor 221 may create a data structure that represents the visual features of the nodes of the spatio-temporal graph may be the input of the trained graph model. A visualisation of an adjacency matrix and visual feature data structure (which represent the spatio-temporal graph) can be seen in Fig. 7 and Fig. 8.
[0107] The machine learning models described in this disclosure (such as the trained graph model) are understood to be models, such as mathematical models, which receive input and generate an output based on the input. The machine learning models may be of an architecture, such as, but not limited to, a neural network, for example. In general, machine learning models are ‘trained’ to learn and recognise patterns in input and provide an output that is a prediction based on the training it has undergone. Training involves updating weights or parameters of the machine learning model, which define the machine learning models, to minimise a loss value, thereby creating a trained machine learning model (in other words, a machine learning model trained to generate an output). This may involve a gradient descent and b ackpropagation method.
[0108] The machine learning models recited in this disclosure may be stored on memory 222 storing the weights that define the model. As such, each of the machine learning models may be referred to as a “memory model,” given that it is defined by parameters (i.e., the weights) that can be stored on computer memory. In some embodiments, the machine learning models may be programmed on an integrated circuit, such as a field-programmable gate array (FPGA) or an NVIDIA processing unit. In such an embodiment, the handling modules (or processor(s) that may perform the method or part thereof) may not retrieve the parameters from memory 222. Instead, an input may be communicated to the integrated circuit, and the integrated circuit may apply the machine learning model to the input and generate an output, which is then communicated to the handling modules (or processor(s) that may perform the method or part thereof).
[0109] Integrated circuits, such as FPGAs, can be used where flexibility, speed, and parallel processing capabilities are desired. In such an embodiment, the integrated circuit may be part of device 220 of system 200 and may be considered as a “processor” or “processing unit”. Other implementations, such as application specific integrated circuits (ASIC) or neuromorphic architectures, are equally useable.
[0110] Applying any of the machine learning models in this disclosure (such as the trained graph model) may also be considered to be executing or evaluating the machine learning model. Applying, executing, or evaluating the machine learning models may involve calling an application programming interface (API) routine to send the input to a server and the server then performs the calculations according to the machine learning model and returns the results. In other examples, applying, executing, or evaluating may involve issuing a command to local hardware, such as a local chip, device, machine learning accelerator (e.g., a USB device designed to efficiently perform machine learning tasks or NVIDIA’ s Deep Learning Accelerator (DLA)), etc., that has the machine learning model stored thereon and provides a command interface to interact with the model. It is also possible to have a local copy of the machine learning model available so that the calculations are performed by the main processor of the local machine. Other local, remote, or distributed implementations (such as cloud computing environments) are equally useable.
[0111] In some examples, the machine learning models described herein may be a neural network. In further examples, these machine learning models may be a neural network comprising one or more convolutional layers. As such, the machine learning models mayperform the methods described herein by creating feature maps using convolutional filters. Such a machine learning model is known as a convolutional neural network (CNN). A CNN is ideal for applications involving images, and the image as it accounts for the positioning and shape of objects captured in the image.
[0112] However, other types of machine learning models are equally applicable here, such as K nearest neighbour, decision tree, support vector machines, regression models and other artificial neural networks, such as long short-term memory (LSTM) networks or deep neural networks. The machine learning models described herein may also be a collective of different models. The machine learning models may also be based on a transformer model, which comprise encoder and / or decoder blocks and predictions by ‘tokenising’ the input (such as an image). As such, the machine learning models may also comprise self-attention. The machine learning models may output a numerical value.
[0113] In the above embodiments, processor 221 may apply a trained graph model to the graph. The trained graph model may be a machine learning model which may receive a graph as input and may output one or more values (e.g., numerical values) representing a prediction or the like based on the information contained in the input graph. The trained graph model may also receive a graph as input and output a different graph that corresponds to a prediction based on the input graph. Similar to other machine learning models, graph models may comprise adjustable weights (parameters) that define the model. These weights may be modified during a training process. The trained graph model may be trained to improve the accuracy of the estimated locations by adjusting the estimated locations based on the received distance measurements, for example. In some examples, the trained graph model is a graph neural network.
[0114] In some embodiments, processor 221 may train a graph model (or another type of machine learning model) to determine the trained graph model. Processor 221 may train the graph model in a supervised manner. For example, a ground truth may be used (such as the actual physiological signal or corresponding physiological measurement) and processor 221 may compare the result of the graph model to this ground truth. Processor 221 may then adjust the parameters defining the trained graph model based on the difference between the result and the ground truth. Processor 221 may train the graph model based on a loss function (e.g., a function based on a difference between a label and the output of the model). It is noted that in some embodiments, the loss function used to train the disclosed graph model may include phase-shift invariance to improve training on poorly synchronized datasets. Data augmentation may also beused to increase the amount of training data, given that training data is often limited, and also to increase the robustness of the model’s prediction. For example, augmentation used during training may include temporal re-scaling of the STGraph / ground-truth signal and temporal flipping.
[0115] However, in other embodiments, processor 221 may train the graph model in an unsupervised manner. In particular, a ground truth may not be used to train the graph model. Instead, processor 221 may use an unsupervised learning technique, such as clustering, to train the graph model. In a sense, by training the graph model in an unsupervised manner, the graph model learns to focus on certain physiological signals and hence, the graph model does not need to be trained in a supervised manner.
[0116] In some embodiments, the trained graph model comprises one or more pooling operations configured to reduce an input graph by reducing a number of nodes of the input graph within a spatial dimension or a temporal dimension. Pooling may be referred to as feature aggregation, which coarsens the feature maps acts to increase the receptive field and decrease computational cost in subsequent layers. Graph pooling may generalise this operation to arbitrary graph structures allowing coarsening of node features over the dynamic node neighbourhood described by the adjacency matrix. Hierarchical pooling methods may also be used, which consider and preserve the intrinsic graph structure, enabling iterative coarsening. It is noted that “input graph” here may refer to the spatio-temporal graph or a graph created by applying one or more operations to the spatio-temporal graph, e.g., the output graph of one pooling operation may be the input graph to a subsequent pooling operation. Processor 221 may also reduce the nodes and / or edges by applying a pooling operation such as, but not limited to, max pooling, average pooling, global max pooling, global average pooling, min pooling, and L2 pooling. In some examples, the trained graph model may comprise one or more up-sampling blocks for the spatial or temporal dimension.
[0117] In some embodiments, processor 221 may reduce the number of nodes of the input graph within the spatial dimension by determining a subset of nodes of the input graph by clustering the nodes of the input graph based on their corresponding positions within a canonical coordinate space. For example, processor 221 may reduce the number of nodes by cluster nodes that are close together e.g., below a predetermined threshold distance in a canonical coordinate space. In the case of a mesh (such as mesh 550 in Fig. 5), each node position in the canonical coordinate space may be based on the midpoint of the respective region in the canonicalcoordinate space. In some examples, processor 221 may cluster the nodes using k-means clustering. However, in other examples, processor 221 may use other clustering methods such as, but not limited to, hierarchical clustering, density -based clustering and DBSCAN.
[0118] With some clustering methods, such as k-means clustering, the number of nodes that are clustered during a pooling operation may be optimised to optimal accuracy and computational performance. As such, in some embodiments, processor 221 may reduce the number of nodes of the input graph by a factor of between 3 and 5. As will be explained in the experimental results section of this disclosure, reducing the number of nodes of the input graph by a factor of between 3 and 5 (say around a factor of 4) was found to provide an optimal trade-off between computational cost and performance.
[0119] In some embodiments, when reducing the number of nodes of the input graph within the spatial dimension, processor 221 may compute an average or a reduction (i.e., an average reduction or maximum reduction) of visual features associated with the nodes of the input graph. For example, if the nodes are clustered during the pooling operation, then the resulting clustered node may be associated with a node feature which is an average reduction or a maximum reduction of the node features of the clustered nodes. In one example, each node may be associated with a visual feature, such as colour of physiological surface 215, which may be represented by an RGB value. As such, the feature of the clustered node may be an average of the RGB value of the nodes that are clustered. Hence, the cluster node may also be associated with a visual feature. However, in other examples, the node features may be aggregated by applying a different mathematical operation or algorithm to the node features. For example, if the edges are weighted based on spatial relationships, then the node features may be aggregated by calculating a weighted sum using the edge values as the weights.
[0120] In some embodiments, the one or more pooling operations are configured to reduce a number of edges of the input graph within the spatial dimension by determining one or more connections between the nodes of the input graph within the spatial dimension. In other words, processor 221 aggregates the connections (i.e., edges) between coarsened nodes based on the existing connections between their contained nodes. As such, the one or more pooling operations may also reduce the input adjacency matrix (which may be then referred to as a coarsened adjacency matrix). For example, where the edges are represented by binary values, a coarsened edge may have a value of “1” if there is a connection between any nodes within coarsened nodes. A similar process may occur if the edges are weighted based on spatial relationships. However,the values of the edges may be averaged, for example. In other examples, the values of the coarsened edges may be calculated in a different manner e.g., maximum reduction.
[0121] In some embodiments, the one or more pooling operations are configured to reduce a number of nodes and / or edges of the input graph within the temporal dimension similar to the pooling operations described in relation to the spatial dimension. For example, processor 221 may apply one or more pooling operations to reduce a number of nodes and / or edges by clustering the nodes, computing an average or a maximum or the like. Processor 221 may also reduce the nodes and / or edges within the temporal dimension by applying a pooling operation such as, but not limited to, max pooling, average pooling, global max pooling, global average pooling, min pooling, and L2 pooling. In some examples, the trained graph model may comprise a temporal pooling operation, one or more temporal up-sampling blocks, node permutation invariant batch normalization, or a combination thereof.
[0122] In some embodiments, the trained graph model comprises one or more graph convolution operations configured to perform a convolution of node features across a spatial dimension at an associated time of an input graph. Similarly, in some embodiments, the trained graph model comprises one or more convolution operations configured to perform a convolution of node features across a temporal dimension for an associated node of the input graph. As such, the trained graph model may extract both spatial and temporal features captured in the spatiotemporal graph. As a result of these convolutions, the trained graph model may be referred as a graph convolutional model or a graph convolutional neural network.Estimating the physiological signal
[0123] Processor 221 estimates 304 the physiological signal of the subject based on the output. In some examples, the output of the trained graph model may correspond to the physiological signal. However, in other examples, processor 221 may apply further processing to the output to estimate 304 the physiological signal. For example, processor 221 may apply an algorithm, mathematical operation or machine learning model to the output to estimate 304 the physiological signal. In some examples, the physiological signal may be a single value corresponding to a physiological signal at one time (e.g., a time with the period of time for which the image data was captured). In other examples, the physiological signal may be more than one value. For example, the physiological signal may be a vector of signal values. The physiological signal may be a vector of length equal to the number of the multiple frames. Hence, in thisexample, each element of the vector may correspond to a physiological signal at a time captured in an individual frame.
[0124] In some embodiments, the trained graph model comprises a prediction head comprising one or more of: an up-sampling block; a max pooling layer and a fully connected layer. The prediction head may aggregate the spatial features and output the estimated physiological signal of length equal to the number of the multiple frames. Further, the prediction head may temporally up-sample the number of nodes in the input graph, then aggregate the remaining spatial dimensions using maximum pooling to enforce node permutation.
[0125] In some embodiments, processor 221 determines a physiological measurement based on the physiological signal. A physiological measurement may also be referred to as a physiological parameter. In essence, the physiological measurement may be considered to be an indicator of the health of subject 210. In some examples, processor 221 may determine a physiological measurement by applying an algorithm, mathematical operation or machine learning model to the estimated physiological signal. The physiological measurement may be one of but not limited to: rate of change in blood volume; heart rate; heart rate variability; respiration rate; depth of anaesthesia; blood pressure; blood flow; and blood oxygenation.Mathematical formulation and example embodiment
[0126] The mathematical formulation of the disclosed method is now presented. An example embodiment of the disclosed method is also presented. However, it is noted that various aspects discussed in relation to this example embodiment are equally applicable to other embodiments of the disclosed method. The disclosed method is not limited to the example embodiment presented herein. This example embodiment was used to determine the experimental results described herein.
[0127] In this example embodiment, subject 210 is a human and physiological surface 215 is the face of subject 210. As such, the spatio-temporal may be referred to as a facial spatiotemporal graph (STGraph), in the example embodiment. However, as noted previously, the disclosed method is not limited to such embodiment. The face may be a convenient physiological surface, given that a physiological surface can be estimated during a video consultation without having to place a camera (such as a webcam) on another surface. Moreover, the face is a primary physiological surface that exhibits features that manifest as spatio-temporal patterns, such assubtle changes in skin colour occur as blood volume pulse. Even further, the face has an inherent spatial structure which may be leveraged to provide accurate physiological signals or physiological measurements.Facial Spatio-Temporal Graph
[0128] Modelling the 3D facial surface. The 3D facial surface is modelled per video frame as a 3D mesh with a fixed topology across frames. The nodes and edges of the facial STGraph, Gf, are defined from the attributes and connections of the 3D mesh faces, as depicted in Fig. 6. Fig. 6 shows the node features Xtand adjacency matrix A(0)of the spatial graph Gzat time step / , which are derived from the attributes and relationships of the triangular 3D mesh faces.
[0129] In the example embodiment, a STGraph that encodes features and relationships across sequence of video frames at times T e {!,..., | T |} is constructed. Concretely, let G= (V, E, X, A) denote an STGraph that spans T. Here, V denotes the node set, where each node vtjG V represents the z -th spatial region or surface element at time t T. The edge set, E, represents adjacency relationships, with each edge e(t i) (t. GE connecting nodes vt iand vt,}. The feature matrix ' G RT V|' contains the features xtiG RCassociated with each node v, capturing properties such as appearance and geometry. Lastly, the adjacency matrix A G R|THVMTHVIX£encodes edge features, modelling pairwise relationships based on E.
[0130] Facial STGraph
[0131] An instantiation of G called Gfis introduced, which encodes facial surface features and geometric structure. The STGraph Gfis designed so that each node vt iencodes the observable features of a temporally consistent facial surface region across T. Next, the specific components of Gf, including V, E, X, and A, are defined. A diagram of the STGraph construction can be seen in Fig. 7.
[0132] Node Regions. The nodes V are derived by modelling the facial surface across T as a fixed topology 3D triangular mesh derived from 3D facial landmarks:M3D= (VM), F(M)) (i)where V(M)GR|T|X|V< )|X3denotes the set of vertex positions, and F(M)G Z|T|X|F < )|X3represents the triangular mesh faces, with each face referencing the index of three vertices. Each node vt tcorresponds to the facial surface region modelled by the triangular mesh face ft iG F1.
[0133] Edge Connectivity. The connectivity, E, (e.g., the edges) are represented using the structure of M3D, based on intra-frame spatial and inter-frame temporal connectivity. For spatial connectivity, to mimic the dense connectivity present in the image-plane, two nodes vt tand vtare considered adjacent if their associated mesh faces share at least one vertex. Specifically, for mesh faces G F(M ):(2)
[0134] For temporal connectivity, to mimic the temporal correspondence of pixels in video, node vt iis considered to be connected to the corresponding node vz+11in the consecutive frame, that is:%),(G7)G E <^^ = 7 / , =^ + 1- 0)
[0135] This may act to constrain the receptive field of graph convolution operations to local facial surface regions whilst maintaining frame-to-frame correspondence.
[0136] Node Features. For each node vt t, the node features xt tare extracted based on the colour, position, and surface normal vectors of the associated mesh face ft iG F(M), hence:=K. Il« ll^]GR9. (4)
[0137] Noting that spatial averaging is particularly useful, the node colour x^ is defined as the mean value of the pixels contained within the projection of ft tonto the image-plane. This maximizes the number of pixels for each xtciwithout introducing correlations between nodes. The node position x is defined as the mean value of the three vertex positions associated with ft iprojected into a pre-defined canonical UV-coordinate space. This ensures x does not varyacross T, providing a symmetry to be exploited by A for computational efficiency. Lastly, the node surface normal x"tis defined as the outwards facing unit vector of ft tbased on the cross product of two edge vectors. It is noted that the computation of xtciuses vectorized implementation on the GPU to enable real-time computation Table 1. These datasets are further discussed in the Experiments section.Avg. Resolution (Dataset Avg. VRAM Use (GiB) Throughput (FPS)H x W xC)PURE 480 x 640 x3 0.35 198 UBFC-rPPG 480x640x3 0.33 206 MMPD 320x240x3 0.32 225 VIPL-HR 500x455x3 0.33 218Table 1: Performance of facial STGraph construction algorithm across datasets on an NVIDIA GeForce RTX 3090 and Xeon Silver 4114 using predefined 3D meshes averaged across 10 runs.
[0138] Edge Features. The edge features a( ) (z,y)are defined between nodes vtjand vt,}for each edge e(t t,}. For spatial edge features, nearby facial regions often exhibit spatially correlated colour changes, and the spatial coherence is incorporated in the edge features by weighting nearby nodes more strongly using a decaying Gaussian weight based on the Euclidean distance between x and x, by:= exP<.~\ IIx' ~xf H2)G R1- (5)
[0139] As x and x do not vary across T this temporal symmetry is used to storeA G R|v|x|v|xl. For the temporal edge features,a,. = 1 to enable standard convolution along the temporal direction. Where no edge connection exists,a,.n= 0.The trained graph model architecture
[0140] As depicted in Fig. 8, the trained graph model in the example embodiment consists of an eight-layer fully-convolutional STGCN feature extraction backbone and an rPPG prediction head. The STGCN blocks perform extraction and modelling of local spatiotemporal features withrespect to the facial surface, as encoded by Gf= X, A), where X G R|T|X|V|XCand A G R|V|X|V|. The STGCN block in layer 1 expands the channel dimension of X from 3 to cb. The STGCN blocks in layers 3, 5, and 7 further expand the channel dimension by a factor of rc. The pooling blocks iteratively coarsen the spatial and temporal structure of Gf. The temporal pooling blocks ( Pool(l)) in layers 1 and 3 each reduce the temporal dimension of X by a factor of rt. The spatial graph pooling blocks (NPool(r>) in layers 2, 4, and 6 each reduce the spatial dimension of X by a factor of rswhilst coarsening the adjacency A), as depicted in Fig. 8. Hence, given Gfthe backbone outputs extracted features(8)G R|T| / r / |V| / ''XCz / c. Subsequently, the rPPG predictor head temporally up-samples and refines(8)through the TUpsample blocks back to length | T | and aggregates the remaining spatial dimensions using maximum pooling thus enforcing node permutation invariance. Lastly, dropout with probability phis applied to the channel dimensions before projecting the features to Y G R|T|using a fully-connected layer.
[0141] ST-GCN blocks. Each ST-GCN block in the network is implemented as follows: First, graph convolution is applied to the node features h(kper frame to capture intra-frame spatial relationships, where k represents the network layer and 7z(0)= X. The graph convolution is implemented by applying a IxT standard 2D convolution, where T represents the temporal kernel size, and then multiplying by the adjacency matrix Aw, which represents the geometric relationships, on the second dimension, where A(0)= A.
[0142] Temporal edges only connect the same node across consecutive frames. Therefore, standard convolution is used across the temporal dimension of the intermediary node features *(k)to model inter-frame relationships, leveraging the static temporal topology of G..Spatial and Temporal Graph Pooling
[0143] The one or more pooling operations (which may be referred to as HSGP blocks) are used to iteratively compress the spatial dimension of the feature maps in an efficient and parameter-less manner as depicted in Fig. 9. For example, the HSGP block in layer k = 5 of the trained graph model coarsens the graph from 213 nodes to 53. The HSGP is described by the composition of the selection, reduction, and connection functions detailed below.
[0144] Fig. 9 shows how HSGP block maps a fixed input graph G characterized by the node features h(kand adjacency matrix A(k)to a fixed coarsened graph G characterized by the pooled node features h'(k)and adjacency matrix A'(k)based on pre-defined node clusters. The HSGP block in layer k = 5 of the trained graph model coarsens the graph from 213 nodes to 53. The specific process of how the HSGP blocks reduce the number of nodes is given below.
[0145] Selection. Selection maps the | V | nodes of G to the | V | coarsened nodes of G, where | V |<| V |. Each coarsened node vt’jV represents a set of original nodes from V, and the node features and attributes are mapped to their corresponding coarsened node. Selection is implemented using k-means clustering on the node positions to enable aggregation of local spatio-temporal features. Specifically, x "'7' is clustered within a canonical UV-coordinate space to avoid bias in the clusters due to variations in positions across frames. The number of output clusters, | V |, is determined by | V |=| V | / / , where rp& U >1 controls the coarsening ratio in each HSGP block. The clustering is pre-computed to maintain a fixed and pre-defined output topology of G across frames, ensuring the validity of the graph structure assumption, while also minimizing run-time computational cost.
[0146] Reduction. Reduction aggregates the set of node features h^A contained within V using a specified function to obtain h'(k). vQ a Q and maximum reduction of the coarsened node features is implemented.
[0147] Connection. Connection resolves the topology of G as described by the coarsened adjacency matrix A' G R|V|X|V|. Connection is implemented by aggregating the connections between coarsened nodes based on the existing connections between their contained nodes. An entry is assigned At'j= 1 if there is a connection between any nodes within coarsened nodes vt'iand v', as encapsulated by A, otherwise = 0. The topology of G described by A is computed upon initialization. When calculating A' for a given HSGP block, xk"‘ ' is averaged per cluster, enabling iterative clustering and pooling. Each HSGP block in the network may then output a coarsened adjacency matrix, A(k), for subsequent layers as shown in Fig. 9.
[0148] In summary, the pooling blocks iteratively coarsen the spatial and temporal features and graph of Gf. For spatial graph pooling blocks, a hierarchical graph pooling is adapted withclustering based on the node positions x. To coarsen the graph from XU), AI)) to (( / )l, A(l)l) in a fixed manner across T k -means clustering of x is performed, to achieve a similar effect to mesh down-sampling. The number of output nodes | V | is determined by | V |=| V 111 rs. The node features are aggregated per cluster using maximum pooling and the node positions per cluster using average pooling. Lastly, the coarsened graph structure A(r>' is resolved by aggregating the connections between and within clusters. This implementation maintains the fixed topology assumption embedded in G,. For temporal pooling blocks, a maximum pooling of X is employed along the temporal dimension with a stride of rt. This acts to enlarge the receptive field of the trained graph model whilst reducing the computational cost in deeper layers.Loss Functions
[0149] In the example embodiment, a loss function that integrates both time-domain and frequency-domain constraints.
[0150] Time-domain Loss
[0151] Phase-shift can be challenging in training the graph model, due to poor synchronization between videos and reference signals. Furthermore, physiological factors such as pulse transit time (PTT) introduce delays which cannot be accounted for during data acquisition. To address these problems and enhance robustness against sub-sequence level phase-shifts, a modified maximum cross-correlation loss redefining the Pearson correlation coefficient Lris used during training as a soft-expectation across shifts<p e:Lr (. K y) = EW(9?)[LrF)] (6)where y is y shifted by some cp, T controls the sharpness of the selection, and the shift distribution rt ( < / >) is given by:exp\rLr(y y)\W(< P) = — - - — • COS exp[7rLr(y^y)\
[0152] During training, T = 1O and < E=| T | to provide moderately sharp selection at cp with the maximum correlation.
[0153] Frequency-domain Loss
[0154] The frequency-domain loss is computed as the cross-entropy between the power spectrum density (PSD) of y and the normal distribution N (fPR( ), cr), where fHRrepresents the PR derived from y following the evaluation protocol described in the Experiments section of this disclosure. The loss for fPR(y) e [30,180] beats per minute (BPM) is calculated and cr = 6 BPM is used. Thus, the overall loss is expressed as:L =U,y)+LC£(N (fm(y),6), PSD(y) (8)Experiments
[0155] Experiments for rPPG-based pulse rate (PR) evaluation are conducted on four publicly available datasets: PURE, UBFC-rPPG, MMPD, and VIPL-HR. PURE and UBFC-rPPG are small-scale datasets containing 59 and 42 videos respectively in relatively constrained conditions. The MMPD dataset is a medium-scale dataset containing a variety of real-world scenarios across 660 videos. The VIPL-HR dataset is a large-scale dataset containing 2,378 videos captured across a variety of samples and video recording devices.Experimental Setup
[0156] Data Processing. For dataset preparation, video frames were resampled to a regular sampling rate using cubic spline interpolation when frame-time data is available. Next, the photoplethysmography (PPG) signal was resampled to a regular sampling rate aligned with the associated video length using cubic spline interpolation. This ensures both video and PPG data have regular and aligned sampling rates, primarily impacting the PURE and VIPL-HR datasets. For construction of Gf, the 3D facial surface M3Dwas modelled using the MediaPipe FaceMesh pipeline resulting in | V|= 852. The MediaPipe FaceMesh pipeline is a tool for detecting and tracking facial landmarks in real-time, which uses a combination of machine learning models to identify and map 3D facial landmarks from a 2D image or video feed.
[0157] X was computed as described above on non-overlapping video segments of length | T |=384 with contiguous detections. A was computed using MediaPipes x ands= 5.00 as described above. For processing, xbjG X was dynamically masked which face away from the camera (6? > 90° based on the orientation between x"tand the camera normal (nc= [0, 0, -1] ). Only the node colour features were retained for training. Next, the temporal difference of X was computed and normalization per node across time was applied. Similarly, the temporal difference of the ground-truth y was computed and normalization across time was applied.
[0158] Training Details. The parameters used to generate experimental results were cb= 8, rc= 2, rt= 2, rs= 4, pt= 0.45, ph= 0.25, employing the loss function described in example embodiment. The disclosed graph model was trained using the Adam optimizer with a OneCycleLR scheduler and a maximum learning rate of 0.009. All datasets were trained for 50-epochs, using a batch size of 4 for PURE and UBFC-rPPG, and 16 for MMPD and VIPL-HR. Where applicable, the model with the lowest validation loss was retained for testing. For data augmentation, Gaussian noise N (0,0.01) was applied to X, then temporal re-sampling was performed to randomly rescale the length of X and y by up to [-20%, 20%], and temporal flipping p = 0.50 ) of X and y was then performed. For all reported results, a fixed seed (0 ) was assigned, aggregated the results across 5 runs of the associated experimental protocols and report the performance.
[0159] Evaluation Details. The pulse rate (PR) per video was extracted from the predicted rPPG signals, using a>c= [0.50,3.00]77z. It was observed that filtering suppresses the fundamental frequency for low PR, thus the extracted PR was refined / PA(.) to fb,.) / 2 when the ratio of the power p fPR / 2) / p fPR) was above 75%. The mean absolute error (MAE), root mean square error (RMSE), and Pearson’s correlation coefficient (r) between fPR(y) and fPR(y) measured in BPM is reported in the results disclosed herein.Intra-dataset Evaluation
[0160] Table 2 highlights the ability of Gfto directly and robustly encode the relevant features. For complex datasets such as MMPD and VIPL-HR, the disclosed method achieves competitiveperformance while using up to 97.3% fewer parameters and 99.4% less computational cost than methods.PURE UBFC-rPPG MMPD VIPL-HR RMSE 2*r RMSE 2*r RMSE 2*r RMSE 2*r MAE. MAE. MAE. MAE. MethodT T T T (BPM) (BPM) (BPM) (BPM) (BPM) (BPM) (BPM) (BPM) Method i 5.77 14.93 0.81 4.06 8.83 0.89 13.66 18.76 0.08 11.4 16.9 0.28 Method 2 3.67 11.82 0.88 4.08 7.72 0.92 12.36 17.71 0.18 11.5 17.2 0.30 Method 3 - - - 1.19 2.10 0,98 - - - 5.30 8.14 0.76 Method 4 -........ 5.12 8.01 0.79 Method 5 1.29 2.01 0.98 2.19 3.12 0.99 - - - 5.02 7.97 0.79 Method 6 0.82 1.31 0,99 0,44 0.67 0.99 - - - 4.93 7.68 0.81 Method 7 0.40 0.92 0,99 0.17 0.21 0.99 - - - 4.52 7,49 0.81 Method s -........ 4.76 7.51 0.84 Method 9 0.83 1.54 0,99 6.27 10.82 0.65 22.27 28.92 -0.03 11.0 13.8 0.11 Method 10 2.10 2.60 0,99 2.95 3.67 0.97 4.80 11.80 0.60 10.8 14.8 0.20 Method 11 2.48 9.01 0.92 1.70 2.72 0.99 9.71 17.22 0.44Method 12 1.10 1.75 0,99 0.50 0.71 0.99 11.99 18.41 0.18 4.97 7.79 0.78 Method 13 -........ 4.88 7.62 0.80 Method 14 - - - 1.14 1.81 0.99 13.47 21.32 0.21Method 15 0,27 0,47 0,99 0.50 0.78 0.99 3,07 6,81 0,86 4.51 7.98 0.78 Method 16 0.23 0.34 0,99 0.50 0.75 0.99 3.16 7.27 0.84 7,49 0.81 Method 17 - - - 4.64 7.37 0.87 - - - 4.70 7.44 0,82 Method 18 -........ 6.69 9.70 0.48 Method 19 0.69 0.98 0,99 0.45 0,58 0.99 -.....Disclosed0.45 1.24 1.00 0.59 1.44 0.99 2.38 5.54 0.91 4.29 8.51 0.76 methodTable 2: Intra-dataset PR estimation performance of models on the PURE, UBFC-rPPG, MMPD, and VIPL-HR datasets. The best and second best results are formatted as bold and underline respectively. Methods 1-2 are signal processing methods, methods 3-8 are STMap-based deep learning methods and methods 9-19 are video-based deep learning methods.Cross-dataset Evaluation
[0161] As shown in Table 3, the disclosed method demonstrates high generalization performance against significantly larger methods when used without temporal pooling, highlighting the importance of fine-grained temporal information. The disclosed methodotherwise exhibits a strong ability to overfit the training dataset despite its size demonstrating poor generalization, although highlighting the ability of G;to directly and robustly encode the facial surface features. Fig. 10 shows a cross-dataset evaluation from UBFC-rPPGMMPD across folds.PURE UBFC- UBFC-rPPG UBFC-rPPG PURE MMPDrPPG PURE MMPD RMSE RMSE RMSE RMSE MAE 2*r MAE 2*r MAE 2*r MAE 2*r Method I I I I (BPM) T (BPM) T (BPM) T (BPM) T (BPM) (BPM) (BPM) (BPM) Method 9 1.21 2.90 0.99 16.92 24.61 0.05 5.54 18.51 0.66 17.50 25.00 0.05 Method 10 0.98 2,48 0.99 13.22 19.61 0.23 8.06 19.71 0.61 10.24 16,54 0.29 Method 11 1.30 2.87 0.99 13.94 21.61 0.20 3.69 13.80 0.82 14.01 21.04 0.24 Method 12 1.44 3.77 0,98 14.57 20.71 0.15 12.92 24.36 0.47 12.10 17.79 0.17 Method 14 2.07 6.32 0.94 14.03 21.62 0.17 5.47 17.04 0.71 13.78 22.25 0.09 Method 15 2.80 - 0.95 14.57 - 0.14 3.83 - 0.83 14.15 - 0.15 Method 20 0.89 1.83 0.99 8.98 14.85 0.51 0,97 3,36 0,99 9.08 15.07 0.53 Method 16 0,95 1.83 0.99 10,44 16,70 0.36 1.98 6.51 0.96 10.63 17.14 0.34 Disclosed2.05 7.23 0.92 13.85 23.72 0,39 7.58 18.63 0.73 19.98 28.12 0.16 methodDisclosedmethod 3.03 8.26 0.90 10.79 18.22 0.36 0.46 1.28 1.00 9,99 17.09 0,44 < Jt= 1)Table 3: Cross-dataset performance PR estimation of models trained on PURE / UBFC-rPPG and tested on PURE / UBFC-rPPG / MMPD. The best and second best results are formatted as bold and underline respectively.Ablation Study
[0162] Gfwas ablated to isolate its impact on performance and demonstrate the benefits of its design. Ablation studies were conducted following intra-dataset evaluation on MMPD.
[0163] Impact of Node Region Definition. As demonstrated in Table 4, spatial averaging reduces pixel-level noise to improve. The definition and computation of xtciG X usingftiGF(M)deploys all available pixels without introducing correlations, inherently addressing the need for region size optimization.Avg. Pixels PerNode Region Definition MAE J, (BPM) RMSE J, (BPM) 2*r T NodePixel (1 x 1) 1.0 3.55 7.19 0.86 Region (3x3) 9.0 3.44 6.97 0.87 Region (5 x 5) 25.0 3.90 9.13 0.79 Region (10x 10) 100.0 3.91 9.03 0.80 Region (15x 15) 225 3.70 8.49 0.80 Mesh Face 17.6 2.38 5.54 0.91Table 4: Ablation on node region colour feature definition xtciwith average number of pixels per node across the MMPD dataset.
[0164] Impact of Node Orientation Masking. X was masked based on 6rto remove nodes which face away from the camera, ensuring each node vt iencodes only the observable colour improves performance as demonstrated in Table 5. Leveraging a 3D facial surface representation, such as M3£), is useful to enable computation of 0rto remove node features representing re-projected information.Node Feature Masking MAE (BPM) RMSE (BPM) 2*r T - 3.02 6.70 0.88 er>90° 2.38 5.54 0.91 er> 60° 2.92 5.95 0.90 er>45° 6.75 13.89 0.55 er>30° 8.12 13.90 0.55Table 5: Ablation on node feature orientation masking using 9r.
[0165] Impact of Edge Connectivity. As demonstrated in Table 6, leveraging a dense spatial connectivity around each node as defined by E is useful for performance, providing improved performance over the edge connectivity method. The performance improvement of bothscenarios over self-connectivity highlights the importance of spatial feature propagation in modelling local spatiotemporal features.Avg. No. Edges (Edge Connectivity MAE (BPM) RMSE (BPM) 2*r T | V|)Self 1.0 3.53 7.27 0.85 Edges 3.9 3.22 7.61 0.84 gray! 50 Vertices 12.7 2.38 5.54 0.91Table 6: Ablation on edge connectivity E based on M with average number of edges.
[0166] Impact of Edge Features. As shown in Table 7, stronger aggregation of nearby nodes improves performance. To contrast, distant nodes are weighted throughs= -5.00 and the degraded performance is observed, including over the binary case. This highlights the usefulness of geometric information for improving feature aggregation.Edge Features 2*2, MAE (BPM) RMSE (BPM) 2*r TBinary 0.00 2.62 5.95 0.91 Growing Weight -5.00 3.64 7.64 0.87 Decaying Weight 5.00 2.38 5.54 0.91Table 7: Ablation on distance-based ec ge features.
[0167] Impact of Surface Representation. Modelling the facial surface as M3provides a temporally consistent and physically meaningful structure to encode the facial surface features. As evidenced in Table 8, M3provides improved performance over alternative facial surface representations, validating our choice. Furthermore, it is observed that using coarser spatial regions degrades performance. To achieve this, G is constructed using common facial surface representations, as depicted in Fig. 11. Fig. 11 shows a visualization of STGraphs constructed from 3D landmarks, 2D landmarks, and bounding boxes. Facial detection was performed using BlazeFace for all methods. The BlazeFace facial detector is a lightweight and efficient face detection model developed by Google Research, which is designed specifically for mobile devices, offering real-time performance with speeds ranging from 200 to over 1000 frames per second (FPS) on high-end devices.
[0168] 2D landmarks were extracted using FAN and define a fixed triangulation scheme. The Face Alignment Network (FAN) is a deep learning-based library designed for detecting facial landmarks in both 2D and 3D coordinates. Then, X and A were constructed as described above without masking due to lacking 3D information. For dynamic and static bounding boxes, the video was cropped using boxes from all and the 0 -th frame respectively. For the video, the full video frames were used. To construct X, the cropped frames were resized to 29x29 and then the result was flattened. A was then constructed by exploiting the image-structure to obtain a dense 3x3 connectivity around each node.Surface2* V MAE (BPM) RMSE (BPM) 2*r T Representation3D Landmarks852 2.38 5.54 0.91 (MediaPipe)2D Landmarks852 3.02 6.70 0.88 (MediaPipe)2D Landmarks109 5.36 9.52 0.73 (FAN)Dynamic841 3.44 7.89 0.82 Bounding BoxesStatic Bounding841 4.26 8.92 0.79 Box (xl.25)Static Bounding841 4.88 9.92 0.73 Box (xl. OO)None (Video) 841 4.81 9.28 0.75Table 8: Ablation on facial surface representation.
[0169] Impact of Graph Structure. Leveraging the graph structure, A, to constrain the receptive field of the convolutions of the trained graph model to localized facial regions is useful for performance, as demonstrated in Table 9. To explore this, A was removed, necessitating replacement of the GCN operations in the trained graph model with CNN operations, for fair comparison, ID CNN operations are used. Without A to define the node spatial structure the performance is sensitive to node ordering due to spatial locality assumptions of CNN kernels.RMSEGraph Structure Conv. Op. Node Layout MAE (BPM) 2*r T (BPM)A GCN Default 2.38 5.54 0.91 - CNN Default 5.59 10.03 0.70Table 9: Ablation on use of grap i structure A.Discussion
[0170] Computational Cost. As shown in Table 10, the disclosed method uses significantly fewer parameters and computational cost compared to other competitive methods, providing strong performance efficiency. This performance was compared to other methods using | T |= 900 frames with respect to their default input sizes. Per frame metrics are reported in Table 10. The disclosed method performs well when including the overhead of MediaPipe in Gfconstruction, though throughput is limited by the CPU inference of MediaPipe.Throughput / Frame Compute / Frame VRAM / Frame Method Params. (M)(MFLOP)^ (MiB)^ (kFPS)TMethod 8 13.98 9.78 16.96 0.09 Method 7 88.12 5,37 46,63 0.75 Method 9 2.23 228.38 15.03 4.75 Method 10 0.77 138.50 25.27 0.87 Method 11 2.23 228.38 12.25 4.94 Method 12 7.43 316.90 6.95 5.99 Method 14 2.16 115.42 16.78 3.30 Method 15 3.33 241.00 7.83 4.33 Method 16 5.14 81.00 2.68 2.82 Disclosed method 0.09 1.36 79.32 0,14Disclosed0,40 24.74 0.08 0,14 method+ G;Table 10: Computational cost using | T |= 900 on an NVIDIA GeForce RTX 3090 and Xeon Silver 4114 averaged across 10 runs. Best and second best results are formatted as bold and underline respectively.
[0171] Interpretability and Visualization. The disclosed method provides robust interpretability, 3D facial landmarks provide feedback about suitable frames and enabling finegrained spatial understanding of signal strengths with respect to the facial surface. Fig. 12 shows a visualization of the activations of the disclosed method in layers I = 3 with respect to the 3D facial mesh for a video from PURE (left) and across the MMPD dataset (right). In Fig. 12, the average activation across T after the STGCN block in layer 3 is visualised and this is related back to the landmark positions. Across the dataset, strong activation can be observed in regions aligning with physiological expectations.Using a sampling operator for remote photoplethysmography
[0172] A further embodiment will now be described using corresponding reference numerals to those of preceding figures where appropriate for corresponding elements.
[0173] Some rPPG pipelines first construct structured regional signals, making the choice of where and how to aggregate spatial signals a central design decision. Yet regions are typically fixed heuristically or tuned using signal to noise ratio (SNR-like proxies, treating sampling as preprocessing and leaving region placement, spatial support, and surface correspondence decoupled from the supervised objective. In this exemplary embodiment, a sampling operator may be trained end-to-end to further advantageously optimise both region placement and measurement footprint under a downstream rPPG loss. The sampling operator may be considered to be a learnable, surface-anchored sampling operator. The disclosed sampling operator may be used for handling / learning unknowns or uncertainty in the input image data, which may be advantageous in scenarios where the physiological surface of the subject (e.g., the head) is moving frequently during initial calibration.
[0174] The sampling operator described below may be complementary to the disclosed method described above as it removes the fixed mesh topology aspect, in some embodiments. For example, instead of relying on a predefined mesh partition, the sampling operator may learn through training where to place the regions, how broad each region should be, and which regions are worth retaining for the task. The spatio-temporal graph may then be constructed over these learned regions e.g., using a Delaunay triangulation of the learned sampling centres.
[0175] As will be discussed below, in some embodiments, through the use of the disclosed sampling operator, regions are parameterised as Gaussians kernels in canonical UV space,anchored to per-frame face geometry via differentiable soft barycentric mapping, and their supports are transported to the image plane using a local UV XY Jacobian (first-order pushforward) to remain deformation-aware under pose and expression. Evaluated as a plug-in frontend across multiple rPPG architectures, training setups, and representation baselines, the disclosed sampling operator yields consistent improvements in accuracy and cross-dataset transfer over heuristic and proxy region schemes.
[0176] As discussed above, rPPG estimates physiological signals from facial video by measuring periodic blood-volume-induced changes in skin appearance. In practice, these variations are extremely weak relative to nuisance factors such as head motion, facial expression, illumination changes, and partial occlusion. Consequently, some methods may extract signals that aggregate local colour dynamics from video regions over time and construct structured representations before applying a temporal model. As those skilled in the art will appreciate, this construction step may play an important role: it can be used to determine what is measured, over what spatial support, and how / whether measurements track the facial surface.
[0177] In some embodiments, processor 221 creates the spatio-temporal graph from the image data by applying a sampling operator to the image data, the sampling operator being configured to calculate multiple visual features of the physiological surface. In particular, in some embodiments, for each of the multiple frames, processor 221 determines a mesh projected onto the respective frame based on anatomical landmarks of the subject captured within the respective frame, as previously described in this disclosure. The mesh may comprise multiple vertices corresponding to the anatomical landmarks. Each frame may have a different mesh and / or different vertices. Processor 221 may create the spatio-temporal graph from the image data by applying a sampling operator to the image data and the mesh. More specifically, the sampling operator may be applied to the plurality of meshes for all frames of the image data. The plurality of meshes may have be ordered according to a temporal sequence similar or equivalent to the temporal sequence of the frames of the image data. The sampling operator may be configured to calculate a visual feature of the physiological surface for each region of the mesh. More specifically, the sampling operator may calculate a visual feature for each region of each of the plurality of meshes.
[0178] In other words, in some embodiments, the construction of these signals (e.g., visual features) can be constructed from the perspective of a sampling operator over the image data (e.g., a video). For example, given a facial video, the sampling operator may produce a set of Ksignals by integrating local appearance within the video over time. The sampling operator may (i) be anchored to the facial surface to enforce temporal correspondence, (ii) specify the spatial support of each measurement, and (iii) be robust to motion and deformation. However, the task-optimal sampling operator may be difficult to prescribe analytically because rPPG signal strength and reliability may be spatially non-uniform and depend on subject attributes and capture conditions. This may create a design bottleneck: the sampling operator that best serves a downstream supervised objective is generally unknown a priori.
[0179] Fixed sampling operators using heuristic regions (e.g., predefined facial regions or patch grids) or proxy-selected regions (e.g., SNR-based pruning) may be employed to address this. However, these choices impose strong priors about where to measure and the support of each measurement, and they optimize surrogates rather than the downstream physiological objectives. Deep models with learned spatial weighting (e.g., attention or masks) may also mitigate this to some extent, but they may operate as dense image- / feature-plane reweighting mechanisms and may not provide an explicit, surface-consistent, compact set of trackable measurement regions whose locations and supports are directly optimized by the rPPG loss. As a result, sampling design remains only weakly coupled to task optimization. More broadly, learned sampling mechanisms such as deformable attention or learned ROI / glimpse selection can adapt where features are aggregated, but they may act in the image or feature plane and learn offsets / supports without enforcing surface-consistent correspondence under non-rigid facial motion. In contrast, rPPG benefits from physiologically grounded measurements that (i) remain anchored to the same facial surface locations over time, (ii) carry an explicit spatial support that deforms with expression and pose, and (iii) yield a compact, budgeted set of regions whose parameters are directly optimised by the rPPG loss rather than an auxiliary reweighting objective.
[0180] To address this, the disclosure provides a sampling operator that may be learnable, surface-anchored sampling operator for facial rPPG that jointly learns where to measure on the face and how large each measurement region should be. In some embodiments, the sampling operator is configured to calculate a distribution for each region of the mesh and the sampling operator is configured to calculate the visual feature for each region based on the respective distribution. For example, each region may be a triangular region defined by three vertices. In some embodiments, given per-frame face geometry (such as a mesh), K sampling regions may be parameterised as distributions (such as Gaussian kernels) defined on a canonical UV facialmesh. One or more of the distributions may be a Gaussian (normal) distribution, Bernoulli distribution, Poisson distribution, log-normal distribution or the like.
[0181] In some embodiments, each of the distributions is parameterised by a mean parameter and a covariance parameter. The mean parameter may represent a location of the distribution in a space (e.g., canonical UV space or the image space (XY space)). The covariance may represent the spread, shape, and orientation of the distribution’s density. In some examples, the mean parameter and covariance parameter are trainable parameters. Hence, the sampling operator itself may be trainable based on the mean parameter and the covariance parameter. The sampling operator may be trained using gradient decent and / or backpropagation, where a known physiological signal associated with training image data may be used as ground truth. In some examples, the sampling operator may be trained in a similar manner as the trained graph model. In some examples, each region of each of the plurality of meshes (i.e., the mesh for each of the frames) may have a distribution parameterised by a mean parameter and a covariance parameter.
[0182] In some embodiments, the mean parameter is a weighted sum of vertices of the mesh, the vertices corresponding to the anatomical landmarks of the subject. In a sense, each vertex of the mesh may contribute to the distribution, where vertices closer to the distribution centre (i.e., the mean) contribute more than vertices further from the mean. In this sense, the distribution may not necessary be confined to a particular region of the mesh. In other words, each distribution may not be necessarily ‘hard’ associated with a region of the mesh but rather have a ‘soft’ association. This may be referred to as anchoring a continuous UV mean to a discrete mesh without a hard association. As such, some distributions associated with a region This assists with sampling, particularly as physiological surfaces are smooth and continuous.
[0183] In some embodiments, weights of the weighted sum are based on a score for each region, the score being indicative of an overlap between the region and the respective distribution. The score may be referred to as an insideness score, representing how much a distribution lies within its associated region. In some examples, this score (i.e., the insideness score) may be calculated using barycentric coordinates, which may represent where the mean of the distribution is within its associated region (which may be a triangular region for some meshes). In some examples, the weights of the weighed sum add up to 1 by construction. In some examples, the weighed sum provides a mean parameter in the image space (e.g. the XY space) rather than in the canonical UV space. In general, the weights may be used as a soft surface mapping to map the canonical UV space to the image space e.g., UV -> XY. In otherwords, a differentiable surface mapping may anchor these distributions (e.g., kernels) to the moving face in each frame. The insideness score shows (or may be used to interpret) what is happening by mapping graph points / regions to what is being used in an analysis or discrimination. This may be advantageous for explainability and / or deployment of the disclosed methods, in comparison to other machine learning model which may be ‘black box’ in nature.
[0184] In some embodiments, the sampling operator is configured to calculate the covariance by calculating a Jacobian based on the weights of the weighted sum. Processor 221 may use a Jacobian-based push-forward to transport the covariance of the distribution so the measurement footprints deform with pose and expression rather than remaining fixed image-plane crops. In other words, the Jacobian may be used to determine how a mean parameter change in the image-space position when a region centre changes in UV space. As the mean parameter (i.e., the region centre) in the XY space may be given as a weighted sum, the Jacobian may be calculated using the weights. Further, in some examples, the Jacobian may be calculated using the vertices of the mesh.
[0185] In some embodiments, the sampling operator is configured to calculate the visual feature by combining pixel values with the associated distribution. For example, the regional traces may be extracted from the image data by differentiably integrating video pixels under each distribution, enabling gradients from the downstream rPPG loss to directly optimize region locations and supports end-to-end. The sampling operator is configured to calculate the visual feature using a weighted sum (e.g., using quadratures or Riemann summation). Hence, in some examples, the sampling operator may perform numerical integration.
[0186] The cost of UV— >image transport and Gaussian quadrature grow linearly with the number of evaluated regions. To bound compute (and reduce redundant regions), in some embodiments, the sampling operator is trained using a top-K gate configured to select top-K regions of the mesh during forward pass of the training. In other words, a sparse subset of regions under a fixed sampling budget may be learned, so inference cost scales with K instead of dense per-pixel reweighting. The top-K gate may also be learnable to optimise the regions that are chosen.
[0187] The sampling operator disclosed herein may be used with rPPG backbones, such as the trained graph model disclosed herein. For example, the sampling operator may be used to create the spatio-temporal graph, which may also be considered to be a spatio-temporal tensor. Thesampling operator disclosed herein may consistently improves both within-dataset performance and cross-dataset generalization over existing heuristic and proxy-driven sampling schemes.
[0188] The disclosed sampling operator may provide the following:• Task-driven sampling formulation, regional signal extraction in rPPG may be cast as a learnable sampling operator anchored to the facial surface, and the sampling operator may itself be optimised end-to-end under a downstream physiological objective.• Differentiable surface-anchored regions. A compact region parameterization in canonical UV space may be used using Gaussian kernels, together with differentiable centre transport (soft surface mapping) and deformation-aware support transport (Jacobian push-forward), enabling surface-consistent, trackable measurement regions across time.• Differentiable signal formation. Region signals may be extracted via differentiable integration of video pixels under each transported kernel, which may provide an explicit signal-formation operator (region integrals) rather than dense feature reweighting and directly coupling region location and footprint to the rPPG loss.• Budgeted region selection and evaluation. A sparse subset of regions under a fixed budget may be learned for efficient inference, via a learnable hard Top- T selection mechanism and demonstrate improved accuracy and cross-dataset generalization across multiple rPPG methods and training settings compared to fixed and proxy -based baselines.Mathematical formulation and example embodiment
[0189] The sampling operator 0^: ( / , X) I— > S may be a learnable, surface sampling operator that converts a video into a fixed budget of K temporally aligned regional signals. 0^ maintains N candidate distributions (e.g., Gaussian kernels) defined on a canonical UV facial mesh. For each forward pass, a hard Top- T gate selects an active set K. For each selected region and each frame, (i) the canonical UV mean is anchored to the mesh via a differentiable soft barycentric map, (ii) the region centre and support is transported through per-frame geometry, and (iii) a region signal is measured via differentiable Gaussian-weighted integration of the image. The parameters of O are optimized end-to-end under the downstream objective.0r / )outputs a fixed-size set of explicit region integrals with surface correspondence.
[0190] Notation. Let I G ^BXTXHXWXCdenote the input video, and let Itcp) be the c -th channel evaluated at continuous image coordinates p G R2via differentiable interpolation. Let X G R3denote the image-space location of vertex / landmark v at time t. LetR2denote the projection of Xp into the image-plane. A fixed canonical UV mesh is given by vertices U G [0, l]Kx2and faces F G {0,..., V - 1 }Ma. For brevity, the batch index B is omitted throughout.
[0191] Learnable surface-to-video sampling operator dy,. Given a video I and per-frame face geometry Xxyon a canonical UV mesh (U, F), Cy maintains N UV Gaussian kernels and selects K per forward pass with a learnable Top- K gate. Soft barycentric anchoring projects each region to image space and deforms its support via the local UV — XY Jacobian. Region signals are computed by differentiable Gaussian-weighted integration of video pixels, back-facing masking, and UV-regularized. The entire operator is differentiable and may be trained end-to-end with a downstream objective.
[0192] Surface- Anchored Deformable Gaussian Kernels
[0193] Canonical UV Gaussian Kernel parameterization. N candidate regions are maintained, each modelled as a 2D Gaussian kernel in canonical UV with mean and Symmetric Positive Definite (SPD) covariance:wv T) 2 _ 2x2FnGK, ~Ln VLn ) ■,LnG K• (9)
[0194] / "vis optimised directly as a free parameter initialized within the canonical domain. For stable SPD parameterization, / "' is explicitly represented via axis scales and a rotation:< „ = softplus(s„ ) + s, sne R2,=’cos£„g-sin$n g~_sin^gcos0„>g’LU: = ^(^,g)diag(< T„i, CT„2),so E"v= C(C)’ >-0 by construction.
[0195] Differentiable surface anchoring. To anchor a continuous UV mean to the discrete mesh without a hard triangle lookup, a soft surface map is computed. For each face f = a, b, c), barycentric coordinatesn.f= Bary( / "'; Ua, Ub, Uc) G R3and an insideness score qnf= minAmaybecomputed.
[0196] The soft assignment over faces istz„ / =softmax / (7„ / / rft),which is aggregated to per-vertex blend weights WnG R1:M 2 Vir.. YK = 1. / =! 1-0 v-1 Q2)where. is the Kronecker delta. Here, Wnsums to 1 by construction and approaches a convex combination of a single face’s vertices at Tb0 when / "v besinside that face.dWnI d / uuxG R1'2is also computed to propagate local deformation.
[0197] Differentiable framewise transport. Given per-frame landmarks Xx, the region centre tracks the face viaV / e -'E X!'- (13)
[0198] To deform the region support coherently with the local surface motion, the local Jacobian of the soft surface map (UV XY) is used and obtained by differentiating (13):j = J^ = yXxy— G R2X2.
[0199] Region support is transported in factored form to obtain an image-plane sampling factor without factorizing a pushed-forward covariance. Specifically, the UV factor forward is pushed by the local Jacobian:5)and the image-plane covariance is constructed when needed as, which is equivalent to the first-order covariance push-forward Z^1=but avoids a per-region decomposition when Ttis required.
[0200] Gaussian-Weighted Image Integration
[0201] A per-region signal is extracted by Gaussian-weighted sampling of each frame, approximated via Gauss-Hermite quadrature. Given transported Gaussian kernels{( / ^, Z^)}, the region signal may be given as a Gaussian-weighted image expectation:^C=E [ / teO)]-(16)
[0202] Similarly, using the reparameterization p =+I^z with z: N (0,1):Ez:N <0, Z)|^Ac+Qy)
[0203] This enables gradients to flow through the sampling coordinates into both the region centre and the support factorvia differentiable interpolation.
[0204] Quadrature-based implementation. From (9), the 2D standard-normal expectation is approximated using a fixed tensor-product Gauss-Hermite rule. Let {^q,wq}®=lbe ID nodes / weights scaled for a standard normal with the standard A / 2 node scaling and 11 n normalization absorbed into w. Then(18)where / te(’) is evaluated using differentiable bilinear interpolation. Stacking over time and the K selected regions yield signalsS e
[0205] Back-facing masking. To avoid gradients flowing through regions that are back-facing, regions are gated using the local surface normal. Per-vertex normals N^zare computed from, F) and blend to each region:fl = VtF NA; IZn =n / Pn Pv=i (19)For example, with a fixed viewing direction d R3, mtn= l|n’ d > cos6(] and gate<—, such that masked regions advantageously contribute neither signal nor gradient.
[0206] Budgeted Region Sampling
[0207] Hard Top- T region budget. While a pool of N candidate regions is maintained, the cost of UV image transport and Gaussian quadrature grows linearly with the number of evaluated regions. To bound compute (and reduce redundant regions), exactly K= N candidates are extracted per forward pass using a learnable Top- T gate. Each candidate n has a learnable score anduring training Gumbel noise g„: Gumbel(0,l) may be injected to encourage exploration:trainingI an, inferenceK= TopK(d, A"), An= l[w GK],20)
[0208] Here,gscales the deterministic scores relative to the stochastic perturbation. Because h is exactly K -sparse, this selection is applied before the transport and integration and those operations are only evaluated for n K thus, sampling cost scales with K rather than N. A consequence is that only selected candidates receive task gradients in a given step; the Gumbel perturbation mitigates this by occasionally activating different candidates early in training.
[0209] Straight-through gradients and per-region gain. The hard Top- T selection is discrete, so a straight-through estimator is used to optimize the selection scores. Let hnbe the hard indicator from (20), which may be given as:ĥn= hn+ ãn− stopgrad(ãn),which evaluates as the binary mask in the forward pass (hn=hn) while passing gradients directly to the selection scores as if hn= anin the backward pass. In addition, a bounded gainγn= γmaxSigmoid(bn) is learned which forms the final gate weight zn= ĥnγn. This gate is applied to the sampled region features,, / (:for n ∈ K. (22)
[0210] This lets task gradients (when a candidate is active) update anto shape which candidates survive under the Top- T budget, and update bnto calibrate the contribution of selected regions. For efficiency, stncis computed only for n ∈ K unselected candidates are skipped and receive no task gradient on that step.
[0211] Geometric Support Constraints
[0212] UV surface containment. Regions are regularised to remain within a canonical UV facial domain represented as half-spaces Aj' x < b, j = 1,..., / / huN.Hhullℓμ(n) = Σ max(0, A*_j μuv_n − b_j),(23)ℓΣ(n) = Σ max(0, A*_j μuv_n − b_j + khull√(A*_j Σuv_n A_j)),■i-' (24)where«=1 (25)
[0213] Training objective. Given video I and landmarks X, the sampling operator produces regional signals S = Φφ(I, X) (e.g., the spatio-temporal graph or spatio-temporal tensor), whichare consumed by a downstream model fϑto predict ŷ = fϑ(S). The sampler parameters φ may be optimised individually or optimised together with parameters ϑ of a downstream model (such as the trained graph model described herein) under a downstream task loss and geometric regularizes:min Ltask(fϑ(Φφ(I, X)), y) + λhullLhull.(26)
[0214] Hence, in some embodiments, the sampling operator and the trained graph model may have be trained simultaneously using a shared loss function.Experiments
[0215] Experimental Setup
[0216] Datasets, protocols, and metrics. The following dataset were used for evaluation of the sampling operator: PURE, UBFC-rPPG, and MMPD to cover both simple and complex environments.
[0217] Implementation.was instantiated on the MediaPipe FaceMesh canonical UV template (468 vertices), maintaining N = 468 candidate regions and selecting a fixed inference budget via a learnable Top- T gate (§ 3.3); unless stated otherwise, K = 213 is used at test time. Per-frame geometry X (i.e., the mesh) was obtained from MediaPipe FaceMesh tracking (2D landmarks for sampling; 3D geometry for back-face gating). Gaussian kernels were anchored with soft barycentric blending ib= 0.05 ), transported with Jacobian push-forward, and integrated via differentiable Gaussian quadrature (tensor-product Gauss-Hermite, Q = 9 nodes per axis) with bilinear interpolation; back-facing regions are masked using a normal-view threshold 3V= 90°. puvwas initialised uniformly within the convex hull of UV landmarks, using cr = (0.02, 0.02) and 0g= 0, and UV containment regularization was applied with &hull= 3 and uu=1 (§ 3.4).was optimized with Adam using a separate optimizer from the downstream estimator (max LR 5 × 10−3). Top- T uses Gumbel perturbations during training withgannealed to 0.1 and deterministic Top- K at inference.
[0218] Comparison to other rPPG Sampling Operators
[0219] The sampling operator disclosed herein was compared against other region-design approaches to assess performance gains. Specifically, the disclosed sampling operator was compared against (i) fixed region discretizations (e.g., image-plane grids or predefined facial partitions), (ii) surface- / landmark-anchored region suites with geometry priors (e.g., orientation masking), and (iii) proxy-selected region subsets (e.g., SNR / SQI- or error-driven selection) from a fixed candidate set. For each rPPG method fϑ, only the sampling operator that maps video to regional signals was changed. For learnable rPPG methods, O was trained jointly end-to-end with f9using the method’s standard supervised rPPG loss; for signal processing methods, f9was kept fixed and only Φφwas optimised by back-propagating a supervised rPPG loss through the fixed estimator to the sampled region signals.
[0220] Intra-dataset evaluation. Results show that learning the sampling operatorimproves in-distribution PR estimation when the downstream estimator fθ(architecture, training recipe, and post-processing) is held fixed and only the sampling operator is swapped. Across PURE, UBFC-rPPG, and MMPD, the disclosed method achieves lower MAE / RMSE and higher correlation than heuristic ROI discretizations (e.g., grids / partitions), fixed surface-anchored suites, and proxy-selected ROI subsets. This pattern is consistent with the sampling operator learning where to measure and how to aggregate (deformable support) directly under the physiological loss while maintaining surface correspondence, rather than committing to a fixed discretization or optimizing SNR-like surrogates. These results support treating the sampling operator (region signal construction) as a learnable, task-coupled component of the pipeline rather than a preprocessing choice.
[0221] Cross-dataset evaluation. Experiments were performed to determine whether the learned operator yields region signals that remain reliable under dataset shift, again changing only the sampling operator while keeping fθfixed. In established source target transfers, the disclosed method reduces cross-domain MAE / RMSE and improves correlation relative to fixed and proxy-driven ROI operators, with the most consistent gains in harder target domains. This suggests that proxy ROI selection overfits to source-specific surrogate quality measures, while optimizing the sampling operator under the true rPPG objective encourages measurement layouts and supports that improve generalization beyond the training distribution.
[0222] Ablations
[0223] The disclosed method was ablated on MMPD (intra-dataset) using BVPNet, changing only the indicated sampler component while keeping the downstream pipeline fixed.
[0224] Impact of learnable placement and support. Results show that the gains come from jointly learning ROI location and footprint, not from either degree of freedom in isolation.Relative to fixed regions (MAE X), learning only placement improves by X BPM and learning only support improves by X BPM, whereas learning both yields the largest improvement (X BPM). This indicates a coupling between where to measure and how to integrate locally, supporting the importance of learnable sampling-operator design under the rPPG objective.
[0225] Impact of learnable selection and weighting. Results show that selection is a secondary gain on top of learned geometry, while scalar weighting has a negligible effect.Relative to the {uv, LUV} baseline (MAE X), learning selection scores a improves MAE to X (X), whereas adding only / changes MAE little (MAE X / X). This indicates that a mainly reallocates a fixed compute budget across already -useful regions, but neither selection nor reweighting can substitute for learning placement / support, which drives the bulk of the improvement.
[0226] Impact of geometric constraints. Table 3 shows that geometric constraints are necessary for the learned sampler to work reliably. Relative to the full model (MAE X), removing back-facing masking increases the error to X (+X BPM), and removing the hull loss increases the error to X (+X BPM). This indicates that these terms prevent degenerate solutions — sampling from occluded / back-facing or off-face regions — that the task loss can otherwise exploit, improving stability and physical validity during optimization.Discussion
[0227] Node Budget
[0228] Baseline region methods may differ in their native region count, so accuracy can reflect the sampling budget rather than the sampling design. Therefore, the region budget K was swept for tuneable baselines and for our learned operator (via the Top- X gate), reporting accuracy and a lightweight runtime indicator (Fig. 1). At matched X, the disclosed samplingoperator yields lower error and often matches much denser grids, implying the gains come from task-optimized region placement / support rather than region count alone.AppendicesA. Overview
[0229] The appendices provide the following additional details and results with regards to the example embodiment and experiments:• Appendix B offers comprehensive description and design considerations of the facial STGraph construction and further details into the HSGP block implementation.• Appendix C provides further specifics into the experimental setup to ensure reproducibility.B. Method
[0230] Facial STGraph Construction
[0231] Node Feature Computation. The node features, X(rgb), for some frame, is computed as the average RGB value of the pixels contained within the projection of the triangular mesh faces onto the frame plane. This ensures video frames can be processed rapidly for training and inference. Below, the implementation is described, which is optimized to run on the GPU in realtime with minimal VRAM usage.
[0232] The computation for a single triangular mesh face z at time t, for the associated frame F ∈ RH×W×Cis nowdescribed. The image-space positions of the vertices associated with the triangular mesh face are obtained, and the position of the triangle () are computed as the mean of the vertex positions. The set of linear equations describing the edge of each triangle are then computed, as depicted in Fig. 13. The green dots show the positions of the triangular mesh face vertices, the blue dot shows the computed position, x(b)f, and the red lines depict the linear equations for each of the edges.
[0233] Next, an image-space mask is constructed which is True for all pixels within the triangle. To achieve this, interior-facing masks are computed for each edge of the triangle using the linear equations, the mask is defined as True for all pixels which are below the line. Themasks are then inverted if the triangles y -position is above the y -position of the line at the corresponding x -position, this process is illustrated in Fig. 14. As the triangle centre is above the line the mask is inverted to ensure it is True for the pixels on the interior of the triangle.
[0234] The interior image-space mask for the triangle is then computed as the intersection of the three interior-facing edge masks, as shown in Fig. 15. This mask is then applied to the frame to obtain the set of pixels, p,, inside the triangle.
[0235] To efficiently obtain these sets of pixels for all triangular mesh faces, p = {p0,...,pv} where ptR ‘ and P =P > the computation is vectorized across all V triangular mesh faces in the frame. It is noted that the number of pixels depends on the area of the facial surface relative to the frame, as the facial surface typically occupies less than the entire frame P < HWC. Vectorization is performed by repeating both the frame, masks, and other intermediary variables across the V triangular mesh faces, this can incur significant VRAM usage. To address this, a fixed size window of shape, Wj∈ RH×W×C×K, isconsidered around each triangular mesh face based on the dimensions of the largest triangle, where Hw« H and Ww« W. This window per frame is extracted and the masks with respect to this window are computed. This drastically reduces the VRAM cost as compared to naive vectorization using the full video frame.
[0236] Furthermore, triangles and their associated sets of pixels p are optionally clustered together based on a pre-computed k -means clustering of the triangle positions within a canonical UV-coordinate space. This clustering is the same as the one employed in the HSGP pooling blocks within the disclosed method, with the number of output clusters determined as a ratio of the input clusters. Larger ratios aggregate a larger number of triangles together, coarsening the spatial resolution of the input facial STGraph. It is noted, if using this clustering, the corresponding clustered adjacency matrix should also be computed.
[0237] Lastly, mean reduction is performed on the sets of pixels, p, to obtain the average RGB value per cluster of triangular mesh faces. This is achieved through scattering the computation on the GPU with mean reduction using torch scatter. It is noted that the computational cost and memory cost of this process scale both with max(Jp) and P. This process is repeated for all frames within a given sample to obtain the node features.
[0238] Components of the trained graph model
[0239] Spatio-Temporal Graph Convolution
[0240] The graph convolution operation within the ST-GCN blocks of the trained graph model is implemented by first applying a x T standard 2D convolution without bias where T represents the temporal kernel size, to the node features h(k)in the k -th network layer, where h(0)= X. The weighted node features are then multiplied by the adjacency matrix A(k)to aggregate adjacent node features. Each spatial graph convolution operation in the ST-GCN blocks is implemented in code as follows:Algorithm 1 Spatial graph convolution pseudo-code implementation. / z*w= Conv2d(Qw, C®-| V® |, l)( / ?w) h(k)= Einsum(“bctv,vw bctw”, ( / z*w, A(k)))where C‘k>represents the number of input channels, C'kthe number of output channels, and | X(k)| the number of nodes in the k -th network layer respectively. It is also noted that∈ RB×C×T×|V|and A(k)∈ RB×|V|×|V|where B represents the batch size. The computational cost in floating point operations (FLOP) of the standard 2D convolution operation (Conv2d) without bias can be computed as:XtOc„,.„ =BC<‘>C®7’|V,> |2
[0241] After the Conv2d operation:h* ∈ RB×C·|V|×T×1. The computational cost of the subsequent batched tensor multiplication, implemented using einsum, can be computed as:FLOP— = BC^'T | V‘> |!(21 V‘> | -1)FLOP_ = 2BC”T | V‘> I5-BC' 'T | V‘> |!
[0242] Hence, the computational cost of the graph convolution operation in the k -th network layer can be derived by considering the combined cost of the Conv2d and einsum operations:FLOPGCN= FLOPConv2d+ FLOPeinsumFFLOPGCN= BC(k-1)C(k)T|V(k)|2+ 2BC(k)T|V(k)|3− BC(k)T|V(k)|2
[0243] Based on this, the einsum operation dominates the computational cost of the graph convolution operation, scaling cubically with |V(k)|. Consequently, balancing the number of nodes and channels throughout the network is useful to manage the overall computational cost of the disclosed method.
[0244] Hierarchical Spatial Graph Pooling Blocks
[0245] The HSGP blocks within the trained graph model returns the pre-computed pooled / clustered adjacency matrices for subsequent GCN operations. As the HSGP blocks are non-learnable and spatial pooling is performed in a static manner, the adjacency matrices (A(k)) can be pre-computed in the k-th layer of the network upon initialization. Specifically, k -means clusters is performed on the node positions defined within a canonical UV-coordinate space as shown in Fig. 16, to avoid bias or variability in the clusters due to variations in the image-space positions across frames.
[0246] Fig. 16 shows a visualization of the clusters of node position which maps the initial |V(0)| = 852 nodes to |V(1)| = 213 coarsened nodes. During the second iteration this clustering maps the |V(1)| = 213 nodes to |V(2)| = 53 coarsened nodes. This occurs in the canonical UV-coordinate space, the white dots indicate the node positions. The number of output clusters is determined as a ratio of the input nodes, |V(k)| / / rpwhere rprp> 1. For a cluster of nodes, their average position is then computed to effectively spatially coarsen the UV-coordinate space landmarks. This process tends to coarsen regions of high node density over sparser regions. This clustering is iteratively applied to progressively coarsening the nodes, this process is depicted in Fig. 16.
[0247] The obtained clusters are used to coarsen the adjacency matrix progressively upon model initialization. During coarsening of the adjacency matrix all internal and external clustered node connections are aggregated and all connections are re-normalized to 1, which is depicted inFig. 17. This process ensures the graph maintains a fixed and pre-defined output topology across frames and enables efficient run-time clustering.C. Experimental Setup
[0248] Hardware & Software Details
[0249] All experiments were conducted on a Linux VM running Ubuntu 20.04, equipped with 8 x NVIDIA Al 00 GPUs, each with 81,920 MiB of VRAM, using NVIDIA driver version 555.42.06 and CUDA version 12.5 and 2 x Xeon Gold 6348 CPUs (56846 Cores). Platforms with different CPUs / GPUs were used for benchmarking and are explicitly stated. All models were implemented in Python 3.12 using PyTorch 2.4.1.
[0250] Data Processing Details
[0251] The 3D facial surface was modelled as a fixed topology mesh per frame using the MediaPipe FaceMesh 3D landmark detection pipeline, with a confidence threshold of 0.45. MediaPipe outputs 468 detected 3D landmarks alongside the pre-defined triangular mesh tessellation protocol, this pipeline includes the BlazeFace facial detector and the FaceMesh 3D landmark detector. The BlazeFace facial detector is a lightweight and efficient face detection model developed by Google Research, which is designed specifically for mobile devices, offering real-time performance with speeds ranging from 200 to over 1000 frames per second (FPS) on high-end devices. The FaceMesh 3D landmark detector is a 3D facial landmark detection model developed by Google as part of their MediaPipe framework, which provides a detailed mapping of facial features by detecting 468 3D landmarks on the face.
[0252] The explicit dependence on detected landmarks for mesh construction can impact the number of samples able to be formed from the dataset videos. It was found that this does not significantly impact the PURE and UBFC-rPPG datasets due to their low number of missing frames. However, it has a larger impact on the MMPD dataset. It was observed that the frames with missing landmarks were typically associated with significant rotation and poor lighting conditions.
[0253] The relative orientation between the i -th node and the defined camera normal ( ncam= [0, 0, 1]T) for the t -th frame was computed following:arccosl — 7 —: -
[0254] The node features that face away from the camera (θt,i> π / 2) were masked, to ensurethe node features represent the observable colour of a surface region across time. The first-order temporal difference of the node features was computed and normalization was applied per node across time to ensure the scale of the data across all nodes is consistent. Following the facial STGraph construction protocol, Gfrepresented by the node features X ∈ R128×852×3and the adjacency matrix A(0)∈ R852×852was constructed.
[0255] Model Details
[0256] To ensure re-producibility in the computed clusters across runs, the k-means clustering algorithm is manually seeded with 0 and clustering on the node positions within a canonical UV-coordinate space is performed to avoid potential bias due to variations in node positions across frames.
[0257] Evaluation Details
[0258] Evaluation Metrics: The mean absolute error (MAE) measured in Beats Per Minute (BPM), root mean square error (RMSE) measured in BPM, mean absolute percentage error (MAPE) measured in (%), Pearson’s correlation coefficient (r) measured in BPM, and the Signal-to-Noise Ratio (SNR) measured in decibels (dB) are reported. The MAE, RMSE, and MAPE are also reported to provide insights into the bias in the estimates of the predicted pulse rates, r is reported to provide insights into the correlation in the estimates of the predicted pulse rates. SNR is reported to provide insights into the frequency domain characteristics of the estimated rPPG signal. The standard error (SE) is also computed appropriately per metric to provide a measure of the statistical accuracy of the estimates.
[0259] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.
Claims
CLAIMS:
1. A method for estimating a physiological signal of a subject based on remote photoplethysmography, the method comprising:receiving image data capturing a physiological surface of the subject across multiple frames having a temporal sequence;creating a spatio-temporal graph from the image data, the spatio-temporal graph representing a spatial structure of the physiological surface across the multiple frames according to the temporal sequence;evaluating a trained graph model, trained to generate an output corresponding to a prediction of the physiological signal of the subject, on the spatio-temporal graph to generate an output; andestimating the physiological signal of the subject based on the output.
2. The method of claim 1, wherein the spatio-temporal graph comprises edges representing the spatial structure of the physiological surface and nodes associated with a visual feature of the physiological surface at a location of the physiological surface related to the respective node.
3. The method of claim 2, wherein the visual feature comprises colour of the physiological surface at the location related to the respective node.
4. The method of any one of the preceding claims, whereinthe method further comprises creating multiple spatial graphs from the image data, each of the multiple spatial graphs representing one of the multiple frames; andcreating the spatio-temporal graph comprises connecting the multiple spatial graphs according to the temporal sequence.
5. The method of claim 4, wherein the method further comprises, for each of the multiple frames, determining a mesh projected onto the respective frame based on anatomical landmarks of the subject captured within the respective frame, wherein creating the multiple spatial graphs is based on the mesh.
6. The method of claim 5, wherein the mesh for at least one of the multiple frames comprises multiple regions and each of the multiple regions is associated with a visual feature corresponding to a spatial average of pixels of the respective frame within the respective region.
7. The method of claim 6, wherein edges of the at least one of the multiple spatial graphs are based on spatially adjacent regions, the spatially adjacent regions comprising a subset of the multiple regions whose boundary polygon shares one or more vertices.
8. The method of claim 7, wherein the edges of the at least one of the multiple spatial graphs are weighted based on the spatially adjacent regions.
9. The method of claim 8, wherein the edges of the at least one of the multiple spatial graphs are weighted using a Gaussian weight based on a distance between the spatially adjacent regions.
10. The method of any one of claims 5 to 9, wherein the mesh comprises a fixed topology for each of the multiple frames, wherein each of the multiple spatial graphs comprises the edges of the at least one of the multiple spatial graphs based on the fixed topology.
11. The method of any one of claims 2 to 10, wherein the edges of the spatio-temporal graph correspond to the edges of each of the multiple spatial graphs.
12. The method of any one of claims 5 to 11, wherein the mesh is a two-dimensional representation of a three-dimensional surface mesh of the physiological surface of the subject.
13. The method of any one of the preceding claims, wherein the trained graph model comprises one or more pooling operations configured to reduce an input graph by reducing a number of nodes of the input graph within a spatial dimension or a temporal dimension.
14. The method of claim 13, wherein reducing the number of nodes of the input graph within the spatial dimension comprises determining a subset of nodes of the input graph by clustering the nodes of the input graph based on their corresponding positions within a canonical coordinate space.
15. The method of claim 13 or 14, wherein reducing the number of nodes of the input graph within the spatial dimension comprises reducing the number of nodes by a factor of between 3 and 5.
16. The method of any one of claims 13 to 15, wherein reducing the number of nodes of the input graph within the spatial dimension comprises computing an average or a maximum of visual features associated with the nodes of the input graph.
17. The method of any one of claims 13 to 16, wherein the one or more pooling operations are configured to reduce a number of edges of the input graph by determining one or more connections between the nodes of the input graph within the spatial dimension.
18. The method of any one of the preceding claims, wherein the trained graph model comprises one or more graph convolution operations configured to perform a convolution of node features across a spatial dimension at an associated time of an input graph.
19. The method of claim 18, wherein the trained graph model comprises one or more convolution operations configured to perform a convolution of node features across a temporal dimension for an associated node of the input graph.
20. The method of any one of the preceding claims, wherein the trained graph model comprises a prediction head comprising one or more of: an up-sampling block, a max pooling layer and a fully connected layer.
21. The method of any one of the preceding claims, wherein the method further comprises determining a physiological measurement based on the physiological signal, the physiological measurement being one of:rate of change in blood volume;heart rate;heart rate variability;respiration rate;depth of anaesthesia;blood pressure;blood flow; andblood oxygenation.
22. The method of any one of the preceding claims, wherein the image data captures a face of the subject.
23. The method of any one of the preceding claims, whereinthe method further comprises, for each of the multiple frames, determining a mesh projected onto the respective frame based on anatomical landmarks of the subject captured within the respective frame; andcreating the spatio-temporal graph from the image data comprises applying a sampling operator to the image data and the mesh, the sampling operator being configured to calculate a visual feature of the physiological surface for each region of the mesh.
24. The method of claim 23, wherein the sampling operator is configured to calculate a distribution for each region of the mesh and the sampling operator is configured to calculate the visual feature for each region based on the respective distribution.
25. The method of claim 24, wherein each of the distributions is parameterised by a mean parameter and a covariance parameter, the mean parameter and covariance parameter being trainable parameters.
26. The method of claim 25, wherein the mean parameter is a weighted sum of vertices of the mesh, the vertices corresponding to the anatomical landmarks of the subject.
27. The method of claim 26, wherein weights of the weighted sum are based on a score for each region, the score being indicative of an overlap between the region and the respective distribution.
28. The method of claim 25 or 26, wherein the sampling operator is configured to calculate the covariance by calculating a Jacobian based on weights of the weighted sum.
29. The method of any one of claims 24 to 28, wherein the sampling operator is configured to calculate the visual feature by combining pixel values with the associated distribution.
30. The method of any one of claims 23 to 29, wherein the sampling operator and the trained graph model have been trained simultaneously using a shared loss function.
31. The method of any one of claims 23 to 30, wherein the sampling operator is trained using a top-K gate configured to select top-K regions of the mesh during forward pass of the training.
32. Software that, when executed by a computer, causes the computer to perform the method of any one of the preceding claims.
33. A system for estimating a physiological signal of a subject based on remote photoplethysmography, the system comprising one or more processors configured to perform the method of claims 1 to 31.