Photoplethysmography signal extraction

The method for determining PPG signals using remote camera observation by dividing images into regions and clustering pixels effectively addresses interference issues, providing accurate and efficient PPG signal extraction for medical imaging.

JP2026512940APending Publication Date: 2026-04-22KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2023-10-25
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing camera-based photoplethysmography (PPG) systems face challenges in accurately extracting PPG signals due to interference from MRI scanners and the need for complex setups, which can introduce noise and require extensive post-processing.

Method used

A method and apparatus for determining PPG signals using remote camera observation by dividing camera images into regions of interest and non-interest, applying dynamic characteristic analysis, and clustering pixels to isolate stable PPG signals, reducing interference and post-processing needs.

Benefits of technology

This approach enables accurate and stable PPG signal extraction with reduced noise and computational resources, suitable for medical imaging and other diagnostic procedures without direct contact, improving signal quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026512940000001_ABST
    Figure 2026512940000001_ABST
Patent Text Reader

Abstract

Apparatus 10 and method for use in camera-based photoplethysmography are disclosed. The apparatus has a camera image input 11 and a subset generator 12 that divides image regions where tissue of interest is presumed to exist into a first plurality of pixel subsets and second regions where irrelevant tissue is presumed to exist into a second plurality of subsets. A signal dynamics analyzer 13 determines the dynamic characteristics of each subset. A clustering processor 14 clusters the first and second plurality of subsets, respectively, based on their dynamic characteristics. A selector 15 selects clusters from the first plurality of subsets based on size and / or morphology, thereby excluding each cluster that substantially corresponds to a cluster in the second plurality of subsets in its associated signal dynamics. Regions of interest defining the selected clusters and / or PPG signals extracted from the regions of interest are provided as output 16.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more particularly, to determining a photoplethysmography (PPG) signal from a subject based on a camera image. More specifically, the present invention relates to an apparatus, method, system, and computer program for use in determining a photoplethysmography signal by camera observation.

Background Art

[0002] Photoplethysmography (PPG) is a technique that uses changes in light interaction, such as light absorption, to measure changes in blood volume in the microvasculature. For example, a simple PPG sensor can have a light source for illuminating a biological tissue such as skin, mucosa, and other translucent tissues, and a photodetector such as a photodiode, a phototransistor, and / or a photomultiplier tube for measuring the amount of light scattered and / or absorbed by the tissue. The detected signal indicates changes in blood volume in the microvasculature, and thus provides a non-invasive method for measuring changes in blood volume (almost) in real time. PPG has a wide range of applications, including monitoring blood oxygen concentration, pulse rate, current cardiac phase, and blood pressure. The information contained in the PPG signal can be used, for example, as a trigger and / or gate during a magnetic resonance imaging (MRI) procedure.

[0003] Furthermore, it is known in the art to use camera observation of a subject, such as a medical image or a patient undergoing a therapeutic procedure, to obtain a useful signal indicating a physiological parameter of the subject. For example, knowledge of the cardiovascular and / or respiratory function of a subject can be useful in processing images obtained in a diagnostic imaging procedure for monitoring the state of the patient, such as health, during the procedure, and / or in guiding or controlling the image acquisition of the performed procedure in a related manner.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Remote photoplethysmography (rPPG), also known as camera PPG or video PPG, is a non-contact PPG signal extraction method that uses a camera to capture images of tissue during examination. Camera-based PPG has several advantages over conventional (contact) PPG, including the ability to capture a wide area of ​​tissue, which is useful when the tissue is not uniformly illuminated, and the ability to simultaneously capture multiple PPG signals from different parts of the tissue, which can improve the accuracy of the results, for example, by averaging them.

[0005] A typical approach to acquiring rPPG signals involves first identifying a region of interest (ROI), such as a patch of skin tissue, in one or more images (e.g., a video stream) acquired by a camera. Next, a PPG is constructed from the dynamic changes in pixel values ​​corresponding to this ROI. This can include various image / signal processing techniques, such as downmixing of color signals and post-processing to clean the signal from undesirable components. From the PPG signal thus obtained, markers (e.g., time points) identifying features of interest, such as the pulse rate and / or one or more specific stages of the cardiac cycle, can be extracted. Such markers can be used by trigger generator systems, for example, to trigger MRI prepulses and / or to control data acquisition windows.

[0006] In MRI, ROI identification can advantageously utilize prior knowledge of the MRI environment, since the environment typically remains substantially the same across different sessions. Illumination level information (e.g., preferring areas that are neither too dark nor too bright to select well-lit skin portions) and pulsation information (i.e., information indicating living skin) are also usually considered.

[0007] The dynamic range of the camera-based PPG signal itself is very small, with a 10% offset relative to the average illumination (average pixel value offset). -3Because the impact of other factors that may affect the detected pixel signal is minimal, careful control of illumination conditions is generally an important consideration for achieving continuous and accurate monitoring. Therefore, the use of a high-quality, stable infrared illumination source is considered advantageous. A simple setup can be used with a single-channel near-infrared (NIR) illumination source and a camera that picks up the NIR signal after modulation due to interaction with the tissue of interest. The signals from different pixels within the ROI region can be combined to construct a PPG signal.

[0008] Nuisance components in a signal can be removed by signal decomposition techniques. This may include disturbances due to the movement of the object; that is, motion signals in the image can be measured and removed from the (raw) PPG signal. Similarly, general light fluctuations can be measured and removed from the PPG signal.

[0009] Exemplary approaches that take the known pulsation characteristics of skin-based PPG signals are sometimes commonly referred to as “living skin models.” In such approaches, a camera image, or a portion thereof pre-selected based on prior knowledge (e.g., an area in the image where a patient is typically placed for a predetermined procedure), is divided into a grid of neighboring rectangular patches to form smaller patches, e.g., compartments or pre-selected portions of the image.

[0010] Subsequently, the average pixel intensity level and spread (e.g., standard deviation) of each patch, and the motion vector (e.g., using an optical flow algorithm) of each patch can be determined. When the average pixel level, spread, and motion all satisfy predetermined conditions, the patch is selected as a uniform, stable patch for further use, for example, to prune patches that do not exhibit these desirable characteristics. For example, the average pixel level condition can correspond to lighting conditions that fit within the dynamic range of the camera system with sufficient margin and match the typical characteristics of the target tissue under such lighting conditions. If the spread and motion are below (respectively) predetermined thresholds, it can be assumed that there is a certain level of stability for determining the signal of interest, for example, that the patch is not excessively affected by various sources of noise and / or error, or by strong motion.

[0011] For patches not removed by the pruning process, the average value as a function of time can be analyzed. After applying the Fourier transform, frequency components representing, for example, the heart rate band, which fall within a predetermined range, can be examined to find spectral peaks within this range. If the strongest peak is sufficiently clear, for example, if the peak's energy is sufficiently high compared to the range average (and / or higher compared to the second strongest peak), the patch is associated with the frequency of this primary peak (the strongest spectral peak). This value is called the patch frequency. Finally, by clustering the patch frequencies, the largest cluster can be selected as the most likely candidate to represent the PPG signal, and the pixels belonging to the patch within this cluster can be used to define the region of interest for extracting the PPG signal after this initial setup / calibration procedure.

[0012] For example, U.S. Patent Application Publication No. 10441173 describes a prior art method for extracting physiological information, such as PPG signals, from remote camera observation data. A data stream containing image data representing an observation region of interest is received. Multiple sub-regions are defined within this region, and a classifier classifies the sub-regions into indicator-type and auxiliary-type regions, the former including a region that at least partially represents the region of interest, and the latter including a reference region. The sub-region(s) classified as the region of interest are then further processed to obtain vital information of interest, such as PPG signals.

[0013] The object of the present invention is to provide good and efficient means and methods for determining PPG signals from a subject (e.g., a patient undergoing a medical imaging procedure) from, for example, a camera image (video stream) of the subject, based on remote observation. [Means for solving the problem]

[0014] In particular, a region of interest particularly suitable for extracting such PPG signals from camera images is determined, and / or this ROI is used to extract and provide such PPG signals.

[0015] An advantage of the embodiments of the present invention is that remote imaging can be used so as to eliminate the need for direct contact with the object and / or so as to allow the device to be positioned substantially away from the object. For example, in a medical setting, it may be preferable to have no obstacles in immediate vicinity of the patient so that procedures can be performed without unnecessary interference.

[0016] Furthermore, in medical imaging, diagnostic imaging techniques can potentially interfere with other electronic devices.

[0017] In magnetic resonance imaging (MRI), the hardware of the MR scanner, such as the high-frequency antenna and / or magnetic field gradient coil, can interfere with the correct operation of other electronic devices. For example, PPG sensors may pick up noise caused by high-frequency signals and / or be affected by strong magnetic fields (and / or magnetic field gradients and / or dynamic magnetic field changes). Similarly, the illumination source used to light the patient's skin for PPG detection can be interfered with by the MR scanner hardware, which can cause illumination changes that interfere with PPG detection. Furthermore, another advantage is that an accurate and / or stable PPG signal can be obtained by reducing and / or avoiding the carrying over of interfering signals, such as random noise and systematic errors, into the PPG output signal inferred from the raw data stream.

[0018] Therefore, the output signal obtained in this way requires less post-processing, such as further cleaning of the signal. This can improve the efficient use of computational resources, such as processing power, memory resources, and power consumption, and / or improve the accuracy and stability of the results derived from the output. For example, many signal cleaning strategies require that the measured disturbance signal and the disturbance in the raw PPG are synchronized, and this condition is not always met. By carefully selecting the pixels from which the signal is derived, a more stable and representative PPG signal can be obtained, in which case the residual interference components are generally smaller and less random.

[0019] An advantage of the embodiments of the present invention is that useful signals can be automatically determined (e.g., by algorithm) from camera observations of the target. Since in many environments a suitable camera system already exists or can be installed with an easy upgrade, the embodiments can be applied to existing systems without complex and / or expensive intervention.

[0020] An advantage of the embodiments of the present invention is that in diagnostic imaging procedures using scanner systems such as MRI, CT, SPECT, and PET scanners, signals usable for triggering and / or system control can be determined from camera observations, thereby advantageously improving the quality of diagnostic images. Such trigger signals can also be used, for example, in radiotherapy and similar (e.g., therapeutic) procedures, and are not necessarily strictly limited to medical imaging applications. Furthermore, other application areas for PPG acquisition are not necessarily excluded.

[0021] An advantage of the embodiments of the present invention is that appropriate region of interest (ROI) selection can be obtained for pixels representing PPG pulsating signals containing such substantial components. In particular, prior knowledge can be advantageously used to prevent undesirable signal components, i.e., interfering signals, noise, and / or other disturbances, from contributing to the extracted PPG signal. For example, a sophisticated pixel screening process can be applied to select appropriate ROIs based on knowledge of desirable versus undesirable regions and / or the characteristics of the pixel signal content. Knowledge of non-skin locations (or, more generally, locations that do not represent the tissue type under observation) can be advantageously used to extract pulsating features. Furthermore, morphological knowledge of the characteristics of promising candidate regions can be applied to ROI selection. Prior knowledge of pulsating characteristics associated with PPGs can also be advantageously used.

[0022] The apparatus, system, method, and computer program product according to embodiments of the present invention achieve the above objectives.

[0023] In a first aspect, the present invention relates to a method for extracting a photoplethysmographic signal indicating the physiology of a subject based on remote camera observation (for example, from a video stream obtained by said remote camera observation), for example, a method for determining a photoplethysmographic signal, and / or a method for determining an image region of interest for use in extracting a photoplethysmographic signal. The method includes acquiring a camera image (for example, a stream of video images) from a camera configured to monitor at least one body part of a subject. The method includes dividing a first region of the camera image representing an image portion in which tissue of interest is presumed to be present into a first subset of pixels, and dividing a second region of the camera image representing an image portion in which tissue of interest is presumed not to be present into a second subset of pixels, the second region being different from the first region.

[0024] The method further includes determining, for each subset of a plurality of first and second subsets (e.g., pruned), at least one dynamic characteristic value that represents the time signal dynamics of the subset (the pixels forming it) in a sequence of camera images.

[0025] This method involves clustering a first set of subsets (e.g., pruned) and a second set of subsets (e.g., pruned) based on at least one dynamic characteristic value in order to group subsets of pixels into clusters of similar signal dynamics. By using (substantially) the same clustering variables (e.g., peak frequency and / or magnitude, or one or more substantially equivalent values) for clustering subsets (e.g., patches) of the first region and for clustering subsets (e.g., patches) of the second region, information from estimated unrelated clusters in the second region can be taken into consideration when selecting one or more related clusters in the first region, as discussed below.

[0026] The method includes a step of selecting at least one cluster from clusters of subsets obtained by clustering a first plurality of subsets based on cluster size and / or cluster form, and the step of selecting a cluster excludes each cluster of the first plurality of subsets that substantially corresponds to the time signal dynamics of the clusters of the second plurality of subsets in the associated time signal dynamics.

[0027] The method further includes providing an output. The output includes a region of interest definition for defining a photoplethysmograph signal extracted from pixels forming the at least one selected cluster and / or a region of interest within the camera image, based on, for example, other sequences of images acquired by a camera.

[0028] In a method according to an embodiment of the invention, the subset generator is configured to identify anatomical key points within a sequence of camera images. The subset generator is configured to use the identified anatomical key points to determine a first plurality of subsets and a second plurality of subsets.

[0029] Anatomical key points are positions relative to a part of the subject's body. They can be, for example, the positions of joints, positions on the face, or indicative of other body parts. Currently, there are commercially available software that can automatically identify anatomical key points. Even video game systems such as Microsoft Xbox incorporate such software capabilities. Neural networks such as convolutional neural networks trained using images labeled with anatomical key points can also be used.

[0030] In the method according to embodiments of the present invention, the subset generator is further configured to determine a first plurality of subsets and a second plurality of subsets by fitting a template to identified anatomical keypoints in a sequence of camera images. The template can be fitted to and / or modified to provide the first plurality of subsets and the second plurality of subsets. This can provide an effective means of identifying the target region to be used for the first plurality of subsets and the second plurality of subsets.

[0031] In the method according to an embodiment of the present invention, the subset generator has a neural network configured to output anatomical keypoints in response to receiving a sequence of camera images as input. As described above, a neural network trained with images labeled with anatomical keypoints can be used. Various types of neural networks can be used. For example, a convolutional neural network, U-Net, ResNet, or a network with fully connected layers can be configured to receive images as input and output the locations of anatomical keypoints.

[0032] In the method according to embodiments of the present invention, at least one dynamic characteristic value has a first-order spectral peak, and the signal dynamics analyzer is configured to determine the first-order spectral peaks of a first subset and a second subset by performing a fast Fourier transform of the pixel values ​​to detect the first-order spectral peak. For each of both the first subset and the second subset, the average characteristic of the subset, such as brightness, contrast, position in the color spectrum, color change, or other characteristics, can be recorded as a function of time using multiple images. This set of values ​​can then be Fourier transformed. The frequency with the largest peak value becomes the first-order spectral peak.

[0033] In the method according to an embodiment of the present invention, the clustering processor is configured to cluster a first plurality of subsets and a second plurality of subsets according to their primary spectral peaks. In this example, each of the first plurality of subsets is clustered according to its primary spectral peak. Each of the second plurality of subsets is also clustered according to its primary spectral peak. This provides a convenient method for grouping or clustering subsets in a manner that can be used to generate a photoplethysmography signal.

[0034] In the method according to embodiments of the present invention, the primary spectral peak is limited to the heart rate band. This can be effective in reducing or eliminating false signals from the photoplethysmography signal.

[0035] In the method according to the embodiment of the present invention, the heart rate range is 5 beats / min to 300 beats / min.

[0036] In the method according to an embodiment of the present invention, the selector is configured to select the largest cluster among a first subset that has a primary spectral peak different from the largest cluster among a second subset. Since the second subset is not used to image the target skin, the largest cluster of the second subset is likely to be a noise signal rather than a photoplethysmography signal. Thus, this example can provide an improved means of reducing errors in the photoplethysmography signal.

[0037] In the method according to embodiments of the present invention, the selector is further configured to ignore clusters having spectral peak magnitudes exceeding a predetermined threshold during the selection process. This can also provide an effective means of reducing errors in the photoplethysmography signal.

[0038] In the method according to embodiments of the present invention, the first region, each second region, can be divided into first, each second, and a plurality of pixel subsets based on predetermined prior knowledge assumptions that the target tissue is unlikely to be present in the camera image, respectively, and a known setup of the camera for the volume in the space in which the target is located is given.

[0039] In the method according to embodiments of the present invention, selecting at least one cluster may include selecting the largest cluster from a first subset which is excluded or not excluded based on the correspondence of a second subset to a cluster.

[0040] In the method according to embodiments of the present invention, selecting at least one cluster may include determining a shape scale for each cluster in a first subset that represents the shape of the image region formed by the pixels (or patches) within the cluster. The selection may be based on the fact that the shape scale of at least one selected cluster matches a predetermined target shape of a region of tissue of interest, for example, an oval or ellipse representing exposed skin of the face or forehead of a subject.

[0041] In the method according to embodiments of the present invention, the selection of at least one cluster may be based on a ranking of clusters in a first subset (e.g., those not excluded based on the aforementioned correspondence to clusters in a second subset), wherein the ranking combines a score based on cluster size and a score based on shape scale. The ranking may optionally include a further score based on the cluster correspondence with clusters in a second subset (e.g., best match of clusters).

[0042] In the methods according to embodiments of the present invention, the first and second regions can be divided into a first plurality of pixel subsets and a second plurality of pixel subsets by forming the partitions of the first and second regions into subsets based on a spatial distance metric and / or based on image codomain distance and / or a combination thereof, so that each subset groups pixels that are spatially and / or numerically close. For example, in embodiments where different pixel subset sizes are considered (e.g., in the optimization strategies discussed below), the grouping of pixels into (first, second) subsets can be switched or the weights shifted from forming more or fewer subsets based on distance in a spatial domain to forming more or fewer subsets based on distance in an image codomain (e.g., pixel value domain).

[0043] In the method according to embodiments of the present invention, dividing a first region and each second region into a plurality of first and second pixel subsets may include image segmentation and / or image processing of images acquired by a camera, and / or at least one image and / or spatial information acquired by a further camera, and / or further spatial information sources (e.g., a 3D camera, a diagnostic imaging device, etc.) to determine the first and second regions.

[0044] In the method according to embodiments of the present invention, dividing the first region into a second region may include image subtraction between an image acquired with an object present and a blank image acquired without an object, and / or a similar image background compensation technique, in order to detect the first region and / or the second region.

[0045] A method according to an embodiment of the present invention may include at least the step of iteratively dividing a first region into a first plurality of pixel subsets (e.g., applying an optimization strategy), such that different iterations correspond to different values ​​of at least one optimization parameter representing at least the size of the pixel subsets. The method may further include evaluating a quality metric based on the pixel subsets obtained in each iteration to select a parameter value that yields a sufficient or optimal value of the quality metric Is, and using the plurality of pixel subsets obtained for the selected parameter value, such that, for example, clustering and thus the region of interest determined therefrom are also based on the subsets corresponding to the parameter selection.

[0046] In the method according to embodiments of the present invention, at least one optimization parameter may further include the size of the temporal observation window used to determine at least one dynamic characteristic value for each subset.

[0047] In the method according to embodiments of the present invention, the quality metric used in the optimization strategy may include or consist of a signal-to-noise ratio.

[0048] In the method according to embodiments of the present invention, determining at least one dynamic characteristic value may include performing a Fourier transform in the time domain to obtain one or more time-frequency characteristics.

[0049] Methods according to embodiments of the present invention may include pruning a first plurality of subsets based on predetermined criteria for one or more values ​​determined for each subset, in order to remove subsets that are unlikely to correspond to homogeneous regions of tissue of interest, and similarly, a second plurality of subsets may be pruned based on predetermined criteria (similar and / or similar). Thus, clustering can be applied to the pruned first plurality of subsets and the pruned second plurality of subsets.

[0050] In the methods according to embodiments of the present invention, the pruning may be performed based on predetermined criteria, the predetermined criteria may have one or more criteria for at least one value indicating movement associated with the subset, such as the average pixel value and / or the spread of pixel values ​​per subset, and / or subsets exhibiting excessive movement, non-uniform pixel value distributions, and / or, on average, pixel values ​​outside a predetermined target range. It will be understood that various alternative measures can be used for means (e.g., generally a measure of the centrality of pixel values ​​across the subset), spread (e.g., generally a measure of the variance of pixel values ​​across the subset), and / or movement (e.g., a measure that correlates with or is associated with movement, such as determined by optical flow).

[0051] The method according to embodiments of the present invention may include further pruning of the first and / or second subsets (which may have already been pruned in the first step described above) based on at least one dynamic characteristic value, before applying the clustering to the further pruned first and / or second subsets. The further pruning may be based on a predetermined criterion of 1 or a number of values ​​for the spectral energy of the largest spectral peak of the Fourier spectrum determined for each subset, and / or a value determined therefrom.

[0052] In a second aspect, the present invention relates to an apparatus for use in extracting photoplethysmographic signals indicating the physiology of a subject based on remote camera observation, for example, an apparatus for extracting photoplethysmographic signals and / or for providing a region of interest definition to be used to extract photoplethysmographic signals. The apparatus has an input unit for receiving camera images from a camera configured to monitor at least one body part of a subject.

[0053] The device includes a subset generator that divides a first region representing an image portion in the camera image in which tissue of interest is presumed to exist into a first subset of pixels, and divides a second region different from the first region representing an image portion in the camera image in which tissue of interest is presumed not to exist into a second subset of pixels.

[0054] The device includes a signal dynamics analyzer for determining, for each subset of a first subset and (for example, independently) for each subset of a second subset, at least one dynamic characteristic value that represents the temporal signal dynamics of the subset in a sequence of camera images received via the input.

[0055] This device has a clustering processor for grouping subsets of pixels into clusters having similar signal dynamics. It clusters a first set of subsets and a second set of subsets based on at least one dynamic characteristic value (e.g., clustering the first set of subsets independently of the second set of subsets, and vice versa). (Image) It is understood that clustering algorithms can generally be thought of as taking into account spatial relationships, such as related dynamic characteristics, as well as grouping subsets that are spatially close to each other together.

[0056] The device has a selector that selects at least one cluster from a first subset of clusters provided by a clustering processor, based on cluster size and / or cluster configuration. The selector is adapted to exclude from the selection each cluster of the first subset that substantially corresponds to the signal dynamics of the clusters of a second subset in the associated signal dynamics.

[0057] The device has an output unit that outputs a region of interest definition that defines pixels forming at least one cluster selected by a selector for use in extracting a photoplethysmography signal from a camera image acquired by a camera, and / or outputs a photoplethysmography signal extracted from the region of interest in the camera image received via the input.

[0058] The apparatus according to embodiments of the present invention may have a camera suitable for PPG signal extraction, such as an infrared camera, a color camera, and / or a camera having a predetermined PPG filter, such as a selective green filter. It is understood that the apparatus may further (optionally) have a light source that illuminates the object (or at least a body part of interest) using light containing appropriate spectral components corresponding to the sensitivity of the camera. The camera may consist of one camera or a combination of multiple cameras.

[0059] An apparatus according to an embodiment of the present invention may have a photoplethysmography signal extractor for extracting a photoplethysmography signal provided as an output from a region of interest in the camera image (sequence of images).

[0060] In an apparatus according to an embodiment of the present invention, the subset generator can be configured to divide a first region into a first set of pixel subsets and a second region into a second set of pixel subsets, based on a known camera setup for the spatial volume in which the object is placed, and based on predetermined prior knowledge assumptions about where tissue of interest may or may not exist. For example, the regions can be defined via a user interface, pre-configured (e.g., stored in configuration memory and / or hardcoded), or retrieved from an external data storage device stored during a camera setup calibration procedure, for example.

[0061] In the apparatus according to embodiments of the present invention, a selector for selecting at least one cluster can be adapted to select the largest cluster that is not excluded from a first subset based on the correspondence of clusters in a second subset.

[0062] In an apparatus according to an embodiment of the present invention, the selector can be adapted to determine a shape scale representing the shape of each cluster in a first subset, and to select at least one cluster based on the fact that the shape scale of the selected cluster matches a predetermined target shape of the region of interest.

[0063] In an apparatus according to an embodiment of the present invention, the selector can be adapted to determine a ranking of clusters in a first subset (i.e., those not excluded based on the correspondence to clusters in a second subset), the ranking combining a first score based on cluster size and a second score based on cluster shape. Thus, the ranking can be used to select, for example, the highest or top N (e.g., 2, 3, ...) highest-ranked clusters to determine the selection. The ranking may optionally include other scores based on the correspondence between the cluster and clusters in a second subset (e.g., best match of clusters).

[0064] In an apparatus according to an embodiment of the present invention, the subset generator can be adapted to divide the first region into a first plurality of pixel subsets and the second region into a second plurality of pixel subsets by forming partitions of the first region and the second region into blocks or patches based on a spatial distance metric and / or into subsets based on image codomain distance and / or a combination thereof, wherein each subset groups together pixels that are spatially and / or numerically close.

[0065] In the apparatus according to embodiments of the present invention, the subset generator can be adapted to determine the first and second regions by performing image segmentation (and / or other appropriate processing) of a reference image and / or reference spatial information received via the input, the reference image and / or reference spatial information being acquired by a camera, another camera, and / or other spatial information source and received via the input (11). For example, additionally or alternatively, the subset generator can be adapted to perform image subtraction between a reference image acquired with the subject present and a blank image acquired without the subject present (e.g., as one of the other appropriate processing or an element thereof), and / or apply image background compensation techniques similar to those used in determining the first and second regions, for example, in combination with image segmentation of the image after irrelevant background has been removed (optionally).

[0066] An apparatus according to an embodiment of the present invention includes an optimizer, in which the subset generator is configured to iteratively divide a first region into a plurality of first pixel subsets for different values ​​of at least one optimization parameter representing the size of at least one pixel subset, and the optimizer is configured to calculate and evaluate a quality metric based on the pixel subsets obtained in each iteration for different values ​​of at least one optimization parameter, and to select a parameter value that yields a sufficient or optimal value of the quality metric. Accordingly, the selected parameter value, and the corresponding divisions of the first and second regions into first and second plurality of subsets, respectively, can be used by the clustering processor and the selector.

[0067] In the apparatus according to an embodiment of the present invention, the optimizer may also use the size of the observation time window used by the signal dynamics analyzer as a component of the at least one optimization parameter.

[0068] In the apparatus according to an embodiment of the present invention, the quality metric calculated by the optimizer may include or consist of a signal-to-noise ratio.

[0069] In the apparatus according to embodiments of the present invention, the signal dynamics analyzer can be adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency characteristics for use in determining the at least one dynamic characteristic value.

[0070] Apparatus according to embodiments of the present invention may have a subset eliminator for pruning a first plurality of subsets based on predetermined criteria for one or more values ​​determined for each subset, and for pruning a second plurality of subsets based on the same and / or other predetermined criteria, in order to remove subsets that are unlikely to correspond to homogeneous regions of the target tissue, and a clustering processor may be adapted to cluster the pruned first plurality of subsets and the pruned second plurality of subsets, respectively.

[0071] In an apparatus according to an embodiment of the present invention, a subset eliminator can be adapted to prune a first and / or second subset based on a predetermined criterion including one or more criteria for the average pixel value per subset, the spread of pixel values ​​per subset, and / or at least one value indicating movement associated with the subset, thereby adapting the subset eliminator to remove pixels exhibiting excessive movement, non-uniform pixel value distributions, and / or pixel values ​​that, on average, fall outside a predetermined target range.

[0072] In the apparatus according to an embodiment of the present invention, the subset eliminator can be further adapted to apply further pruning of a first and / or second subset based on at least one dynamic characteristic value, and the clustering processor can be adapted to cluster the further pruned first subset and the further pruned second subset, respectively.

[0073] In the apparatus according to embodiments of the present invention, the subset eliminator can be adapted to apply this further pruning, which is based on at least one predetermined criterion for the spectral energy of the largest spectral peak of the Fourier spectrum determined for each subset, and / or a value determined therefrom.

[0074] In a third aspect, the present invention relates to a diagnostic imaging system having an examination zone, the system comprising a camera for acquiring images from a subject when the subject is being examined while positioned in the examination zone, and an apparatus according to an embodiment of the present invention operably connected to the camera to receive camera images from the camera as input.

[0075] In a fourth aspect, the present invention relates to a computer program that performs a method according to an embodiment of the present invention when executed by a computing device.

[0076] The independent and dependent claims describe specific and preferred features of the invention. Features of dependent claims may be combined with features of the independent claims and other dependent claims as appropriate, and are not necessarily only expressly described in the claims. [Brief explanation of the drawing]

[0077] [Figure 1] A diagram illustrating a method according to an embodiment of the present invention. [Figure 2] A diagram showing an apparatus according to an embodiment of the present invention. [Modes for carrying out the invention]

[0078] The drawings are schematic and not limiting. Components in the drawings are not necessarily represented to scale. The present invention is not necessarily limited to the specific embodiments shown in the drawings.

[0079] Detailed description of the embodiment

[0080] Notwithstanding the exemplary embodiments described below, the present invention is limited only by the appended claims, which are expressly incorporated into this detailed description, and each claim, and each combination of claims permitted by the dependent structures defined by the claims, forms a distinct embodiment of the present invention.

[0081] When used in the claims, the term “comprising” is not limited to, but does not exclude, additional features, components, or steps, such as those described below. Therefore, it expresses the presence of the features mentioned without excluding the presence or addition of one or more features.

[0082] This detailed description presents various specific details. Embodiments of the present invention can be implemented without these specific details. Furthermore, well-known features, components, and / or steps are not necessarily described in detail for the sake of clarity and brevity of this disclosure.

[0083] In a first aspect, the present invention relates to a method for determining a photoplethysmography (PPG) signal indicating the physiology of a subject based on remote camera observation, and / or a method for determining an image region of interest (ROI) for extracting such a PPG signal by means of a prior art method of receiving an ROI (e.g., as an image map, as a mask image, or as a corresponding masked image in a video stream) and subsequently determining the PPG. Such a method for extracting a PPG signal is well known in the art, provided that a suitable definition of the available image content (e.g., a mask or other suitable description of the area of ​​interest to be used) is provided, and such a method has not been discussed in detail, but it is understood that a method according to an embodiment of the present invention may have such a well known signal extraction step to obtain and provide a high-quality PPG signal (e.g., having a good signal-to-noise ratio and / or being robust to motion and / or other artifacts) based on the area of ​​interest determined according to an embodiment of the present invention.

[0084] As is well known in this field, photoplethysmography (PPG) is an advantageous, low-cost optical technique for detecting microvascular blood volume changes associated with heart rate dynamics. It can advantageously provide non-invasive imaging, for example, using camera images. The raw PPG signal contains the pulsatile component (i.e., the "AC" component) synchronized with the cardiac cycle, i.e., the PPG signal of interest to be extracted. Frequencies that do not show the signal of interest can be readily removed by filtering, for example, irrelevant slow ("DC" and / or lower frequency) components, such as components due to average illumination conditions, and / or potential changes due to slow motion, respiration, thermoregulation (e.g., especially when using images in the infrared region), and / or other such factors. Higher frequencies outside the physiological range of interest, which may substantially relate to noise, can be filtered out as well. Therefore, once a set of pixels in (and each) video frame from which a signal should be extracted—that is, a set of pixels corresponding to the region of interest definition provided by the method according to the embodiment—is established, a useful PPG signal can be extracted from the live video stream provided by camera surveillance by applying conventional signal processing techniques known in the art (e.g., machine learning-based filtering and / or more advanced techniques).

[0085] Figure 1 shows an exemplary method 100 according to an embodiment. This method includes acquiring a camera image, such as an image stream (115), and includes capturing a video or live stream from a camera configured to monitor at least one body part of an object. The camera image may be acquired by directly observing the object, for example, the object may be in the direct line of sight of the camera, or it may be observed indirectly, for example, using one or more reflective surfaces in the optical path. Lenses and / or other optical elements may be optionally used to optimize the camera's field of view, for example, to magnify the relevant body part, to focus the image, and / or for similar purposes.

[0086] The camera can be an infrared camera (e.g., an NIR camera operating in the near-infrared range, e.g., a camera sensing in the 800-850 nm range), thereby enabling the extraction of the PPG signal without requiring adjustment of (visible) illumination conditions and / or being affected by ambient light (in the visible spectrum). Since the human eye is not sensitive to infrared light, the subject being monitored is not affected by bright light that is unpleasant, intrusive, and / or uncomfortable (or painful). However, this embodiment is not limited thereto. For example, it is also possible to extract the PPG signal from imaging in the visible light region. For example, the light absorption of (de)oxygenated hemoglobin is stronger in the (near) infrared wavelength range than in visible light (e.g., red light), and infrared light can penetrate deeper into the skin (e.g., the effect of melanin in the skin is not as pronounced). Furthermore, infrared cameras based on semiconductor pixel detectors, for example, can have higher sensitivity in the NIR range (e.g., 800-850 nm) than wavelengths above 900 nm, which can result in a higher signal-to-noise ratio and / or better performance when the body part of interest is not ideally illuminated. Other optical ranges (other than IR or NIR) can also be used for PPG signal detection, as it has been observed that video images in the red or green wavelength range can also provide sufficient information to detect changes in intensity used.

[0087] This method may include acquiring camera images from a camera configured to monitor a body part of a subject during a medical and / or diagnostic examination or intervention. Examinations may include, for example, magnetic resonance imaging, computed tomography, positron emission tomography, or single-photon emission computed tomography. Examples of interventions include treatments and related procedures such as surgery and radiotherapy. Embodiments may also be applied to general subject or patient monitoring in, for example, intensive care units, patient wards, or non-medical settings such as nursing homes. In particular, the PPG signal determined by this method may be used to monitor the health and / or condition of a subject and / or to control systems or devices used in examinations or interventions. The body part of interest may include, but is not limited to, the face or its parts, such as the forehead (for example, skin from another body part, such as the arms, chest, legs, or abdomen, may be used).

[0088] The method includes dividing at least one first region in a camera image, for example in a reference image, into a subset of pixels, for example, patches (101). The subset of pixels, for example, patches, can form a partition of the first region (for example, such that the entire first region is covered by a collection of patches that do not overlap). However, the latter is not necessarily required, for example, some overlap between patches is not necessarily excluded.

[0089] Furthermore, in addition to step 101 of dividing the first region into patches, according to embodiments of the present invention, a second region in the camera image is similarly divided into a subset of pixels (e.g., patches). The first and second regions are different image regions, for example, typically separate (and non-empty) parts of an image, for example, there is no overlap between the first and second regions, each region is not equal to the whole image, and each region is not empty. However, there may be slight overlaps, such as pixels of a shared boundary contour between the two regions, for example, negligible. For example, the intersection between the two regions may be less than 10%, preferably less than 5%, more preferably less than 1%, for example substantially 0%, relative to the aggregate of both regions in terms of size, when evaluating the ratio of the number of pixels in the intersection to the number of pixels in the aggregate.

[0090] The first area represents image content in which the organization of interest is presumed to exist, while the second area represents image content in which the organization of interest is presumed not to exist (for example, unlikely to be found or unlikely to be found).

[0091] The combination of the first and second regions can form an entire image; for example, the two regions can partition an image or cover the entire image. However, this is not always the case. For example, background image content (e.g., areas of the image where the subject is unlikely to be present) can be ignored, while the first region can define a zone of the image where the tissue of interest (e.g., exposed skin) is likely to be visible (e.g., in any appropriate form such as well-lit), and the second region can define a zone where the tissue may not be clearly visible (however, the image content is not necessarily unrelated to the subject itself). Such a zone where the tissue of interest may not be clearly visible may include parts of the subject's body covered by clothing, glasses, blankets, etc., in which case a clear view is likely to be found where bodily features (eyes, nose, hair, fingers, nails, navel, nipples, skin folds, etc.) may be potentially obstructed by the instrument and / or cause errors when inferring the PPG signal.

[0092] A more conventional "biological skin" approach, in which image patches are analyzed to obtain useful signal content, can be extended according to embodiments of the present invention by analyzing a first region having potentially relevant image content and a second region having potentially irrelevant image content. Thus, in addition to analyzing candidate biological skin patches, patches of non-skin regions (a priori known, i.e., presumed) are also taken into consideration.

[0093] For this initial region segmentation step, a reference image can be used, which may be a first image at the start of the procedure, a random image from a (e.g., short) sequence of images, an image selected based on a quality metric (e.g., to eliminate the effects of initial stabilization of the camera acquisition process, e.g., autofocus, gain calibration, etc.), and / or a set of several images constructed from multiple (short subsequences) of camera images, e.g., by averaging, etc. However, since the object and camera can be considered substantially static (i.e., preferably the camera is fixed and preferably the movement of the object is avoided), the selection of such a reference image is not necessarily required, and for example, segmentation into patches can be applied, at least roughly, to image frames acquired any later (and / or earlier) within the same session, and can also be considered applicable to any such session if based on entirely predetermined assumptions (e.g., always using the camera in the same spatial configuration and always positioning the object in the same manner relative to the camera with respect to the session).

[0094] A first region, and / or a fixed set of patches covering the first region, can be predetermined and defined, for example, as a function of the locations where the tissue of interest (of the subject's body) is estimated to be approximately present. A second region, and / or a fixed set of patches covering the second region, can be predetermined and defined, for example, as a function of the locations in the image where the tissue of interest is estimated not to be present. Similarly, patches can be defined in a fixed manner, regardless of the session and / or subject to which the procedure relates, for example, as predetermined sections of a predetermined first region (second region) in a camera image. In other words, a predetermined or case-by-case (independently) determined mapping of camera pixel coordinates (e.g., image grid) can be used to determine the first and second regions and / or patches of each region without requiring a reference image. Region definitions and / or (region-by-region) patch definitions can be established independently of a particular camera image and can therefore be considered applicable to any acquired image, for example, by definition, to each of the single image frames in a sequence acquired to extract PPG information. Region definitions and / or patch definitions may be based entirely on prior knowledge of the camera (and subject) setup, and / or may be determined independently (for example, based on another camera system and / or spatial data acquisition system).

[0095] This is particularly advantageous in applications where the camera is largely fixed, or where the camera is positioned, oriented, focused, zoomed, and / or otherwise configured using a predetermined reference frame and / or parameters (e.g., systems using fixed positioning and / or motorized / automatic configuration in a scanner room), and where the object is generally positioned, oriented, etc., in the same or generally similar manner with respect to the camera or a reference frame having a known relationship with the camera's reference frame (perhaps taking into account the aforementioned parameters such as the orientation and position of the automatically controlled camera).

[0096] For example, in an exemplary application of use in a medical imaging room (or other medical facility where a patient is to be monitored), the camera can view a known area of ​​the room (or more precisely, a cone in space), and the patient can be positioned in a known location within the room (taking into account patient posture for parameters and / or procedures of the medical imaging system, such as patient table translation), thereby making it easy to infer an approximate area in the camera image where the tissue of interest is located, and similarly, a second area where the tissue is unlikely to be present (or is presumed not to be imaged clearly enough) can also be determined.

[0097] Alternatively, the first region (and / or a patch set dividing the first region) and / or the second region (and / or a patch set dividing the second region) can be determined by image segmentation or similar image processing techniques. For example, simple, e.g., two-component segmentation can be performed to separate the body and / or body part and / or tissue of interest (e.g., exposed skin) (first region) from its background (second region). Segmentation of three or more components can be used to detect skin (or other tissue of interest, i.e., the first region), (a) irrelevant parts of the body (second region), and (a) background image regions (e.g., a room and / or equipment; i.e., forming a third component and / or some other components that can be completely ignored). Other segmentation techniques are not necessarily excluded, for example, expected image content is used to predefine several components that can be identified visually (and algorithmically), at least one of which (combinations are not excluded) can be associated with a first region, and at least one of which (or more) can be associated with a second region. For example, multiple (i.e., two or more) image segments can be expected based on assumed pixel properties (e.g., intensity values), and explicitly taking into account such non-uniform second regions and / or (optional, further ignored) background regions (by accommodating additional segmentation image segments) can, to the advantage, improve the quality of the obtainable image segmentation.

[0098] As another example, using a blank image (without any subject) of the same environment, a first region can be determined by background subtraction, and a second region can be defined as the remainder (possibly including normalization procedures and / or more advanced techniques). Combinations of background subtraction and segmentation, and / or other image processing techniques can also be used. For example, a “blank” background image can be subtracted to remove static image components, and segmentation can be performed on the result to identify a first region consisting of tissue of interest (e.g., exposed skin) and a second region consisting of other non-static and / or variable elements in the image, such as body, hair, eyes, coverings of equipment (e.g., tubes, cables, breathing masks, etc.) and / or other such irrelevant image content, in particular, where some dynamic changes in image content can be expected in this second region, in contrast to the static content removed by image subtraction. It is understood that the image subtraction described above is more complex than simple arithmetic matrix subtraction, for example, it can be achieved by any equivalent approach that takes image normalization into account and / or in general applies any appropriate background compensation technique.

[0099] Image segmentation (or other appropriate image processing) is not necessarily limited to images acquired by the camera system, nor does it necessarily have to use only the images used for subsequent PPG extraction to be performed. For example, other cameras can be configured to provide better image contrast and / or other advantageous characteristics, which allows the subject (more generally, the body part or tissue area of ​​interest) to be numerically separated from unrelated parts of the subject, such as clothing or covered parts of the body. For example, a camera used for PPG signal extraction can be optimized to provide good sensitivity for detecting subtle PPG signals of a subject (e.g., in the infrared spectrum and / or another specific part of the spectrum, e.g., the green light band), but is not necessarily ideal for image segmentation, for example, for identifying the subject tissue, such as skin, from other image content. Therefore, as long as the segmentation results can be converted to a reference frame of the camera used for PPG signal extraction, another camera (or camera) more suitable for this purpose, such as a color camera or a depth camera, can be used. Even if the deterministic relationship between camera reference frames is not known a priori, it will be understood that other techniques for transforming the segmentation map can be used, for example, by minimizing the mutual information between images acquired by a PPG camera system and images acquired by a further camera system used for image segmentation (and / or to determine the first and second regions by similar processing techniques). Similarly, the information source for inferring the first and second regions by segmentation is not necessarily limited to a camera system in the narrowest sense, but can include or consist of any suitable means for acquiring spatial information, including, for example, a medical imaging system that may be intended for use during PPG monitoring of a subject.

[0100] While segmentation is mentioned, it should be understood that this does not necessarily exclude more advanced approaches, such as fitting a geometric model of the object (or related body parts or parts of the object) to image data and / or evaluating a machine learning model trained to determine appropriate (approximate) segmentation (i.e., generally descriptions of first and second regions, and / or direct descriptions of the patches forming each of the said regions) as output.

[0101] In the “biological skin” procedure and / or similar procedures, it is typically assumed that the pulsating component of interest, i.e., the PPG signal, is prominent from the measurement noise, i.e., sufficiently detectable and identifiable in the acquired camera pixel signal (e.g., after appropriate processing). In particular, averaging across different pixels of each patch (or, more generally, applying appropriate summarization calculations (such as a measure of statistical centrality)) can sufficiently reduce the sensor's quantization noise and / or other noise components, allowing usable signal content (and associated patches) to be identified from noisy content in irrelevant and / or inappropriate patches where the PPG signal is too weak, contaminated by strong noise, confused by other irrelevant signals, or simply absent. Furthermore, as will be described in detail below, techniques for analyzing the signal dynamics (changes over time) are typically applied, resulting in a sufficiently long observation time required to detect the signal component of interest. For example, an appropriate window size for Fourier analysis must be selected, but one that does not significantly impair temporal resolution or increase sensitivity to errors.

[0102] Instead of using predetermined (and potentially arbitrary) patch sizes in the step of dividing the first (and second) region into patches, the patch size and / or observation time window (e.g., Fourier window) used for analyzing signal dynamics can be optimized, for example, to obtain a good signal-to-noise ratio (SNR) according to embodiments of the present invention. For example, temporal variations that may be due to illumination conditions (e.g., when natural illumination or low-quality light sources are used), environmental stability, and / or motion may be present in the raw pixel signal. Furthermore, physiological differences (intra-subject and / or inter-subject) such as natural heart rate variability may affect the optimal selection of patch size (and / or time window) between sessions. Thus, intrinsic differences (differences related to heart rate cycles) and extrinsic differences (e.g., illumination conditions) at different times and / or in different environments between sessions, for example, between different applications of the method (e.g., used to detect PPG signals from different patients), may lead to (e.g., optimal or optimized) parameter selection being session-dependent. Therefore, such parameter optimization according to the embodiment can advantageously improve accuracy, robustness, stability, and / or generally, the quality of the processed PPG signal output.

[0103] The optimal, or at least appropriate or good, selection of the observation time window (such as the FFT window) and the patch size largely depend on the expected heart rate and heart rate variability in each individual case, for example, but unfortunately these may be unknown (or clearly unknown) beforehand. Indeed, the goal of the method according to the embodiment may be to accurately measure and / or monitor this heart rate and / or other cardiac parameters. In some situations, a larger patch size may be desirable, for example, when shortening the FFT window or reducing noise by (spatial) averaging, but a patch size that is too large can also be disadvantageous. The probability of finding a uniform patch (e.g., one corresponding to substantially exclusively exposed skin area) may decrease as the patch size increases. Therefore, it will be understood that there may be an optimal value between these extremes, which is not known beforehand and will vary from case to case.

[0104] When referring to patch size, it is understood that for a particular selection of patch size parameters, the patch size is not necessarily strictly constant across all patches. In other words, for a particular patch size parameter, there may be some variation in patch size. For example, the step of dividing a first region (and / or a second region) into patches can be implemented in a manner that allows for some variation, and the size parameter, for instance, more accurately determines the average (or median, etc.) patch size used (in the iterative process of that parameter selection).

[0105] Even when a fixed (relatively simple) method is used to divide a region into patches with predetermined deterministic (uniform) patch sizes, such as dividing the region with an orthogonal block pattern of a given block size, irregular boundaries between the first and second regions can result in some edge regions having fewer (or more) pixels than typical, more centrally located (block-like) patches.

[0106] However, a uniform patch size (per iteration, for a corresponding specific selection of patch sizes) can be (arbitrarily) achieved, for example, when processing techniques that rely on (or are sensitive to) a uniform patch size (number of pixels per patch) are used. For example, the (first / second) region definitions can be adjusted to fit all patch sizes considered without splitting patches near the boundaries, and the region can be defined at a resolution (scale) corresponding to the least common multiple of the different patch scales being evaluated. As another example, the first (or second) region may not necessarily correspond precisely to the target organization (or irrelevant image content), and for example, the region may be merely the best guess based on available prior information, so each (uniformly sized) patch (e.g., block) can be assigned to either the first or second region without major problems or significant consequences, based on a majority vote.

[0107] Therefore, the PPG signal and noise components (and / or, generally, interference contributions) of interest may vary between sessions, for example, due to differences in the pulsity of at least the cardiac response of each subject. In the method according to embodiments of the present invention, the patch size (e.g., used in steps 101, 111 for dividing the region) and / or the time window size for analyzing the time variation (e.g., the FFT window size) are selected (123), optimized and / or adjusted, for example, based on the acquired image data (directly or indirectly), so as to be large enough to avoid or reduce adverse effects associated with averaging over a wider region and / or analyzing the signal over longer time frames, for example, while exceeding the noise floor. Such parameter selection 123 can be implemented in a variety of ways.

[0108] For example, the step of dividing at least a first region into patches (101) (and optionally, the step of dividing a second region into patches (111), and / or one or more other steps described below) can be repeated for different values ​​of at least one optimization parameter representing the size of the pixel subset (125), for example, the step can be repeated for different patch sizes. Other parameters, such as the time window size for signal dynamics analysis (e.g., Fourier analysis), can also be adjusted (e.g., in addition to size optimization). After each iteration, the fitness of the tested parameter selection is indicated by calculating a quality metric based on the acquired image data organized into the set of patches of that parameter selection (directly or indirectly), for example, based on the intermediate output of the steps performed in each iteration and / or based on the output of the region of interest obtained in the iteration.

[0109] For example, the entire procedure can be performed in each iteration, for instance, until the region of interest is determined, and the quality metric can be determined directly from the region of interest acquired for a particular parameter selection (and / or from the PPG signal extracted from the video stream acquired by the camera using its ROI mask). Thus, as indicated by the quality metric, the best region of interest may be selected from multiple trials with different parameter selections, and / or the parameters may be adjusted for use in each subsequent iteration based on the output of one or more previous iterations. The quality metric may include, but is not limited to, an estimate of the signal-to-noise ratio (SNR) of the extracted PPG signal for the acquired ROI, and / or may become more complex by considering various factors such as ROI size, SNR, signal contrast, and / or one or more known measures of signal quality (which are general), and / or (in particular) PPG signal quality.

[0110] As an advantageous and efficient approach, a simpler quality metric can be used that does not require the execution of all steps, and such a quality metric can be calculated, for example, based on intermediate results, for example based on patches obtained only in the domain partitioning step, for example based on patches after applying the pruning step described later, or based on the dynamic parameter estimation (e.g., Fourier analysis) step described later. It will be understood that the domain partitioning (domain partitioning 101 and / or 111) step and / or pruning step can be performed relatively efficiently compared to the more computationally intensive signal dynamics analysis (Fourier analysis) and / or steps that depend on it. Thus, advantageously, the quality metric can be calculated efficiently based on the pruned patch set (first domain or first and second domains) without considering dynamics. However, for example, especially if the time analysis window should also be optimized, the quality metric can also be based on the intermediate output of other steps that take into account (multiple) characteristics of the signal dynamics determined for (e.g., selected) patches. For example, signal dynamics can be characterized relatively efficiently using, for instance, dedicated Fourier transform hardware, multiprocessors, and / or GPUs, and consequently, the use of such information in (e.g., iterative) parameter optimization is not necessarily prohibited.

[0111] Furthermore, in the optimization step, instead of a high-quality time dynamics analysis that may be used in a later step of the method described later, and instead of calculating the entire Fourier spectrum (in detail), for example, a simplified or approximate analysis of time dynamics can be used by detecting signal components at only specific frequencies, or only some specific frequencies and / or only specific spectral bands, or only some spectral bands. For example, such limited information is sufficient to calculate (or use as an element in) a quality metric, and is sufficient to give at least a rough indication of signal quality with respect to the expected time dynamics. To illustrate this, the spectral power of a predetermined frequency (or frequency band) can be determined quickly and / or easily, and this frequency (band) is estimated to contain at least some cardiac cycle-related activity (e.g., based on mean heart rate), while in the more detailed analysis discussed below (e.g., after parameter selection), the peak frequency and peak signal intensity can be determined in detail by more demanding calculations.

[0112] Variations of this approach can easily be attempted, such as normalizing the estimated spectral power (or other appropriate value) of the relevant frequency (band) by similarly calculated values ​​for frequencies or bands where no relevant (i.e., cardiac-related) activity is expected.

[0113] In another exemplary approach, scale-space analysis can be used, for example, where a signal is downsampled (in the time dimension) by various factors (e.g., a scale pyramid of powers of 2), and as a result, the mean, standard deviation, and / or other such measures of the signal (e.g., already averaged per patch) can be used as surrogates for signal dynamics (e.g., dependent on the corresponding time scale) in different (approximate) frequency bands.

[0114] The quality metric used in parameter selection (e.g., optimization) 123 can generally represent the fitness of the parameter selection, e.g., the patch scale and / or time analysis window in which it is calculated. Thus, the parameter selection corresponding to a higher (or highest) quality metric value can also be selected for use in determining the region of interest, and as a result, the PPG signal of interest can ultimately be extracted. Obviously, the quality metric can also be defined as a cost function, which is entirely equivalent, for example, by defining it such that a lower value indicates a better parameter selection.

[0115] If a sufficient number of patches (e.g., at least a predetermined number) (generally subsets of pixels, i.e., “actual,” spatially connected patches, and / or “virtual,” aggregated patches, as described below) render consistent signal intensity within an expected (e.g., predetermined) range, the quality metric may have, for example, a maximum value (or, conversely, the corresponding cost function may have a minimum value). Signal intensity can be estimated as an approximate indicator of PPG signal intensity, for example, based on the mean pixel value and / or deviation, based on the maximum spectral peak amplitude value, and / or based on the intermediate results of the methods discussed further below, or it can be estimated directly, for example, by extracting a (e.g., short-duration) PPG signal using parameter selection and directly estimating the signal-to-noise ratio of the PPG signal for that parameter selection.

[0116] The expected predetermined range of signal intensity can be determined by the expected PPG signal characteristics, taking physiology into consideration. This can be achieved, but is not limited to, by testing the signal intensity values ​​of each patch against a normal distribution, calculating the Z-score of the mean value (e.g., maximum spectral peak intensity) associated with each patch against a normal distribution with parameters (mean and standard deviation) corresponding to the predictive distribution (physiologically expected signal), and summing (or averaging, or combining in another appropriate way) the obtained Z-scores over all patches within a region (first region).

[0117] Therefore, the total score increases as more patches fit the expected curve. Many alternatives can be considered, for example, a test that compares the set of (pixel-based) values ​​within a patch to a known (i.e., estimated) distribution instead of the patch mean (e.g., comparing the sample to a target distribution using a Student's t-test or F-test). Other types of distributions can also be used for the test (this may include empirical distributions, such as histogram-based empirical distributions). Furthermore, various alternatives to the patch-by-patch or pixel-by-pixel characteristics on which such scoring and / or tests are based can be considered, such as signal intensity, pixel intensity, signal-to-noise ratio, spectral peak size, spectral peak size ratio, etc.

[0118] The use of pixel intensity values ​​(or average intensity values ​​per patch) has the advantage of allowing the quality metric to be calculated by performing only the step of dividing a first region (and optionally a second region) for each specific parameter selection under test. Thus, this can be performed very efficiently (specifically, across multiple parameter test iterations). In another exemplary approach, a pruning step could be performed in each iteration to remove irrelevant patches, and / or a step could be performed to characterize the signal dynamics so that, for example, the signal dynamics per patch can be taken into consideration in the parameter selection process (e.g., for calculating the quality metric).

[0119] In a relatively simple approach, the region of interest can be determined for each parameter space point over a range of parameters (e.g., patch size) or over a parameter grid of parameters (e.g., time window and patch size), and the best region of interest can be used as the output and / or for PPG signal extraction, i.e., a simple line / grid search optimization can be used. Parameter points can be sampled equidistantly along lines or across grids (e.g., in a Cartesian grid scan), or vary (e.g., sampled more densely in regions of the parameter space where the optimal selection as a whole is more likely to be obtained). The search strategy can optionally take into account the characteristics and / or specific properties of the parameters, and for example, multiple solution search and / or hierarchical search can be used to find the optimal (or at least well-tuned) parameters. When two or more parameters are being optimized, it will be understood that the combination can be optimized jointly (e.g., as a parameter vector for optimization) or separately (e.g., in nested optimization procedures).

[0120] More advanced techniques may use, for example, estimates of the parameter gradient and / or the Jacobian matrix of the parameter to use results from one or more previous iterations to determine the next sampling point.

[0121] In yet another strategy, the initial region of interest can be estimated using the first run, or several initial runs, from which the heart rate can be estimated (e.g., using ROI to extract the PPG signal over a short time window), and then this information can be used to obtain a better estimate of appropriate parameter selection, from which the procedure can be restarted to obtain a better ROI definition. This can be repeated multiple times if deemed necessary, or it can be repeated for only one refinement step. It will also be understood that signal noise can also be considered, for example, using estimates of heart rate and noise (and / or other quality measures that can be derived from the observed PPG output and / or ROI) to search for appropriate parameter selection in the lookup table by empirical relationships, etc. Alternatively, a machine learning strategy can be used to propose optimal values ​​for the parameter space based, for example, the results of a previous run (e.g., the ROI obtained from the previous parameter selection, i.e., an iterative approach) and / or an input vector describing the details of the use case (e.g., patient demographic data, room conditions, time, information about the procedure to which PPG extraction is applied (e.g., an identifier for the medical imaging procedure), and / or other data that can be implicitly used to infer appropriate parameter selection by an algorithm trained on a suitable example).

[0122] The embodiment can typically divide the first / second region into adjacent continuous regions, for example, to form a section of a single connected image patch, but several modifications and / or alternatives can be contemplated. For example, instead of dividing the region into a single connected area (a patch in a more conventional sense), or in addition thereto, alternatively (or in addition thereto), different pixels can be grouped together if their pixel values ​​are similar, instead of being spatially close to each other. Spatial grouping and value region grouping of pixels can also be combined to form “patches,” where each patch consists of pixels that are relatively close to each other in value space (i.e., of the same or similar pixel values), and does not necessarily require a predetermined shape (e.g., a block, square, rectangle, etc.) or connectivity.

[0123] For example, a (spatial) distance metric (not necessarily in a Cartesian coordinate system, but such as the Manhattan norm, etc., but not limited to these examples) can be combined with a (value-space) distance metric (such as the absolute difference between pixel values, etc., but not limited to these examples), for example by weighted summation, and the (first or second) region can be divided into pixel sets to optimize the combined metric, thereby obtaining, for example, a target number of pixels per set while minimizing the metric.

[0124] Determining patches for each region using only value-space (image codomain) distance, or in combination with spatial (image domain) distance (generally grouping pixels into multiple subsets), is particularly advantageous when applied to the parameter optimization strategies described above, but is not necessarily limited to this. Increasing patch size, for example, by increasing the number of pixels per subset "patch" (at least on average), increases the likelihood of forming heterogeneous patches, and, for example, it becomes less likely that a single pixel group (subset or patch) will contain only the tissue of interest (e.g., skin). Therefore, grouping pixels into patches using only spatial domain has limitations when the patch size is large, because beyond a certain size, the quality of the patch deteriorates as the desired signal, such as skin-related signals, becomes contaminated by pixels with poor or irrelevant signal content assigned to the same patch.

[0125] When the patch size (i.e., subset size) is large, it may be desirable to switch from a pixel combination based on spatial adjacency to another metric based on pixel output values ​​(e.g., an intensity-based metric). Such a switch can be done discretely, but it can also be adjusted continuously by changing the weights of the combined spatial distance and value distance metric.

[0126] Alternatively, patches can be formed based on spatial neighbors, but if they exceed a predetermined size threshold, for example, larger patches can be formed by combining smaller (spatial neighbor-based) patches based on an alternative value-space metric. For example, four primitive patch blocks can be selected based on the similarity of pixel values, without necessarily requiring spatial proximity. The latter can enable efficient implementation; for example, native (smaller) patches can be formed by dividing the at-hand image region into blocks with very low computational cost (e.g., forming a Cartesian grid across the region), while these primitive patches can be combined based on their intensity into a subset of the desired (larger) target size, which can also be done with reasonable computational cost (e.g., considering the simple computations performed on a relatively small number of entities, i.e., primitive patches). For example, but not limited to, a suitable combination of patches can be found by using only calculations on a relatively small NxN matrix (or only its upper / lower triangular portions, considering matrix symmetry) of the absolute average (averaged per primitive patch) pixel intensity differences between pairs of primitive patches, where N represents a relatively small number of primitive patches.

[0127] When referring to exemplary intensity-based metrics, it should be noted that this generally refers to any appropriate value-space (e.g., codomain) metric. For example, a camera may be configured to collect vector data (e.g., color images), and a value-space similarity metric may be evaluated based on vector pixel values, and therefore not necessarily limited to scalar comparisons of intensity values. Also, similarity measures are not necessarily limited to considering only measures of centrality (e.g., mean intensity), but may (alternatively or additionally) be comparison measures of (one) other properties (one) of the set of pixels, e.g., standard deviation, signal-to-noise ratio (in a spatial sense, e.g., mean intensity relative to standard deviation), higher-order statistical moments, texture metrics, and / or any other spatially descriptive measures. For example, for each pixel (or "primitive patch"), a vector (e.g., the primitive patch described above) may be formed with different components representing the pixel value and / or the pixel value in the local (spatial) neighborhood of the pixel, and pixels (or primitive patches) may be clustered based on the similarity of this descriptive vector.

[0128] As described above, when applying parameter tuning strategies, it can be particularly advantageous to discretely or continuously switch between forming spatial domain neighborhoods (patches in the narrowest sense) and value domain neighborhoods (patches in the sense of a generalized subset). It will also be understood that continuously changing combinations of both types of metrics, such as weighted sums, may be more suitable for some optimization algorithms and / or that separate switching points between metrics can be deferred to higher size thresholds and / or that the parameter optimization can be avoided by missing a portion of the parameter space where it would be possible to obtain a parameter optimization if the metric switching were performed in a more preferable continuous form in this case. However, it should also be noted that, as described above, when the target size requires it (e.g., when it exceeds a certain target size threshold), the approach of joining primitive (based only on spatial neighborhoods) patches to larger patches based on intensity (generally distance in the image codomain or related space) has the advantage of being easier to implement and / or more efficient. Furthermore, it should be noted that by using clustering techniques, such as those known in the field of image processing, to group pixels (or primitive patches) based on their value-domain characteristics, it is possible to ensure a certain degree of spatial cohesion of the resulting patches (generally subsets) even when the spatial neighbor metric is no longer explicitly considered (for example, when it is implicitly implemented and considered by the clustering algorithm).

[0129] To illustrate the approach discussed above, consider a scenario where we evaluate patches of multiple patch sizes, e.g., 10x10, 15x15, 20x20 (or equivalent approximate pixel counts). Assume that further increasing the patch size leads to the unfavorable situation where (almost) all patches (in the first region's partition) contain a significant amount of non-skin area. Thus, the naive approach prohibits further increasing the patch size (based on spatial neighborhoods) and fails to yield the desired result, for example, finding useful ROIs (or higher-quality ROIs than those found in the 20x20 test case). Instead of increasing the patch area to, for example, 40x40, smaller patches may be combined into larger virtual patches, for example, four 20x20 patches may be combined into one virtual (potentially spatially separated) patch totaling 40x40 pixels. Such combinations of patches can be obtained by clustering techniques, for example, by extracting mean level and standard deviation features across the patches (and possibly by averaging / calculating over the entire observation period, e.g., over a predetermined time window). Next, the selection of an appropriate scale (patch size) can be done by choosing a size that allows a sufficient number of (real or virtual) patches to render consistent signal intensity and / or another appropriate quality metric within the expected range (defined by physiology).

[0130] After step 101, which divides the first region into patches, and optionally after optimizing the patch size and / or other parameters, the patches of the first region can be pruned based on predetermined criteria of one or more values ​​determined for each patch (102). This may include selecting patches that have average pixel values, spread and / or motion values ​​(e.g., magnitude, motion vector components and / or one or more metrics representing motion) in a predetermined (corresponding) range and / or satisfy predetermined criteria. Motion values ​​such as the magnitude of motion, the vector components of the displacement vector and / or other such metrics representing motion can be determined, for example, by applying an optical flow algorithm (e.g., using two or more images acquired by the camera at different time points, e.g., a short dynamic sequence).

[0131] Therefore, a set of patches can be pruned to select a subset of patches that are substantially static in position (little or no detectable movement), substantially well-illuminated, and / or substantially homogeneous (e.g., have a low standard deviation of pixel values ​​forming the patch). Where we refer to mean, spread, standard deviation, etc., it will be understood that a similar effect can be obtained by various similar statistical measures of centrality and / or variance, e.g., median, quartiles, variance, etc. Similarly, other simple statistics (or other representative values) per patch, such as higher-order statistical moments, skewness, kurtosis, texture and / or pattern-based values, spatial gradient, and / or other similar values, can be used to locally (per patch) prune patches that exhibit characteristics unsuitable for the desired tissue of interest and / or reflect undesirable imaging conditions.

[0132] Furthermore, after step 101 of dividing the second region into patches, the patches of the second region may also be pruned 112 based on predetermined criteria relating to one or more values ​​determined for each patch. These criteria may be the same as or different from those for pruning the patches of the first region. This means that several variations can be considered in the implementation. For example, the values ​​of interest may be calculated over the entire image (but per patch) or over the set of the first and second regions (but per patch). By applying some recording, flagging, or other appropriate approach, it is easy to track which patch corresponds to which region. Depending on the details of the embodiment (e.g., the hardware used, the selection of values ​​used for pruning, and / or other such factors), it may be more efficient to calculate the values ​​of interest over all patches or separately (and possibly in parallel) for each region. Even if the same criteria are not applied to the patches of the first and second domains, in some situations it may be more efficient to calculate the values ​​related to the criteria (of the combined set of criteria) across all patches, which can be achieved, for example, by taking advantage of parallelization techniques (e.g., graphics processing units (GPUs) and / or general-purpose graphics processing units (GP GPUs)), in which case only different comparisons, thresholding, and / or other criterion evaluations need to be applied to the patches of the two different domains.

[0133] Pruning patches in the first and / or second regions 102, 112 may include selecting patches that have average pixel values, spread and / or motion values ​​(e.g., magnitude, motion vector components and / or one or more metrics representing motion) in a predetermined (corresponding) range and / or satisfy predetermined criteria. Motion values ​​such as the magnitude of motion, the vector components of the displacement vector and / or other such metrics representing motion can be determined, for example, by applying an optical flow algorithm (e.g., using two or more images acquired by the camera at different time points, e.g., a short dynamic sequence).

[0134] Therefore, a set of patches can be pruned (region by region) to select patches that are substantially static in position (little or no detectable motion), substantially bright, and / or substantially homogeneous (e.g., have a low standard deviation of pixel values ​​forming the patch). Where we refer to mean, spread, standard deviation, etc., it will be understood that a similar effect can be obtained by various similar statistical measures of centrality and / or variance, e.g., median, quartiles, variance, etc. Similarly, other simple statistics (or other representative values) per patch, such as higher-order statistical moments, skewness, kurtosis, texture and / or pattern-based values, spatial gradient, and / or other similar values, can be used to locally (patch by patch) prune patches that exhibit characteristics unsuitable for the desired tissue of interest and / or reflect undesirable imaging conditions.

[0135] For example, patches in which (large) motion is detected are removed (pruned) in both the first and second regions, so that, for example, (sufficiently) static (i.e., not moving) patches are selected. Patches in which (considerable) change (across pixel positions within the patch) is detected can be removed (pruned) in both the first and second regions to preserve homogeneous patches. Patches that are excessively or under-illuminated can be removed in the first region, and optionally in the second region as well, for example, to select patches that fall within the appropriate range of the camera system's dynamic range. These pruning criteria can be applied by logical OR operations, for example, a patch can be removed if one removal criterion is met, or, equally, retention criteria can be applied by logical AND operations, for example, a patch can be selected to be preserved for further processing only if all selection criteria (such as the logical negation of the aforementioned removal criteria) are met. Criteria for combining the results of multiple criteria (e.g., by logical operations, by majority vote) and / or values ​​for evaluating such criteria, and / or other choices of strategies are also possible according to the embodiment and can be determined, for example, by a person skilled in the art by relying on knowledge of the art and direct practice, testing, and / or simulation.

[0136] The method further includes determining one or more dynamic characteristic values, e.g., characteristics, that indicate the temporal dynamics for each patch. Thus, these values ​​are based on multiple camera images corresponding to different points in time, e.g., a sequence (time series) of camera images. These dynamic characteristic values ​​are determined for patches in the first region (109) and also for patches in the second region (119). In particular, the characteristics determined for each patch (109, 119) may be the same (e.g., of the same type, realized in the same way, and / or corresponding to the same physical quantity), for example, the same dynamic characteristic values ​​as those for patches in the first region can be determined for patches in the second region (119). This allows the results obtained for the patch-by-patch dynamic characteristics in the second region to be taken into consideration when processing the same characteristics obtained in the first region, as will be discussed later.

[0137] To improve efficiency, these dynamic characteristic values ​​can be determined (109, 119) after pruning (102, 112) the patches in the first and second regions, respectively, to avoid calculating the dynamic characteristic values ​​of patches that are to be removed (pruned). However, this is not always necessary. For example, by using parallelization techniques, the dynamic characteristics of all patches can be calculated quickly and efficiently (still patch by patch) even if the values ​​obtained for the pruned / selected / scheduled patches are substantially ignored. The method includes clustering the patches in the first region (e.g., patches not removed by the pruning step) based on one or more dynamic characteristic values ​​(representing the dynamic characteristics of the patches) (103), and similarly clustering the patches in the second region (113) based on the dynamic characteristic values. This clustering 103, 113 can optionally be performed on unpruned patches (i.e., only considering patches not removed by pruning). Clustering can be performed separately (e.g., independently) on the patches in the first and second regions.

[0138] Herein, it should be noted that, according to some embodiments, it is possible to apply clustering to all patches, for example, to remove clusters containing removed (pruned) patches, and / or to reduce such clusters by removing patches removed from those clusters. Therefore, the order of the pruning step and the clustering step does not necessarily have to be predetermined; for example, pruning may be performed after clustering. However, it will be understood that performing pruning before clustering (and limiting clustering to accepted patches) may be a simpler strategy, (typically) more efficient, less prone to errors, and / or more robust.

[0139] Furthermore, to increase efficiency, for example, both clustering steps 103 and 113 can be combined. However, each cluster determined by the clustering steps preferably corresponds to only one of either the first or second region, and for example, a single cluster cannot simultaneously contain patches from both the first and second regions. Therefore, when combining clustering steps 103 and 113 (but not limited to this), for example, by clustering a combined set of patches from both regions (preferably after pruning), mixed clusters may be obtained as part of the clustering result, regardless of whether they are assigned to the first or second region, for example, such mixed clusters may contain patches from both the first and second regions. Since mixed clusters should preferably be avoided, each mixed cluster can be simply removed (i.e., the cluster is further ignored), split into two clusters (one for each of the first and second regions), and / or more complex strategies can be devised to handle this. For example, a clustering strategy can penalize mixed clusters, for instance, by an iterative approach, so that it converges to a set of clusters without such mixed clusters. Nevertheless, it should be noted that performing the clustering of patches in the first region and the clustering of patches in the second region independently of each other (preferably after pruning in both cases) may generally be a simpler solution. Both clustering steps 103, 113 can be performed in parallel, for example, to keep the computational load low.

[0140] For example, after pruning based on the average pixel values ​​of the patch (e.g.), the spread and / or movement of the pixel values ​​of the patch, and one or more (temporal) frequency characteristics of each patch can be determined (steps 109, 119). Thus, the aforementioned dynamic characteristics may, in particular, include (a) temporal frequency characteristics, e.g., Fourier components and / or values ​​derived therefrom.

[0141] Fourier analysis (e.g., transformation from the time domain to the time-frequency domain) can be performed for each patch in steps 109 and 119. This can be applied (optionally) to patches that have not yet been removed in the pruning step, for example, to reduce the load on computational / processing resources for efficiency. For example, for patches that were not removed by the pruning step based on average pixel values ​​(and / or alternative values), spread (and / or alternative values), and / or motion, the average pixel values ​​(of the patch in question) as a function of time can be analyzed by applying a Fourier transform such as the Discrete Fourier Transform (DFT) or the Fast Fourier Transform (FFT).

[0142] In this example, the Fourier transform is based on the average pixel values ​​as a function of time (e.g., as the input domain for the transform), but it will be understood that other appropriate values ​​can be used to summarize the pixel content of the entire patch domain (FT input, e.g., FFT) (e.g., central behavior represented by a statistical centrality measure). Alternatively, a single pixel or a small group of pixels can be selected as representative of the entire patch, and each pixel of the patch can also be (Fourier) analyzed. If multiple Fourier spectra are calculated for the same patch (e.g., multiple pixels or multiple pixel subgroups), the resulting frequency spectra (of the same patch) can be combined in the frequency domain, for example by averaging (e.g., averaging of spectral powers), however, aggregation of data for the entire patch in the time domain may be preferable to aggregation of data in the frequency domain, for example, from the standpoint of efficiency (e.g., only one Fourier transform per patch is needed by using a single time series Fourier transform of the average pixel values).

[0143] Frequency analysis, such as Fourier analysis, can be performed on a sample over a predetermined time window, for example, appropriately selected and optimized as parameters as described above. In particular, the time window should be long enough to cover the frequencies of interest, but short enough for practicality (and / or clarity or specificity in the time domain). Those skilled in the art are very familiar with the appropriate configuration of Fourier analysis (e.g., time window and / or sampling rate) and related considerations. For example, the time window may be, but is not limited to, a range of 0.5s to 30s, for example, a range of 1s to 5s. In particular, the Fourier analysis can be configured and / or pre-processed and / or post-processed (e.g., using low-pass filtering, band-pass filtering, and / or high-pass filtering) to cover the spectral cardiac band, for example, the spectral cardiac band in which the PPG signal of interest is presumed to have been found. For example, a band-pass filter of, for example, 0.4Hz to 4Hz can be used to correspond to the heart-rhythmic signal component under typical heart rate conditions (e.g., 24 to 240 bpm). However, the embodiments are not necessarily limited thereto, and for example, appropriate filter parameters can be determined by routine experiments, simulations, and / or simple design techniques, taking into account the knowledge of those skilled in the art.

[0144] It will be understood that by applying various alternatives to Fourier analysis, such as wavelet analysis, time-scale analysis, and multi-resolution analysis, the same or similar effects, i.e., the dynamic characteristics of the camera signal associated with each patch, can be represented. Therefore, embodiments may include the use of such direct alternatives instead of Fourier analysis, and / or Fourier analysis may be combined with other such techniques.

[0145] Frequency analyses, e.g., Fourier spectra 109, 119 determined for use when clustering patches based on their dynamic characteristics 103, 113, can also be used for further steps in pruning. This may include determining (for each patch) the spectral energy of the strongest peak (and / or its predetermined band, e.g., heart rate frequency band) in the Fourier spectrum, and / or the ratio of the spectral energy of this strongest peak to the spectral energy of a second strongest peak. This peak spectral energy and / or this ratio can be applied as a further pruning criterion to remove patches with weak or indistinguishable maximum values ​​(e.g., multiple peaks instead of a single clearly defined strong peak) by removing patches where this peak spectral energy and / or ratio is less than a predetermined threshold (122). Such further pruning 122 may be applied to patches in both the first and second regions, or only to patches of interest in the first region. The further pruning criteria may be the same or different for the first and second regions.

[0146] For example, before proceeding to the selection of regions of interest discussed below, such further pruning steps can be used, for example, as additional verification steps to consider the correspondence between the signal intensity and / or spectral characteristics / features (signal dynamics characteristics or features) of each patch and the desired / expected characteristics of the PPG signal. Possible criteria include thresholding (one-sided or two-sided) of the peak spectral energy and / or the ratio of the peak spectral energy to the average spectral energy (and / or the spectral energy of the second largest peak). For example, patches with weak peaks are removed, and / or patches with indistinct peaks (compared to the average and / or the second largest peak) are also removed. Furthermore, if the amplitude of the spectral peak is too high, it can be considered an unsuitable candidate for further consideration (for example, the thresholding process can be carried out considering not only a lower limit but also an arbitrary upper limit). The predetermined thresholds can take into account predetermined patch sizes and predetermined observation windows (FFT windows) of predetermined length, and can be readily determined by those skilled in the art, for example, based on selected operating parameters.

[0147] As an addition or alternative, such metrics (e.g., peak signal intensity, peak frequency, etc.) can also be used to determine dynamic characteristics for further use in clustering, as discussed below, for example, further pruning steps (binary removal) can be replaced or complemented by using the underlying metrics in the clustering steps(s) discussed further below. For example, a metric (e.g., probability) can be associated with each patch based on peak amplitude, based on peak frequency, or a combination thereof, in order to construct clusters from the patches in the steps below.

[0148] Therefore, in addition to or without further pruning, the peak frequency of the strongest peak, or the peak frequency and peak magnitude, can be used for clustering steps 103, 113. The peak magnitude can be expressed as a ratio to the second largest peak, a ratio to the mean over the spectral band under consideration (heart rate band), and / or another suitable expression. The frequency and magnitude can be combined in a weighted scalar combination, for example, the fitness of the frequency (e.g., probability) and the fitness of the magnitude can be combined to make a combined fitness (probability, likelihood, etc.) for clustering patches, or they can be used together as vector values ​​for clustering.

[0149] Clustering 103, 113 can be based on the frequency of the strongest spectral peak determined for each patch (e.g., at least) (for patches not removed by the initial and / or further pruning steps). However, alternatives can also be used, for example, by using multi-resolution time-scale analysis (e.g., using a hierarchical pyramid of scale analysis) to determine the time scale at which the strongest signal is detected (e.g., the highest power or intensity, magnitude scale, etc.), and associating this (most prominent) time-resolution scale value with the patches to be used as a clustering variable. Clustering based on multiple variables (e.g., associating vector-value variables with the patches to be used as a criterion for clustering) is not necessarily ruled out (e.g., combining peak frequency and peak time scale).

[0150] For example, using the peak frequency (or a suitable equivalent, e.g., the scale in the scale hierarchy with the highest observed energy, the time period corresponding to the peak frequency, the wavelength, etc.), it is possible to determine clusters of patches where pixels pulsate at substantially the same or similar frequencies (relative to other patches in the cluster). As mentioned above, modifiers of signal intensity (for this frequency component) can also be included in the clustering variability.

[0151] To cluster patches in the first region and to cluster patches in the second region, the same or similar clustering variables (e.g., peak frequency and / or amplitude) can be used to select one or more clusters of related patches in the first region using information from unrelated clusters in the second region, as described later.

[0152] The method includes step 104 of selecting at least one cluster from the clusters of patches obtained by step 103 of clustering patches of a first region based on size and / or shape. An output is then provided (130). The region of interest (ROI) definition may be provided as an output defining the pixels that form the selected one or more clusters, which can be used, for example, by a dedicated PPG signal extraction algorithm, and / or the PPG signal may be determined from the pixel values ​​in the selected one or more clusters (in a series of images acquired by the camera) (105). For example, the PPG signal may be determined based on the dynamic change (e.g., spatial average) of substantially only the pixels of the selected clusters (105). In other words, the PPG signal can be determined from a sequence of camera images based on a received input representing (or including) the ROI definition, by using the selected clusters and computing a robust PPG signal (from pixels corresponding to the ROI in a sequence of camera images) directly by a further step 105 of the method according to embodiments of the present invention, or by providing the ROI to an external algorithm as is known in the art.

[0153] However, step 104 of selecting clusters includes excluding clusters obtained for the first region that correspond to clusters obtained by clustering patches of the second region 113. In particular, step 104 of selecting clusters excludes from the selection any clusters in the first subsets whose associated temporal signal dynamics substantially correspond to the temporal signal dynamics of clusters in the second subsets, for example, if the (absolute) difference (or any appropriate distance metric) between at least one dynamic characteristic value of a pair of clusters being evaluated (one from the first region and the other from the second region) is less than a predetermined threshold. For example, for each cluster in the first region, a distance (generally, for example, a metric based on dynamic characteristics) to each cluster in the second region can be determined, and if the minimum distance thus obtained is below a threshold, this cluster in the first region can be removed. Optionally, only the largest clusters in the second region may be considered for this exclusion step, to avoid removing large clusters in the first region that match small clusters in the second region, e.g., floating patches or a few floating patches of the tissue of interest (e.g., skin). In other words, the size criterion can be applied to clusters in the second region to avoid filtering out clusters in the first region based on matches to small clusters in the second region that could correspond to small areas of the tissue of interest that were incorrectly estimated to be irrelevant (by the pre-assignment of corresponding image regions to the second region).

[0154] To select clusters based on size (104), the largest cluster, or a predetermined number of largest clusters, can be selected from the clusters obtained by clustering (103) the patches of the first region after excluding the corresponding clusters. The size of each cluster for selecting the largest clusters can be determined based on the number of patches, the number of pixels, the corresponding region, or another similar measure representing the spatial extent of the cluster. Thus, clusters are selected that advantageously contain a large number of pixels (as much as possible) but do not match the dynamic characteristics encountered for the second region. Then, using the pixels belonging to the selected one or more clusters, a region of interest can be defined from which a robust and accurate PPG signal can be extracted.

[0155] Instead of directly selecting the largest cluster (of the patches obtained for the first relevant region), the largest cluster unrelated to any irrelevant (e.g., non-skin) frequencies (or alternative dynamic characteristics used for clustering) is selected; that is, clusters corresponding to vibrations and / or dynamics of pixel signals (e.g., pixel intensity) related to non-skin (or more generally, unrelated to the tissue type of interest) are ignored, i.e., excluded from selection.

[0156] For example, the step (104) of selecting a cluster from a cluster of patches obtained by clustering (103) the patches of a first region may include the step of pruning the cluster of patches of the first region such that its dynamic characteristic value (e.g., peak frequency) corresponds to the dynamic characteristic value (e.g., peak frequency) of the cluster obtained by clustering (113) the patches of a second region.

[0157] In an exemplary approach, peak frequencies (or their alternative values) are averaged across patches within each cluster so that a single peak frequency is associated with a cluster, and cluster peak frequencies (or their alternative values) obtained in both the first and second regions are removed (i.e., the corresponding clusters in the first region are pruned, or simply such clusters are ignored). Peak frequencies can be compared (between clusters in the first and second regions, respectively) simply by thresholding their absolute difference (e.g., requiring it to be less than a predetermined threshold for comparison). Alternatives can easily be contemplated. For example, the standard deviation of peak frequencies within a cluster (on the patches forming the cluster) may be considered, and the two clusters may be compared by, for example, Student's t-test, F-test, etc. Optionally, smaller clusters in the second region can be ignored, and a criterion may be imposed, for example, that the matching cluster frequency corresponds to a cluster of at least substantial size (e.g., a predetermined size threshold) in the second region. Therefore, in order to be excluded from selection, a (large) cluster in the first region must have a frequency match with a cluster in the second region that is also at least a substantial size. After removing the clusters in the first region by the corresponding clusters in the second region (with respect to the dynamic parameters used for clustering, such as peak frequency), the largest cluster in the first region under consideration, or a set of the top ranks of the largest cluster (top two, top three, etc.), can be selected to determine the region of interest (i.e., the region used to extract the PPG signal).

[0158] For simplicity and efficiency, the corresponding clusters within the first and second domains are described as being determined based on the dynamic characteristics used for clustering (e.g., peak frequency), although it will be understood that this is not necessarily required. The same or similar effect can also be obtained (additionally or alternatively) by comparing the temporal dynamics of the clusters in the first and second domains using another appropriate measure, such as the cross-correlation of the mean signals over time (alternative values ​​such as the entire spatial range in each of the cluster pairs being compared, or centrality), statistical tests for comparing distributions, mutual information or other information-theoretic metrics, and / or another appropriate measure for comparing the scalars or vectors representing each of the cluster pairs being compared (e.g., time series, Fourier spectra, and / or measures for comparing other characteristics of the dynamic behavior).

[0159] Furthermore, selection 104 may also consider the morphology of each cluster. In particular, each cluster found in a tissue of interest, e.g., a first region of skin (presumably relevant), is associated with a measure based on the shape of the image region formed by the pixels (or patches) within the cluster (i.e., the collection of patches within the cluster, and / or the combined set of pixels forming the cluster), for example, to rate the cluster for its consistency with the expected shape of the region of tissue of interest. Thus, for each cluster that was not eliminated by correspondence with matching clusters of irrelevant image content (i.e., a second region), a measure (e.g., a probability index) can be determined 121 indicating whether (and / or to what extent) the shape of the image region formed by the pixels (or patches) within the cluster matches the expected shape.

[0160] For example, the forehead in question can form a target region of tissue of interest (skin) corresponding to a connected set of roughly elliptical patches. Thus, in this example, the scale 121 that can be determined for each cluster (e.g., a first region not removed by a previously applied pruning step) can represent the cluster's correspondence to the ellipse. Equivalently, the scale can represent the deviation from there, i.e., the scale can be a substantially monotonic relationship that increases or equivalently decreases as the shape of the image region formed by the pixels (or patches) within the cluster better matches the intended shape.

[0161] For example, an oval or ellipse can be fitted to a patch (or, equivalently, a set of pixels) forming a cluster, and an error value (of the shape) can be assigned to patches (or pixels) that extend beyond the fitted shape and to missing patches within the shape. Thus, if the collection of patches / pixels (i.e., clusters) forms a line segment (for example), it is unlikely to represent a desired feature, such as a face or part thereof, such as a forehead, in detail. Furthermore, alternatively or additionally, the scale (e.g., error value) can take into account the fitted parameters of the fitted ellipse, such as the major and minor axes (and / or their ratio) of the fitted ellipse, and prioritize the region of the fitted ellipse (or oval) that falls within a range corresponding to the overall shape that is expected to be obtained. For example, a fitted ellipse with a very large ratio of maximum axis to minimum axis will match a line segment more closely to the desired target shape (in this example) of an ellipse that is (approximately) closer to a circular form.

[0162] Based on the expected shape of the target region of interest, many modifications can be readily conceived. For example, to target a portion of the skin of the forearm, a roughly rectangular area can be used as a criterion for defining the shape scale by fitting a rectangular or rounded rectangle (for example, at any angle with respect to the camera image coordinates) into the cluster. Thus, in this example, the selection may actively prefer more elongated clusters based on prior knowledge (e.g., intended use for extracting PPG signals from camera observations of the forearm).

[0163] Accordingly, clusters are selected based on size and / or morphological probability, e.g., a combination of both, while corresponding clusters in a second region (presumably irrelevant to the tissue of interest) (based on dynamics, e.g., peak Fourier frequency) can be removed. Thus, embodiments of the present invention can improve upon prior art "biological skin" selection methods by considering information obtained in (one or more) regions of an image where tissue of interest may not be found (i.e., signal dynamics in the second region described above), and optionally by also considering the morphological characteristics (e.g., shape criteria) of the selected clusters.

[0164] Different contributions to the selection (e.g., ranking) step, such as those based on cluster size, cluster shape, and correspondence with unrelated clusters in a second domain, can be combined in various ways. For example, different types of metrics can be weighted and combined to obtain a ranking metric in order to select the most promising one or more clusters. Alternatively, binary exclusion from selection can be applied to one or more metrics, such as removing clusters with corresponding dynamic properties in the clusters of the second domain, and / or removing clusters that are too small or too large, and / or removing clusters that do not correspond to a desired shape. For example, a combination of these metrics can be used, for example, some of these metrics for binary selection and one or more for ranking selection after binary selection has been applied. The metrics used for binary selection and ranking selection may be different or overlapping. For example, the initial elimination step may be based on these criteria with relatively broad thresholds, and a weighted combination of these metrics may be used for the final ranking selection.

[0165] In a second aspect, the present invention relates to an apparatus for use in extracting a photoplethysmographic signal indicating the physiology of an object based on remote camera observation, for example, an apparatus that provides such a PPG signal and / or a region of interest adapted to select pixels in an image frame provided by a camera for extracting such a PPG.

[0166] Referring to Figure 2, an exemplary apparatus 10 according to an embodiment of the present invention is schematically shown.

[0167] The device 10 has an input unit 11 for receiving camera images from a camera configured to monitor at least one body part of a subject. The device 10 may optionally have a camera 19, or the device may be adapted to connect to such an (external) camera via, for example, a wired, wireless, or indirect connection (e.g., retrieving images stored on an intermediate device such as a streaming server). The device may optionally have a light source (or multiple light sources) for illuminating the subject (the relevant body part or part thereof). The camera may be an infrared camera, a monochrome camera operating in the visible wavelength range or a portion thereof, e.g., a green light spectral band, a color camera, and / or a multispectral camera. Similarly, the light source may emit light in at least the overlapping portion of the spectrum to which the camera is sensitive, such as an infrared light source, a green light source, or a broad-spectrum (e.g., white) light source.

[0168] The device 10 includes a subset generator 12 for dividing a first region representing an image portion in a camera image in which the tissue of interest is presumed to exist into a first subset of pixels (e.g., patches), and a second region representing an image portion in a camera image in which the tissue of interest is presumed not to exist into a second subset of pixels (e.g., patches).

[0169] The subset generator 12 can also be adapted to divide each of the first and second regions into a plurality of first and second pixel subsets, based on predetermined prior knowledge assumptions about the likelihood of tissue of interest existing in each region, given a known camera setup for the spatial volume in which the object is located. Such prior knowledge can be, for example (but not limited to), hardcoded or hardwired within the device, stored in appropriate configuration memory, and / or received via a user interface or interface to an external controller.

[0170] The subset generator 12 is configured to divide the first region into a first set of pixel subsets and the second region into a second set of pixel subsets by forming partitions of the first or second region into blocks or patches based on a spatial distance metric, thereby grouping spatially close pixels in each subset. Additionally or alternatively, each region can be partitioned into subsets based on image codomain distance, for example, by grouping pixels that are close to each other in terms of their pixel values. A combination of spatial-based distance and codomain-based distance can also be used, and / or based on such combinations, for example, each subset can group together pixels that are spatially and numerically close (in a balanced combination).

[0171] The subset generator 12 can be adapted to perform image segmentation and / or processing of a reference image (and / or generally, reference spatial information) received via the input 11, for example, from a camera, from another camera, and / or a (general-purpose) spatial information source (e.g., a 3D surface scanner, a tomographic imaging system, etc.) to determine the first and second regions.

[0172] The subset generator 12 can be adapted to perform image subtraction between a reference image acquired when the target is present and a blank image acquired when the target is absent.

[0173] The apparatus may also have a subset eliminator 18 for pruning a first set of subsets based on predetermined criteria for one or more values ​​determined for each subset, in order to remove subsets that are unlikely to correspond to a uniform area of ​​the target tissue. The subset eliminator 18 may further be adapted to prune a second set of subsets based on the same and / or other predetermined criteria. The clustering processor 14, further described below, may be adapted to cluster the pruned first set of subsets and the pruned second set of subsets, respectively, as determined by the subset eliminator.

[0174] The subset eliminator 18 can be adapted to prune a first and / or second subset based on predetermined criteria including one or more criteria for the mean pixel value per subset, the spread of pixel values ​​per subset, and / or for at least one value indicating movement associated with the subset, so that the subset eliminator is adapted to remove pixels exhibiting excessive movement, non-uniform pixel value distributions, and / or, on average, pixel values ​​outside a predetermined target range. It will be understood that various alternatives and / or equivalents can be used for median values ​​(e.g., generally a measure of centrality), spread (e.g., generally a measure of variance, e.g., variability or standard deviation), and movement (e.g., motion vector components, magnitude of motion vectors, magnitude of squared motion vectors, etc.).

[0175] The device includes a signal dynamics analyzer 13 that determines, for each subset of a first subset and for each subset of a second subset, at least one dynamic characteristic value that indicates the temporal signal dynamics of a subset in a sequence of camera images received via the input 11.

[0176] The signal dynamics analyzer 13 can be adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency characteristics to be used in determining at least one dynamic characteristic value.

[0177] The subset eliminator 18 can be adapted to apply further pruning (e.g., already first time-pruned, see above) of a first and / or second subset based on at least one dynamic characteristic value. Thus, the clustering processor 14 can be adapted to cluster the further selected first subsets and the further pruned second subsets, respectively.

[0178] For example, further pruning may be based on at least one predetermined criterion for the spectral energy of the largest spectral peak of the Fourier spectrum determined for each subset, and / or on a value determined therefrom, for example, subsets whose spectral energy falls below a certain threshold or falls outside a predetermined value range are removed.

[0179] The device includes a clustering processor 14 for grouping a subset of pixels into clusters having similar signal dynamics by clustering a first subset and a second subset, respectively, based on at least one dynamic characteristic value.

[0180] The device has a selector 15 that selects at least one cluster from a first subset of clusters provided by a clustering processor 14 based on cluster size and / or cluster configuration. The selector is adapted to exclude from the selection each cluster of the first subset that substantially corresponds to the signal dynamics of the clusters of a second subset in terms of their associated signal dynamics.

[0181] Referencing the cluster size-based selection, the selector 15 can be adapted to select one or more maximum clusters from the first subset that are not excluded, based on their correspondence to clusters of the second subset.

[0182] In the case of selection based on cluster morphology, the selector may be adapted to determine a shape scale representing the shape of the image region formed by pixels (or patches) within a cluster for each cluster of a first subset (or, to make preliminary selections from those clusters, for example, after applying a rough first selection criterion, for example, after excluding clusters that are too small, and / or after excluding based on correspondence with clusters in a second subset), and to select at least one cluster (i.e., the result used in the output) based on the fact that the shape scale of the selected cluster matches a predetermined target shape of the region of tissue of interest.

[0183] The selector 15 can further be adapted to determine a ranking of clusters in the first subset that are not excluded based on the correspondence with clusters in a second subset, and this ranking is used to determine the selected clusters. The ranking can combine a first score based on cluster size and a second score based on shape scale, for example, to combine selection based on morphology and size. Exclusion criteria (correspondence with clusters in a second subset) can be introduced into this ranking, and it will be understood that exclusions are imposed on the ranking by penalty, for example. In other words, exclusions may be hard exclusions performed preventively, or they may be imposed by soft requirements, such as lowering the ranking of clusters whose dynamic characteristics match to clusters in a second domain that are presumed to be irrelevant. It should also be noted that this can avoid or reduce restrictions imposed by criteria that are strictly imposed on the matching between clusters in the first subset and clusters in a second subset. This can be achieved, for example, by requiring that the absolute difference between peak frequencies be smaller than a predetermined threshold, or alternatively by including this absolute difference (or a similar measure, etc.) in the ranking measure with appropriate weighting, but particularly large clusters with the expected shape may be selected even if their dynamic characteristics are relatively close (but not "too close") to the clusters in the second set.

[0184] The device has an output 16 for outputting a region of interest definition that defines pixels forming at least one cluster selected by the selector 15, for use in extracting a photoplethysmography signal from a camera image acquired by the camera, or alternatively or additionally, for outputting a photoplethysmography signal extracted from a region of interest in a camera image received via input 11.

[0185] Accordingly, the apparatus may optionally further include a photoplethysmography signal extractor 21 for extracting a photoplethysmography signal to be provided as an output from a region of interest in a series of camera images received from the camera via input 11.

[0186] The device may also have an optimizer 17. The subset generator 12 can be adapted to iteratively divide a first region into a first set of pixel subsets (and optionally a second region into a second set of subsets) for each different value of at least one optimization parameter representing the size of a pixel subset. The optimizer 17 can be adapted to calculate and evaluate a quality metric based on the pixel subsets obtained in each such iteration for each different value of at least one optimization parameter, so that the optimizer can select a parameter value from which a sufficient (e.g., above a predetermined threshold, or after a predetermined number of iteration steps, or up to a predetermined parameter precision) or optimal value of the quality metric is obtained. The parameter values ​​thus selected, as well as the corresponding divisions into subsets of the first and second regions, can then be used by the clustering processor 14 and the selector 15, for example, to determine a region of interest.

[0187] The optimizer 17 may also use the size of the observation time window used by the signal dynamics analyzer 13, for example, the Fourier time window parameter, as a component of at least one optimization parameter (for example, in addition to the subset size to be optimized).

[0188] The optimizer 17 can, for example, calculate a quality metric that includes or is composed of the signal-to-noise ratio.

[0189] For example, the apparatus may have a processor, computer, or similar general-purpose computing device in combination with software adapted to perform the method according to embodiments of the present invention, as described above. Alternatively or additionally, the apparatus may have dedicated hardware designed to perform the method or one or more steps according to embodiments of the present invention. For example, such dedicated hardware may include application-specific integrated circuits and / or configurable hardware such as field-programmable gate arrays.

[0190] In a third aspect, the present invention relates to a diagnostic imaging system, such as a magnetic resonance imaging system or a computed tomography system. The system has an examination zone and a camera for monitoring an object to be examined while it is located within the examination zone. The system also has a device according to an embodiment of the second aspect of the present invention, which is operably connected to the camera, for example, to receive camera images from camera 19 via input 11, and which receives camera images from the camera as input.

[0191] As is known in the art, the apparatus can be adapted to image a target by magnetic resonance imaging, computed tomography, positron emission tomography, and / or single-photon emission computed tomography. Alternatively, in other embodiments, the present invention relates to a similar system for performing medical interventions, such as a radiotherapy system, having, for example, a camera system and apparatus according to an embodiment of a second embodiment of the present invention. Other embodiments may have a system for monitoring a patient in an intensive care unit during a surgical intervention, and / or generally for monitoring a patient's health in a medical or medical situation.

[0192] For example, the system may be a magnetic resonance imaging system, but the principles of the present invention can also be equally applied to systems for different diagnostic imaging modalities and / or for use in other medical or non-medical situations.

[0193] A magnetic resonance imaging system according to an embodiment may have a primary magnet assembly that defines an inspection zone, for example, the inspection zone may be formed by a volume in which magnetic field conditions suitable for magnetic resonance imaging are substantially generated and controlled by the magnet assembly. Thus, the inspection zone may correspond to (at least the usable portion of) a volume enclosed by the system's magnet bore (but not limited to, for example, the principles of the present invention apply equally to open bore systems and other, less frequently used magnet assembly configurations).

[0194] The object being examined, such as a patient, can be placed on a patient treatment table within the examination zone when the system is in use. The primary magnet assembly may have magnet windings, such as coaxial (e.g., superconducting) windings, to generate a fixed, uniform magnetic field within the examination zone. The examination zone may be a cylindrical volume surrounded by these magnet windings.

[0195] The system may have a reconstructor for reconstructing magnetic resonance images, such as tomographic MRI images, from magnetic resonance signals acquired by the system in use. The reconstructed images may be provided via output for viewing, processing, or storage.

[0196] Auxiliary devices, such as RF T / R head coils, can be placed in the inspection zone to acquire magnetic resonance signals from the target head during use. Other auxiliary coil configurations can be used to acquire signals from other body parts or for different use cases, but typically the signals can also be received by a receiving coil integrated within the housing of the primary magnet assembly.

[0197] The system includes, for example, a camera or a camera assembly having multiple cameras. The camera system is adapted to acquire information from the subject being examined, for example, to obtain vital signs, movement, pain indicators, etc.

[0198] For example, the camera may be mounted near one of the entries in the inspection zone. For example, the camera may be integrated into or mounted on the flange of the MR bore (for example, so that the usable free bore diameter is not affected or is only minimally reduced, and / or interference with the operation of the MR system is avoided or minimized). For example, (optional) illumination lights may also be provided in or on this flange (but not limited to this).

[0199] The system may have a display for showing images of the inside of the inspection zone acquired by a camera (raw and / or after appropriate processing). This allows the operator to visually monitor the objects within the inspection zone.

[0200] Images acquired by the camera can be provided to an apparatus according to an embodiment, for example, an image processing apparatus for performing a method according to an embodiment. Thus, the image processor may be, or may have, an apparatus 10 according to an embodiment of a second aspect of the present invention. The image processor (i.e., apparatus 10) is adapted to process image information acquired by the camera system, perform image analysis to acquire information from the patient, and in particular to extract a photoplethysmography (PPG) signal from a region of interest in the video stream acquired by the camera, in which case this ROI is acquired as described above.

[0201] Respiratory and / or cardiac phase information (and / or more general information indicating motion), such as the PPG signal generated by the device, and / or information derived therefrom and / or enhanced therefrom, can be provided to the reconstructor to correct the magnetic resonance signal acquired for motion and / or to apply motion correction to the reconstructed magnetic resonance image. For example, a cardiac trigger signal can be determined based on the PPG signal.

[0202] Accordingly, the signals provided by the device 10 can be used to gate data acquisition by a system, such as an MRI system, or other system for performing a diagnostic imaging procedure. Alternatively, or additionally, a treatment system can be used to control the delivery of a treatment procedure. Additionally, or alternatively, the signals can be used to sort, match, select, and / or annotate image data acquired by a diagnostic imaging system to display extracted PPG information (or information derived therefrom, such as cardiac phase) alongside diagnostic images acquired at substantially corresponding times.

[0203] In the magnetic resonance imaging system according to embodiments of the present invention, the camera system may have one or more light sources. Embodiments that rely on passive illumination for imaging are not necessarily excluded, but it will be understood by those skilled in the art that when light conditions are better controlled and active illumination is used, (camera) imaging may be more effective.

[0204] The light source and / or camera can be positioned outside the examination zone, or on or near its edge region. This can simplify the setup of the magnetic resonance imaging system (for example, avoiding or reducing interference with the system's RF and magnetic field operations) and provide more flexibility in bore width within the examination zone. For example, in a cylindrical bore system, both the camera and the light source (or either one thereof) can be positioned on the flange of the bore enclosure at one end of the bore, leaving the other end virtually free, for example, enabling unobstructed access to the examination zone (for bringing the patient and / or auxiliary equipment into the examination zone) and reducing the potential claustrophobic effects on the subject, and thus potential discomfort, while being imaged by the system.

[0205] In a magnetic resonance imaging system according to an embodiment of the present invention, the camera (or a group of cameras) can be adapted to process (e.g., substantially exclusively photosensitive) light in (e.g., within a narrow) infrared wavelength range and outside the visible wavelength range.

[0206] In a magnetic resonance imaging system according to an embodiment of the present invention, the camera (or a group of cameras) can be adapted to operate in the visible wavelength range (for example, to be substantially exclusively sensitive), and can be adapted to be sensitive to, for example, a broad white light spectrum or a portion thereof, for example, a color band, for example, green light. The camera can be adapted to acquire monochrome information, or it may be a color camera adapted to detect, for example, different color components, for example, red, green, and blue components (but not limited to these), preferably independently and substantially simultaneously. The camera can also be adapted to detect relatively large (e.g., more than three) spectral components, and can be, for example, a multispectral camera.

[0207] A light source can emit light with a spectrum suitable for the camera; for example, broadband white light can provide illumination for a monochrome or color camera operating in the visible range. Similarly, an infrared light source can be used to emit infrared light in the spectral range to which an infrared camera is sensitive. The spectra of the light source and the camera do not necessarily have to be identical or closely related; for example, the spectrum of the light source may be broader, as long as there is sufficient overlap with the spectrum to which the camera is sensitive. The camera may be, for example, a digital camera having an array of pixel photodetectors.

[0208] In a fourth aspect, the present invention relates to a computer program for performing a method according to an embodiment of the first aspect of the present invention when performed by a computing device. For example, the computer program may have machine-interpretable instructions for instructing a computing device, such as a computer, to implement (i.e., perform) the method of the embodiment.

[0209] Any other features or details of the apparatus, system, and / or computer program according to embodiments of the present invention that are described above will be apparent from the above-provided description of the method according to embodiments of the present invention and / or vice versa; that is, regardless of the context (the embodiments and / or aspects of the present invention described), the details and features described above can be applied equally to all different aspects of the present invention with necessary modifications.

Claims

1. A device used to extract photoplethysmographic signals that indicate the physiological state of a subject based on remote camera observation, An input that receives camera images from a camera configured to monitor at least one body part of a target, A subset generator that divides a first region representing an image portion in the camera image in which the tissue of interest is presumed to exist into a first subset of pixels, and divides a second region different from the first region representing an image portion in the camera image in which the tissue of interest is presumed not to exist into a second subset of pixels, A signal dynamics analyzer that determines at least one dynamic characteristic value representing the temporal signal dynamics of a subset in a sequence of camera images received via the input, for each subset of the first plurality of subsets and for each subset of the second plurality of subsets, A clustering processor that clusters the first subset and the second subset based on the at least one dynamic characteristic value, and groups the subsets of pixels into clusters having similar signal dynamics. A selector that selects at least one cluster from a first subset of clusters provided by the clustering processor based on cluster size and / or cluster configuration, wherein the selector excludes from the selection each cluster of the first subset whose associated signal dynamics substantially correspond to the signal dynamics of the clusters of the second subset of clusters; Outputs a region of interest definition that defines pixels forming at least one cluster selected by the selector, for use in extracting a photoplethysmography signal from a camera image acquired by the camera, and / or outputs a photoplethysmography signal extracted from the region of interest in the camera image received via the input. A device having.

2. The apparatus according to claim 1, wherein the subset generator is configured to identify anatomical keypoints in the sequence of camera images, and the subset generator is configured to determine a first subset and a second subset using the identified anatomical keypoints.

3. The apparatus according to claim 2, wherein the subset generator is further configured to determine the first subset and the second subset by fitting templates to the identified anatomical keypoints in the sequence of camera images.

4. The apparatus according to claim 2 or 3, wherein the subset generator has a neural network that outputs the anatomical keypoints in response to receiving a sequence of camera images as input.

5. The apparatus according to any one of claims 1 to 4, wherein the at least one dynamic characteristic value has a first-order spectral peak, and the signal dynamics analyzer determines the first-order spectral peak of a first subset and a second subset by performing a fast Fourier transform of the pixel values ​​to detect the first-order spectral peak.

6. The apparatus according to claim 5, wherein the clustering processor clusters the first subset and the second subset according to the primary spectral peaks.

7. The apparatus according to claim 5 or 6, wherein the primary spectral peak is limited to the heart rate band.

8. The apparatus according to claim 7, wherein the heart rate range is 5 beats / min to 300 beats / min.

9. The apparatus according to any one of claims 5 to 8, wherein the selector is configured to select the largest cluster among the first subsets having a primary spectral peak different from the largest cluster among the second subsets.

10. The apparatus according to claim 9, wherein the selector is further configured to ignore clusters having spectral peaks exceeding a predetermined threshold during the selection process.

11. The apparatus according to any one of claims 1 to 10, further comprising a photoplethysmography signal extractor for extracting a photoplethysmography signal provided via the output from the region of interest in the sequence of camera images received via the input.

12. The selector is configured to select one or more maximum clusters from the first subsets that are not excluded based on the correspondence with the clusters of the second subsets, and / or The selector is configured to determine a shape scale for each cluster in the first subset that represents the shape of the image region formed by the pixels and / or patches within the cluster, and to select at least one cluster based on the fact that the shape scale of the selected cluster matches a predetermined target shape of the region of interest, and / or The apparatus according to any one of claims 1 to 11, wherein the selector is configured to determine the ranking of the first subset of clusters, the ranking is used to determine the selection, and the ranking combines a first score based on the size of the clusters and a second score based on the shape scale.

13. The apparatus according to any one of claims 1 to 12, wherein the subset generator divides each of the first and second regions into a first plurality of pixel subsets and a second plurality of pixel subsets by forming each of the portions of the first region and the second region into subsets based on a spatial distance metric and / or based on image codomain distance and / or a combination thereof, and each subset is configured to group together pixels that are spatially and / or numerically close.

14. The apparatus according to any one of claims 1 to 13, wherein the subset generator is adapted to determine the first and second regions by performing image segmentation and / or processing of a reference image and / or reference spatial information received via the input, and / or is adapted to perform image subtraction between a reference image acquired when the object is present and a blank image acquired when the object is not present, and the reference image and / or the reference spatial information is acquired by the camera, another camera, and / or another spatial information source and received via the input.

15. The apparatus according to any one of claims 1 to 14, wherein the apparatus comprises an optimizer, the subset generator is adapted to iteratively divide the first region into a plurality of first pixel subsets for different values ​​of at least one optimization parameter representing the size of the pixel subset, and the optimizer is adapted to calculate and evaluate a quality metric based on the pixel subsets obtained in each iteration for different values ​​of the at least one optimization parameter in order to select the parameter value for which a sufficient or optimal value of the quality metric is obtained.

16. The apparatus according to any one of claims 1 to 15, wherein the optimizer further uses the size of the observation time window used by the signal dynamics analyzer as a component of the at least one optimization parameter, and / or the quality metric calculated by the optimizer has or is composed of a signal-to-noise ratio.

17. The apparatus according to any one of claims 1 to 16, wherein the signal dynamics analyzer is adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency characteristics for use in determining the at least one dynamic characteristic value.

18. The apparatus has a subset eliminator that prunes the first plurality of subsets based on a predetermined criterion of one or more values ​​determined for each subset, and prunes the second plurality of subsets based on the same and / or other predetermined criterion, so that the apparatus further removes subsets that are less likely to correspond to homogeneous regions of the target tissue, wherein the clustering processor is adapted to cluster the pruned first plurality of subsets and the pruned second plurality of subsets, respectively.

19. The apparatus according to claim 18, wherein the subset eliminator is adapted to prune the first and / or second subsets based on the predetermined criteria which include one or more criteria which include the average pixel value for each subset, the spread of pixel values ​​for each subset, and / or at least one value which indicates movement associated with the subset, and as a result the subset eliminator is adapted to remove pixel subsets exhibiting excessive movement, non-uniform pixel value distributions, and / or pixel values ​​which, on average, fall outside a predetermined target range.

20. The apparatus according to claim 18 or 19, wherein the subset eliminator is further adapted to apply further pruning of the first and / or second subsets based on the at least one dynamic characteristic value, and the clustering processor is adapted to cluster the further pruned first subsets and the further pruned second subsets, respectively.

21. A method for extracting photoplethysmographic signals that indicate the physiological state of a subject based on remote camera observation, The steps include acquiring a camera image from a camera configured to monitor at least one body part of a target, The steps include: dividing a first region representing an image portion in the camera image in which tissue of interest is presumed to exist into a first subset of pixels; and dividing a second region, different from the first region, representing an image portion in the camera image in which tissue of interest is presumed not to exist into a second subset of pixels; For each subset of the first plurality of subsets and each subset of the second plurality of subsets, the step of determining at least one dynamic characteristic value that represents the temporal signal dynamics of the subset in the sequence of camera images, The steps include clustering the first subsets and the second subsets based on at least one dynamic characteristic value so as to group subsets of pixels into clusters of similar signal dynamics, A step of selecting at least one cluster from a subset cluster obtained by the step of clustering the first subset based on cluster size and / or cluster configuration, wherein the step of selecting is to exclude from the selection each cluster of the first subset that substantially corresponds in the associated signal dynamics to the signal dynamics of the cluster of the second subset, A step of providing an output, wherein the output has a region of interest definition that defines pixels forming at least one selected cluster for use in extracting a photoplethysmography signal from a camera image acquired by the camera, and / or a photoplethysmography signal extracted from the region of interest in the camera image acquired by the camera. A method of having.

22. The method according to claim 21, wherein the first region and the second region are divided into a first pixel subset and a second plurality of pixel subsets, respectively, based on predetermined prior knowledge assumptions about the likelihood of the presence and non-prevalence of the target tissue.

23. A diagnostic imaging system having an examination zone, A camera placed in the aforementioned inspection zone to acquire images from the subject being inspected, As input, the device according to any one of claims 1 to 20, which is operablely connected to the camera to receive a camera image from the camera, A system that has

24. A computer program for performing the method described in claim 21 or 22 when executed by a computing device.