Photoplethysmography signal extraction
By automatically determining the region of interest in the camera image and clustering processing, the problem of difficulty in extracting PPG signals during medical imaging is solved, and more accurate and stable PPG signal detection is achieved.
Patent Information
- Application Number
- CN202380076818.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-25
- Publication Date
- 2025-06-13
AI Technical Summary
During medical imaging, camera-based photoplethysmography (PPG) signals are difficult to accurately extract in an interfering environment, especially during magnetic resonance imaging (MRI), where noise and strong magnetic fields of the device interfere with the detection of PPG signals.
By automatically determining the region of interest (ROI) from the camera image, using prior knowledge and dynamic feature analysis, the image region is segmented into a subset containing and not containing the tissue of interest, and clustering is performed to select a subset of pixels suitable for extracting the PPG signal.
This method can improve the accuracy and stability of PPG signals while reducing interference signals, reduce post-processing requirements, and improve the utilization efficiency of computing resources.
Smart Images

Figure CN120153409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and, more particularly, to determining a photoplethysmogram (PPG) signal from an object based on camera images. More specifically, the present invention relates to an apparatus, method, system, and computer program product for determining a photoplethysmogram signal by camera observation. Background Art
[0002] Photoplethysmography (PPG) is a technique that uses changes in the interaction of light (e.g., light absorption) to measure changes in blood volume in the microvascular system. For example, a simple PPG sensor can consist of a light source for irradiating living tissue such as skin, mucous membranes, and other semi-transparent tissues, and a light detector such as a photodiode, phototransistor, and / or photomultiplier tube to measure the amount of light scattered and / or absorbed by the tissue. The detected signal indicates changes in blood volume in the microvascular system and thus provides a non-invasive way to measure (near) real-time changes in blood volume. PPG has a wide range of applications, including monitoring blood oxygen, pulse rate, current cardiac phase, and blood pressure. For example, the information contained in the PPG signal can be used for triggering and / or gating during a magnetic resonance imaging (MRI) procedure.
[0003] In addition, it is known in the art to use camera observations of an object (e.g., a patient undergoing a medical imaging or treatment procedure) to obtain useful signals indicative of the physiological parameters of the object. For example, knowledge of the cardiovascular and / or respiratory function of an object is useful for monitoring the (e.g.) health status of a patient during surgery, to consider when processing images obtained during a diagnostic imaging procedure and / or to guide or control the image acquisition of the procedure being performed in a relevant manner.
[0004] Remote photoplethysmography (rPPG) (also known as camera-PPG or video-PPG) is a method for non-contact PPG signal extraction that uses a camera to capture images of the tissue under study. Camera-based PPG has several advantages over conventional (contact) PPG, including the ability to capture a larger tissue area, which can be helpful in cases where the tissue is not uniformly illuminated, and / or the ability to simultaneously capture multiple PPG signals from different parts of the tissue, e.g., to improve the accuracy of the results by averaging.
[0005] In typical methods of obtaining rPPG signals, first, a region of interest (ROI), such as a patch of skin tissue, is identified in one or more images (e.g., a video stream) obtained by a camera. Based on the dynamic changes in the pixel values corresponding to the ROI, a PPG is then constructed. This can involve various image / signal processing techniques, such as downmixing and post-processing of color signals to clean the signal from unwanted components. Then, features of interest, such as the pulse rate of the object and / or markers (e.g., time points) that identify one or more specific phases in the cardiac cycle, can be extracted from the PPG signal thus obtained. For example, a trigger generator system can use such markers to trigger MRI pre-pulses and / or control the data acquisition window.
[0006] In MRI, ROI identification can advantageously use prior knowledge about the MRI environment, since the environment generally remains substantially the same across different sessions. Illumination level information (e.g., favoring regions that are neither too dark nor too bright in order to select well-illuminated skin patches) and pulsatility information (i.e., indicating living skin) are also typically considered.
[0007] The dynamic range of the camera-based PPG signal itself is very small, i.e., on the order of about 10 -3 relative to the average illumination (average pixel value offset), and also small for other factors that can affect the detected pixel signal, making carefully controlled illumination conditions often an important consideration for achieving continuous and accurate monitoring. Therefore, using a high-quality, stable infrared illumination source can be considered advantageous. After modulation through interaction with the tissue of interest, a simple setup can use a single-channel near-infrared (NIR) illumination source and a camera that picks up the NIR signal. The signals of different pixels in the ROI region can be combined to construct the PPG signal.
[0008] Interference components in the signal can be removed through signal decomposition techniques. This can include interference caused by the movement of the object, i.e., the motion signal in the image can be measured and removed from the (raw) PPG signal. Similarly, general light variations can be measured and removed from the PPG signal.
[0009] Illustrative methods that employ the known pulsatile properties of skin-based PPG signals can generally be referred to as "living skin models". In such methods, a camera image or a portion thereof (such as the region in the image where the patient is typically positioned for a predetermined procedure) preselected based on prior knowledge is divided into small patches, e.g., a grid of rectangular patches adjacent to their neighborhoods, to form a partition of the image or a preselected portion thereof.
[0010] Then, for each patch and the motion vector of each patch, an average pixel intensity level and spread (e.g., standard deviation) can be determined (e.g., using an optical flow algorithm). When the average pixel level, spread, and motion all meet predetermined conditions, the patch is selected as a uniform non-moving patch for further use, e.g., thus pruning patches that do not exhibit these desired properties. For example, the average pixel level condition can correspond to an illumination condition that falls within a comfortable margin of the dynamic range of the camera system and is consistent with the typical characteristics of the tissue of interest under such illumination conditions. When the spread and motion are below (corresponding) predetermined thresholds, a certain level of stability for determining the signal of interest can be assumed, e.g., the patch will not be overly affected by various noise sources and / or errors, nor by strong motion.
[0011] For patches not retained by the pruning process, the average value as a function of time can then be analyzed. After applying the Fourier transform, the frequency components located within a predetermined range (e.g., representing the heart rate band) can be examined to find spectral peaks within that range. If the strongest peak is sufficiently distinct, e.g., the energy at the peak is sufficiently high relative to the average value of the range (and / or relative to the second strongest peak), the patch is associated with the frequency of that dominant peak (the strongest spectral peak). This value can be referred to as the patch frequency. Finally, by clustering the patch frequencies, the largest cluster can be selected as the most likely candidate representing the PPG signal, i.e., the pixels of the patches belonging to that cluster can be used to define the region of interest for extracting the PPG signal after this initial setup / calibration process.
[0012] For example, US10441173 B2 describes a prior art method for extracting physiological information (e.g., PPG signal) from remote camera observation data. A data stream including image data representing an observation area of an object of interest is received. A plurality of sub-regions within the region are defined, and a classifier classifies the sub-regions into indicative type regions and auxiliary type regions, where the former includes the region(s) that at least partially represent the object of interest, and the latter includes the reference region(s). Then, the sub-region(s) classified as the region of interest are further processed to obtain important information of interest, such as the PPG signal. SUMMARY OF THE INVENTION
[0013] The aim of embodiments of the present invention is to provide good and effective means and methods for determining a PPG signal from an object (e.g., a patient undergoing a medical imaging process) based on remote observation of, for example, a video stream of a camera image of the object. In particular, a region of interest that is particularly suitable for extracting such a PPG signal from a camera image is determined, and / or the ROI is used to extract and provide such a PPG signal.
[0014] Advantages of embodiments of the present invention are that remote imaging can be used, e.g., such that direct contact with the object is not required and / or such that the device can be placed at a considerable distance from the object. For example, in a medical environment, it may be preferred to keep the area immediately adjacent to the patient free of obstacles, e.g., such that procedures can be performed without unnecessary hindrance.
[0015] In addition, in medical imaging, diagnostic imaging techniques may interfere with other electronic devices.
[0016] In magnetic resonance imaging (MRI) in particular, the hardware of the MR scanner (e.g., radiofrequency antennas and / or magnetic field gradient coils) may interfere with the correct operation of other electronic devices. For example, a PPG sensor may pick up noise due to radiofrequency signals and / or be affected by strong magnetic fields (and / or field gradients and / or dynamic field changes). Similarly, an illumination source used to illuminate the patient's skin for PPG detection may be interfered with by the MR scanner hardware, causing illumination variations that may interfere with PPG detection.
[0017] Other advantages are that an accurate and / or stable PPG signal can be obtained, e.g., by reducing and / or avoiding carrying interference signals (such as random noise or systematic errors) into the PPG output signal inferred from the original data stream.
[0018] Thus, the output signal so provided will require less post-processing, e.g., further cleaning of the signal. This can improve the efficient use of computing resources (e.g., processing power, memory resources, power consumption, etc.), and / or can improve the accuracy and stability of the results derived from the output. For example, many signal cleaning strategies require that the measured perturbation signal and the perturbation in the original PPG be synchronous, and this condition is not always met. By carefully selecting the pixels from which the signal is derived, a more stable and representative PPG signal can be obtained, where the residual interference components can typically be (relatively) smaller and less random.
[0019] Advantages of embodiments of the present invention are that useful signals can be automatically determined (e.g., by an algorithm) based on camera observations of the object. In many environments, a suitable camera system may already exist or can be installed with a simple upgrade, such that the embodiments can be applied to existing systems without complex and / or expensive intervention.
[0020] Advantages of embodiments of the present invention are that signals can be determined based on camera observations, which can be used to trigger and / or system control during diagnostic imaging processes using a scanner system (such as an MRI, CT, SPECT, or PET scanner), which can advantageously improve diagnostic image quality. Such a trigger signal can also be used, for example, in radiotherapy and similar (e.g., therapeutic) processes, and is not necessarily strictly limited to medical imaging applications. In addition, other application areas of PPG acquisition are not excluded either.
[0021] Advantages of embodiments of the present invention are that a good selection of the region of interest (ROI) of pixels representing the PPG pulsation signal can be obtained, which has such a substantial component. In particular, prior knowledge can be advantageously used to prevent unwanted signal components (i.e., interference signals, noise, and / or other perturbations) from contributing to the extracted PPG signal. For example, a refined pixel screening process can be applied based on knowledge of the attributes and / or pixel signal content of the desired region and the undesired region to select a suitable ROI. Knowledge of non-skin locations (or more generally, locations that do not represent the type of tissue of interest to be monitored) can be advantageously used to extract pulsation features. In addition, morphological knowledge of the characteristics of promising candidate regions can be applied to ROI selection. Prior knowledge of the nature of the pulsations associated with PPG can also be advantageously used.
[0022] Devices, systems, methods, and computer program products according to embodiments of the present invention achieve the above object.
[0023] In a first aspect, the present invention relates to a method for extracting a photoplethysmography signal indicative of the physiology of an object based on remote camera observations (e.g., from a video stream obtained from the remote camera observations), e.g., a computer-implemented method, e.g., a method for determining a photoplethysmography signal and / or for determining an image region of interest for extracting a photoplethysmography signal. The method includes acquiring camera images (e.g., a video image stream) from a camera configured to monitor at least one body part of an object. The method includes dividing a first region into a first plurality of pixel subsets and dividing a second region into a second plurality of pixel subsets, the first region representing an image portion of the camera image where it is assumed that tissue of interest exists, and the second region being different from the first region and representing an image portion of the camera image where it is assumed that tissue of interest does not exist.
[0024] The method further includes determining, for each subset of the (e.g., trimmed) first plurality of subsets and the second plurality of subsets, at least one dynamic eigenvalue indicative of the temporal signal dynamics of the subset (the pixels forming the subset) in a sequence of camera images.
[0025] The method includes clustering a first plurality of subsets (e.g., trimmed) and a second plurality of subsets (e.g., trimmed) based on the at least one dynamic eigenvalue respectively, so as to group subsets of pixels into clusters of similar signal dynamics. By using (substantially) the same clustering variables (e.g., peak frequency and / or amplitude, or one or more values substantially equivalent thereto) to cluster subsets (e.g., patches) of a first region and to cluster subsets (e.g., patches) of a second region, information from putative non-related clusters in the second region can be considered when selecting one or more relevant clusters in the first region, as described below.
[0026] The method includes selecting at least one cluster from the clusters of subsets obtained by clustering the first plurality of subsets based on cluster size and / or cluster morphology, wherein the step of selecting the (one or more) clusters excludes from the selection each cluster in the first plurality of subsets whose associated temporal signal dynamics substantially correspond to the temporal signal dynamics of a cluster in the second plurality of subsets.
[0027] The method further includes providing an output. The output includes a region of interest definition to define the pixels forming the selected at least one cluster and / or a photoplethysmogram signal extracted from a region of interest in the camera image based on another image sequence acquired by the camera, for example.
[0028] In a method according to an embodiment of the present invention, a subset generator is configured to identify anatomical key points in a sequence of camera images. The subset generator is configured to use the identified anatomical key points to determine a first plurality of subsets and a second plurality of subsets.
[0029] The anatomical key points are positions of partitions of the body of a reference object. For example, they can indicate the positions of joints, positions on the face, or other body partitions. There are currently commercial software capable of automatically identifying anatomical key points. Even video game systems such as the Microsoft Xbox have such software capabilities built in. Neural networks can also be used, such as a convolutional neural network trained using images labeled with anatomical key points.
[0030] In a method according to an embodiment of the present invention, the subset generator is further configured to determine a first plurality of subsets and a second plurality of subsets by fitting a template to the identified anatomical key points in a sequence of camera images. The template can be fitted and / or deformed to the anatomical key points to provide the first plurality of subsets and the second plurality of subsets. This can provide an effective means for identifying regions of an object for the first plurality of subsets and the second plurality of subsets.
[0031] In the method according to an embodiment of the present invention, the subset generator includes a neural network configured to output anatomical key points in response to receiving a sequence of the camera images as input. As described above, a neural network trained with images labeled with anatomical key points can be used. Various types of neural networks can be used. For example, a convolutional neural network, a U-Net, a ResNet, or a network with fully connected layers can be configured to receive an image as input and output the positions of the anatomical key points.
[0032] In the method according to an embodiment of the present invention, at least one dynamic eigenvalue includes a main spectral peak, and wherein the signal dynamic analyzer is configured to determine the main spectral peaks of the first plurality of subsets and the second plurality of subsets by performing a fast Fourier transform of the pixel values to detect the main spectral peak. For each subset among both the first plurality of subsets and the second plurality of subsets, multiple images can be used to record the average attributes (such as brightness, contrast, position in the chromatogram, color change, or other attributes) of the subset over time. Then, a Fourier transform can be performed on the sequence of values. The frequency with the maximum peak will be the main spectral peak.
[0033] In the method according to an embodiment of the present invention, the clustering processor is configured to cluster the first plurality of subsets and the second plurality of subsets based on the main spectral peaks. In this example, each subset in the first plurality of subsets is clustered according to its main spectral peak. Each subset in the second plurality of subsets is also clustered according to its main spectral peak. This provides an easy way to group or cluster the subsets in a manner that can be used to generate a photoplethysmogram signal.
[0034] In the method according to an embodiment of the present invention, the main spectral peak is restricted within the heart rate frequency band. This can effectively reduce or eliminate the error signals in the photoplethysmogram signal.
[0035] In the method according to an embodiment of the present invention, the heart rate frequency band is between 5 beats per minute and 300 beats per minute.
[0036] In the method according to an embodiment of the present invention, the selector is configured to select the largest cluster in the first plurality of subsets, which has a different main spectral peak from the largest cluster in the second plurality of subsets. Since the second plurality of subsets is not used to image the skin of the object, the largest cluster in the second plurality of subsets may be a noise signal rather than a photoplethysmogram signal. Therefore, this example can provide an improved means for reducing the error in the photoplethysmogram signal.
[0037] In the method according to an embodiment of the present invention, the selector is further configured to ignore the clusters with a spectral peak amplitude higher than a predetermined threshold during the selection process. This can also provide an effective means for reducing the error in the photoplethysmogram signal.
[0038] In the method according to an embodiment of the present invention, considering a known setting of the volume in the space where the camera is to be positioned relative to the object, the first region (and the second region respectively) can be divided into the first plurality of pixel subsets and the second plurality of pixel subsets based on a predetermined prior knowledge assumption of where (and respectively where not) the tissue of interest may be present in the camera image.
[0039] In the method according to an embodiment of the present invention, selecting at least one cluster may include selecting, from the first plurality of subsets, one or more largest clusters that are not excluded based on the correspondence with the clusters in the second plurality of subsets.
[0040] In the method according to an embodiment of the present invention, selecting at least one cluster may include determining, for each cluster in the first plurality of subsets, a shape metric representing the shape of the image region formed by the pixels (or patches) in the cluster. The selection may be based on the shape metric of at least one selected cluster being consistent with a predetermined target shape of the region of interest, such as an oval or elliptical shape indicating the exposed skin of the face or forehead of the object.
[0041] In the method according to an embodiment of the present invention, selecting at least one cluster may be based on the ranking of the clusters in the first plurality of subsets (e.g., not excluded based on the correspondence with the clusters in the second plurality of subsets), where the ranking combines a score based on the size of the cluster and a score based on the shape metric. Optionally, the ranking may further include a further score based on the correspondence (e.g., best match) of the cluster with the clusters in the second plurality of subsets.
[0042] In the method according to an embodiment of the present invention, the first region (and the second region respectively) can be divided into the first (and second respectively) plurality of pixel subsets such that each subset groups pixels that are close together in space and / or value by forming the partition of the first region (and the second region respectively) into blocks or patches based on a spatial distance metric and / or forming the partition of the first region (and the second region respectively) into subsets based on an image co-domain distance and / or a combination of a spatial distance metric and an image co-domain distance. For example, in an embodiment considering different pixel subset sizes (e.g., in the optimization strategy discussed below), grouping the pixels into (first and second respectively) subsets may switch or shift in weight from forming more or only subsets based on distances in the spatial domain to forming more or only subsets based on distances in the image co-domain (e.g., the pixel value domain).
[0043] In the method according to an embodiment of the present invention, partitioning the first region (and the second region separately) into the first (and second respective) plurality of pixel subsets may include image segmentation and / or image processing of an image acquired by a camera and / or at least one image and / or spatial information acquired by another camera and / or another spatial information source (e.g., 3D camera, diagnostic imaging device, …) to determine the first region and the second region.
[0044] In the method according to an embodiment of the present invention, partitioning the first region (and the second region separately) into the first (and second respective) plurality of pixel subsets may include image subtraction between an image acquired in the presence of an object and a blank image acquired in the absence of an object, and / or similar image background compensation techniques, in order to detect the first region and / or the second region.
[0045] The method according to an embodiment of the present invention may include at least performing the step of repeatedly partitioning the first region into the first plurality of pixel subsets (e.g., to apply an optimization strategy) such that different iterations correspond to different values of at least one optimization parameter that at least represents the size of the pixel subsets. The method may further include evaluating a quality metric based on the pixel subsets obtained in each iteration in order to select a parameter value that obtains a sufficient or optimal value of the quality metric, and using the plurality of pixel subsets obtained for the selected parameter value in a further step of the method, e.g., such that clustering and thus the region of interest determined therefrom is based on the subsets corresponding to the parameter selection.
[0046] In the method according to an embodiment of the present invention, at least one optimization parameter may further include the size of a time observation window for determining at least one dynamic eigenvalue for each subset.
[0047] In the method according to an embodiment of the present invention, the quality metric used in the optimization strategy may include or consist of a signal-to-noise ratio.
[0048] In the method according to an embodiment of the present invention, determining at least one dynamic eigenvalue may include performing a Fourier transform in the time domain to obtain one or more time-frequency features.
[0049] The method according to an embodiment of the present invention may include pruning the first plurality of subsets based on a predetermined criterion for one or more values determined on a subset-by-subset basis in order to reject subsets having a low likelihood of a homogeneous region corresponding to the tissue of interest, and may further include pruning the second plurality of subsets based on the (same and / or similar) predetermined criterion. Thus, clustering may be applied to the pruned first plurality of subsets and the pruned second plurality of subsets.
[0050] In the method according to an embodiment of the present invention, the pruning may be performed based on the predetermined criteria, where the predetermined criteria may include the average pixel value for each subset and / or the spread of the pixel values of each subset and / or one or more criteria for at least one value indicating the motion associated with the subset, so as to reject subsets showing excessive motion, non-uniform pixel value distribution, and / or pixel values having an average outside a predetermined target range. It should be understood that various alternative metrics for the average value (e.g., typically a centrality metric of the pixel values over the subset), spread (e.g., typically a dispersion metric of the pixel values over the subset), and / or motion (e.g., a metric related or associated with movement, such as determined by optical flow) may be used.
[0051] The method according to an embodiment of the present invention may include further pruning the first and / or second plurality of subsets based on the at least one dynamic eigenvalue before applying the clustering to the first and / or second plurality of subsets for further pruning (e.g., which may have been initially pruned by the steps described above). The further pruning may be based on predetermined (one or more) criteria for the spectral energy of the maximum spectral peak of the Fourier spectrum determined for each subset and / or on values determined therefrom.
[0052] In a second aspect, the present invention relates to a device for extracting a photoplethysmogram signal indicative of the physiology of an object based on remote camera observation, such as a device for extracting a photoplethysmogram signal and / or for providing a definition of a region of interest for extracting a photoplethysmogram signal. The device includes an input for receiving a camera image from a camera configured to monitor at least one body part of an object.
[0053] The device includes a subset generator for dividing a first region into a first plurality of pixel subsets and dividing a second region into a second plurality of pixel subsets, where the first region represents an image portion of the camera image where it is assumed that the tissue of interest exists, and the second region is different from the first region and represents an image portion of the camera image where it is assumed that the tissue of interest does not exist.
[0054] The device includes a signal dynamic analyzer for determining at least one dynamic eigenvalue for each subset of the first plurality of subsets and (e.g., independently) for each subset of the second plurality of subsets, where the dynamic eigenvalue indicates the temporal signal dynamics of the subset in a sequence of camera images received via the input.
[0055] The device includes a clustering processor that is configured to cluster a first plurality of subsets and a second plurality of subsets respectively based on at least one dynamic eigenvalue (e.g., clustering the first plurality of subsets independently of clustering the second plurality of subsets, and vice versa) to group pixel subsets into clusters of similar signal dynamics. It should be understood that (image) clustering algorithms generally can also be considered to take into account spatial relationships, e.g., grouping subsets that are spatially close together, such as being similar in the associated (one or more) dynamic features.
[0056] The device includes a selector that is configured to select at least one cluster from the clusters in the first plurality of subsets provided by the clustering processor based on cluster size and / or cluster morphology. The selector is adapted to exclude from the selection each cluster in the first plurality of subsets whose associated signal dynamics substantially correspond to the signal dynamics of the clusters in the second plurality of subsets.
[0057] The device includes an output unit that is configured to: output a region of interest definition that defines the pixels forming at least one cluster selected by the selector for extracting a photoplethysmogram signal from a camera image acquired by a camera; and / or output a photoplethysmogram signal extracted from a region of interest in the camera image received via an input unit.
[0058] The device according to an embodiment of the present invention may include the camera, e.g., a camera suitable for PPG signal extraction, such as an infrared camera, a color camera, and / or a camera having a predetermined PPG filter (e.g., a selective green filter). It should be understood that the device may also include (optionally) a light source to illuminate an object (or at least the body part of interest) with light including suitable spectral components corresponding to the sensitivity of the camera. The camera may include a single camera or a combination of multiple cameras.
[0059] The device according to an embodiment of the present invention may include a photoplethysmogram signal extractor that is configured to extract a photoplethysmogram signal from a region of interest in the (sequence of) camera images and provide it as an output.
[0060] In the device according to an embodiment of the present invention, the subset generator may be adapted to divide a first region (and each of the second regions respectively) into a first plurality of pixel subsets (and a second plurality of pixel subsets) based on a predetermined prior knowledge assumption of where the tissue of interest may (respectively may not) be present, given a known setting of the volume of the space where the camera is located relative to the object. For example, the region may be defined via a user interface, may be pre-configured (e.g., stored in a configuration memory and / or hard-coded), or may be retrieved from an external data storage device, e.g., a location stored during the calibration process of the camera settings.
[0061] In a device according to an embodiment of the present invention, a selector for selecting at least one cluster may be adapted to select the largest (one or more) clusters from a first plurality of subsets that are not excluded based on the correspondence with clusters in a second plurality of subsets.
[0062] In a device according to an embodiment of the present invention, the selector may be adapted to determine a shape metric representing the shape of each cluster in the first plurality of subsets and to select the at least one cluster based on the shape metric of the selected cluster being consistent with a predetermined target shape of a region of interest tissue.
[0063] In a device according to an embodiment of the present invention, the selector may be adapted to determine a ranking of the clusters in the first plurality of subsets (e.g., not excluded based on the correspondence with clusters in the second plurality of subsets), where the ranking combines a first score based on the size of the cluster and a second score based on the shape metric of the cluster. Thus, the ranking can be used to determine a selection, e.g., selecting the highest-ranked cluster or the top N (e.g., 2, 3, …) highest-ranked clusters. Optionally, the ranking may further include a further score based on the correspondence (e.g., best match) of the cluster with clusters in the second plurality of subsets.
[0064] In a device according to an embodiment of the present invention, a subset generator may be adapted to divide a first region (and a second region, respectively) into the first plurality of pixel subsets (and the second plurality of pixel subsets, respectively) by forming partitions of the first region (and the second region, respectively) into blocks or patches based on a spatial distance metric and / or by forming partitions of the first region (and the second region, respectively) into subsets based on an image co-domain distance and / or a combination of a spatial distance metric and an image co-domain distance, such that each subset groups together pixels that are close in space and / or value.
[0065] In a device according to an embodiment of the present invention, the subset generator may be adapted to perform image segmentation (and / or other suitable processing) of a reference image and / or reference spatial information received via an input unit to determine the first region and the second region, where the reference image and / or reference spatial information is received via the input unit (11) and acquired by the camera, another camera, and / or another spatial information source. For example, additionally or alternatively, the subset generator may be adapted to perform image subtraction (e.g., as the other suitable processing or an element thereof) between a reference image acquired in the presence of an object and a blank image acquired in the absence of an object, and / or apply a similar image background compensation technique, for determining the first region and the second region, e.g., (optionally) in combination with image segmentation of an image after removing an irrelevant background.
[0066] Devices according to embodiments of the present invention may include an optimizer, wherein the subset generator is adapted to repeatedly divide a first region into the first plurality of pixel subsets for different values of at least one optimization parameter representing the size of (at least) a pixel subset, and wherein the optimizer is adapted to calculate and evaluate a quality metric based on the pixel subsets obtained in each iteration for different values of the at least one optimization parameter in order to select a parameter value that obtains a sufficient or optimal value of the quality metric. Thus, the selected (one or more) parameter values, as well as the corresponding divisions of the first region and the second region into the first plurality of subsets and the second plurality of subsets, can then be used by the clustering processor and the selector.
[0067] In a device according to an embodiment of the present invention, the optimizer may also use the size of the time observation window used by the signal dynamic analyzer as a component of the at least one optimization parameter.
[0068] In a device according to an embodiment of the present invention, the quality metric calculated by the optimizer may include or consist of a signal-to-noise ratio.
[0069] In a device according to an embodiment of the present invention, the signal dynamic analyzer may be adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency features for determining the at least one dynamic eigenvalue.
[0070] Devices according to embodiments of the present invention may include a subset eliminator for pruning the first plurality of subsets according to a predetermined criterion based on one or more values determined on a subset-by-subset basis in order to reject subsets having a low likelihood of corresponding to a homogeneous region of the tissue of interest, and for pruning the second plurality of subsets based on the same and / or additional predetermined criteria, and wherein the clustering processor may be adapted to cluster the pruned first plurality of subsets and the pruned second plurality of subsets respectively.
[0071] In a device according to an embodiment of the present invention, the subset eliminator may be adapted to prune the first plurality of subsets and / or the second plurality of subsets based on a predetermined criterion, the predetermined criterion including one or more of the following criteria: the average pixel value of each subset, the spread of the pixel values of each subset, and / or at least one value indicating the motion associated with the subset, such that the subset eliminator is adapted to reject pixel objects showing excessive motion, uneven pixel value distributions, and / or pixel values having an average outside a predetermined target range.
[0072] In a device according to an embodiment of the present invention, the subset eliminator may also be adapted to apply further pruning of the first plurality of subsets and / or the second plurality of subsets based on at least one dynamic eigenvalue, and wherein the clustering processor is adapted to cluster the further pruned first plurality of subsets and the further pruned second plurality of subsets respectively.
[0073] In an apparatus according to an embodiment of the present invention, the subset eliminator may be adapted to apply the further pruning, wherein the further pruning is based on at least one predetermined criterion of the spectral energy of the maximum spectral peak of the Fourier spectrum determined for each subset and / or based on a value determined therefrom.
[0074] In a third aspect, the present invention relates to a diagnostic imaging system having an examination area, wherein the system comprises: a camera for acquiring an image from an object when the object is being examined in the examination area; and an apparatus according to an embodiment of the present invention, operatively connected to the camera to receive the camera image as an input from the camera.
[0075] In a fourth aspect, the present invention relates to a computer program product for performing a method according to an embodiment of the present invention when executed by a computing device.
[0076] The independent and dependent claims describe specific and preferred features of the present invention. The features of the dependent claims may be combined with the features of the independent claims and the features of other dependent claims considered appropriate, and need not be only as explicitly stated in the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 A method according to an embodiment of the present invention is shown.
[0078] Figure 2 An apparatus according to an embodiment of the present invention is shown.
[0079] The drawings are schematic and not restrictive. The elements in the drawings need not be shown to scale. The present invention need not be limited to the specific embodiments of the present invention shown in the drawings. DETAILED DESCRIPTION
[0080] Although exemplary embodiments are described below, the present invention is limited only by the appended claims. The appended claims are hereby expressly incorporated into this detailed description, where each claim and each combination of claims permitted by the dependency structure defined by the claims form a separate embodiment of the present invention.
[0081] The word "comprising" used in the claims is not limited to the features, elements or steps described hereinafter and does not exclude additional features, elements or steps. Thus, it clearly states the presence of the recited features without excluding the further presence or addition of one or more features.
[0082] In this detailed description, various specific details are presented. Embodiments of the present invention may be carried out without these specific details. In addition, for the sake of clarity and conciseness of this disclosure, well-known features, elements and / or steps need not be described in detail.
[0083] In a first aspect, the present invention relates to a method for determining a photoplethysmogram (PPG) signal indicative of the physiology of an object based on remote camera observations, and / or for determining an image region of interest (ROI) to extract such a PPG signal from prior art methods such as by receiving the ROI (e.g., as an image map, as a mask image, or as a corresponding mask image in a video stream) for subsequent determination of the PPG. Since methods for extracting the PPG signal are well known in the art once a suitable definition of the available image content is provided (e.g., a mask or other suitable description of the region of interest to be used), such methods are not discussed in detail, but it should be understood that the method according to an embodiment of the present invention may include such well-known signal extraction steps in order to obtain and provide a high-quality PPG signal (e.g., having a good signal-to-noise ratio and / or robustness against movement and / or other artifacts) based on the region of interest determined according to an embodiment of the present invention.
[0084] As is known in the art, photoplethysmography (PPG) is a favorable low-cost optical method for detecting changes in the microvascular blood volume associated with the heartbeat dynamics. It can advantageously provide non-invasive measurements, such as using camera imaging. The raw PPG signal includes a pulsatile component (i.e., the "AC" component) synchronized with the cardiac cycle, i.e., the PPG signal of interest to be extracted. Frequencies that do not indicate the signal of interest can be simply removed by filtering, e.g., to remove irrelevant slow ("DC" and / or lower frequency) components, such as those due to average illumination conditions and / or changes that may be caused by slow movement, breathing, temperature tuning (e.g., especially when using imaging in the infrared domain) and / or other such factors. Higher frequencies outside the physiological range of interest (e.g., which may be substantially associated with noise) can be equivalently filtered out. Thus, once the set of pixels in (each) video frame from which the signal is to be extracted has been established, i.e., the set of pixels corresponding to the region of interest definition provided by the method according to the embodiment, conventional signal processing techniques known in the art (e.g., filtering and / or more advanced techniques, such as machine learning-based) can be applied to extract a useful PPG signal from the real-time video stream provided by camera monitoring.
[0085] Figure 1Illustrates an illustrative method 100 according to an embodiment. The method includes acquiring 115 camera images (e.g., an image stream) from a camera, such as capturing video or a live stream, the camera being configured to monitor at least one body part of an object. The camera images can be obtained by directly observing the object (e.g., the object can be in the direct line of sight of the (one or more) cameras) or indirectly (e.g., using one or more reflective surfaces in the optical path). Lenses and / or other optical elements can optionally be used to optimize the field of view of the camera, such as magnifying the (one or more) relevant body parts, focusing the image, etc.
[0086] The camera can be an infrared camera (e.g., operating in the near-infrared range NIR, such as a camera sensitive in the range of 800 nm to 850 nm), such that a PPG signal can be extracted without the need to adjust (visible) lighting conditions and / or without being affected by ambient light (in the visible spectrum). Since the human eye is insensitive to infrared light, the object being monitored is not affected by unpleasant, disturbing, and / or uncomfortable (or even painful) bright light. However, the embodiment is not limited thereto. For example, a PPG signal can also be extracted from imaging in the visible light range. For example, the light absorption of (de)oxygenated hemoglobin is stronger in the (near) infrared wavelength range than in (e.g., red) visible light, and infrared light can penetrate deeper into the skin (e.g., the effect of melanin in the skin is less significant). In addition, for example, an infrared camera based on a semiconductor pixel detector can have higher sensitivity in the NIR range (e.g., 800 - 850 nm) than at wavelengths above 900 nm, which can result in a higher signal-to-noise ratio and / or better operation when the body part of interest is not ideally illuminated. Other light ranges (besides IR or NIR) can also be used for PPG signal detection, as it has been observed that, for example, video imaging in the red or green wavelength ranges can also provide sufficient information to detect slight changes in the intensity used.
[0087] The method can include acquiring camera images from a camera configured to monitor a body part of an object during a medical and / or diagnostic examination or intervention. For example, the examination can be a magnetic resonance imaging, computed tomography, positron emission tomography, or single photon emission computed tomography examination. Examples of interventions include treatments and related procedures such as surgery and radiotherapy. The embodiment can also be applied to general object or patient monitoring, such as in an intensive care unit, a patient ward, or a non-medical environment, such as a nursing home. In particular, the PPG signal determined by the method can be used to monitor the health and / or status of the object and / or to control a system or device for the examination or intervention. The (one or more) body parts of interest can include the face or a part thereof, such as the forehead, but are not limited thereto (e.g., the skin of another body part, such as the arm, chest, leg, abdomen,... can also be used).
[0088] The method includes dividing 101 at least one first region in a camera image (e.g., in a reference image) into pixel subsets, e.g., patches. The pixel subsets (e.g., patches) may form a partition of the first region (e.g., such that the entire first region is covered by the union of non-overlapping patches). However, the latter need not be necessary, e.g., some overlap between patches need not be excluded.
[0089] In addition, in addition to the step of dividing 101 the first region into patches, according to an embodiment of the present invention, a second region in the camera image is also divided into pixel subsets (e.g., patches). The first region and the second region are different image regions, e.g., typically separate (and non-empty) parts of the image, e.g., such that there is no overlap between the first region and the second region, each region is not equal to the entire image, and each region is not empty. However, there may be some non-substantial overlap, e.g., which is negligible to some extent, such as pixels of a shared boundary contour between the two regions. For example, with respect to the union of the two regions in terms of size, the intersection between the two regions may be less than 10%, preferably less than 5%, even more preferably less than 1%, e.g., substantially 0%, e.g., when evaluating the ratio of the number of pixels in the intersection to the number of pixels in the union.
[0090] The first region represents image content where it is assumed that the tissue of interest is present, while the second region represents image content where it is assumed that the tissue of interest is not present (e.g., with a lower probability or less likely to be found).
[0091] The union of the first region and the second region may form the entire image, e.g., the two regions may divide the image or may completely cover the image. However, this need not be the case. For example, background image content (e.g., an image region where an object is less likely to be present) may be ignored, while the first region may define an image area where the tissue of interest (e.g., exposed skin) may be visible (in a suitable form, e.g., sufficiently illuminated, etc.), and the second region may define an area where the tissue may not be clearly visible (but the image content need not be independent of the object). Such areas where the tissue of interest is less likely to be clearly visible may include parts of the body of an object covered by clothing, glasses, a blanket, etc., where a clear view may be blocked by the device, and / or where body features are more likely to be found as error sources when inferring the PPG signal, such as eyes, nose, hair, fingers, nails, navel, nipples, skin folds, and / or other such features.
[0092] Thus, according to an embodiment of the present invention, by analyzing a first region having image content that may be relevant and a second region having image content that may be irrelevant, a more conventional "living skin" method can be extended, in which patches of an image are analyzed to obtain useful signal content. Thus, in addition to analyzing candidate active skin patches, patches of (prior known, i.e., assumed) non-skin regions are also considered.
[0093] For this initial region partitioning step, a reference image can be used, which can be the first image at the start of the process, a random image from an (e.g., short) image sequence, an image selected based on a quality metric (e.g., to exclude effects due to the initial stability of the camera acquisition process, e.g., autofocus, gain calibration, and / or others), and / or can be based on the aggregation of several images constructed, e.g., by averaging, from multiple (short subsequence) camera images. However, since the object and the camera can be considered to be substantially static (i.e., the camera is preferably fixed and preferably the object's movement is avoided), the selection of such a reference image may not necessarily be required. For example, partitioning into patches can be considered to be at least roughly applicable to any subsequent (and / or previous) acquired image frames in the same session, and if based entirely on predetermined assumptions (e.g., when the camera is always used in the same spatial configuration and the object is always positioned relative to the camera used for the session in the same way), even applicable to any such session.
[0094] The first region and / or a fixed set of patches covering the first region can be predetermined, e.g., can be defined as a function of the position where it is assumed that the tissue of interest (of the object's body) is likely to be present. The second region and / or a fixed set of patches covering the second region can also be predetermined, e.g., defined as a function of the position in the (one or more) images where it is assumed that the tissue of interest is absent. Similarly, the patches can be defined in a fixed manner, e.g., a predetermined partitioning of a predetermined first region (and each of the second regions) of the camera image, regardless of the session and / or object involved in the process. In other words, a predetermined or situationally (independently) determined mapping of camera pixel coordinates (e.g., an image grid) can be used to determine the first region and the second region and / or the patches of each region, without the need for a reference image. The region definition and / or the patch definition (for each region) can be established independently of the (one or more) specific camera images and can thus be considered applicable to any acquired image, e.g., by definition, applicable to each individual image frame in the sequence obtained by extracting PPG information therefrom. The region definition and / or the patch definition can be based entirely on prior knowledge of the camera (and object) settings, and / or can be determined independently (e.g., based on another camera system and / or spatial data acquisition system).
[0095] This may be particularly advantageous in applications where the camera is typically fixed or positioned, oriented, focused, zoomed, and / or otherwise configured using a predefined reference frame and / or parameters (e.g., a fixed arrangement in a scanner room and / or a system using an electrified / automated configuration), and the object is typically positioned, oriented, etc. in the same or generally similar manner relative to the camera or relative to a reference frame having a known relationship to the camera's reference frame (possibly taking into account parameters such as automatically controlled camera orientation, position, etc. as described above).
[0096] For example, in an exemplary application used in a medical imaging room (or other medical facility where a patient is to be monitored), the camera can view a known area of the room (or more precisely, a cone in space), and the patient can be placed in a known position in the room (possibly taking into account parameters of the medical imaging system such as patient table translation and / or the patient's pose for the procedure), such that the general area where the tissue of interest lies in the camera's image can be easily inferred, and similarly, a second area where the tissue may not be present (or may not be imaged clearly enough) can also be determined.
[0097] Alternatively, the first region (and / or a set of patches partitioning the first region) and / or the second region (and / or a set of patches partitioning the second region) can be determined by image segmentation or similar image processing techniques. For example, a simple (e.g., two-component) segmentation can be performed to separate the body and / or one or more body parts and / or tissue of interest (e.g., exposed skin) (the first region) of the object from its background (the second region). A three or more component segmentation can be used to detect the skin (or other tissue of interest, i.e., the first region), the irrelevant parts of the body (the second region), and the background image regions (e.g., the room and / or the instrument; i.e., forming a third component and / or multiple other components that can be completely ignored). Other segmentation methods are not excluded. For example, in cases where the expected image content is for predefined multiple components that can be visually (and algorithmically) discriminated, at least one (not excluding combinations) can be associated with the first region, and at least one (or more) can be associated with the second region. For example, multiple image segments (i.e., more than two) can be expected based on assumed pixel attributes (e.g., intensity values, etc.), and explicitly considering such non-uniform second regions and / or (optional, further ignored) background regions (by accommodating additional segmented image patches) can advantageously improve the quality of the obtainable image segmentation.
[0098] As another example, a blank image (with no objects) of the same environment can be used to determine a first region by background subtraction, and a second region can be defined as its complement (which may include a normalization process and / or more advanced methods). Combinations of background subtraction and segmentation and / or other image processing techniques can also be used. For example, a "blank" background image can be subtracted to remove the static image components, and segmentation can be performed on the result to identify a first region composed of tissue of interest (e.g., exposed skin) and a second region composed of other non-static and / or variable elements in the image, such as covered body parts, hair, eyes, instruments (e.g., tubes, cables, breathing masks, etc.), and / or other such irrelevant image content, especially in cases where some dynamic changes in the image content may be expected in this second region (compared to the static content removed by image subtraction). It should be understood that the image subtraction mentioned above can be implemented by any equivalent method more complex than just simple arithmetic matrix subtraction, such as considering the normalization of the image and / or generally applying any suitable background compensation technique.
[0099] Image segmentation (or other suitable image processing) need not be limited to images obtained by a camera system, or need not (only) use the images for subsequent PPG extraction to be performed. For example, another camera can be configured to provide better image contrast and / or other advantageous features that allow the object (more generally, the body part or tissue region of interest) to be numerically separated from the non-relevant parts of the object (e.g., covered body parts by fabric). For example, the (one or more) cameras used for PPG signal extraction can be optimized to provide good sensitivity for detecting subtle PPG signals of interest (e.g., in the infrared spectrum and / or another specific part of the spectrum, such as the green light band), but may not be ideal for segmenting the image (e.g., discerning tissue of interest, such as skin, from other image content). Thus, another camera (or cameras) more suitable for this purpose can be used, such as a color camera or a depth camera, and to some extent, the segmentation result can be transformed into the reference frame of the camera used for PPG signal extraction. It should be understood that even when the deterministic relationship between camera reference frames is not known a priori, other techniques can be used to transform the segmentation map, e.g., by minimizing the mutual information between the image obtained by the PPG camera system and the image obtained by another camera system used for image segmentation (and / or for determining the first and second regions by similar processing techniques). Similarly, the information source used to infer the first and second regions, for example, by segmentation, need not be limited to a camera system in the narrowest sense, but can include any suitable device for acquiring spatial information or be composed of any suitable device for acquiring spatial information, including, for example, a medical imaging system, e.g., the medical imaging system can be intended to be used during PPG monitoring of an object.
[0100] Although reference has been made to segmentation, it should also be understood that this need not exclude more advanced methods, such as fitting a geometric model of an object (or a related body part or part of an object) to the image data and / or evaluating a trained machine learning model to determine a suitable (approximate) segmentation (i.e., generally, descriptors of the first and second regions, and / or directly determining descriptors of the patches forming each of the regions) as an output.
[0101] In the "in vivo skin" process and / or similar processes, it is generally assumed that the pulsatile component of interest (i.e., the PPG signal) stands out from the measurement noise, or in other words, it is sufficiently detectable and distinguishable in the acquired camera pixel signals (e.g., after suitable processing). In particular, the averaging operation on different pixels of each patch (or, generally, applying a suitable aggregation calculation, such as a measure of statistical centrality) should sufficiently reduce the sensor quantization noise and / or other noise components so as to be able to distinguish the available signal content (and the patch associated therewith) from the noise content in irrelevant and / or unsuitable patches, where the PPG signal is too weak, contaminated by heavy noise, confused by other irrelevant signals, or simply absent. Furthermore, as discussed in detail below, techniques for analyzing the dynamics (variation over time) of the signal are generally applied, such that a sufficiently long observation time is required to detect the signal component of interest. For example, it may be necessary to select a suitable window size for Fourier analysis, without losing too much time resolution and without increasing the sensitivity to errors too much.
[0102] Different from using a predetermined (and possibly arbitrary) patch size in the (one or more) steps of dividing the first region (and the second region) into patches, according to embodiments of the present invention, the patch size and / or the time observation window (e.g., Fourier window) for analyzing the signal dynamics can be optimized, for example, in order to obtain a good signal-to-noise ratio (SNR). For example, temporal fluctuations may exist in the raw pixel signals, which can be attributed to illumination conditions (e.g., when using natural lighting or a low-quality light source), environmental stability, and / or motion. In addition, physiological differences (such as natural heart rate variability) (within and / or between objects) can affect the optimal choice of the patch size (and / or time window) from one session to another. Therefore, at different times and / or in different environments, the intrinsic (heartbeat cycle-related) and extrinsic (e.g., illumination conditions) differences between sessions (e.g., between different applications of the method, such as for detecting PPG signals for different patients) may result in the choice of (e.g., optimal or optimized) parameters being session-dependent. Thus, the optimization of such (one or more) parameters according to embodiments can advantageously improve the accuracy, robustness, stability, and / or generally the quality of the processed PPG signal output.
[0103] For example, the optimal or at least good or suitable selection of the time observation window (e.g., FFT window selection) and the selection of the patch size can strongly depend on the heart rate and heart rate variability expected for each individual case, which, unfortunately, may be a priori unknown (or not accurately known). In fact, the goal of the method according to an embodiment can well be to accurately measure and / or monitor this heart rate and / or other cardiac parameters. Although in some cases, a larger patch size may be preferred, for example when using a shorter FFT window and / or reducing noise by (spatial) averaging, too large a patch size can also be disadvantageous. The probability of finding a uniform patch (e.g., substantially exclusively corresponding to the exposed skin area) decreases as the patch size increases. Therefore, it should be understood that there may be an optimum between these extremes, which is not a priori known and depends on each specific case.
[0104] In the case of the reference patch size, it should be understood that for a particular choice of the patch size parameter, the patch size does not have to be strictly constant over all patches. In other words, for a particular patch size parameter, some variation in the patch size is possible. For example, the step of dividing the first region (and / or the second region) into patches can be implemented in a way that allows some variation, e.g., such that the size parameter more accurately determines the average (or median,...) patch size to be used (in an iteration of the parameter selection).
[0105] Even when using a fixed (and relatively simple) method for dividing a region into patches with a predetermined and deterministic (uniform) patch size, e.g., partitioning the region by a Cartesian block pattern of a given block size, the irregular boundaries of the first / second region can still result in some edge regions containing fewer (or more) pixels than a typical more central (block-shaped) patch.
[0106] However, if desired, e.g., when using a processing technique that depends on (or is sensitive to) a uniform patch size (number of pixels per patch), a uniform patch size can be (optionally) enforced (for each iteration, for the corresponding particular choice of the patch size). For example, the (first / second) region definition can be adjusted to fit all patch sizes to be considered without splitting any patch near the boundary, e.g., by defining the region at a resolution (scale) corresponding to the least common multiple of the different patch scales to be evaluated. In another example, since the first (and second respective) regions do not have to exactly correspond to the tissue of interest (corresponding to irrelevant image content respectively), e.g., the regions can be merely a best guess based on available prior information, each (uniform-sized) patch (e.g., block) can also simply be assigned to the first region or the second region as a whole, e.g., based on majority voting, without too many problems and / or substantial consequences.
[0107] Thus, since in the method according to an embodiment of the present invention, the PPG signal of interest and the noise component (and / or generally, the interference contribution) can vary between sessions, for example, at least due to differences in the cardiac response pulsatility of each subject, 123 (e.g., optimize and / or tune) the size of the patches for analyzing temporal variations (e.g., the step of partitioning regions 101, 111) and / or the time window size (e.g., FFT window size) can be selected (directly or indirectly) based on the acquired image data, for example, so as to be above the noise floor but small enough to avoid or reduce the adverse effects associated with averaging over a larger region and / or analyzing the signal over a longer time frame. Such parameter selection 123 can be implemented in various ways.
[0108] For example, for different values of at least one optimization parameter representing the size of a subset of pixels, the step of partitioning 101 at least the first region into moments (and optionally also partitioning 111 the second region into patches and / or one or more additional steps as discussed below) can be repeated 125, e.g., repeating one or more steps for different patch sizes. Other parameters (e.g., in addition to size optimization) can also be adjusted, such as the time window size for signal dynamic analysis (e.g., Fourier analysis). After each iteration, a quality metric is calculated (directly or indirectly) based on the acquired image data organized in the set of patches for this parameter selection, e.g., based on the intermediate output of one or more steps performed in (e.g., each) iteration and / or based on the region of interest output thus obtained in the iteration, so as to indicate the fitness of the tested parameter selection.
[0109] For example, substantially the entire process can be performed in each iteration, e.g., determining the points of the region of interest, and the quality metric can be directly determined from the obtained region of interest for a particular parameter selection (and / or the PPG signal extracted from the video stream acquired by the camera using this ROI mask). Thus, the best region of interest indicated by the quality metric can be selected from multiple trials with different parameter selections, and / or the parameters can be adjusted based on the output of one or more previous iterations for each subsequent iteration. Without being limited thereto, the quality metric can include an estimate of the signal-to-noise ratio (SNR) of the extracted PPG signal for the obtained ROI, and / or can be more complex, e.g., considering various factors such as ROI size, SNR, signal contrast, and / or one or more known metrics of (general) signal quality and / or (specifically) PPG signal quality.
[0110] In an advantageously efficient method, simpler quality metrics can be used that do not require all steps to be performed, e.g., quality metrics that can be easily calculated based on intermediate results, e.g., based on the tiles obtained only by the (one or more) region partitioning steps, e.g., based on the tiles after applying the pruning steps discussed below, or based on the dynamic parameter estimation (e.g., Fourier analysis) steps discussed below. It should be understood that the (one or more) region partitioning (region partitioning 101 and / or 111) steps and / or the (one or more) pruning steps can be performed relatively efficiently compared to more computationally intensive signal dynamic analysis (Fourier analysis) and / or steps that depend thereon. Thus, advantageously, the quality metric can be efficiently calculated based on the (one or more) set of pruned tiles (first region or first and second regions) without considering the dynamics. However, e.g., especially when the time analysis window is also to be optimized, the quality metric can also be based on the intermediate output of other steps, e.g., considering one or more characteristics of the signal dynamics determined for the (e.g., pruned) tiles. For example, the signal dynamics can be characterized relatively efficiently, e.g., using dedicated Fourier transform hardware, multi-processors, and / or GPUs, such that using this information in (e.g., iterative) parameter optimization need not be prohibitive.
[0111] Furthermore, in the optimization step, instead of using a high-quality time dynamic analysis as used in other steps of the method as can be discussed below, e.g., different from fully (detailedly) calculating the Fourier spectrum, a simplified or approximate analysis of the time dynamics can be used as an alternative, e.g., by detecting one or more signal components only for specific frequencies or several specific frequencies and / or only for specific spectral bands or several spectral bands. For example, such limited information may be sufficient to calculate the quality metric (or be used as an element in its calculation), e.g., at least giving a rough indication of the signal quality with respect to the expected time dynamics. To illustrate this, the spectral power of a predetermined frequency (or frequency band) can be determined quickly and / or easily, where it is assumed that the frequency (frequency band) contains at least some cardiac cycle-related activity (e.g., based on the average heart rate), while in the more detailed analysis discussed below (e.g., after parameter selection), the peak frequency and peak signal intensity can be determined in detail by more demanding calculations.
[0112] Variations of this method can be easily envisioned, such as normalizing the spectral power (or other suitable values) of the pre-assumed relevant frequency (frequency band) by similar calculated values for frequencies or frequency bands expected to have no relevant (i.e., heart-related) activity.
[0113] In another illustrative method, scale space analysis can be used, e.g., where the signal is downsampled (in the time dimension) by various factors (e.g., a scale pyramid of powers of two), such that the mean, standard deviation, and / or other such measures (e.g., where each patch has been averaged) of the signal can be used as surrogates for the signal dynamics in different (approximate) frequency bands (e.g., depending on the corresponding time scale).
[0114] Quality metrics for parameter selection (e.g., optimization) 123 can generally express the goodness of fit of the parameter selection (e.g., patch scale and / or temporal analysis window) for which the quality metric is computed. Thus, the (one or more) parameter selections corresponding to the higher (or highest) (one or more) quality metric values can be selected for further use in determining the region of interest from which the PPG signal of interest can ultimately be extracted. Clearly, the quality metric can alternatively be equivalently defined as a cost function, e.g., such that lower values indicate more appropriate parameter selections.
[0115] For example, when a sufficient (e.g., at least some predetermined) number of patches (generally, subsets of pixels, i.e., "real" spatially connected patches and / or "virtual" aggregated patches, as described below) exhibit consistent signal intensity within an expected (e.g., predetermined) range, the quality metric can have a maximum value (or equivalently, the corresponding cost function can have a minimum value). The signal intensity can be estimated as an approximate measure of the PPG signal intensity, e.g., based on the average pixel value and / or deviation, based on the maximum spectral peak amplitude value, and / or based on intermediate results of methods discussed further below, or can be estimated directly, e.g., by extracting the PPG signal (e.g., a short string) using the parameter selection at hand to directly estimate the signal-to-noise ratio of the PPG signal for that parameter selection.
[0116] Given physiology, the expected predetermined range of signal intensity can be determined from the expected PPG signal properties. This can be achieved, but is not limited to, testing the signal intensity values of each patch for a normal distribution, e.g., calculating the Z-score of the mean (e.g., signal intensity, e.g., maximum spectral peak amplitude) associated with each patch relative to a normal distribution with parameters (mean and standard deviation) corresponding to the expected distribution (physiologically expected signal), and summing (or averaging, or combining in another suitable way) the Z-scores obtained over all patches in a region (a first region).
[0117] Thus, a total score is obtained, which increases as more patches better fit the expected curve. Many alternatives can be considered. For example, the samples can be compared with the target distribution, e.g., by using a Student's t-test or an F-test, and the set of (pixel-based) values in the patch (instead of, e.g., the mean value of the patch) can be tested against a known (i.e., assumed) distribution. Other types of distributions (potentially including empirical distributions, e.g., histogram-based distributions) can also be used to test against it. In addition, various alternatives of per-patch or per-pixel features on which such scoring and / or testing is based can be considered, such as signal intensity, pixel intensity, signal-to-noise ratio, spectral peak amplitude, spectral peak amplitude ratio, etc.
[0118] The use of pixel intensity values (or averaged per-patch intensity values) can advantageously allow for the calculation of a quality metric after only the step of partitioning the first region (and optionally the second region) needs to be performed for each specific parameter selection to be tested. Thus, this can be performed very efficiently (especially in multiple parameter test iterations). In another illustrative method, a pruning step can also be performed for each iteration to reject irrelevant patches, and / or the step of calculating the (one or more) features of the signal dynamics can also be performed, e.g., such that the signal dynamics of each patch can be considered during the parameter selection process (e.g., to calculate the quality metric).
[0119] In a relatively simple method, the region of interest can be determined for each parameter space point over a range of parameters (e.g., patch size) or on a parameter grid (e.g., time window and patch size), and the optimal region of interest can be used as the output and / or for PPG signal extraction, i.e., a simple line / grid search optimization can be used. The parameter points can be sampled equidistantly along the line or on the grid (e.g., in a Cartesian grid scan), or the parameter points can vary (e.g., sampled more heavily in regions of the parameter space where the best choice is more likely to be found overall). The search strategy can optionally consider the attributes and / or specific properties of the (one or more) parameters, e.g., using a multi-resolution and / or hierarchical search to find the best (or at least adequately tuned) (one or more) parameters. It should be understood that when more than one parameter is optimized, the combination can be optimized jointly (e.g., as a parameter vector to be optimized) or separately (e.g., in a nested optimization process).
[0120] More advanced techniques can use the results from one or more previous iterations to determine the next sampling point, e.g., using an estimate of the (one or more) parameter gradients and / or the Jacobian matrix of the parameters.
[0121] In yet another strategy, the first run or the first few runs can be used to estimate an initial region of interest from which the heart rate can be estimated (e.g., using an ROI for extracting the PPG signal within a short time window), after which this information can be used to obtain a better estimate of a suitable parameter selection, and the process can be restarted from that parameter selection to obtain a better ROI definition. If deemed necessary, this can be repeated multiple times or only used for one refinement step. It should be understood that signal noise can also be considered, e.g., using the heart rate and an estimate of the noise (and / or other quality metrics derivable from the observed PPG output and / or the ROI), to look up a suitable parameter selection in a look-up table via an empirical relationship, etc. Alternatively, a machine learning strategy can be used, e.g., to propose an optimal parameter space solution based on the output of previous runs (e.g., the ROI obtained for previous parameter selections, i.e., in an iterative method) and / or an input vector describing details of the use case (e.g., patient demographics, room conditions, time of day, information about the process for applying PPG extraction (e.g., identifier of a medical imaging process)) and / or other data that can implicitly be used to infer a suitable parameter selection via an algorithm trained on appropriate examples.
[0122] Although embodiments can generally divide the first region / second region into adjacent contiguous regions, e.g., forming a partition of a single-connected image patch, some variations and / or alternatives can be envisioned. For example, instead of or in addition to dividing the region into single-connected regions (patches in the more conventional sense), alternatively (or additionally), when different pixels are similar in pixel value rather than being spatially close together, the different pixels can also be grouped together. The spatial grouping and value range grouping of pixels can also be combined to form "patches", where each patch consists of pixels that are spatially close together, without the need for a predefined shape (e.g., block, square, rectangle,...) or connectivity, and consists of pixels that are relatively close together in value space (i.e., having the same or similar pixel values).
[0123] For example, a (spatial) distance metric (not necessarily Cartesian, e.g., Manhattan norm; but not limited to these examples) can be combined, e.g., via a weighted sum, with a (value space) distance metric (e.g., the absolute value of the difference in pixel values, but not limited to this), and the (first and second regions respectively) regions can be divided into sets of pixels to optimize the combined metric, e.g., in order to obtain a target number of pixels for each set while minimizing the metric.
[0124] When applied to the parameter optimization strategy as described above, it can be particularly advantageous to use the value space (image co-domain) distance (either exclusively or in combination with the space (image domain) distance) to determine the patches for each region (typically, pixels are grouped into multiple subsets), but it is not necessarily limited to this. For increasing patch sizes, e.g., an increase in the number of pixels per subset "patch" (at least on average), the likelihood of forming uneven patches increases, e.g., the likelihood that a single pixel group (subset or patch) contains only the tissue of interest (e.g., skin) becomes smaller. Thus, grouping pixels into patches only in the spatial domain is naturally limited for larger patch scales, because above a certain size, the quality of the patch deteriorates due to the desired (e.g., skin-related) signal being contaminated by pixels assigned to the same patch with poor or irrelevant signal content.
[0125] For larger patch sizes (i.e., subset sizes), it can be preferably switched from pixel combinations based on spatial adjacency to another metric based on the pixel output values, e.g., an intensity-based metric. This switch can be a discrete switch, but it can also be continuously tuned by changing the weights of the combined metric of the spatial and value distances.
[0126] Alternatively, patches can be formed based on the spatial neighborhood, but for example, above a predetermined size threshold, larger patches can be formed by combining smaller (spatially neighborhood-based) patches based on an alternative value space metric, e.g., four blocks of the original patches are selected based on the similarity of pixel values without the need for spatial proximity. The latter can allow for an efficient implementation. For example, the native (smaller) patches can be formed by dividing the image region at hand into blocks with a very low computational cost (e.g., forming a Cartesian grid over the region), and these original patches can be combined into subsets of the desired (larger) target size based on their intensities, which can also be performed at a reasonable computational cost (e.g., given the simple calculations performed on a relatively small number of entities, i.e., the original patches). For example, but not limited to, the appropriate combination of patches can be found by using only the calculation of a relatively small NxN matrix (or even only its upper / lower triangular part considering matrix symmetry) of the absolute average (average per original patch) pixel intensity difference between pairs of original patches, where N represents a relatively low number of original patches.
[0127] In the context of referring to illustrative intensity-based metrics, note that this can generally refer to any suitable value space (e.g., co-domain) metric. For example, a camera can be configured to collect vector data (e.g., a color image), and can evaluate a value space similarity metric based on vector pixel values, and thus is not limited to a scale comparison of intensity values. Additionally, the similarity metric need not be limited to considering only centrality metrics (e.g., average intensity), but can also use (alternatively or additionally) comparison metrics of (one or more) other properties of a set of pixels, such as standard deviation, signal-to-noise ratio (in a spatial sense, e.g., average intensity versus standard deviation), higher-order statistical moments, texture metrics, and / or any other spatial descriptive metric. For example, for each pixel (or "primitive patch"), a vector can be formed that includes different components indicating one pixel value and / or multiple pixel values in the local (spatial) neighborhood of the pixel (e.g., the primitive patch as discussed above), and the pixels (or primitive patches) can be clustered based on the similarity of the descriptive vector.
[0128] When applying the parameter tuning strategy as described above, it can be particularly advantageous to switch discretely or continuously from forming a spatial domain neighborhood (patch in the narrowest sense) to forming a value domain neighborhood (patch in the generalized sense of a subset). It should also be understood that a combination of continuous variations of the two types of metrics (e.g., via a weighted sum) may be more suitable for some optimization algorithms and / or can avoid losing a part of the parameter space, where the parameter optimum would be found if the discrete switching point between the metrics were postponed to a higher size threshold and / or the metric switching were performed in a more preferred continuous form in such a case. However, it should also be noted that methods such as those discussed above, where when the target size is required (e.g., exceeding a certain target size threshold), primitive (space-only neighborhood-based) patches are joined into larger patches based on intensity (generally, based on distance in the image co-domain or related space), can be easier to implement and / or advantageously more efficient. It should also be noted that using clustering techniques known in the field of image processing, for example, to group pixels (or primitive patches) together based on the value domain properties of the pixels (or primitive patches) may be sufficient to ensure a certain degree of spatial cohesion of the obtained patches (usually subsets), even when spatial neighborhood metrics are no longer explicitly considered (e.g., implicitly implemented and considered by the clustering algorithm).
[0129] To illustrate the method discussed above, consider a scenario of evaluating multiple patch sizes, such as patches of 10x10, 15x15, and 20x20 (or their equivalent approximate number of pixels). Assume the adverse case where a further increase in patch size would mean that (almost) all patches would include a substantial amount of non-skin area (for partitioning of the first region). In a naive approach, further increase in patch size (based on spatial neighborhood) would thus be prohibited, e.g., no desirable result of finding a useful ROI (or a higher-quality ROI than that found for the 20x20 test case) would be presented. Different from increasing the patch area to, e.g., 40x40, smaller patches can be combined into larger virtual patches, e.g., four 20x20 patches can be combined into one virtual (possibly spatially separated) patch of a total of 40x40 pixels. Such combination of patches can be obtained through clustering techniques, e.g., by extracting features of the average level and standard deviation on the patches (possibly averaged / computed over the entire observation period, e.g., also within a predetermined time window). Then, a suitable scale (patch size) can be selected by choosing a sufficient number of (real or virtual) patches that exhibit a consistent signal intensity within the (physiologically defined) expected range and / or another suitable quality metric.
[0130] After the step of partitioning 101 the first region into patches, and optionally after optimizing the patch size and / or other parameters, the patches of the first region can be pruned 102 based on a predetermined criterion of one or more values determined on a per-patch basis. This can include selecting patches having an average pixel value, extent, and / or motion value (e.g., amplitude, motion vector components, and / or one or more metrics representing motion) within a predetermined (corresponding) range and / or satisfying a predetermined criterion. For example, motion values such as motion amplitude, vector components of the displacement vector, and / or other such metrics representing motion can be determined by applying an optical flow algorithm (e.g., using two or more images obtained by a camera at different time points, e.g., a short dynamic sequence).
[0131] Thus, the set of patches can be pruned to select a subset of patches that are substantially stationary in position (little or no detectable movement), substantially well-lit, and / or substantially uniform (e.g., having a low standard deviation of the pixel values forming the patch). In the context of referring to mean, extent, standard deviation, etc., it should be understood that the same effect can be obtained through various similar statistical metrics of centrality and / or dispersion (e.g., median value, quartile value, variance value, etc.). Similarly, other simple per-patch statistics (or other representative values) can be used, such as higher-order statistical moments, skewness, kurtosis, texture, and / or pattern-based values, spatial gradients, and / or other such values, to prune patches that exhibit properties inconsistent with the desired tissue of interest and / or reflect an undesired imaging state, i.e., on a local (per-patch) basis.
[0132] In addition, after the step of dividing the second region 101 into patches, the patches of the second region can also be trimmed 112 based on a predetermined criterion of one or more values determined on a per-patch basis. These criteria can be the same as or different from the criteria used to trim the patches of the first region. This means that some variations can be considered in the implementation. For example, alternatively, the values of interest can be calculated over the entire image (but still based on each patch) or over the union of the first and second regions (but for each patch). By applying some bookkeeping, tagging, or other suitable methods, it is easy to track which patch corresponds to which region. Depending on the details of the embodiment (such as the hardware used, the selection of the values for trimming, and / or other such factors), it may be more efficient to calculate the values of interest either over all patches or separately (possibly in parallel) for each region. It should also be noted that even if the same criteria are not applied separately to the patches of the first and second regions, in some cases, it may still be more efficient to calculate the values (of the combined set of criteria) over all patches, for example, by advantageously using parallelization techniques (such as a graphics processing unit GPU and / or a general-purpose graphics processing unit GPGPU), such that only different comparison, thresholding operations, and / or other criterion evaluators need to be applied to the patches of the two different regions.
[0133] Trimming 102, 112 the patches of the first region and / or the second region can include selecting patches having an average pixel value, an extent, and / or a motion value (such as amplitude, motion vector components, and / or one or more metrics representing motion) within a (one or more) predetermined (corresponding) range and / or satisfying a predetermined criterion. The motion values, such as motion amplitude, vector components of the displacement vector, and / or other such metrics representing motion, can be determined, for example, by applying an optical flow algorithm (such as using two or more images obtained by a camera at different time points, such as a short dynamic sequence).
[0134] Thus, a set of patches can be trimmed to select (for each region) those patches that are substantially stationary in position (little or no detectable movement), substantially well-lit, and / or substantially uniform (e.g., having a low standard deviation of the pixel values forming the patch). In the case of referring to the mean, extent, standard deviation, etc., it should be understood that the same effect can be obtained by various similar statistical metrics of centrality and / or dispersion (such as the median value, quartile value, variance value, etc.). Similarly, other simple per-patch statistics (or other representative values), such as higher-order statistical moments, skewness, kurtosis, texture, and / or pattern-based values, spatial gradients, and / or other such values, can be used to trim patches that exhibit properties inconsistent with the desired tissue of interest and / or reflect undesired imaging conditions, i.e., on a local (per-patch) basis.
[0135] For example, patches in which (substantial) motion is detected may be rejected (trimmed off) in both the first region and the second region, e.g. to select (substantially) static (i.e. non-moving) patches. Blocks in which (substantial) changes (in pixel positions within patches) are detected may be rejected (trimmed off) in both the first region and the second region, e.g. to retain uniform patches. For the first region, and optionally for the second region, over- or under-illuminated patches may be rejected, e.g. in order to select patches that fall within a suitable range of the dynamic range of the camera system. These trimming criteria may be applied by logical OR operations, e.g. satisfying one rejection criterion may be sufficient to reject a patch, or, quite equivalently, the retention criteria may be applied by logical AND operations, e.g. patches to be retained are only selected for further processing when all selection criteria are satisfied (e.g. the logical negation of the preceding rejection criteria). Depending on the embodiment, other choices of criteria and / or values for evaluating such criteria and / or other choices of strategies for combining the results of multiple criteria (e.g., by logical operations, by majority voting) are also possible, for example, as the skilled person will be able to decide by relying on knowledge in the art and straightforward implementation, testing and / or simulations.
[0136] The method further comprises determining, on a per-patch basis, one or more dynamic feature values indicative of temporal dynamics, such as features. This / these value(s) are thus based on a plurality of camera images corresponding to different points in time, such as a sequence of camera images (a temporal sequence). These dynamic feature values are determined 109 for patches of the first region, and these dynamic feature values are determined 119 for patches of the second region. In particular, the features determined 109, 119 for each patch may be the same (e.g. of the same type, implemented in the same way and / or corresponding to the same physical quantity), e.g. the same dynamic feature values may be determined 119 for patches of the second region as for patches of the first region. This allows taking into account the results obtained for the same features in the second region when processing the dynamic features of each patch obtained in the first region, as will be explained below.
[0137] For efficiency reasons, these dynamic feature values may be determined 109, 119 after pruning 102, 112 the patches of the first and second regions, respectively, e.g., in order to avoid computing the dynamic feature value(s) for rejected (pruned) patches. However, this need not be the case. For example, parallelization techniques may be used to quickly and efficiently compute the dynamic feature(s) for all patches (but still on a per-patch basis), even though the values obtained for the pruned / to-be-pruned patches are substantially ignored.
[0138] The method includes clustering 103 the patches of the first region (e.g., the patches not retained by the pruning step) based on one or more dynamic feature values (representing the dynamic features of the patches), and similarly, clustering 113 the patches of the second region based on the (one or more) dynamic feature values. The clustering 103, 113 can be optionally (only) performed on the patches not pruned off (i.e., specifically considering the patches not rejected by pruning). The clustering can be performed separately (e.g., independently) for the patches of the first region and the second region.
[0139] Also, it should be noted that according to some embodiments, it would be acceptable to apply the clustering to all patches and, for example, reject the clusters including the rejected (pruned-off) patches and / or reduce such clusters by removing the rejected patches from the clusters. Thus, the order of the pruning and clustering steps need not be predetermined. For example, the pruning can be performed after the clustering. However, it should also be understood that performing the pruning before the clustering (and restricting the clustering to the accepted patches) would be a simpler strategy, potentially (generally) more efficient, potentially less error-prone, and / or potentially more robust.
[0140] Furthermore, for example, for efficiency, the two clustering steps 103, 113 can be combined. However, each cluster determined by the (one or more) clustering steps preferably specifically corresponds to the first region or the second region. For example, such that a single cluster cannot include patches from both the first region and the second region at the same time. Thus, if the clustering steps 103 and 113 are combined (but not limited to this), for example, by clustering the combined set of patches from the two regions (preferably after pruning has been applied), and regardless of their assignment to the first region or the second region, a mixed cluster can be obtained as part of the clustering result. For example, where such a mixed cluster includes patches from both the first region and the second region. Since mixed clusters are preferably avoided, each mixed cluster can be simply rejected (i.e., the cluster is further ignored), can be split into two clusters (one cluster for each of the first region and the second region), and / or more complex strategies can be designed to handle this. For example, the clustering strategy can penalize the mixed clusters in order to, for example, converge to a set of clusters without such mixed clusters through an iterative method. However, it should be noted that performing the clustering of the patches in the first region and the clustering of the patches in the second region independently of each other (preferably after pruning, in both cases) is generally a simpler solution. The two clustering steps 103, 113 can be performed in parallel, for example, to keep the computational burden low.
[0141] For example, after (e.g.) trimming based on the average pixel value of a block, the spread of the pixel values of the block, and / or motion, one or more (temporal) frequency features of each block can be determined (in steps 109, 119). Thus, the aforementioned (one or more) dynamic features can particularly include (one or more) temporal frequency features, such as Fourier components and / or values derived therefrom.
[0142] In steps 109, 119, Fourier analysis can be performed on a per-block basis (e.g., specifically transforming from the time domain to the time-frequency domain). For efficiency, this can (optionally) be applied only to blocks that have not been rejected in the trimming step, e.g., to reduce the burden on computational / processing resources. For example, for blocks that have not been retained by the trimming step based on the average pixel value (and / or alternative), spread (and / or alternative), and / or motion, the average pixel value (of the block at hand) as a function of time can be analyzed by applying a Fourier transform (e.g., a discrete Fourier transform (DFT), e.g., a fast Fourier transform (FFT)).
[0143] Although in this example the Fourier transform is based on the average pixel value as a function of time (e.g., as the input domain of the transform), it should be understood that other suitable values can be used to summarize the pixel content over the area of the block (as the input to the FT, e.g., FFT), such as the central behavior represented by a statistical centrality measure. Alternatively, a single pixel or a small group of pixels can be selected to represent the entire block, or even each pixel of the block can be (Fourier) analyzed. When calculating multiple Fourier spectra for the same block (e.g., for multiple pixels or multiple sub-groups of pixels), the obtained spectra (for the same block) can be combined in the frequency domain, e.g., by averaging (e.g., averaging the spectral power), e.g., for efficiency, even though data aggregation on the block in the time domain may be a better alternative than data aggregation in the frequency domain (e.g., using the Fourier transform of a single time series of average pixel values such that only one Fourier transform needs to be performed per block).
[0144] Frequency analysis, such as Fourier analysis, can be performed on a sample within a predetermined time window, which is, for example, appropriately selected (e.g., optimized for the parameters discussed above), in particular long enough to cover the frequencies of interest and short enough for practicality (and / or to maintain some qualitative or specificities in the time domain). Those skilled in the art are well familiar with the appropriate configuration of Fourier analysis (e.g., time window and / or sampling rate) and the considerations associated therewith. For example, the time window can be in the range of 0.5 s to 30 s, e.g., in the range of 1 s to 5 s, but is not limited thereto. In particular, the Fourier analysis can be configured and / or pre-processed and / or post-processed (e.g., using low-pass filtering, band-pass filtering, and / or high-pass filtering) to cover, for example, the spectral cardiac band in which the PPG signal of interest is assumed to be found. For example, a band-pass filter of, for example, 0.4 Hz to 4 Hz can be used, which corresponds to the heart rate synchronous signal component under typical heart rate conditions (e.g., 24 - 240 bpm). However, the embodiments are not necessarily limited thereto, and for example, given the knowledge of those skilled in the art, appropriate filter parameters can be determined by conventional experiments, simulations, and / or direct design techniques.
[0145] It should be understood that various alternatives to Fourier analysis, such as wavelet analysis, time-scale analysis, multi-resolution analysis, etc., can be applied to achieve the same or similar effects, i.e., to express the (one or more) dynamic features of the camera signal associated with each patch. Thus, the embodiments can include using such (one or more) direct alternatives instead of Fourier analysis, and / or Fourier analysis can be combined with another or multiple such techniques.
[0146] The frequency analysis (e.g., Fourier spectrum) used to determine 109, 119 the (one or more) dynamic feature clusters 103, 113 of the patches can also be used for further pruning steps. This can include determining the spectral energy of the strongest peak (and / or its predetermined frequency band, such as the heart rate band) of the Fourier spectrum (for each patch), and / or the ratio of the spectral energy of this strongest peak to the spectral energy of the second strongest peak. This peak spectral energy and / or this ratio can be applied 122 as a further pruning criterion, e.g., to reject patches with weak or non-distinct (e.g., multiple peaks instead of a single well-defined strong peak) maxima, for example, by rejecting patches for which the peak spectral energy and / or ratio is less than a predetermined threshold. Such further pruning 122 can be applied to the patches in both the first region and the second region, or only to the patches of interest in the first region. The further pruning criteria can be the same for the first region and the second region, or can be different.
[0147] For example, prior to the selection of the region of interest discussed further below, such a further pruning step 122 can be used, for example as an additional verification step, to consider the signal strength of each patch and / or the correspondence of one or more spectral attributes (one or more signal dynamic features) with the desired / expected attributes of the PPG signal. Potential criteria can include thresholds (on one or both sides) for peak spectral energy and / or the ratio of peak spectral energy to average spectral energy (and / or the spectral energy of the second peak). For example, patches with weak peaks can be rejected and / or patches with non-distinct peaks (relative to the average and / or the second largest peak) can be rejected. Additionally, if the amplitude of the spectral peak is too high, it can also be considered an unlikely candidate for further consideration (e.g., the threshold operation can be based not only on a lower limit but optionally also consider an upper limit). The predetermined threshold can take into account the predetermined patch size and the predetermined length of the observation window (FFT window), and can be easily determined by a person skilled in the art based on the selected operating parameters.
[0148] Additionally or alternatively, such metrics (e.g., peak signal strength, peak frequency, …) can also be used to determine one or more dynamic features for further use in clustering, as discussed below. For example, the further pruning step (binary rejection) can be replaced or supplemented by using one or more base metrics in one or more of the clustering steps discussed further below. For example, a metric (e.g., probability) can be associated with each patch based on peak amplitude, based on peak frequency, or based on a combination thereof, to build clusters based on the patches in the following steps.
[0149] Thus, in addition to or in the absence of such further pruning, the peak frequency of the strongest peak or the peak frequency and peak amplitude can then be used for one or more of the clustering steps 103, 113. The peak amplitude can optionally be expressed as a ratio relative to the second largest peak, relative to the average over the considered spectral band (heart rate band), and / or another suitable expression. The frequency and amplitude can be combined in a weighted size combination. For example, the fitness of the frequency (e.g., probability) and the fitness of the amplitude can be combined into a combined fitness (probability, likelihood, …) for clustering the patches, or they can be used together as vector values for clustering.
[0150] Clusters 103, 113 can be based on (e.g., at least) the frequency of the strongest spectral peak, as determined for each patch (e.g., for patches not rejected by the initial pruning step and / or further pruning steps). However, alternatives can also be used, such as determining the time scale at which the strongest signal is detected (e.g., a measure of the highest power or intensity, amplitude...), e.g., using multi-resolution time scale analysis (e.g., a hierarchical pyramid using scale analysis) and associating the scale value of this (most prominent) time resolution with the patch to be used as a clustering variable. Clustering based on multiple variables also need not be excluded, e.g., associating vector-valued variables with the patches that serve as the basis for clustering, e.g., combining peak frequency and peak time scale.
[0151] For example, the peak frequency (or a suitable equivalent, e.g., the scale in the scale level with the highest observed energy, the time period corresponding to the peak frequency, the wavelength,...) can be used to determine clusters of patches in which the pixels pulsate at substantially the same or similar frequency (relative to other patches in the cluster). As described above, a qualifier of the signal intensity (for that frequency component) can also be included in the clustering variable.
[0152] By clustering the patches of the first region and the patches of the second region using the same or similar clustering variables (e.g., peak frequency and / or amplitude), information from non-relevant clusters in the second region can be used to select one or more clusters of relevant patches in the first region, as described below.
[0153] The method includes selecting 104 at least one cluster from the clusters of patches obtained by the step of clustering 103 the patches of the first region based on size and / or morphology. Then, an output is provided 130. An area of interest (ROI) definition can be provided as the output, which defines the pixels forming the selected one or more clusters, e.g., for use by a dedicated PPG signal extraction algorithm, and / or a PPG signal can be determined 105 based on the pixel values in the selected (one or more) clusters (in a sequence of another image acquired by the camera). For example, the PPG signal can be determined 105 based on the dynamic changes (e.g., spatial averaging thereon) substantially only in the pixels of the selected (one or more) clusters. In other words, the selected (one or more) clusters can be used to directly calculate a robust PPG signal (from the pixels corresponding to the ROI in the sequence of camera images) either through another step 105 of the method according to an embodiment of the present invention or by providing the ROI to an external algorithm (known in the art) to determine the PPG signal based on the received input representing (or including) the ROI definition according to the sequence of camera images.
[0154] However, the step of selecting 104 one or more clusters includes excluding those clusters found for the first region that correspond to the clusters found by clustering 113 the chunks of the second region. In particular, the step of selecting 104 one or more clusters excludes from the selection each cluster in the first plurality of subsets whose time signal dynamics in its associated time signal dynamics substantially correspond to the time signal dynamics of a cluster in the second plurality of subsets, e.g., for which the (absolute) difference (or any suitable distance metric) between at least one dynamic eigenvalue of the cluster pair under evaluation is below a predetermined threshold, one from the first region and the other from the second region. For example, for each cluster in the first region, the distance to each cluster in the second region can be determined (typically, e.g., based on a metric of dynamic features), and if the minimum distance thus found is below the threshold, the cluster in the first region can be rejected. Optionally, for this exclusion step, only the largest clusters in the second region can be considered, e.g., to avoid rejecting large clusters in the first region that match small clusters in the second region (e.g., one or several stray patches of the tissue of interest (e.g., skin)). In other words, a size criterion can be applied to the clusters in the second region so as to avoid filtering out clusters in the first region based on their match with small clusters in the second region, which small clusters in the second region can correspond to small regions of the tissue of interest that are wrongly assumed to be irrelevant (by a priori assigning the corresponding image regions to the second region).
[0155] To select 104 one or more clusters based on size, after excluding the corresponding clusters, the largest cluster or a predetermined number of the largest clusters can be selected from the clusters found by clustering 103 the chunks of the first region. The size of each cluster for selecting the one or more largest clusters can be determined based on the number of chunks, the number of pixels, the corresponding area, or another similar metric representing the spatial extent of the cluster. Thus, clusters are selected that advantageously contain a large number (as far as possible) of pixels but do not match the dynamic features encountered in the second region. Then, the pixels belonging to the selected one or more clusters can be used to define the region of interest from which a robust and accurate PPG signal can be extracted.
[0156] Instead of directly selecting the largest cluster (of the chunks found for the first relevant region), the largest cluster that is not associated with any irrelevant (e.g., non-skin) frequency (or alternative dynamic feature for clustering) is selected; i.e., clusters corresponding to the oscillations and / or dynamics (e.g., pixel intensity) of the pixel signals associated with non-skin (or more generally, not with the tissue type of interest) are ignored, i.e., excluded from the selection.
[0157] For example, the step of selecting 104 clusters from the clusters of patches obtained by clustering 103 the patches of the first region may include pruning, from these clusters of patches of the first region, the clusters that correspond, on their (one or more) dynamic feature values (e.g., peak frequency), to the (one or more) dynamic feature values (e.g., peak frequency) of the clusters obtained by clustering 113 the patches of the second region.
[0158] In an illustrative method, the peak frequency (or its substitute) is averaged over the patches in each cluster to associate a single peak frequency with the cluster, and the cluster peak frequencies (or substitutes) found for both the first region and the second region can be removed (i.e., the corresponding clusters in the first region are pruned, or simply ignored). The peak frequencies (between the clusters in the first region and the second region, respectively) can be compared by simply thresholding the absolute difference of the peak frequencies (e.g., requiring the difference to be less than a predetermined threshold for matching). Alternative approaches can be readily envisioned. For example, the standard deviation of the peak frequencies in the cluster (over the patches forming the cluster) can be considered, e.g., to compare two clusters via a Student's t-test, an F-test, etc. Optionally, small clusters in the second region can be ignored, e.g., a criterion can be imposed that the matching cluster frequencies correspond to clusters in the second region having at least a substantial size (e.g., a predetermined size threshold). Thus, for rejection from the selection, it may be required that the (large) clusters in the first region match in frequency clusters in the second region that also have at least a substantial size. After removing the clusters in the first region that correspond to clusters in the second region (with respect to the dynamic parameter used for clustering, e.g., peak frequency), the largest cluster or the set of the largest clusters ranked highest (e.g., the top 2, top 3, …) among the clusters in the first region that are still under consideration can be selected to determine the region of interest, i.e., for extracting the PPG signal.
[0159] However, for simplicity and effectiveness, the corresponding clusters in the first region and the second region are described as being determined based on the dynamic feature used for clustering (e.g., peak frequency), it should be understood that this is not strictly necessary. The same or a similar effect can also be (additionally or alternatively) obtained by comparing the temporal dynamics of the clusters in the first region and the second region via another suitable metric, e.g., using the cross-correlation over time of the average signal (over the spatial extent of each cluster in the pair of clusters being compared; or an alternative value of centrality, etc.), using statistical tests to compare distributions, using mutual information or another information-theoretic metric, and / or using another suitable metric to compare scalars or vectors representing each cluster in the pair of clusters being compared, e.g., metrics for comparing other qualities of time series, Fourier spectra, and / or dynamic behavior.
[0160] In addition, selection 104 can also consider the morphology of each cluster. In particular, each cluster found in a first region (presumed relevant) of the tissue of interest (e.g., skin) can be associated with a metric based on the shape of the image region formed by the pixels (or patches) in the cluster (i.e., the set of patches in the cluster, and / or the set of sets of combined pixel sets forming the cluster), e.g., in order to rate the cluster to be consistent with the expected shape of the region of the tissue of interest. Thus, for each cluster not rejected by the correspondence with a matching cluster of unrelated image content (i.e., in the second region), a metric 121, e.g., a probability index, can be determined that indicates whether (and / or to what extent) the shape of the image region formed by the pixels (or patches) in the cluster is consistent with the expected shape.
[0161] For example, the forehead of an object can form a target region of the tissue of interest (skin), which will correspond to a connected set of patches of an approximately elliptical shape. Thus, in this example, the metric 121 that can be determined for each cluster (in the first region; e.g., not rejected by a previously applied pruning step) can represent the correspondence of the cluster to the elliptical shape. Equivalently, the metric can express the deviation therefrom, i.e., when the shape of the image region formed by the pixels (or patches) in the cluster better matches the expected shape, the metric can be in a substantially monotonic relationship that increases or equivalently decreases.
[0162] For example, an oval or ellipse can be fitted to the patches (or equivalently, the set of pixels) forming the cluster, and (shape) error values can be assigned to the patches (individual pixels) that extend beyond the fitted shape and the missing patches inside the shape. Thus, if the set of patches / pixels (i.e., the cluster) forms (e.g.) a line segment, it is less likely to detail the desired features, such as the face or a part thereof, e.g., the forehead. Additionally or alternatively, the metric (e.g., the error value) can consider the parameters of the fit, such as the major and minor axes (and / or their ratio) of the fitted ellipse, e.g., in order to preferably fit an elliptical (or oval) region within the range corresponding to the expected overall shape to be found. For example, compared to the desired target shape of an ellipse that is closer to (roughly) a circular form (in this example), an ellipse with a very large ratio of the maximum axis to the minimum axis will be more consistent with a line segment.
[0163] Based on the expected shape of the target region of interest, many variations can be easily envisioned. For example, to target a part of the skin on the forearm, an approximately rectangular region can be used as a standard for defining the shape metric, e.g., by fitting a rectangle or a rounded rectangle to the cluster (e.g., at an arbitrary angle with respect to the camera image coordinates). Thus, in this example, the selection can actively prefer more elongated clusters based on prior knowledge (e.g., the expected use for extracting the PPG signal from a camera observation of the forearm).
[0164] Thus, clusters can be selected based on size and / or morphological probabilities (e.g., a combination of both), while rejecting corresponding clusters in a second region (presumed unrelated to the tissue of interest) based on dynamics (e.g., peak Fourier frequency). Thus, embodiments of the present invention can improve the prior art "in vivo skin" selection method by considering information obtained in (one or more) regions of an image where the tissue of interest is unlikely to be found (i.e., the signal dynamics in the second region discussed above), and optionally also by considering the morphological attributes (e.g., shape criteria) of the selected clusters.
[0165] (e.g., based on correspondence with non-relevant clusters in the second region, based on cluster size, and based on elements of the cluster shape), the different contributions to the selection (e.g., ranking) step can be combined in various ways. For example, different types of metrics can be weighted and combined to obtain a ranking metric for selecting the most promising one or more clusters. Alternatively, binary exclusion from the selection can be applied to one or more metrics, e.g., rejecting clusters having corresponding dynamic attributes to those in the clusters of the second region and / or rejecting clusters that are too small or too large and / or rejecting clusters that do not correspond to the desired shape. Combinations thereof can also be used, e.g., using some of these metrics for binary selection and one or more metrics for selection by ranking after applying the binary selection. The metrics for binary selection and ranking selection can be different or can overlap. For example, the first rejection step can be based on these criteria with a relative tolerance threshold, and a weighted combination of these metrics can be used for the final selection by ranking.
[0166] In a second aspect, the present invention relates to a device for extracting a photoplethysmography signal indicative of an object's physiology based on remote camera observation, e.g., to provide such a PPG signal and / or to provide a region of interest suitable for selecting pixels in an image frame provided by the camera from which to extract such a PPG.
[0167] Reference Figure 2 , schematically shows an illustrative device 10 according to an embodiment of the present invention.
[0168] Device 10 includes an input unit 11 for receiving a camera image from a camera of at least one body part configured to monitor a subject. Device 10 may optionally include a camera 19, or the device may be adapted to be connected to such an (external) camera, for example, via a wired, wireless, or indirect (e.g., retrieving an image stored on an intermediate device (e.g., a streaming server)) connection. The device may also optionally include a light source (or light sources) for illuminating the subject (one or more relevant body parts thereof). The camera may include an infrared camera, a monochromatic camera operating in the visible wavelength range or a part thereof, such as the green spectral band, a color camera, and / or a multispectral camera. Similarly, the light source may emit light in at least an overlapping part of the spectrum to which the camera is sensitive, such as an infrared light source, a green light source, a broad-spectrum (e.g., white) light source, …
[0169] Device 10 includes a subset generator 12 for dividing a first region into a first plurality of pixel subsets (e.g., patches), and a second region into a second plurality of pixel subsets (e.g., patches), where the first region represents an image portion in the camera image where tissue of interest is assumed to be present, and the second region will be different from the first region and represents an image portion where tissue of interest is assumed not to be present in the camera image.
[0170] The subset generator 12 may also be adapted to divide the first region and the second region into the first plurality of pixel subsets and the second plurality of pixel subsets, respectively, based on respective predetermined prior knowledge assumptions about where tissue of interest may be present and absent, given a known setting of the volume in the space in which the given camera will be positioned relative to the subject. Such prior knowledge may, for example (but is not limited to), be hard-coded or hard-wired in the device, may be stored in a suitable configuration memory, and / or may be received via a user interface or an interface to an external controller.
[0171] The subset generator 12 may be adapted to divide the first region and the second region into the first plurality of pixel subsets and the second plurality of pixel subsets, respectively, by forming partitions of the first region and the second region into blocks or patches based on a spatial distance metric such that each subset groups together pixels that are spatially close. Additionally or alternatively, the respective regions may be divided into subsets based on an image co-domain distance, e.g., by grouping pixels whose pixel values are close to each other. A combination of spatial and co-domain distances may also be used, e.g., such that and / or based on their combination, each subset groups together pixels that are close in space and value (a balanced combination thereof).
[0172] The subset generator 12 may be adapted to perform image segmentation and / or processing of a reference image (and / or generally, reference spatial information) received via the input unit 11 (e.g., from a camera, from another camera, and / or from a (general) spatial information source such as a 3D surface scanner, a tomographic imaging system, …) to determine a first region and a second region.
[0173] The subset generator 12 may be adapted to perform image subtraction between a reference image acquired when an object is present and a blank image acquired when the object is absent.
[0174] The device may further include a subset eliminator 18 for pruning the first plurality of subsets based on a predetermined criterion of one or more values determined on a subset-by-subset basis so as to reject subsets having a low likelihood of a homogeneous region corresponding to the tissue of interest. The subset eliminator 18 may also be adapted to prune the second plurality of subsets based on the same and / or additional predetermined criteria. The clustering processor 14, discussed further below, may be adapted to cluster the pruned first plurality of subsets and the pruned second plurality of subsets determined by the subset eliminator, respectively.
[0175] The subset eliminator 18 may be adapted to prune the first plurality of subsets and / or the second plurality of subsets based on a predetermined criterion including one or more of the following criteria: the average pixel value of each subset, the spread of the pixel values of each subset, and / or at least one value indicative of the motion associated with the subset, such that the subset eliminator is adapted to reject objects showing excessive motion of pixels, non-uniform pixel value distribution, and / or pixel values with an average outside a predetermined target range. It should be understood that various alternatives and / or equivalents may be used for the mean (e.g., generally, a centrality measure), the spread (e.g., generally, a dispersion measure such as variance or standard deviation), and the motion (e.g., motion vector components, motion vector magnitude, square of the motion vector magnitude).
[0176] The device includes a signal dynamics analyzer 13 for determining at least one dynamic eigenvalue for each subset of the first plurality of subsets and each subset of the second plurality of subsets, the dynamic eigenvalue indicating the temporal signal dynamics of the subset in a sequence of camera images received via the input unit 11.
[0177] The signal dynamics analyzer 13 may be adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency features for determining at least one dynamic eigenvalue.
[0178] The subset eliminator 18 may be adapted to apply further pruning of the first plurality of subsets and / or the second plurality of subsets based on at least one dynamic eigenvalue (e.g., having been pruned for the first time, see above). Thus, the clustering processor 14 may be adapted to cluster the further pruned first plurality of subsets and the further pruned second plurality of subsets, respectively.
[0179] For example, further pruning can be based on at least one predetermined criterion of the spectral energy of the maximum spectral peak of the Fourier spectrum determined for each subset and / or based on a value determined therefrom, such as rejecting a subset whose spectral energy is below a certain threshold or outside a predetermined value range.
[0180] The device includes a clustering processor 14 for clustering a first plurality of subsets and a second plurality of subsets respectively based on at least one dynamic eigenvalue, so as to group subsets of pixels into clusters of similar signal dynamics.
[0181] The device includes a selector 15 for selecting at least one cluster from the clusters in the first plurality of subsets provided by the clustering processor 14 based on the cluster size and / or the cluster morphology. The selector is adapted to exclude from the selection each cluster in the first plurality of subsets whose associated signal dynamics substantially correspond to the signal dynamics of the clusters in the second plurality of subsets.
[0182] Referring to the selection based on the cluster size, the selector 15 may be adapted to select the largest one or more clusters from the first plurality of subsets that are not excluded based on their correspondence to the clusters in the second plurality of subsets.
[0183] For the selection based on the cluster morphology, the selector may be adapted to determine, for each cluster in the first plurality of subsets (or for a preliminary selection from those clusters, such as after applying a rough first selection criterion (e.g., eliminating too small clusters), and / or after exclusion based on the correspondence to the clusters in the second plurality of subsets), a shape metric representing the shape of the image region formed by the pixels (or patches) in the cluster, and select at least one cluster (i.e., the result to be used in the output) based on the shape metric of the selected cluster(s) being consistent with a predetermined target shape of the region of interest tissue.
[0184] Selector 15 may also be adapted to determine a ranking of clusters in the first plurality of subsets that are not excluded based on the correspondence with the clusters in the second plurality of subsets, where the ranking is used to determine the selected cluster(s). The ranking may combine a first score based on the size of the cluster and a second score based on a shape metric, for example, in order to combine the selection based on morphology and size. It should be understood that the exclusion criteria (corresponding to the clusters in the second plurality of subsets) may also be incorporated into the ranking, for example, such that the exclusion is imposed by a penalty on the ranking. In other words, the exclusion may be a hard exclusion that is pre-emptively enforced, or it may be imposed by a soft requirement, for example, to reduce the ranking of a cluster that matches a cluster in a second region that is assumed to be irrelevant in terms of its dynamic characteristics. It should also be noted that this can avoid or reduce the limitations resulting from the criteria strictly imposed on the matching between the clusters in the first plurality of subsets and the second plurality of subsets. Although this may be achieved, for example, by requiring that the absolute difference between peak frequencies is less than a predetermined threshold, by alternatively including the absolute difference (or a similar metric or equivalent) into the ranking metric (e.g., appropriately weighted), even a particularly large cluster with an expected shape can be selected if it is relatively close (but not "too close") to a cluster in the second group in terms of its dynamic characteristics.
[0185] The device includes an output unit 16 to output a region of interest definition that defines pixels forming at least one cluster selected by selector 15, for extracting a photoplethysmogram signal from a camera image acquired by a camera, or alternatively or additionally, to output a photoplethysmogram signal extracted from a region of interest in a camera image received via input unit 11.
[0186] Thus, the device may also optionally include a photoplethysmogram signal extractor 21 to extract a photoplethysmogram signal from a region of interest in a (sequence of) camera images received from a camera via input unit 11, which is provided as an output.
[0187] The device may also include an optimizer 17. The subset generator 12 may be adapted to repeatedly divide the first region into a first plurality of pixel subsets (and optionally, also divide the second region into a second plurality of subsets) for different values of at least one optimization parameter representing the size of the pixel subsets. The optimizer 17 may be adapted to calculate and evaluate a quality metric based on the pixel subsets obtained for different values of at least one optimization parameter in each such iteration, such that the optimizer can select a parameter value that obtains a sufficient value of the quality metric (e.g., higher than a predetermined threshold, or after a predetermined number of iteration steps, or reaching a predetermined parameter accuracy) or an optimal value of the quality metric. Then, the cluster processor 14 and selector 15 may use the (one or more) parameter values so selected and divide the first region and the second region into subsets accordingly, for example, which may be used to determine the region of interest.
[0188] The optimizer 17 may also use the size of the time observation window (such as the Fourier time window parameter) used by the signal dynamic analyzer 13 as a component of at least one optimization parameter (for example, in addition to the subset size to be optimized).
[0189] For example, the optimizer 17 may calculate a quality metric that includes or consists of a signal-to-noise ratio.
[0190] For example, the device may include a processor, a computer, or a similar general-purpose computing device, in combination with software adapted to execute the method according to an embodiment of the present invention, as described above. According to an embodiment of the present invention, the device may alternatively or additionally include dedicated hardware designed to execute the method or its (one or more) steps. For example, such dedicated hardware may include application-specific integrated circuits and / or configurable hardware, such as field-programmable gate arrays.
[0191] In a third aspect, the present invention relates to a diagnostic imaging system, such as a magnetic resonance imaging system or a computed tomography system. The system has an examination area and includes a camera for monitoring an object when the object is located in the examination area. The system also includes a device according to an embodiment of the second aspect of the present invention, which is operably connected to the camera so as to receive a camera image as an input from the camera, for example, from camera 19 via input unit 11.
[0192] As is known in the art, the system may be adapted to image an object by magnetic resonance imaging, computed tomography, positron emission tomography, and / or single-photon emission computed tomography. Alternatively, in another aspect, the present invention relates to a similar system for performing medical interventions, such as a radiotherapy system, for example, including a camera system and a device according to an embodiment of the second aspect of the present invention. Other embodiments may include systems for monitoring a patient during a surgical intervention in an intensive care unit and / or generally for monitoring the health of a patient in a medical or paramedical environment.
[0193] For example, the system may be a magnetic resonance imaging system. However, the principles of the present invention may equally apply to systems for different diagnostic imaging modalities and / or for another medical or non-medical context.
[0194] A magnetic resonance examination system according to an embodiment may include a main magnet assembly that defines an examination area. For example, the examination area may be formed by a volume in which the magnetic field conditions created and controlled by the magnet assembly are substantially suitable for magnetic resonance imaging. Thus, the examination area may correspond to the volume (at least the available part) surrounded by the magnet bore of the system (but is not limited thereto. For example, the principles of the present invention equally apply to open-bore systems and other less frequently used magnet assembly configurations).
[0195] An object to be examined (e.g., a patient) can be positioned on a patient bed in the examination area during the use of the system. The main magnet assembly can include magnet windings, such as coaxial (e.g., superconducting) windings, to generate a fixed and uniform magnetic field in the examination area. The examination area can be a cylindrical volume surrounded by these magnet windings.
[0196] The system can include a reconstructor to reconstruct (one or more) magnetic resonance images, such as tomographic MRI images, based on the magnetic resonance signals acquired by the system in use. The reconstructed images can be provided via an output unit for viewing, processing, or storage.
[0197] During use, an auxiliary instrument, such as an RF T / R head coil, can be placed in the examination area to acquire magnetic resonance signals from the head of the object. Other auxiliary coil configurations can be used to acquire signals from other body parts or for different use cases, and the signals can generally also be received by a receiver coil already integrated in the housing of the main magnet assembly.
[0198] The system includes a camera or a camera assembly, such as including multiple cameras. The camera system is adapted to obtain information from the object being examined, such as to obtain vital signs, movement, pain indicators, etc.
[0199] For example, the camera can be mounted near an entrance of the examination area. For example, the camera can be integrated in or mounted on the flange of the MR bore (e.g., such that the available free bore diameter is not affected or only minimally reduced, and / or to avoid or minimize interference with the operation of the MR system). For example, an (optional) illumination lamp can also be provided in or on the flange (but not limited thereto).
[0200] The system can include a display to display the image inside the examination area acquired by the camera (raw and / or after appropriate processing). This enables the operator to visually monitor the object in the examination area.
[0201] The image acquired by the camera can be provided to a device according to an embodiment, such as provided to an image processor for performing a method according to an embodiment. Thus, the image processor can be or can include the device 10 according to an embodiment of the second aspect of the present invention. The image processor (i.e., the device 10) is adapted to process the image information acquired by the camera system and perform image analysis to obtain information from the patient, in particular to extract a photoplethysmogram (PPG) signal from a region of interest in the video stream acquired by the camera, wherein the ROI is obtained as described above.
[0202] Respiratory and / or cardiac phase information (and / or more general information indicative of motion), such as a PPG signal generated by the device and / or information derived from and / or enhanced thereby, may be provided to the reconstructor to correct for motion of the acquired magnetic resonance signals and / or to apply motion correction to the reconstructed magnetic resonance images. For example, a cardiac trigger signal may be determined based on the PPG signal.
[0203] Thus, the signal provided by device 10 may be used to gate the data acquisition by the system (e.g., an MRI system or other system for performing a diagnostic imaging procedure). Alternatively or additionally, a treatment system is used to control the delivery of a treatment procedure. Alternatively or additionally, the signal may be used by a diagnostic imaging system to classify, sort, select, and / or annotate the acquired image data, e.g., to show the extracted PPG information (or information derived therefrom, such as cardiac phase) beside the (one or more) diagnostic images acquired at substantially corresponding times.
[0204] In a magnetic resonance imaging system according to an embodiment of the invention, the camera system may also include one or more light sources. Although embodiments relying on passive illumination for imaging are not necessarily excluded, those skilled in the art will understand that when active illumination is used, the illumination conditions can be better controlled and (camera) imaging can be more efficient.
[0205] The light source and / or the camera may be located outside the examination area, or on or near its edge region. This may simplify the configuration of the magnetic resonance imaging system (e.g., avoid or reduce interference with the RF and magnetic field operation of the system), and may provide more free bore width in the examination area. For example, for a cylindrical bore system, both the camera and the light source (or either of them individually) may be located at the flange of the bore housing at one end of the bore, which may leave the other end substantially free, e.g., to allow unobstructed access to the examination area (for bringing the patient and / or auxiliary instruments into the examination area), and reduce the potential claustrophobic effect on the subject during imaging by the system, and thus reduce possible discomfort.
[0206] In a magnetic resonance imaging system according to an embodiment of the invention, the camera (or cameras) may be adapted to operate (e.g., be substantially exclusively sensitive to) light in the (e.g., narrow) infrared wavelength range and outside the visible wavelength range.
[0207] In a magnetic resonance imaging system according to an embodiment of the present invention, a camera (or cameras) may be adapted to operate in the visible wavelength range (e.g., sensitive to a broad white light spectrum) or a portion thereof (e.g., a color band, e.g., green light) (e.g., substantially exclusively sensitive). The camera may be adapted to acquire monochromatic information, or may be a color camera, e.g., adapted to preferably independently and substantially simultaneously detect different color components, such as red, green, and blue components (but not limited thereto). The camera may also be adapted to detect a relatively large (e.g., more than three) spectral components, e.g., may be a multispectral camera.
[0208] (One or more) light sources may emit light in a spectrum suitable for the camera, e.g., broadband white light may provide illumination for a monochromatic or color camera in the visible range for operation. Similarly, an infrared light source may be used to emit infrared light in a spectral range sensitive to an infrared camera. It should be understood that the spectra of the light source and the camera need not be the same or even closely related, e.g., the spectrum of the light source may be wider, to an extent that there is sufficient overlap with the spectrum sensitive to the camera. The camera may be a digital camera, e.g., including an array of pixel light detectors.
[0209] In a fourth aspect, the present invention relates to a computer program product for performing a method according to an embodiment of the first aspect of the present invention when executed by a computing device. For example, the computer program product may include machine-interpretable instructions to direct a computing device (e.g., a computer) to implement (i.e., execute) the method of the embodiment.
[0210] Given the description provided above related to the method according to an embodiment of the present invention, other features of the device, system, and / or computer program product according to an embodiment of the present invention or details of the features described above will be clear, and / or vice versa, i.e., the details and features discussed above, regardless of their context (the described embodiments and / or aspects of the present invention), may be equivalently applied to all different aspects of the present invention, with necessary modifications.
Claims
1. An apparatus (10) for extracting a photoplethysmography signal indicative of the physiology of an object based on remote camera observation, the apparatus comprising: - an input unit (11) configured to receive camera images from a camera (19) configured to monitor at least one body part of the object; - a subset generator (12) configured to divide a first region into a first plurality of pixel subsets and a second region into a second plurality of pixel subsets, the first region representing an image portion of the camera image where tissue of interest is assumed to be present, the second region being different from the first region and representing an image portion of the camera image where tissue of interest is assumed not to be present; - a signal dynamics analyzer (13) configured to determine, for each subset of the first plurality of subsets and for each subset of the second plurality of subsets, at least one dynamic eigenvalue indicative of the temporal signal dynamics of the subset in a sequence of the camera images received via the input unit (11); - a clustering processor (14) configured to cluster the first plurality of subsets and the second plurality of subsets respectively based on the at least one dynamic eigenvalue so as to group the pixel subsets into clusters of similar signal dynamics; - a selector (15) configured to select at least one cluster from the clusters of the first plurality of subsets provided by the clustering processor (14) based on cluster size and / or cluster morphology, wherein the selector is adapted to exclude from the selection each cluster of the first plurality of subsets whose associated signal dynamics substantially correspond to the signal dynamics of a cluster of the second plurality of subsets, and - an output unit (16) configured to output a region of interest definition that defines the pixels forming the at least one cluster selected by the selector (15) for extracting a photoplethysmography signal from camera images acquired by the camera, and / or configured to output a photoplethysmography signal acquired from the region of interest in the camera images received via the input unit (11).
2. The apparatus according to claim 1, wherein, the subset generator is configured to identify anatomical key points in the sequence of camera images, wherein the subset generator is configured to use the identified anatomical key points to determine the first plurality of subsets and the second plurality of subsets.
3. The apparatus according to claim 2, wherein, the subset generator is further configured to determine the first plurality of subsets and the second plurality of subsets by fitting a template to the identified anatomical key points in the sequence of camera images.
4. The apparatus according to claim 2 or 3, wherein, the subset generator includes a neural network configured to output the anatomical key points in response to receiving the sequence of camera images as input.
5. The apparatus according to any one of the preceding claims, wherein, The at least one dynamic eigenvalue includes a main spectral peak, and wherein the signal dynamic analyzer is configured to determine the main spectral peaks of the first plurality of subsets and the second plurality of subsets by performing a fast Fourier transform of the pixel values to detect the main spectral peak.
6. The apparatus according to claim 5, wherein, the clustering processor is configured to cluster the first plurality of subsets and the second plurality of subsets according to the main spectral peak.
7. The apparatus according to claim 5 or 6, wherein, the main spectral peak is limited within a heart rate frequency band.
8. The apparatus according to claim 7, wherein, the heart rate frequency band is between 5 beats per minute and 300 beats per minute.
9. The apparatus according to any one of claims 5 to 8, wherein, the selector is configured to select the largest cluster among the first plurality of subsets having a main spectral peak different from the largest cluster among the second plurality of subsets.
10. The apparatus according to claim 9, wherein, the selector is further configured to ignore clusters having a spectral peak amplitude higher than a predetermined threshold during the selection process.
11. The apparatus according to any one of the preceding claims, comprising a photoplethysmogram signal extractor (21) to extract the photoplethysmogram signal to be provided via the output unit (16) from the region of interest in the sequence of camera images received via the input unit (11).
12. The apparatus according to any one of the preceding claims, wherein, the selector (15) is adapted to select one or more largest clusters from the first plurality of subsets that are not excluded based on correspondence with the clusters among the second plurality of subsets, and / or wherein the selector (15) is adapted to determine a shape metric representing the shape of the image region formed by the pixels and / or patches in the cluster for each cluster among the first plurality of subsets, and is adapted to select the at least one cluster based on the shape metric of the selected cluster being consistent with a predetermined target shape of the region of the tissue of interest, and / or wherein the selector (15) is adapted to determine a ranking of the clusters of the first plurality of subsets, wherein the ranking is used to determine the selection, and the ranking combines a first score based on the size of the cluster with a second score based on the shape metric.
13. The apparatus according to any one of the preceding claims, wherein, the subset generator (12) is adapted to divide the first region and the second region into the first plurality of pixel subsets and the second plurality of pixel subsets respectively by partitioning the first region and the second region into blocks or patches based on a spatial distance metric and / or based on an image co-domain distance and / or based on a combination of a spatial distance metric and an image co-domain distance, such that each subset groups together pixels that are close in space and / or value.
14. The apparatus according to any one of the preceding claims, wherein, The subset generator (12) is adapted to perform image segmentation and / or processing on a reference image and / or reference spatial information received via the input unit (11) to determine the first region and the second region, and / or is adapted to perform image subtraction between a reference image acquired in the presence of the object and a blank image acquired in the absence of the object, wherein the reference image and / or the reference spatial information acquired by the camera, another camera, and / or another spatial information source is received via the input unit (11).
15. The apparatus according to any one of the preceding claims, wherein, the apparatus includes an optimizer (17), and the subset generator (12) is adapted to repeatedly divide the first region into the first plurality of pixel subsets for different values of at least one optimization parameter representing the size of the pixel subsets, and the optimizer (17) is adapted to calculate and evaluate a quality metric based on the pixel subsets obtained in each iteration for different values of the at least one optimization parameter in order to select a parameter value that obtains a sufficient or optimal value of the quality metric.
16. The apparatus according to any one of the preceding claims, wherein, the optimizer (17) also uses the size of the time observation window used by the signal dynamic analyzer (13) as a component of the at least one optimization parameter, and / or wherein the quality metric calculated by the optimizer (17) includes or consists of a signal-to-noise ratio.
17. The apparatus according to any one of the preceding claims, wherein, the signal dynamic analyzer (13) is adapted to perform a Fourier transform in the time domain to obtain one or more time-frequency features for determining the at least one dynamic eigenvalue.
18. The apparatus according to any one of the preceding claims, comprising a subset eliminator (18) for pruning the first plurality of subsets based on a predetermined criterion for one or more values determined on a subset-by-subset basis to reject subsets having a low likelihood of a homogeneous region corresponding to the tissue of interest, and the subset eliminator for pruning the second plurality of subsets based on the same and / or additional predetermined criteria, and wherein, the clustering processor (14) is adapted to cluster the pruned first plurality of subsets and the pruned second plurality of subsets respectively.
19. The apparatus according to claim 18, wherein, the subset eliminator (18) is adapted to prune the first plurality of subsets and / or the second plurality of subsets based on the predetermined criterion, the predetermined criterion including one or more of the following criteria: the average pixel value of each subset, the spread of the pixel values of each subset, and / or at least one value indicating the motion associated with the subset, such that the subset eliminator is adapted to reject objects with pixels showing excessive motion, non-uniform pixel value distribution, and / or pixel values with an average outside a predetermined target range.
20. The apparatus according to claim 18 or 19, wherein, The subset eliminator (18) is further adapted to apply further pruning to the first plurality of subsets and / or the second plurality of subsets based on the at least one dynamic eigenvalue application, and wherein the clustering processor (14) is adapted to cluster the further pruned first plurality of subsets and the further pruned second plurality of subsets respectively.
21. A method (100) for extracting a photoplethysmogram signal indicative of the physiology of an object based on a remote camera observation, the method comprising: - acquiring (115) camera images from a camera configured to monitor at least one body part of the object; - partitioning (101) a first region into a first plurality of pixel subsets and partitioning (111) a second region into a second plurality of pixel subsets, the first region representing an image portion of the camera image where tissue of interest is assumed to be present, the second region being different from the first region and representing an image portion of the camera image where the tissue of interest is assumed not to be present; - determining (109, 119) at least one dynamic eigenvalue for each of the first plurality of subsets and the second plurality of subsets, the at least one dynamic eigenvalue indicative of the temporal signal dynamics of the subset in a sequence of the camera images; - clustering (103, 113) the first plurality of subsets and the second plurality of subsets respectively based on the at least one dynamic eigenvalue so as to group the pixel subsets into clusters of similar signal dynamics; - selecting (104) at least one cluster from the clusters of subsets obtained by the step of clustering (103) the first plurality of subsets based on cluster size and / or cluster morphology, wherein the step of selecting (104) the at least one cluster excludes from the selection each cluster of the first plurality of subsets whose associated signal dynamics substantially correspond to the signal dynamics of a cluster of the second plurality of subsets; and - providing (130) an output, wherein the output includes an interest region definition to define the pixels forming the selected at least one cluster for extracting a photoplethysmogram signal from the camera images acquired by the camera, and / or the output includes a photoplethysmogram signal extracted (105) from the interest region in the camera images acquired by the camera.
22. The method according to claim 21, wherein given a known setting of the volume of the space in which the camera will be positioned relative to the object, the first region and the second region are partitioned (101, 111) into the respective first plurality of pixel subsets and the second plurality of pixel subsets based on a predetermined prior knowledge assumption of where the tissue of interest may be present.
23. A diagnostic imaging system having an examination area (11), the system comprising: - a camera (19) for acquiring images from an object when positioned in the examination area for examination, - an apparatus according to any one of claims 1 to 20, operatively connected to the camera to receive camera images as input from the camera (19).
24. A computer program product for performing the method according to any one of claims 21 to 22 when run on a computing device.
Citation Information
Patent Citations
System and method for extracting physiological information from remotely detected electromagnetic radiation
US10441173B2