Detecting vital signs from image
By adjusting the gain of the optical sensor signal according to the color channel distribution properties of the human skin area and other areas in video image processing, the accuracy and reliability problems of vital sign detection under different lighting conditions and environments in the prior art are solved, and higher detection accuracy and sensitivity are achieved.
Patent Information
- Application Number
- CN202380072403.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-11
- Publication Date
- 2025-05-27
AI Technical Summary
When detecting human vital signs based on images, it is difficult for the prior art to adapt to different lighting conditions and environments, resulting in low accuracy and reliability of detection.
By receiving video images of multi-color channels, identifying the human skin areas and other areas, calculating the color channel distribution attributes, and adjusting the gain of the optical sensor signal based on the brightness attributes and reference color channel distribution attributes to match specific lighting conditions and environments.
It improves the accuracy and reliability of vital sign detection, reduces the requirements for manual adjustment and environmental control, adapts to different lighting conditions and environments, and enhances the sensitivity of detection and the convenience of operation.
Smart Images

Figure CN120051236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the detection of a person's vital signs based on one or more images, and in particular but not exclusively, to the detection of pulse and / or respiratory vital signs based on images of a captured person's face, chest, and / or skin area. Background Art
[0002] In many scenarios, the detection of vital signs based on images of a person has become a very advantageous and useful tool.
[0003] Image-based vital sign monitoring attempts to measure vital signs based on visual cues captured from an image of a person. Such vital signs can specifically include pulse or respiratory attributes. Application areas include but are not limited to emergency wards and department triage, (infectious disease) screening applications, and telehealth solutions, etc.
[0004] Different methods and algorithms for extracting vital sign data from images are known in the art. For example, remote photoplethysmography rPPG is a technique for determining a blood volume pulse signal based on an image of skin tissue, typically captured using a custom or off-the-shelf camera. As another example, respiratory attributes such as a respiratory signal and / or respiratory rate can be estimated from an image by detecting, for example, changes in the image caused by chest / abdomen movement.
[0005] Such vital sign monitoring is very advantageous in many scenarios because it can allow for the actual and non-invasive monitoring and detection of vital signals.
[0006] To provide optimal detection of vital signs, it is important that the image has suitable attributes for a particular detection / estimation. For example, this may require careful setting of the capture attributes of the camera to ensure proper capture operation. In particular, adjustment of the settings and calibration of the optical and sensor parameters is crucial in order to achieve high-quality imaging of a person and, in particular, to ensure that the captured image is suitable for detecting vital signs. In particular, adjusting the camera attributes and settings prior to digitization of the sensor signal may be necessary to ensure that the captured image allows for an accurate and reliable determination of vital signs. Such considerations can include adjusting the camera attributes and settings to prevent cropping, effectively utilize the dynamic range, achieve an improved signal-to-noise ratio, adjust the relationship between different color channels, allow for effective detection of visual attributes suitable for estimating vital signs, etc.
[0007] Setting and parameter adjustment can be particularly challenging for optimizing the detection of vital signs, as it may be different from the optimal settings and parameters for seeking to provide an image that is as visually realistic as possible and optimized for the perception of the human visual system. The optimal settings for detecting vital signs can be different. For example, if the detection is based on an image that exaggerates the visual attributes that vary with the pulse, it may be more reliable to detect a person's pulse.
[0008] Accordingly, an improved method for determining vital signs from a captured image would be advantageous. In particular, a method that allows for improved operation, increased flexibility, improved determination of vital signs, improved representation of visual attributes that allow for the determination of vital signs, reduced sensitivity to changes and lighting conditions in the capture environment, reduced complexity, convenient implementation, reduced memory and / or storage device requirements, and / or improved performance and / or operation would be advantageous. SUMMARY OF THE INVENTION
[0009] Accordingly, the present invention seeks to preferably alleviate, mitigate, or eliminate one or more of the above disadvantages, either alone or in any combination.
[0010] According to one aspect of the present invention, there is provided an apparatus for estimating a person's vital signs from a video image, the apparatus comprising: a receiver arranged to receive a video image comprising a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; a first detector arranged to determine a first image region in at least one image of the video image, the first image region corresponding to a person's skin region; a second detector arranged to determine a second image region in at least one image of the video image, the second image region corresponding to a region of the person that is different from the first image region; a determiner arranged to determine a set of color channel distributions for pixels of the first image region; a processor arranged to provide a skin color indication indicative of the person's skin color; a reference processor arranged to provide a reference color channel distribution property for the skin color; a gain processor arranged to determine a gain for the optical sensor signal based on a brightness property determined for the second image region and based on a comparison of the reference color channel distribution property with the properties of the set of color channel distributions; a gain controller arranged to control a gain value applied to the optical sensor signal to match the determined gain; and a vital sign determiner arranged to estimate a person's vital sign property from the video image.
[0011] In many embodiments, the present invention may allow for improved determination of vital signs based on images. The method may allow for improved adjustment for specific environmental conditions, particularly lighting conditions. The method may in particular allow for improved adjustment for a specific purpose of detecting vital signs rather than, for example, merely optimizing the authenticity of the captured video image. The method may facilitate operation in many scenarios and may, for example, allow for improved user convenience as the requirements for manual adjustment and / or accurate control of the capture environment can be reduced. In many scenarios, the method may allow the capture operation to focus particularly on visual attributes suitable for determining vital signs, thus generally enhancing the reliability and accuracy of detection.
[0012] In many embodiments, the method may provide improved camera / capture control. The sensitivity of the capture operation to changes in lighting conditions and the type of person for determining vital signs can generally be significantly reduced.
[0013] The color channels may specifically be the red, green, and blue channels, and the video image may be an RGB image.
[0014] The first image region and the second image region may generally be in the same image of the video image, but in some embodiments and scenarios may be in different images. The video image may be a frame of a video signal.
[0015] A set of attributes of the color channel distribution may include a set of values, each value for one of the plurality of color channels. The reference color channel distribution attribute may include a set of values, each value for one of the plurality of color channels.
[0016] In many embodiments, the vital signs may be the pulse rate and / or the respiration rate. The second image region may be an image region corresponding to the chest region of a person.
[0017] A set of attributes of the color channel distribution may be a set of one or more (pixel) percentile values for the color channels and may specifically be a set of one or more pixel values corresponding to a given percentile of the pixel values of the color channels.
[0018] The device may be a camera device.
[0019] According to an optional feature of the present invention, the device may further include: a gain circuit arranged to apply a gain value to optical sensor signals from at least two color channel optical sensors; and an image generator coupled to the gain circuit and arranged to generate a video image based on the optical sensor signals received from the gain circuit.
[0020] In many embodiments, the method can provide improved operation and / or convenient implementation, and can particularly provide improved capture of images for vital sign determination.
[0021] According to an optional feature of the present invention, the reference color channel distribution property includes a set of reference intensity values for different color channels, and the property of the color channel distribution includes a set of determined intensity values for a set of color channel distributions.
[0022] This can provide improved operation and / or convenient implementation in many embodiments. In particular, it can generally provide effective adjustment of the capture operation for determining vital signs.
[0023] The determined intensity value can specifically be a determined pixel value, and can specifically be the intensity value / pixel value for a given percentile of the distribution.
[0024] According to an optional feature of the present invention, the determined intensity value is the color channel pixel value for a given percentile of a set of color channel distributions.
[0025] This can provide improved operation and / or convenient implementation in many embodiments.
[0026] According to an optional feature of the present invention, the gain processor is arranged to determine the gain for the first color channel according to the difference between the determined intensity value for the first color channel among a plurality of color channels and the reference intensity value for the first color channel value.
[0027] This can provide improved operation and / or convenient implementation in many embodiments. In particular, it can generally provide effective adjustment of the capture operation for determining vital signs.
[0028] According to an optional feature of the present invention, the gain processor is arranged to set the gain such that the relationship between the determined intensity values for different color channels matches the relationship between the reference intensity values for different color channels.
[0029] This can generally allow particularly effective color channel balancing (e.g., white balance), which can be optimized for a specific algorithm for determining vital signs.
[0030] According to an optional feature of the present invention, the brightness property is the brightness for the second image region generated by applying the determined gain to the second image region, and the gain processor is arranged to set the gain such that the brightness property indicates that the brightness for the second image region meets the standard.
[0031] This can provide effective adjustment. The brightness standard can specifically include the requirement that the brightness is within a given brightness range.
[0032] According to an optional feature of the present invention, the gain processor is arranged to determine a common gain component based on a luminance attribute and a relative gain component based on a comparison between a reference color channel distribution attribute and the attributes of a set of color channel distributions, the common gain component being the gain applied to all color channels among a plurality of color channels, and the relative gain component being the gain component applied to individual color channels among the plurality of color channels.
[0033] This can provide improved operation and / or convenient implementation in many embodiments. It can not only allow effective adjustment of the capture operation, but also allow practical yet flexible implementation.
[0034] According to an optional feature of the present invention, the gain processor is arranged to control the exposure attribute of at least one color channel optical sensor based on at least one of a luminance attribute and a comparison between a reference color channel distribution attribute and the attributes of a set of color channel distributions.
[0035] This can generally allow improved adjustment and can, for example, allow adjustment over a more diverse set of capture scenarios and lighting conditions. It can generally improve image quality and allow, for example, a higher signal-to-noise ratio of the optical sensor signal and / or the captured image.
[0036] This can provide improved operation and / or convenient implementation in many embodiments.
[0037] According to an optional feature of the present invention, the gain processor is arranged to determine and store a first color channel distribution for a third image region of an image for which a skin region has been detected; and to determine a gain for an optical sensor signal for a first image for which no skin region has been detected based on the first color channel distribution and a second color channel distribution, the second color channel distribution being determined for the third image region of the first image.
[0038] This can provide improved operation in many embodiments and scenarios. It can generally allow appropriate adjustment even for images / times in which no skin region can be detected in a video image. It can generally ensure that the capture operation is adapted to provide improved conditions for detecting an image region corresponding to a skin region. For example, it can ensure that the capture operation at times when no skin (e.g., via a face detector) is detected results in a captured image suitable for performing a skin detection operation (e.g., specifically, face detection).
[0039] According to an optional feature of the present invention, the gain processor is arranged to set the gain to reduce the difference between the attributes of the first color channel distribution and the second color channel distribution.
[0040] This can provide improved operation and / or convenient implementation in many embodiments.
[0041] According to an optional feature of the present invention, the third image region is the complete image region.
[0042] According to an optional feature of the present invention, the gain processor is arranged to determine an initial gain for the optical sensor signal based on a comparison of the complete image intensity value with a predetermined reference intensity value; the complete image intensity value is the intensity value of a predetermined percentile of the color channels among a plurality of color channels.
[0043] This can allow for improved initial adjustment and operation, such as enabling the capture of images suitable for initial skin area detection and the like.
[0044] According to one aspect of the present invention, there is provided a method for estimating a person's vital signs from an image, the method comprising: receiving a video image including a plurality of color channels, each color channel representing digitization of an optical sensor signal from a color channel optical sensor; determining a first image region in at least one image of the video image, the first image region corresponding to the skin region of the person; determining a second image region in at least one image of the video image, the second image region corresponding to a region of the person different from the first image region; determining a set of color channel distributions for pixels of the first image region; providing a skin color indication indicating the skin color of the person; providing a reference color channel distribution attribute for the skin color; determining a gain for the optical sensor signal based on the luminance attribute determined for the second image region and based on a comparison of the reference color channel distribution attribute with the attributes of the set of color channel distributions; controlling the gain value applied to the optical sensor signal to match the determined gain; and estimating a vital sign attribute of the person from the video image.
[0045] These and other aspects, features, and advantages of the present invention will become apparent and will be elucidated with reference to the embodiments described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Embodiments of the present invention will be described by way of example only with reference to the accompanying drawings, wherein
[0047] Figure 1 Examples of elements of a camera according to some embodiments of the present invention are shown;
[0048] Figure 2 Examples of elements of a camera according to some embodiments of the present invention are shown;
[0049] Figure 3 Examples of elements of a camera according to some embodiments of the present invention are shown;
[0050] Figure 4Shows some elements of a possible arrangement of a processor for implementing elements of an audio device according to some embodiments of the present invention. Detailed Description
[0051] Figure 1 Shows an example of a camera device arranged to capture an image of a person and determine the person's vital signs based on the captured image. The vital signs that can be detected by the device can include the pulse rate and the respiration rate.
[0052] Detect vital signs based on a video sequence of images / frames, and can specifically detect vital signs based on changes occurring in the images. Thus, an optical camera can be positioned to capture a person / patient, and the resulting video images can be analyzed and compared to each other to detect vital signs. For example, photoplethysmography can be used to determine the volume changes in blood circulation.
[0053] To achieve accurate and reliable detection of vital signs, it is important that the capture of the images is appropriately adjusted to provide images with suitable optical properties for detecting visible cues that can allow the detection of vital signs. This adjustment of the capture operation can be crucial for the accuracy and reliability that can be achieved in determining vital signs.
[0054] In Figure 1 example, the camera device includes a plurality of optical sensors 101 that generate optical sensor signals for different color channels. Thus, at least two optical sensors are arranged to capture visible light with different bandwidths / different frequency sensitivities. In many embodiments, the camera device includes three optical sensors that capture spectra corresponding to the red, green, and blue (RGB) color channels, respectively. Thus, the optical sensors 101 generate optical sensor signals corresponding to RGB color signals having RGB color channels.
[0055] The sensors are coupled to a gain circuit 103 that is arranged to apply a gain to the optical sensor signals. The gain can be implemented, for example, as a set of gain / amplifiers that apply a variable gain to each analog optical sensor signal, where the variable gain can be dynamically set, for example, via a suitable control signal.
[0056] The gain-adjusted optical sensor signal is fed to an image generator 105, which is arranged to generate an image based on the gain-adjusted optical sensor signal. The image is in particular a color image having color channels corresponding to the color channels of the sensor / optical sensor signal, and thus in a specific example an RGB color image. The image generator 105 may generate color channel values for one color channel by digitizing one of the optical sensor signals from one of the optical sensors. The image generator 105 may in particular include an analog-to-digital converter and perform a process of generating a video image based on digital samples of the optical sensor signal. The signal path, gain adjustment, and image generation of such optical sensors are used in many optical cameras and can provide images with high-quality capture.
[0057] In the described method, the camera device further includes a processor 107 to which the generated image is fed. The processor 107 is arranged to determine vital signs based on the received video image. It is also arranged to adjust the capture operation and functions so as to provide an image that is particularly suitable for detecting vital signs. In particular, the processor 107 is arranged to determine the gain for the optical sensor signal and control the gain circuit 103 to apply the determined gain. The gain may be determined separately for each color channel, and different gains may be applied to different optical sensor signals. It should be understood that the camera device may be in one housing, or may have components divided between multiple housings, or may in fact have components located in other equipment pieces.
[0058] The processor 107 may apply a specific method that will be referred to Figure 2 to be described, Figure 2 An example of the elements of the processor 107 is shown.
[0059] The processor 107 includes a receiver 201 that receives the video image from the image generator 105. Thus, the receiver 201 receives a multi-color-channel captured image including a plurality of color channels, and in particular, the image may be an RGB video image.
[0060] The video image is fed to a first detector 203 and a second detector 205, which are arranged to detect different image regions in the received image(s).
[0061] The first detector 203 is arranged to detect a first image region corresponding to a person's skin region in the received image / frame (at least one image / at least one frame). Thus, the first detector 203 may analyze the image to detect a region corresponding to the region where a person's skin has been captured. In many embodiments, the first detector 203 may in particular perform face detection to detect the face of the person for whom vital signs are to be determined.
[0062] It should be understood that many different methods and algorithms are known and can be used to detect skin regions.
[0063] For example, in some embodiments, the first image region can be determined as a substantially predetermined region corresponding to a fixed position of the user relative to the camera. For example, the camera can capture a scene where a person is arranged to sit on a chair facing the camera and, for example, the head is fixed by a physical headrest. In such a scene, the face of the person will tend to always be in substantially the same position in the image, and the first image region can be determined as a predetermined region corresponding to the position of the face.
[0064] As another example, the first image region can include a face detection algorithm. It should be understood that many different face detection algorithms are known and any suitable method can be used without prejudice to the present invention. After detecting the position of the face in the image, the first detector 203 can continue to generate the first image region as a region corresponding to the detected face position.
[0065] Thus, in many embodiments, the first image region and the skin region can correspond to the facial region of a person. However, it should be understood that in other embodiments, the first image region and the skin region can alternatively or additionally be used for another body part. For example, in some embodiments, it may be required that a person place an arm in a specific position (e.g., supported by a restraint), which results in this being captured at a predetermined position in the image. As another example, a person may be required to partially undress and can be positioned such that the upper body of the user fills most of the image, and the first image region / skin region can be determined as the largest region where the visual attributes of the image are substantially similar. For example, segmentation based on image attributes (e.g., color and brightness) can be performed on the image, and the first image region can be determined as the largest segmentation.
[0066] In many embodiments, an appropriate part of the body / image segmentation can be tracked. For example, a person in bed can look around, and a face tracker can be arranged to always focus on the face for camera control. As another example, a person can move around in a room, and the camera can track the person (and the face). The camera control can be arranged to be unaffected by other environmental factors (e.g., for the example of a person walking in front of a window with sunlight, the camera would typically darken the entire video stream, thereby losing the ability to monitor vital signs).
[0067] The second detector 205 is arranged to detect a second image region in the (one or more) images. It may for example detect a predefined region or may detect an image region based on image attributes. In many embodiments, the second image region may be determined as the image region for determining vital signs. For example, the processor 107 may be arranged to detect a pulse rate based on visual changes in a person's chest region, and thus may select the second image region to correspond to the person's chest region.
[0068] In some embodiments, the second image region may be determined as a predefined region in a video image, for example when the position of the person relative to the camera is strictly controlled. In other embodiments, it may be detected based on, for example, markers or other recognizable visual features in the image. For example, markers having different visual attributes may be attached to the person being imaged. The markers may be regions around the body, for example, from which vital signs are determined. In such a case, the second detector 205 may detect the outline of the marker by detecting a specific visual attribute (e.g., a specific pattern) in the image. The second image region may then be determined as the region within the determined outline.
[0069] As yet another example, the second image region may be determined in some regions as the region having a predefined relationship with the first image region. For example, the second image region may be determined as a predefined region located a predefined distance below the detected face region and having a predefined outline. This may for example allow the determination of the chest region based on the face detection that determines the first image region.
[0070] As another example, the second image region may be determined by using a body part detector to identify the chest region.
[0071] The first detector 203 is coupled to a distribution determiner 207, which is arranged to determine a color channel distribution for the pixels of the first image region. The color channel distribution may generate, for each color channel, a distribution of the pixel values for that color channel. For a given color channel, the color channel distribution may indicate the occurrence frequency / ratio of the pixel values of the pixels belonging to the first image region. The color channel distribution for each color channel may indicate the occurrence or frequency of the pixel values in the first image region for each (or groups of) possible pixel values. For example, a histogram may be determined for each color channel, which reflects the distribution of the pixel values for that color channel.
[0072] The distribution determiner 207 may be arranged to generate a color channel distribution that reflects, for a given color channel and threshold, the percentile / ratio of pixel values below the threshold (or equivalently above the threshold). Equivalently, the distribution determiner 207 may be arranged to generate a color channel distribution that reflects, for a given color channel and percentile, the threshold for which that percentile of pixel values is below the threshold (or equivalently above the threshold).
[0073] The camera device further includes a skin color processor 209, which is arranged to provide an indication of the skin color of the person being imaged and to determine a vital sign for that skin color indication. The skin color indication indicates the skin color of the person and, in particular, indicates visual skin color attributes. In many embodiments, the skin color processor 209 may store a plurality of skin color categories, and the skin color indication may indicate the category that is considered to be appropriate for the particular person being imaged.
[0074] In many cases, it may be assumed that the ratio between the average RGB values for skin tissue is similar across skin light types, as found by De Haan, G., Jeanne, V. in "Robust pulse rate from chrominance-based rPPG". As will be described in detail later, such an assumption can be used to control the capture to provide images that are particularly suitable for detecting vital signs.
[0075] In some embodiments, the skin color processor 209 may be coupled to or include a user input and may determine an appropriate skin color indication in response to the user input. For example, in some embodiments, a set of images corresponding to different skin type categories stored by the skin color processor 209 may be presented to a person. The person may then select the image that is considered to most closely reflect the skin color of the person being imaged, and the skin color indication may be set to indicate the corresponding category.
[0076] In some embodiments, the detection of skin color may be determined, for example, based on an analysis of the captured video image and, in particular, the first image region. The average pixel value may be determined, for example, for each color channel and compared to a set of reference average pixel values for each skin color category. The skin color category having the closest match to the reference average value may then be indicated by the skin color indication as the skin color of the person. Although such detection may sometimes be very inaccurate because it is based on an initial image that may not have been captured under fully adjusted or optimized capture conditions, it is typically sufficient because the difference between the reference average values may be quite large.
[0077] The skin color processor 209 is coupled to a reference processor 211, which is arranged to provide reference color channel distribution attributes for the indicated skin color. Thus, based on the received skin color indication, the reference processor 211 can determine the reference distribution attributes for each color channel. In different embodiments, the attributes can be different attributes (usually separately for each color channel) representing attributes or characteristics of the distribution. For example, the attributes can be an indication of mean, variance, percentile, mode, median. Specifically, in many embodiments, the reference attributes can be a set of pixel values corresponding to a given (e.g., predetermined) percentile of the distribution of pixel values in the color channel.
[0078] The reference color channel distribution attributes can specifically include a set of reference intensity values for different color channels. For example, pixel values can be provided for each color channel. The reference pixel values can indicate, for example, the preferred pixel values for a given percentile of the pixel value distribution in the first image region.
[0079] In many embodiments, the reference processor 211 can store a set of reference attributes, such as reference intensity values, for each of the possible skin color categories. For example, the reference attributes can be predetermined values stored in a memory, for example, during the manufacturing process. Thus, in some embodiments, the predetermined stored reference values can be associated with each possible skin color category. For example, during the development and design of a method for determining vital signs, appropriate values may have been found during testing and experimentation. The reference attributes may have been specifically optimized to produce the desired and as best as possible detection performance for vital sign determination.
[0080] The reference processor 211 is coupled to a gain processor 213, which is arranged to determine the gain for the gain circuit 103 based on the reference color channel distribution attributes provided for a specific skin color and based on the color channel distribution of the pixels in the first image region. The gain processor 213 can specifically adjust the gain for the color channels to reduce the difference between the reference color channel distribution attributes and the attributes of the color channel distribution. Thus, the gain processor 213 can be arranged to determine the gain to reduce the difference between the determined attributes of the color channel distribution and the reference color channel distribution attributes for the specifically indicated skin color of a person. Specifically, it can seek to set the gain to reduce the difference between the reference intensity values for each color channel and the determined intensity values for a given percentile of each color channel distribution.
[0081] For example, the reference color channel distribution property may indicate the reference pixel / intensity values for each color channel. The reference pixel value may indicate, for example, the preferred pixel value for a given percentile (e.g., the ninetieth percentile) of the pixels. The gain processor 213 may determine, for each color channel, the pixel value at the ninetieth percentile of the specifically determined color channel distribution for the first image region. The gain for each color channel may then be set to the ratio between the preferred pixel value and the determined ninetieth percentile pixel value. Thus, if the pixel values of the individual pixels in the first image region are scaled by the determined gain value, it will result in the first image region having the ninetieth percentile value corresponding to the preferred value. Therefore, applying the determined gain to the gain circuit 103 will result in the captured image having the desired properties.
[0082] However, the gain processor 213 is also arranged to consider not only the first image region when determining the gain, but the gain determination also depends on the properties of the second image region. The gain processor 213 is specifically arranged to determine the luminance property for the second image region and set the gain in response to that luminance property. In many embodiments, the luminance property may be a gain-compensated luminance property and may reflect the luminance for the second image region that would result if the determined gain had been applied to the optical sensor signal, and thus that luminance would result in a future image unless the capture environment changes. For example, the luminance property may represent the luminance of the second image region that would result from setting the gain determined by the gain processor 213 for the optical sensor signal and in the absence of any other change in the capture or lighting conditions.
[0083] In many cases, the luminance property may be determined for the second image region and then scaled by a factor depending on the determined gain. For example, for a given gain setting, a predetermined function may be evaluated to provide a gain that reflects how the luminance is modified by the gain change (from the gain used when capturing the image to the currently considered gain). The luminance property may then be scaled by this value to provide the gain-compensated luminance.
[0084] The luminance property may indicate the luminance level for the second image region and specifically may be an indication of the maximum, average, median, percentile (e.g., ninetieth percentile) luminance level for the second image region. The luminance property generally indicates the luminance level as the luminance per unit area (e.g., per unit pixel) relative to the luminance range. For example, the range may have the lowest luminance level corresponding to all color channel values having the minimum value (e.g., RGB value (0,0)) and the highest luminance level corresponding to all color channel values having the maximum value (e.g., for 8-bit values, RGB value (255,255,255)).
[0085] As a low-complexity example, a luminance value can be calculated for each pixel of the second image region according to a suitable luminance metric (e.g., by combining weighted individual color channel values, where the weighting is performed by factors reflecting the relative contribution / importance of each color channel in determining the vital signs). Then, an average value, mean value, median value, or, for example, a given percentile (e.g., the ninetieth percentile) luminance value can be determined.
[0086] Then, the gain processor 213 can adjust the gain considering the luminance attribute. For example, the previously described gain determination can be applied to set the ninetieth percentile color channel pixel value to a preferred value. However, instead of blindly using these gain values, they can be made to meet the requirement that the luminance attribute for the second image region must satisfy a given criterion. For example, the luminance range can be given by a predetermined minimum luminance level and a predetermined maximum luminance level. The gain values for the individual color channels can be determined as previously described. A corresponding luminance gain (appropriately weighting the gains of the individual color channels to reflect their influence on luminance) can be determined to generate a gain-compensated luminance attribute. For example, the luminance attribute (e.g., the ninetieth percentile luminance) for the second image region can be multiplied by the luminance gain, and the resulting value can be compared with the luminance range. If the modified luminance is below the predetermined range, the gain value can be adjusted (increased) until the modified luminance exceeds the predetermined minimum luminance value. Similarly, if the modified luminance is above the predetermined range, the gain value can be adjusted (decreased) until the modified luminance is below the predetermined maximum luminance value.
[0087] In many embodiments, the operation can be linear and generally can be achieved by scaling the gain determined based on the color channels of the first image region to ensure that the luminance attribute of the second image region is within the desired range, where the scaling is the same for all color channels. Thus, the determined gain can simply be multiplied by a given factor that is the same for all color channels. In many cases, the scaling factor can simply be the ratio between the predetermined minimum luminance value and the determined luminance value (if below the minimum luminance value), or it can be the ratio between the predetermined maximum luminance value and the determined luminance value (if above the maximum luminance value).
[0088] The gain processor 213 is coupled to the gain controller 215, which is fed the determined gain and is arranged to control the gain circuit 103 (and, in some cases, the exposure attribute of the sensor 101) to apply the determined gain. For example, the gain processor 213 can be arranged to set the amplification level of the gain amplifier of the gain circuit 103 to match the determined gain.
[0089] Therefore, the method can provide a feedback mechanism that regulates the capture operation to provide an image with color channel balance, which can be carefully controlled to provide favorable attributes for detecting vital signs. This is achieved by specific consideration of skin color, allowing not only effective and reliable adjustment and calibration, but also specific adjustment and potential optimization for an individual person. In addition, a favorable color channel balance is achieved while providing adjustment and ensuring that the imaging of a given image area commonly used to determine vital signs is suitable for and conducive to such detection. The method can advantageously consider different regions of a person and ensure that imaging is completed through combined adjustment / calibration such that a favorable color balance (which can be considered a white balance operation in many cases) and brightness adjustment can be achieved and optimized for vital sign detection.
[0090] The camera device includes a vital sign detector 217 that is arranged to detect vital signs based on video images. After an initial gain setting, the video images can be specifically adjusted and possibly optimized for a particular vital sign detection algorithm as previously described. For example, the color channels can be individually adjusted to highlight features that are particularly suitable for detecting specific attributes. As an example, in some cases, the changes in the captured image of a person due to the pulse may be more obvious in some color channels than in others. For example, depending on skin color, the optical changes of a given illumination of a person can be different in different color channels. The method described can weight the color channels with higher variation relative to those with lower variation.
[0091] The vital sign detector 217 is fed with video images and performs a suitable image-based vital sign detection / algorithm.
[0092] For example, the vital sign detector 217 can perform a photoplethysmography algorithm to determine the volume change of blood circulation based on the video images. Such measurements can be specifically used to determine the pulse rate. Specifically, it can be assumed that the effect of the volume change on the spectrum of light (light absorption by blood) is the same among people; its direction in the RGB color space changes little among people. Given the knowledge of the pulse color direction, the measurement of the pulse signal can be considered a least squares optimization problem. Since the skin is actually a person-dependent filter that changes the light source spectrum, a short-time averaging operation is used to normalize the color signal before the least squares optimization to compensate for the effect of the skin.
[0093] It is assumed that the influence of volume change on the spectrum of light (light absorbed by blood) is the same among humans; its direction in the RGB color space varies little among humans. Given the knowledge of the pulse color direction, the measurement of the pulse signal can be considered as a least squares optimization problem. Since the skin is actually a human-dependent filter that changes the light source spectrum, a short-time averaging operation is used to normalize the color signal before the least squares optimization to compensate for the influence of the skin.
[0094] As another example, the vital sign detector 217 can determine the respiration signal by cross-correlating a one-dimensional representation of an image (or a selected region of interest of the image) with a one-dimensional representation of an earlier image. This one-dimensional representation can be any function that captures the texture in the image or region of interest (e.g., by calculating the average of all pixel values in a row). The maximum correlation is the derivative of the chest position, and when integrated, it forms a measurement of the chest expansion (respiration signal) during respiration. And, all derivatives of pulse and respiration, such as: heart rate interval, respiratory effort, etc.
[0095] As another example, the vital sign detector 217 can determine the pulse transit time based on pulse signals extracted from the skin at two different body positions (or the skin at one position and the ECG). This can be used as a surrogate for blood pressure (see, for example, https: / / pubmed.ncbi.nlm.nih.gov / 28324936 / ).
[0096] In different embodiments, different algorithms and functions can be used to determine the gain based on reference color channel distribution properties, properties of a set of color channel distributions, and luminance properties.
[0097] As described, in many embodiments, the gain setting can be determined mainly based on considerations of color channel intensity values and the distribution of color channel intensity values in individual color channels relative to reference intensity values for the color channels. The color channel intensity values can be represented by color channel pixel values (e.g., each of the RGB values).
[0098] In some embodiments, the gain can be determined according to each value. For example, the gain processor 213 can include a look-up table (LUT) that has as inputs the luminance value determined for the second image region, a given statistical property for each color channel that describes the distribution of each color channel (e.g., the ninetieth percentile pixel value), and the reference intensity values for different color channels. The LUT can provide stored values for a set of gain values for the color channels based on these input values. Thus, the LUT can directly provide an appropriate gain setting.
[0099] Such a LUT can be determined, for example, by performing a large number of tests during the design phase and determining gain values that provide effective vital sign detection for different parameters, settings, and scenarios. The LUT can be populated with the gain values determined in this way. In other examples, the gain values can be determined based on, for example, partial measurements and / or analytical evaluations.
[0100] In some embodiments, such a method can be used instead of populating the LUT to determine the functional relationship between different parameters and appropriate gain values, and the gain processor 213 can be arranged to apply such a function to determine the appropriate gain.
[0101] In many embodiments, the gain processor 213 can be arranged to determine an individual gain value for each color channel based on a reference intensity value and a determined intensity value for the color channel. The gain processor 213 can determine the gain for the first color channel based on the difference between the determined intensity value for the first color channel and the reference intensity value for the first color channel value. For example, as previously described, in some embodiments, the gain can be determined to reflect the difference between the determined intensity value and the reference intensity value for a given color channel. Specifically, the gain for a given color channel can be determined as the ratio between the reference intensity value and the determined intensity value. For example, the reference intensity value can indicate the preferred ninety - percentile color channel value, and the determined intensity value can indicate the measured ninety - percentile color channel value for the first image region. Determining the gain as the ratio between these values will adjust the capture operation such that the measured ninety - percentile value for the image region is set to the preferred value.
[0102] Furthermore, applying this method individually to all color channels will result in all color channels having the desired scaling. Applying this method to all color channels will further achieve a match between the relationship (and specifically the ratio) of the determined intensity values for different color channels and the relationship (and specifically the ratio) of the reference intensity values for different color channels. This can provide particularly advantageous operation in many scenarios and generally can allow emphasizing a particular color channel or, in fact, can allow setting the white balance of the image as desired (including as natural as possible).
[0103] In fact, in many embodiments, gain can generally be generated to maintain the same color channel relationship as the reference intensity value without having to set the individual color channels to the reference intensity values. For example, based on the luminance attribute of the second image region, the common gain component for all gains can be changed, for example, to ensure that the luminance attribute remains within a given range. This may result in individual color channel attributes being different from the reference values, but the ratio / relationship between the color channels is maintained.
[0104] In many embodiments, an individual gain component for a color channel is determined based on the determined intensity value and a reference intensity value, and a common gain component for all color channels is determined based on a luminance attribute. For example, a gain value for an individual color channel can be determined based on the reference intensity value and the determined intensity value of the individual color channel value as described above. Then, the gain processor 213 can determine a gain compensation luminance attribute for these gains, which corresponds to the measured luminance value for the second image region scaled proportionally to the determined gains. If the resulting gain compensation luminance value is within a given range, the gain value can be maintained, otherwise a common scaling factor for all color channels can be applied to the determined gains to cause the gain compensation luminance value to fall within the given range.
[0105] In some embodiments, the gain processor 213 is arranged to control an exposure attribute of the optical sensor based on a comparison of the luminance attribute and / or a reference color channel distribution attribute with an attribute of a set of color channel distributions. The exposure attribute can specifically be a shutter speed / duration and / or an aperture. Thus, in some embodiments, the gain controller 215 can also be arranged to control the capture operation of the optical sensor. These exposure attributes can control how much light is captured for each image. For example, increasing the shutter duration and / or the aperture causes more light to reach the sensor and thus results in a greater value of the optical sensor signal and correspondingly higher pixel values.
[0106] In fact, generally, changing the effect of exposure control has the same effect as changing the gain / amplification applied to the optical sensor signal, and in fact it can be modeled as such. The signal path of the capture operation can correspondingly be Figure 3 illustrated by an example of, where the gain 201 represents the effect of changing the exposure parameters of the sensor 201. The gain 205 represents the main gain common to all color channels, and the gain 207 represents the gain for each individual color channel.
[0107] Thus, for a given set of determined gains, these gains can be divided into a common gain and an individual gain for each color channel. For example, the common gain is applied to all optical sensor signals and can be set, for example, to the gain value determined for one of the color channels. Then the individual gain for that color channel can be set to 1, and the individual gains for the other color channels can be set to the ratio between the determined gain for that color channel and the common gain.
[0108] The common gain can be achieved by controlling the exposure and setting the main gain. The exact method can depend on the specific preferences and requirements of individual embodiments. For example, in some embodiments, it may be desirable to set the exposure attributes to maximize light capture (high aperture, long exposure) in order to maximize the signal-to-noise ratio. Then, the remaining common gain can be achieved by the main gain. In other examples, it may be desirable to set the shutter speed and / or aperture to a desired value (or within a certain range) to provide a clear image. Then the main gain can be set again to produce the desired overall common gain.
[0109] As a specific example, the following method can be used to implement the determined gain:
[0110] 1. The minimum gain can be implemented using the exposure and the main gain, i.e., the common gain is set to the lowest determined color channel gain.
[0111] 2. The gain controller 215 can first maximize the exposure, i.e., it can set the shutter speed and / or aperture to the maximum value to improve the signal-to-noise ratio. Then the remaining contribution to the common gain is implemented by appropriately setting the main gain. In many embodiments, the main gain can be set only with a rather coarse granularity, such that the nearest quantization step can be set. The following steps can be performed to further fine-tune the capture operation:
[0112] 3. The gain controller 215 can determine the corresponding total gain for a given color channel for the determined exposure and main gain settings. This can be specifically done for a single color channel and typically for the green channel, as it has the highest contribution to luminance and is most important for blood volume pulse measurement. Since the main gain is typically coarse, the resulting multiplication factor will usually not match the desired multiplication factor for the green channel.
[0113] 4. To fine-tune the gain for the green channel, fine-tuning of the exposure attributes can be performed. Thus, the shutter speed and / or aperture can be adjusted to more accurately achieve the overall desired gain for the green channel.
[0114] 5. Then the individual color channel gains for the red and blue channels can be determined for the given exposure and main gain settings that have been determined for the green channel.
[0115] The above methods can be applied to all images in many embodiments. However, in many embodiments, the camera device can be arranged to distinguish between images in which a skin region is detected / can be detected and images in which no skin region is detected / cannot be detected. Specifically, for embodiments in which face detection is used to determine a first image region, the system can be arranged to distinguish between images in which face detection is successful and images in which face detection is unsuccessful.
[0116] The gain processor 213 may be specifically arranged in such an embodiment to determine the color channel distribution for an image in which a skin region is detected, and store these color channel distributions for use in determining the gain for an image in which no skin region is detected. The determined and stored color channel distributions are typically for a region different from the first and second regions, and are specifically typically larger. In many embodiments, the gain processor 213 determines the color channel distribution for the entire image in which a skin region has been detected, and stores this full image color channel distribution.
[0117] When processing an image in which no skin region has been detected, the gain processor 213 may continue to retrieve the stored color channel distributions, and may further determine the color channel distribution for the current image (also referred to as a skinless image for brevity). The color channel distribution for the skinless image (also referred to as the current color channel distribution for brevity) may be determined for the same image regions for which the stored color channel distributions were determined, and may typically be determined for the full image (i.e., the image regions for which the current color channel distribution can be determined as the full image region).
[0118] As a low complexity example, if the comparison indicates that the stored color channel distribution and the current color channel distribution are similar enough (according to any suitable difference / similarity criterion), the gain processor 213 may continue to directly reuse the gain determined for the image of the stored color channel distribution (and thus these gains may have been stored together with the color channel distribution). If the distributions are not similar enough, the gain processor 213 may, for example, continue to alternatively apply a nominal or default gain value.
[0119] In many embodiments, the gain processor 213 may be arranged to modify the gain determined for the stored color channel distribution to reflect the difference between this distribution and the current color channel distribution. The compensation may specifically cause the resulting difference between the color channel distributions to be reduced.
[0120] For example, the gain processor 213 may determine the given (e.g., ninetieth) percentile intensity / pixel value for each color channel of both distributions. The compensation factor for each color channel may then be determined as the ratio between the (ninetieth) percentile value of the stored color channel distribution and the (ninetieth) percentile value of the current color channel distribution. The gain for each color channel may then be determined as the stored gain for the stored color channel distribution multiplied by the compensation factor. This may result in setting the gain that achieves the same (ninetieth) percentile value, which may be applicable to many practical scenarios where the capture environment varies and the time interval during which face detection occurs is relatively small.
[0121] As a specific example, in some embodiments, the gain processor 213 may be arranged to determine the full-image percentile intensity values for all images, color channels, and percentiles (0% to 100% for each color channel). These values may be stored (i.e., the color channel distribution given by the percentile values is stored). If then a situation occurs where no skin region is detected, the gain processor 213 may retrieve the stored color channel distribution of the last image in which a skin region was detected. Then, the percentiles that are equal to or below the highest target percentile for the first image region may be determined. Then, this percentile may be used as the new target value for the channel having the highest intensity value for the given percentile.
[0122] This method may, for example, allow the capture operation to remain unchanged if no skin region is detected, unless there is a significant lighting change, in which case the gain processor 213 will seek to reproduce the conditions as they were just before the failure to detect the skin region.
[0123] In some embodiments, a particular problem may be the startup process and how to set the initial capture attributes. In particular, it may be highly desirable to set the initial gain value such that a suitable image is captured and, in many embodiments, such that accurate skin / facial detection is enabled.
[0124] In some embodiments, the gain processor 213 may be arranged to determine the initial gain based on a consideration of the attributes of the full image and thus does not perform or require skin or facial detection. The gain processor 213 may analyze the full image to determine intensity values / attributes. For example, the gain processor 213 may determine a given percentile intensity value for each color channel. For example, the pixel / intensity value for the ninety-fifth percentile of the entire image may be determined for each color channel. The determined pixel / intensity value may then be compared with a predetermined reference intensity value, which may specifically be set to reflect the preferred intensity value for the ninety-fifth percentile. The predetermined reference intensity value may, for example, represent the value of the ninety-fifth percentile, which has been found in suitable trials and analysis to generally provide a suitable exposure setting, thereby producing an image that is particularly likely to be suitable for skin region and / or facial detection. In some embodiments, a common initial gain for all color channels may be determined and specifically may be set to the ratio between the predetermined reference intensity value and the determined percentile intensity value.
[0125] In some embodiments, percentile intensity values can be determined for only one of the color channels. For example, one color channel may be more important for face detection, and exposure can be optimized for that color channel. However, in many embodiments, values can be determined for all color channels, and the percentile intensity value to be compared with a predetermined reference intensity value can be selected as the highest value for each color channel. Thus, the initial gain can be adjusted such that a given percentile value for the highest-intensity color channel is set to a preferred level. This can allow for capturing images suitable for skin / face detection in many scenarios.
[0126] As a specific example, the gain processor 213 can perform an initial full-image exposure control to achieve initial face detection and tracking. The initial camera control can control the maximum measured full-image nth percentile towards a (configurable) target value.
[0127] Figure 4 FIG. 7 is a block diagram showing an example processor 400 according to an embodiment of the present disclosure. The processor 400 can be used to implement one or more processors that implement the device or its elements as described above. The processor 400 can be of any suitable processor type, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (where the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (where the ASIC has been designed to form a processor), or a combination thereof.
[0128] The processor 400 can include one or more cores 402. The cores 402 can include one or more arithmetic logic units (ALUs) 404. In some embodiments, in addition to or instead of the ALU 404, the cores 402 can include a floating point logic unit (FPLU) 406 and / or a digital signal processing unit (DSPU) 408.
[0129] The processor 400 can include one or more registers 412 communicatively coupled to the cores 402. The registers 412 can be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 412 can be implemented using static memory. The registers can provide data, instructions, and addresses to the cores 402.
[0130] In some embodiments, the processor 400 may include one or more levels of cache memory 410 communicatively coupled to the core 402. The cache memory 410 may provide computer-readable instructions for execution to the core 402. The cache memory 410 may provide data for processing by the core 402. In some embodiments, the computer-readable instructions may be provided to the cache memory 410 from local memory (e.g., local memory attached to the external bus 416). The cache memory 410 may be implemented with any suitable cache memory type, e.g., metal-oxide semiconductor (MOS) memory such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.
[0131] The processor 400 may include a controller 414, which may control inputs to the processor 400 from other processors and / or components included in the system and / or outputs from the processor 400 to other processors and / or components included in the system. The controller 414 may control the data paths in the ALU 404, FPLU 406, and / or DSPU 408. The controller 414 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 414 may be implemented as discrete gates, FPGAs, ASICs, or any other suitable technology.
[0132] The registers 412 and the cache 410 may communicate with the controller 414 and the core 402 via internal connections 420A, 420B, 420C, and 420D. The internal connections may be implemented as buses, multiplexers, cross switches, and / or any other suitable connection technology.
[0133] The input and output of the processor 400 may be provided via the bus 416, which may include one or more conductive lines. The bus 416 may be communicatively coupled to one or more components of the processor 400, such as the controller 414, the cache 410, and / or the registers 412. The bus 416 may be coupled to one or more components of the system.
[0134] The bus 416 can be coupled to one or more external memories. The external memories can include a read-only memory (ROM) 432. The ROM 432 can be a masked ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memories can include a random access memory (RAM) 433. The RAM 433 can be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memories can include an electrically erasable programmable read-only memory (EEPROM) 435. The external memories can include a flash memory 434. The external memories can include a magnetic storage device such as a disk 436. In some embodiments, the external memories can be included in the system.
[0135] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention can optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.
[0136] In this application, any reference to one of the terms "responsive to", "based on", "dependent on", and "in accordance with" should be considered a reference to the term "responsive to / based on / dependent on / in accordance with". Any term should be considered a disclosure of any other term, and the use of only a single term should be considered a shorthand notation that includes other alternatives / terms.
[0137] Although the present invention has been described in connection with some embodiments, the present invention is not intended to be limited to the specific forms set forth herein. Rather, the scope of the present invention is limited only by the claims. Additionally, although features may appear to be described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0138] In addition, although listed individually, multiple devices, elements, circuits or method steps can be implemented by, for example, a single circuit, unit or processor. In addition, although individual features may be included in different claims, these features may be advantageously combined, and inclusion in different claims does not mean that the combination of features is not feasible and / or unfavorable. Moreover, including a feature in a claim category does not imply a limitation on the category, but rather indicates that the feature is equally applicable to other claim categories when appropriate. In addition, the order of features in the claims does not imply that the features must work in any particular order, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in that order. On the contrary, the steps may be performed in any suitable order. In addition, singular references do not exclude multiples. Therefore, references to "one", "one", "first", "second", etc. do not exclude multiples. The figure marks in the claims are provided only as clarification examples and should not be interpreted as limiting the scope of the claims in any way.
Claims
1. An apparatus for estimating a person's vital signs based on a video image, the apparatus comprising: a receiver (201) arranged to receive a video image including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; a first detector (203) arranged to determine a first image region in at least one image of the video image, the first image region corresponding to the person's skin region; a second detector (205) arranged to determine a second image region in at least one image of the video image, the second image region corresponding to a region of the person different from the first image region; a determiner (207) arranged to determine a set of color channel distributions for pixels of the first image region; a processor (209) arranged to provide a skin color indication indicating the person's skin color; a reference processor (211) arranged to provide reference color channel distribution attributes for the skin color; a gain processor (213) arranged to determine a gain for the optical sensor signal based on a brightness attribute determined for the second image region and based on a comparison of the reference color channel distribution attributes with the attributes of the set of color channel distributions; a gain controller (215) arranged to control a gain value applied to the optical sensor signal to match the determined gain; and a vital sign determiner (217) arranged to estimate the person's vital sign attributes based on the video image.
2. The apparatus according to claim 1, further comprising: a gain circuit (103) arranged to apply the gain value to optical sensor signals from at least two color channel optical sensors; and an image generator (105) coupled to the gain circuit (103) and arranged to generate the video image based on the optical sensor signals received from the gain circuit (103).
3. The apparatus according to any one of the preceding claims, wherein, the reference color channel distribution attributes include a set of reference intensity values for different color channels, and the attributes of the color channel distribution include a set of determined intensity values for the set of color channel distributions.
4. The apparatus according to claim 3, wherein, the determined intensity values are color channel pixel values at a given percentile for the set of color channel distributions.
5. The apparatus according to claim 3 or 4, wherein, the gain processor (213) is arranged to determine a gain for the first color channel based on a difference between a determined intensity value for the first color channel among the plurality of color channels and a reference intensity value for the first color channel value.
6. The apparatus according to any one of claims 3 to 5, wherein, the gain processor (213) is arranged to set the gain such that a relationship between determined intensity values for different color channels matches a relationship between reference intensity values for different color channels.
7. The apparatus according to any one of claims 3 to 5, wherein, the luminance attribute is the luminance for the second image region generated by applying the determined gain to the second image region, and the gain processor (213) is arranged to set the gain such that the luminance attribute indicates that the luminance for the second image region meets a criterion.
8. The apparatus according to any one of the preceding claims, wherein, the gain processor (213) is arranged to determine a common gain component based on the luminance attribute and to determine a relative gain component based on the comparison of the reference color channel distribution attribute with the attributes of the set of color channel distributions, the common gain component being a gain applied to all of the plurality of color channels, and the relative gain component being a gain component applied to an individual color channel of the plurality of color channels.
9. The apparatus according to any one of the preceding claims, wherein, the gain processor (213) is arranged to control an exposure attribute of at least one color channel optical sensor based on at least one of the luminance attribute and the comparison of the reference color channel distribution attribute with the attributes of the set of color channel distributions.
10. The apparatus according to any one of the preceding claims, wherein, the gain processor (213) is arranged to determine and store a first color channel distribution for a third image region of an image for which a skin region has been detected; and to determine a gain for the optical sensor signal for a first image in which no skin region has been detected based on the first color channel distribution and a second color channel distribution, the second color channel distribution being determined for the third image region of the first image.
11. The apparatus according to claim 10, wherein, the gain processor (213) is arranged to set the gain to reduce the difference between the attributes of the first color channel distribution and the second color channel distribution.
12. The apparatus according to claim 10 or 11, wherein, the third image region is a full image region.
13. The apparatus according to any one of the preceding claims, wherein, the gain processor (213) is arranged to determine an initial gain for the optical sensor signal based on a comparison of a full image intensity value with a predetermined reference intensity value; the full image intensity value being an intensity value for a predetermined percentile of the color channels of the plurality of color channels.
14. A method of estimating a person's vital signs from an image, the method comprising: receiving a video image including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; determining a first image region in at least one image of the video image, the first image region corresponding to a skin region of the person; determining a second image region in at least one image of the video image, the second image region corresponding to a region of the person different from the first image region; determining a set of color channel distributions for pixels of the first image region; providing a skin color indication indicative of the skin color of the person; Provide a reference color channel distribution attribute for the skin color; Determine a gain for the optical sensor signal based on the luminance attribute determined for the second image region and based on a comparison of the reference color channel distribution attribute with the attributes of the set of color channel distributions; Control the gain value applied to the optical sensor signal to match the determined gain; and Estimate the vital sign attribute of the person based on the video image.
15. A computer program product comprising computer program code modules which, when the program is run on a computer, are adapted to carry out all the steps of claim 14.