Detecting vital signs from images

The apparatus and method enhance vital sign detection by adapting camera settings based on image region analysis and color channel distributions, addressing inefficiencies in existing systems and improving detection accuracy and simplicity.

JP2025536196APending Publication Date: 2025-11-05KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517528
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-11
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing image-based vital sign detection systems face challenges in optimizing camera settings for accurate and reliable determination of vital signs, particularly in varying lighting conditions and for different skin tones, leading to inefficiencies and increased complexity.

Method used

An apparatus and method that adapt camera settings by determining specific image regions, calculating color channel distributions, and adjusting gains to match reference characteristics, ensuring optimal image capture for vital sign detection, including pulse and respiratory rates.

Benefits of technology

Improves the accuracy and reliability of vital sign detection by reducing sensitivity to environmental variations and simplifying implementation, while optimizing image capture for specific detection algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536196000001_ABST
    Figure 2025536196000001_ABST
Patent Text Reader

Abstract

The apparatus estimates a person's vital signs from a video image having color channels representing digitized optical sensor signals. A first detector 203 detects a first image region that is a skin region of the person, and a second detector 205 detects a second image region that corresponds to a different region of the person from the first image region. A determiner 207 determines a color channel distribution of pixels in the first image region. A processor 209 provides a skin tone index, and a reference processor 211 provides a reference color channel distribution characteristic for skin tone. A gain processor 213 determines a gain of the optical sensor signal depending on a luminance characteristic of the second image region and a comparison of the reference distribution characteristic with the color channel distribution characteristic. A gain controller 215 controls a gain value applied to the optical sensor signal accordingly, and a vital sign determiner 217 estimates the person's vital sign characteristic from the image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to detecting a person's vital signs from one or more images, and in particular, but not exclusively, to detecting pulse and / or respiratory vital signs from images of a person's face, chest and / or skin areas. [Background technology]

[0002] Image-based vital sign detection of a person has become a very advantageous and useful tool in many settings.

[0003] Image-based vital signs monitoring seeks to measure vital signs from visual cues captured by an image of a person. Such vital signs specifically include pulse rate and respiratory characteristics. Applications include, but are not limited to, emergency medical services and medical triage, (infectious disease) screening applications, and telemedicine solutions.

[0004] Various approaches and algorithms for extracting vital sign data from images are known in the art. For example, remote photoplethysmography (rPPG) is a technique used to determine blood volume pulse signals from images of skin tissue, typically captured using custom or commercially available cameras. As another example, respiratory signals and / or respiratory characteristics such as respiratory rate can be estimated from images using, for example, detection of changes in the image due to chest / abdominal movement.

[0005] Such vital signs monitoring is highly advantageous in many settings because it allows practical, non-intrusive monitoring and detection of vital signals.

[0006] For optimal detection of vital signs, it is important that the image has characteristics suitable for the specific detection / estimation. For example, the camera's capture characteristics must be carefully configured to ensure proper capture operation. In particular, the configuration and calibration of optical and sensor parameters is crucial to achieve high-quality imaging of a person, particularly to ensure that the captured image is suitable for vital sign detection. In particular, adapting the camera's characteristics and settings prior to digitization of the sensor signal is essential to ensure that the captured image allows for accurate and reliable determination of vital signs. Such considerations include adapting the camera's characteristics and settings to prevent clipping, effectively utilize the dynamic range, achieve an improved signal-to-noise ratio, adapt the relationship between different color channels, and efficiently detect visual characteristics suitable for vital sign estimation.

[0007] Optimizing for vital sign detection can be particularly challenging because the settings and parameters may differ from optimal settings and parameters that are optimized for perception by the human visual system and attempt to provide the most visually realistic image possible. Optimal settings for vital sign detection may differ. For example, detecting a person's pulse may be more reliable if the detection is based on an exaggerated image of the visual characteristics that change with the pulse.

[0008] Accordingly, improved approaches for determining vital signs from captured images would be advantageous, particularly approaches that allow for improved operation, increased flexibility, improved determination of vital signs, improved representation of visual characteristics that enable vital signs to be determined, reduced sensitivity to variations in the imaging environment and lighting conditions, reduced complexity, easier implementation, reduced memory and / or storage requirements, and / or improved performance and / or operation. Summary of the Invention [Problem to be solved by the invention]

[0009] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0010] According to one aspect of the present invention, there is provided an apparatus for estimating vital signs of a person from video images, the apparatus comprising: a receiver configured to receive video images including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; a first detector configured to determine a first image region in at least one of the video images, the first image region corresponding to a skin region of the person; a second detector configured to determine a second image region in the at least one of the video images, the second image region corresponding to a region of the person different from the first image region; and a color filter configured to filter pixels of the first image region. a determiner configured to determine a set of color channel distributions; a processor configured to provide a skin tone index, the skin tone index being indicative of a skin tone of the person; a reference processor configured to provide a reference color channel distribution characteristic of skin tone; a gain processor configured to determine a gain of the optical sensor signal dependent on the luminance characteristics determined for the second image region and dependent on a comparison of the reference color channel distribution characteristic with a characteristic of the set of color channel distributions; a gain controller configured to control a gain value applied to the optical sensor signal to match the determined gain; and a vital sign determiner configured to estimate a vital sign characteristic of the person from the video image.

[0011] The present invention, in many embodiments, can improve image-based vital sign determination. The approach can enable improved adaptation to specific environmental conditions, particularly lighting conditions. The approach can particularly enable improved adaptation for the specific purpose of detecting vital signs, rather than simply optimizing the realism of the captured video image, for example. The approach can, in many scenarios, facilitate operation and improve user convenience, for example, by reducing the need for manual adaptation and / or precise control of the imaging environment. In many scenarios, the approach can enable capture operations that specifically focus on visual characteristics suitable for determining vital signs, thereby typically improving the reliability and accuracy of detection.

[0012] This approach can provide improved camera / capture control in many embodiments: the sensitivity of capture operations to variations in lighting conditions and the different people whose vital signs are being determined can often be significantly reduced.

[0013] The color channels may specifically be red, green, and blue color channels, and the video image may be an RGB image.

[0014] The first image region and the second image region may typically be in the same image of a video signal, but in some embodiments and scenarios may be in different images. The video image may be a frame of a video signal.

[0015] The color channel distribution set characteristic comprises a set of values, each value being for one color channel of the multiple color channels. The reference color channel distribution characteristic can comprise a set of values, each value being for one color channel of the multiple color channels.

[0016] The vital signs, in many embodiments, are pulse rate and / or respiratory rate. The second image region is an image region corresponding to the person's chest.

[0017] The characteristic of the set of color channel distributions can be a set of one or more (pixel) percentile values ​​of the color channels, and specifically, a set of one or more pixel values ​​corresponding to predetermined percentiles of pixel values ​​of the color channels.

[0018] The device may be a camera device.

[0019] According to an optional feature of the invention, the apparatus may further include a gain circuit configured to apply gain values ​​to optical sensor signals from the at least two color channel optical sensors; and an image generator coupled to the gain circuit and configured to generate a video image from the optical sensor signals received from the gain circuit.

[0020] This approach may, in many embodiments, provide improved operation and / or ease of implementation, particularly in capturing images for vital signs determination.

[0021] In accordance with an optional feature of the invention, the characteristics of the reference color channel distributions include a set of reference intensity values ​​for the different color channels, and the characteristics of the color channel distributions include a set of determined intensity values ​​for the set of color channel distributions.

[0022] This can provide improved operation and / or easier implementation in many embodiments, particularly typically providing efficient adaptation of capture operations to determine vital signs.

[0023] The determined intensity values ​​may in particular be determined pixel values, in particular intensity values / pixel values ​​for a given percentile of the distribution.

[0024] In accordance with an optional feature of the invention, the determined intensity values ​​are color channel pixel values ​​for predetermined percentiles of a set of color channel distributions.

[0025] This can result in optimal operation and / or ease of implementation in many embodiments.

[0026] According to an optional feature of the invention, the gain processor is configured to determine a gain for a first color channel of the plurality of color channels dependent on a difference between the determined intensity value for the first color channel and a reference intensity value for the first color channel value.

[0027] This can provide improved operation and / or easier implementation in many embodiments, particularly typically providing efficient adaptation of capture operations to determine vital signs.

[0028] According to an optional feature of the invention, the gain processor is configured to set the gains such that a relationship between the determined intensity values ​​of the different color channels matches a relationship between reference intensity values ​​of the different color channels.

[0029] This can allow for particularly efficient color channel balancing (e.g., white balancing), which can typically be optimized to the particular algorithm used to determine the vital signs.

[0030] According to an optional feature of the invention, the luminance characteristic is a luminance for the second image region resulting from applying the determined gain to the second image region, and the gain processor is configured to set the gain to indicate a luminance for the second image region for which the luminance characteristic satisfies the criterion.

[0031] This can provide efficient adaptation.The luminance criteria can include, among other things, a requirement that the luminance be within a predetermined luminance range.

[0032] According to an optional feature of the invention, the gain processor is configured to determine a common gain component dependent on a luminance characteristic and to determine relative gain components dependent on a comparison of a reference color channel distribution characteristic with a characteristic of the set of color channel distributions, the common gain component being a gain applied to all color channels of the plurality of color channels and the relative gain components being gain components applied to individual color channels of the plurality of color channels.

[0033] This may result in improved operation and / or easier implementation in many embodiments, allowing for efficient adaptation of capture operations as well as practical yet flexible implementations.

[0034] According to an optional feature of the invention, the gain processor is configured to control an exposure characteristic of the at least one color channel optical sensor depending on at least one of a luminance characteristic, a comparison between a reference color channel distribution characteristic and a characteristic of the color channel distribution set.

[0035] This typically allows for improved adaptation, e.g., to a wider variety of imaging scenarios and lighting conditions, which can often improve image quality, e.g., allowing for a higher signal-to-noise ratio for the optical sensor signal and / or captured image.

[0036] This can result in improved operation and / or ease of implementation in many embodiments.

[0037] According to an optional feature of the invention, the gain processor is configured to determine and store a first color channel distribution for a third image region of the image in which a skin region is detected, and to determine a gain of the optical sensor signal for the first image in which no skin region is detected depending on the first color channel distribution and the second color channel distribution determined for the third image region of the first image.

[0038] This can improve operation in many embodiments and scenarios. It allows appropriate adaptation to be performed even for images / times where typically no skin regions can be detected in the video image. It can often ensure that capture operations are adapted to provide improved conditions for detecting image regions corresponding to skin regions. For example, it can ensure that capture operations for time periods where no skin is detected (e.g., by a face detector) are such that the captured images are suitable for performing skin detection operations (e.g., face detection specifically).

[0039] According to an optional feature of the invention, the gain processor is configured to set the gain to reduce a difference between characteristics of the first color channel distribution and characteristics of the second color channel distribution.

[0040] This can result in improved operation and / or ease of implementation in many embodiments.

[0041] In accordance with an optional feature of the invention, the third image area is a full image area.

[0042] According to an optional feature of the invention, the gain processor is configured to determine an initial gain of the optical sensor signal dependent on a comparison of the full image intensity value with a predetermined reference intensity value, the full image intensity value being a predetermined percentile intensity value of a color channel of the plurality of color channels.

[0043] This allows for improved initial adaptation and performance, for example by allowing better images to be captured for early skin area detection.

[0044] According to one aspect of the present invention, there is provided a method for estimating a person's vital signs from images, the method including: receiving video images including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; determining a first image region in at least one of the video images, the first image region corresponding to a skin region of the person; determining a second image region in the at least one of the video images, the second image region corresponding to a different region of the person from the first image region; determining a set of color channel distributions for pixels of the first image region; and providing a skin tone index, the skin tone index indicative of a skin tone of the person; providing a reference color channel distribution characteristic for skin tone; determining a gain of the optical sensor signal dependent on a luminance characteristic determined for the second image region and dependent on a comparison between the reference color channel distribution characteristic and a characteristic of the set of color channel distributions;

[0045] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. [Brief explanation of the drawings]

[0046] An embodiment of the present invention will now be described, by way of example only, with reference to the drawings. [Figure 1] FIG. 1 illustrates an example of elements of a camera according to some embodiments of the present invention. [Figure 2] FIG. 1 illustrates an example of elements of a camera according to some embodiments of the present invention. [Figure 3]FIG. 1 illustrates an example of elements of a camera according to some embodiments of the present invention. [Figure 4] 1A and 1B illustrate some elements of a possible configuration of a processor for implementing elements of an audio device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0047] Figure 1 shows an example of a camera device configured to capture an image of a person and determine the person's vital signs from the captured image, including pulse rate and respiratory rate.

[0048] Vital signs can be detected from a video sequence of images / frames, particularly based on variations occurring in the images. Thus, an optical camera can be positioned to capture a person / patient, and the resulting video images can be analyzed and compared to each other to detect vital signs. For example, photoplethysmography can be used to determine volumetric changes in blood circulation.

[0049] To achieve accurate and reliable detection of vital signs, it is important that the capture of the image is appropriately adapted to provide an image with optical properties suitable for detecting visual cues that may enable the vital signs to be detected. Such adaptation of the capture operation is important to the accuracy and reliability that can be achieved in determining the vital signs.

[0050] In the example of Figure 1, the camera device has multiple optical sensors 101 that generate optical sensor signals for different color channels. Thus, at least two optical sensors are configured to capture visible light at different bandwidths / frequency sensitivities. In many embodiments, the camera device includes three optical sensors that capture spectra corresponding to red, green, and blue (RGB) color channels, respectively. Thus, the optical sensors 101 generate optical sensor signals corresponding to RGB color signals with the RGB color channels.

[0051] The sensors are coupled to a gain circuit 103 configured to apply a gain to the optical sensor signals. The gain can be implemented, for example, as a series of gain / amplifiers that apply a variable gain to each analog optical sensor signal, which can be dynamically set, for example, via a suitable control signal.

[0052] The gain-adjusted optical sensor signal is provided to an image generator 105 configured to generate an image from the gain-adjusted optical sensor signal. The image is specifically a color image having color channels corresponding to the color channels of the sensor / optical sensor signal, and thus, in this specific example, an RGB color image. The image generator 105 can generate color channel values ​​for one color channel by digitizing one of the optical sensor signals from one of the optical sensors. The image generator 105 can specifically include an analog-to-digital converter and perform processing to generate a video image from the digital samples of the optical sensor signal. Such a signal path of optical sensor, gain adjustment, and image generation is used in many optical cameras and can provide high-quality captured images.

[0053] In this approach, the camera device further includes a processor 107 that is supplied with the generated image. The processor 107 is configured to determine vital signs based on the received video image. It is further configured to adapt capture operations and functions to provide images particularly suited for detecting vital signs. In particular, the processor 107 is configured to determine a gain for the optical sensor signal and control the gain circuit 103 to apply the determined gain. The gain can be determined separately for each color channel, and different gains can be applied to different optical sensor signals. It should be understood that the camera device may be within a single housing, have components split among multiple housings, or indeed have components located in other devices.

[0054] The processor 107 may apply a particular approach that is described with reference to FIG. 2, which shows an example of elements of the processor 107.

[0055] The processor 107 includes a receiver 201 that receives a video image from the image generator 105. In this manner, the receiver 201 receives a multi-color channel captured image consisting of multiple color channels, and specifically, the image may be an RGB video image.

[0056] The video image is fed to a first detector 203 and a second detector 205 configured to detect different image regions within the received image.

[0057] The first detector 203 is configured to detect, in the (at least one) received image / frame, a first image region corresponding to a skin region of a person. Thus, the first detector 203 can analyze the image to detect a region corresponding to an area capturing the skin of the person. Specifically, in many embodiments, the first detector 203 can perform face detection to detect the face of the person whose vital signs are to be determined.

[0058] It will be appreciated that many different approaches and algorithms are known and can be used to detect skin regions.

[0059] For example, in some embodiments, the first image region can be determined as a substantially predetermined region corresponding to a fixed position of a user relative to the camera. For example, the camera can capture a scenario in which a person is positioned to sit in a chair facing the camera, with the head fixed, for example, by a physical head bracket. In such a scenario, the person's face tends to always be in substantially the same position in the image, and the first image region can be determined as a predetermined region corresponding to the position of the face.

[0060] As another example, the first image region may include a face detection algorithm. Many different face detection algorithms are known, and it will be appreciated that any suitable approach may be used without detracting from the present invention. After detecting the location of a face within the image, the first detector 203 may proceed to generate the first image region as a region corresponding to the location of the detected face.

[0061] Thus, in many embodiments, the first image region and skin region may correspond to a facial region of the person. However, it will be understood that in other embodiments, the first image region and skin region may alternatively or additionally be for another body part. For example, in some embodiments, a person may be required to place their arms in a particular position (e.g., supported by a restraint), so that this is captured at a predetermined location in the image. As another example, a person may be required to be partially undressed, and the user's upper body may be positioned so that it fills most of the image, and the first image region / skin region may be determined as the largest region of the image with substantially similar visual characteristics. For example, segmentation based on image characteristics (such as color or brightness) may be performed on the image, and the first image region may be determined as the largest segment.

[0062] In many embodiments, appropriate parts of the body / image segments can be tracked. For example, a face tracker can be configured to focus on a person's face for camera control even as they look around in bed. As another example, a camera can track a person (and their face) as they move around the room. Camera control can be configured to be unaffected by other environmental factors (e.g., if a person walks in front of a sunny window, the camera would typically dim the entire video stream, thereby losing the ability to monitor vital signs).

[0063] The second detector 205 is configured to detect a second image region within the image. For example, it can detect a predetermined region or can detect an image region based on image characteristics. In many embodiments, the second image region can be determined as an image region used to determine vital signs. For example, the processor 107 can be configured to detect a pulse rate from visual changes in a person's chest, and therefore the second image region can be selected to correspond to the person's chest.

[0064] In some embodiments, the second image region can be determined as a predetermined region within the video image, for example, when the position of the person relative to the camera is tightly controlled. In other embodiments, the second image region can be detected based on, for example, a marker or other recognizable visual feature within the image. For example, a marker with distinct visual characteristics can be attached to the person being imaged. The marker can, for example, surround an area of ​​the body from which vital signs are determined. In this case, the second detector 205 can detect the outline of the marker by detecting a specific visual characteristic (e.g., a specific pattern) within the image. The second image region can then be determined as the area within the determined outline.

[0065] As yet another example, the second image region can be determined as a region that has a predetermined relationship to the first image region, depending on the region. For example, the second image region can be determined as a predetermined region that is located a predetermined distance below the detected face region and has a predetermined contour. This allows, for example, a chest region to be determined from the face detection that determines the first image region.

[0066] As another example, the second image region can be determined by identifying the chest using a body part detector.

[0067] The first detector 203 is coupled to a distribution determiner 207 configured to determine a color channel distribution of pixels of the first image region. The color channel distribution may generate, for each color channel, a distribution of pixel values ​​for that color channel. The color channel distribution may indicate, for a given color channel, the occurrence / proportion of pixel values ​​of pixels belonging to the first image region. The color channel distribution may indicate, for each color channel, the occurrence or frequency of pixel values ​​in the first image region for each (or group) of possible pixel values. For example, for each color channel, a histogram may be determined that reflects the distribution of pixel values ​​for that color channel.

[0068] The distribution determiner 207 can be configured to generate, for a given color channel and threshold, a color channel distribution that reflects the percentile / proportion of pixel values ​​that are below (or alternatively, equivalently, above) that threshold. Equivalently, the distribution determiner 207 can be configured to generate, for a given color channel and percentile, a color channel distribution that reflects the threshold for which that percentile of pixel values ​​are below (or alternatively, equivalently, above) the threshold.

[0069] The camera device further comprises a skin tone processor 209 configured to provide a skin tone indicator of the person whose image is being captured and whose vital signs are being determined. The skin tone indicator is indicative of the person's skin tone, and more particularly, visual skin tone characteristics. In many embodiments, multiple skin tone categories can be stored by the skin tone processor 209, and the skin tone indicator can indicate a category that is deemed appropriate for the particular person being captured.

[0070] In many cases, an assumption can be made that the ratio between the mean RGB values ​​of skin tissue is similar across skin phototypes, as found by De Haan, G., Jeanne, V., "Robust pulse-rate from chrominance-based rPPG." As detailed below, such an assumption can be used to control capture to provide images that are particularly suitable for vital sign detection.

[0071] In some embodiments, the skin tone processor 209 may be coupled to or include a user input, and an appropriate skin tone index may be determined in response to the user's input. For example, in some embodiments, a person may be presented with a set of images corresponding to different skin type categories stored by the skin tone processor 209. The person may then select the image that they believe best reflects the skin tone of the person being photographed, and the skin tone index may be set to indicate the corresponding category.

[0072] In some embodiments, skin tone detection can be determined, for example, from an analysis of the captured video image, particularly the first image region. An average pixel value can be determined, for example, for each color channel and compared to a set of reference average pixel values ​​for each of the skin tone categories. A skin tone index can then indicate that the skin tone category with the most matching reference average value is the person's skin tone. While such detection can sometimes be quite inaccurate because it is based on initial images that may not have been captured when the capture process was fully adapted or optimized, the difference between the reference average values ​​will usually be sufficient.

[0073] The skin tone processor 209 is coupled to a reference processor 211 configured to provide reference color channel distribution characteristics for the indicated skin tones. Thus, based on the received skin tone indicators, the reference processor 211 can determine reference distribution characteristics for each of the color channels. In different embodiments, the characteristics can be different characteristics that represent characteristics or features of the distributions (typically for each color channel separately). For example, the characteristics can be indicators of the mean, variance, percentile, mode, or median. Specifically, in many embodiments, the reference characteristics can be a set of pixel values ​​that correspond to a predetermined (e.g., predetermined) percentile of the distribution of pixel values ​​in the color channels.

[0074] The reference color channel distribution characteristics may specifically include a set of reference intensity values ​​for each color channel. For example, a pixel value may be provided for each color channel. The reference pixel value may, for example, indicate a preferred pixel value for a predetermined percentile of the distribution of pixel values ​​in the first image region.

[0075] The reference processor 211, in many embodiments, can store a set of reference characteristics, such as reference intensity values, for each possible skin tone category. The reference characteristics can be predetermined values ​​stored in memory, for example, during the manufacturing process. Thus, a predetermined, stored reference value can be associated with each possible skin tone category in some embodiments. Suitable values ​​can be found, for example, through testing and experimentation during the development and design of an approach to determining vital signs. The reference characteristics can be specifically optimized to provide the best possible detection performance desired for determining vital signs.

[0076] The reference processor 211 is coupled to a gain processor 213, which is configured to determine gains for the gain circuit 103 based on reference color channel distribution characteristics provided for a particular skin tone and the color channel distributions for pixels of the first image region. The gain processor 213 can specifically adapt the gains of the color channels to reduce the difference between the reference color channel distribution characteristics and the characteristics of the color channel distribution. Thus, the gain processor 213 can be configured to determine gains to reduce the difference between the determined color channel distribution characteristics and the reference color channel distribution characteristics for a particular skin tone indicated for a person. Specifically, the gains can be set to reduce the difference between the reference intensity values ​​for each color channel and the intensity values ​​determined for a given percentile for each of the color channel distributions.

[0077] For example, the reference color channel distribution characteristics may indicate a reference pixel / intensity value for each color channel. The reference pixel value may indicate, for example, a preferred pixel value for a given percentile (e.g., the 90th percentile) of pixels. The gain processor 213 may determine, for each color channel, a 90th percentile pixel value for a particular determined color channel distribution for the first image region. Then, for each color channel, a gain may be set to the ratio between the preferred pixel value and the determined 90th percentile pixel value. Thus, when the pixel values ​​of individual pixels in the first image region are scaled by the determined gain value, the first image region will have a 90th percentile value corresponding to the preferred value. Therefore, applying the determined gain to the gain circuit 103 results in a captured image with the desired characteristics.

[0078] However, the gain processor 213 is further configured not only to consider the first image region when determining the gain, but also to depend on the characteristics of the second image region. Specifically, the gain processor 213 is configured to determine a luminance characteristic of the second image region and set the gain according to this luminance characteristic. In many embodiments, the luminance characteristic can be a gain-compensated luminance characteristic that reflects the luminance of the second image region that would result if the determined gain were applied to the optical sensor signal, and thus for future images unless the capture environment changes. For example, the luminance characteristic can represent the luminance of the second image region that would result if the gain of the optical sensor signal were set as determined by the gain processor 213 and there were no other changes in capture or lighting conditions.

[0079] In many cases, the luminance characteristic can be determined for the second image region and scaled by a factor that depends on the determined gain. For example, for a given gain setting, a given function can be evaluated to provide a gain that reflects how the luminance would be changed by a change in gain from the gain used when the image was captured to the gain currently being considered. The luminance characteristic can then be scaled by this value to provide a gain-compensated luminance.

[0080] The luminance characteristic may indicate the luminance level of the second image region, specifically the maximum luminance, mean luminance, median luminance, or percentile (e.g., 90th percentile) luminance level of the second image region. The luminance characteristic typically indicates the luminance level as luminance per area (e.g., per pixel) for a luminance range. For example, the range may have a minimum luminance level corresponding to all color channel values ​​having a minimum value (e.g., RGB value of (0,0,0)) and a maximum luminance level corresponding to all color channel values ​​having a maximum value (e.g., RGB value of 8-bit value (255,255,255)).

[0081] As a low-complexity example, a luminance value can be calculated for each pixel of the second image region according to an appropriate luminance scale (e.g., by combining individual color channel values ​​according to weighting by a coefficient reflecting the relative contribution / importance of each color channel in determining the vital sign), and then a mean luminance value, a median luminance value, or, for example, a predetermined percentile (e.g., 90(th) percentile) luminance value can be determined.

[0082] The gain processor 213 can then adapt the gain taking into account the luminance characteristics. For example, the gain determination described above can be applied to set the 90th percentile color channel pixel value to a preferred value. However, rather than blindly using these gain values, the luminance characteristics of the second image region can be subject to a requirement that they meet a predetermined standard. For example, a luminance range can be given by a predetermined minimum luminance level and a predetermined maximum luminance level. Gain values ​​for each color channel can be determined as described above. Corresponding luminance gains can be determined to generate a gain-compensated luminance characteristic (by appropriately weighting the gains of each color channel to reflect their impact on luminance). For example, the luminance characteristic of the second image region (e.g., the 90th percentile luminance) can be multiplied by the luminance gain and the resulting value can be compared with the luminance range. If the modified luminance is below the predetermined range, the gain value can be adapted (increased) until the modified luminance exceeds a predetermined minimum luminance value. Similarly, if the modified luminance is above the predetermined range, the gain value can be adapted (decreased) until the modified luminance is below a predetermined maximum luminance value.

[0083] In many embodiments, the operation is linear, and gain adaptation to ensure the luminance characteristics of the second image region are within a desired range can typically be achieved by scaling the determined gain based on the color channels of the first image region, with the scaling being the same for all color channels. Thus, the determined gain can simply be multiplied by a predetermined factor that is the same for all color channels. In many cases, the scale factor can simply be the ratio between a predetermined minimum luminance value and the determined luminance value (if less than the minimum luminance value) or the ratio between a predetermined maximum luminance value and the determined luminance value (if greater than or equal to the maximum luminance value).

[0084] The gain processor 213 is coupled to a gain controller 215 that is supplied with the determined gain, and the gain controller 215 is configured to control the gain circuit 103 (and possibly the exposure characteristics of the sensor 101) to apply the determined gain. For example, the gain processor 213 can be configured to set an amplification level of a gain amplifier of the gain circuit 103 to match the determined gain.

[0085] This approach can then provide a feedback mechanism to adjust the capture operation to provide an image with a color channel balance that can be carefully controlled to provide advantageous characteristics for detecting vital signs. This is achieved by specifically considering skin tone, thereby enabling not only efficient and reliable adaptation and calibration, but also person-specific adaptation and potential optimization. Furthermore, advantageous color channel balance is achieved while simultaneously providing adaptation and ensuring that imaging of predetermined image regions typically used to determine vital signs is appropriate and advantageous for such detection. This approach can advantageously consider different regions of a person, and the combined adaptation / calibration can ensure that both advantageous color balance (which can often be thought of as a white balance operation) and luminance adaptation are achieved, resulting in imaging optimized for vital sign detection.

[0086] The camera device includes a vital sign detector 217 configured to detect vital signs based on the video image. After an initial gain setting, the video image can be specifically adapted and possibly optimized for a particular vital sign detection algorithm, as described above. For example, color channels can be individually adapted to emphasize features that are particularly suited to detecting certain characteristics. As an example, in some cases, fluctuations in a captured image of a person due to their pulse may be more noticeable in some color channels than in others. For example, depending on skin tone, a person's optical changes to a given illumination may differ in each color channel. The described approach can weight color channels with higher variability relative to color channels with lower variability.

[0087] The video image is provided to a vital signs detector 217, which performs a suitable image-based vital signs detection / algorithm.

[0088] For example, the vital signs detector 217 can execute a photoplethysmography algorithm to determine volumetric changes in blood circulation from video images. Such measurements can be used, in particular, to determine pulse rate. Specifically, the effect of volumetric variations on the spectrum of light (light absorption by blood) can be assumed to be identical across people, and its direction in the RGB color space shows little variation across people. Measuring the pulse signal can be considered a least-squares optimization problem, given knowledge of the pulse color direction. Because skin is actually a person-dependent filter that alters the spectrum of the light source, a short-term averaging operation is used to normalize the color signal before the least-squares optimization to compensate for the skin effect.

[0089] The effect of volume changes on the light spectrum (light absorption by blood) is assumed to be identical between individuals, and its direction in the RGB color space shows little variation between individuals. Measuring the pulse signal can be considered as a least-squares optimization problem, given knowledge of the color direction of the pulse. Because skin is actually a person-dependent filter that alters the spectrum of the light source, a short-term averaging operation is used to normalize the color signal before the least-squares optimization to compensate for the skin effect.

[0090] As another example, the vital signs detector 217 can determine the respiration signal by cross-correlating a one-dimensional representation of an image (or a selected region of interest on that image) with that of a previous image. This one-dimensional representation can be any function that captures the texture within the image or region of interest (e.g., by calculating the average over all pixel values ​​in a row). The maximum correlation is the derivative of chest position, which, when integrated, forms a measure of chest expansion during respiration (the respiration signal). It also includes all derivatives of pulse and respiration (e.g., beat-to-beat interval, respiratory effort, etc.).

[0091] As another example, the vital signs detector 217 can determine pulse transit time based on pulse signals extracted from the skin at two different body locations (or the skin at one location and an ECG), which can be used as a surrogate for blood pressure (see, e.g., https: / / pubmed.ncbi.nlm.nih.gov / 28324936 / ).

[0092] Different algorithms and functions for determining the gain based on the reference color channel distribution characteristics, the characteristics of the set of color channel distributions, and the luminance characteristics may be used in different embodiments.

[0093] As described, in many embodiments, gain settings can be determined primarily by considering color channel intensity values ​​and their distribution in an individual color channel relative to a reference intensity value for that color channel. The color channel intensity values ​​can be represented by the pixel values ​​of the color channel, e.g., each of the RGB values.

[0094] In some embodiments, the gain can be determined as a function of each value. For example, the gain processor 213 can have a look-up table (LUT) that has as inputs the luminance values ​​determined for the second image region, predetermined statistical characteristics for each color channel describing each color channel distribution (e.g., 90th percentile pixel values), and a reference intensity value for each color channel. The LUT can provide stored values ​​for a set of gain values ​​for the color channels based on these input values. In this way, the LUT can directly provide the appropriate gain setting.

[0095] Such an LUT can be determined, for example, during the design phase, by running multiple tests to determine gain values ​​that provide efficient vital sign detection for different parameters, settings, and scenarios. The LUT can be populated with the gain values ​​thus determined. In other examples, the gain values ​​can be determined based on, for example, partial measurements and / or analytical evaluations.

[0096] In some embodiments, such an approach can be used to determine a functional relationship between each parameter and an appropriate gain value instead of inputting it into a LUT, and the gain processor 213 can be configured to apply such a function to determine the appropriate gain.

[0097] In many embodiments, the gain processor 213 can be configured to determine individual gain values ​​for each of the color channels based on a reference intensity value and the determined intensity value for that color channel. The gain processor 213 can determine the gain for a first color channel depending on the difference between the determined intensity value for the first color channel and the reference intensity value for the first color channel. For example, as described above, in some embodiments, the gain can be determined to reflect the difference between the determined intensity value and the reference intensity value for a given color channel. Specifically, the gain for a given color channel can be determined as a ratio between the reference intensity value and the determined intensity value. For example, the reference intensity value may indicate a preferred 90th percentile color channel value, and the determined intensity value may indicate a measured 90th percentile color channel value for a first image region. By determining the gain as a ratio between these values, the capture operation is adapted so that the measured 90th percentile value for the image region is set to the preferred value.

[0098] Furthermore, when this approach is applied to all color channels individually, all color channels will have the desired scaling. Furthermore, when this approach is applied to all color channels, it is achieved that the relationship, specifically the ratio, between the determined intensity values ​​for each color channel matches the relationship, specifically the ratio, between the reference intensity values ​​for each color channel. This can provide particularly advantageous operation in many scenarios, typically allowing for the emphasis of a particular color channel, or indeed for the white balance of an image to be set as desired (as natural as possible).

[0099] Indeed, in many embodiments, the gains can generally be generated to maintain the same color channel relationships as the reference intensity values, without necessarily setting each individual color channel to a reference intensity value. For example, based on the luminance characteristics of the second image region, a common gain component of all gains can be varied to ensure that the luminance characteristics remain within a given range. As a result, although the individual color channel characteristics differ from the reference values, the ratios / relationships between the color channels can be maintained.

[0100] In many embodiments, individual gain components are determined for each color channel based on the determined intensity values ​​and reference intensity values, and a common gain component for all color channels is determined based on a luminance characteristic. For example, gain values ​​for each color channel can be determined as described above based on the reference intensity values ​​and the determined intensity values ​​for each color channel value. The gain processor 213 can then determine, for these gains, gain-compensated luminance characteristics corresponding to measured luminance values ​​for the second image region scaled in proportion to the determined gains. If the resulting gain-compensated luminance values ​​are within a predetermined range, the gain values ​​can be maintained; otherwise, a scale factor common to all color channels can be applied to the determined gains so that the resulting gain-compensated luminance values ​​fall within the predetermined range.

[0101] In some embodiments, the gain processor 213 is configured to control the exposure characteristics of the optical sensor depending on the luminance characteristics and / or depending on a comparison between the reference color channel distribution characteristics and the characteristics of the set of color channel distributions. The exposure characteristics can be, specifically, shutter speed / duration and / or aperture. Thus, in some embodiments, the gain controller 215 can also be configured to control the capture operation of the optical sensor. These exposure characteristics can control how much light is captured for each image. For example, increasing the shutter time and / or aperture allows more light to reach the sensor, resulting in a larger optical sensor signal and a correspondingly higher pixel value.

[0102] In fact, the effect of changing the exposure control typically has the same effect as changing the gain / amplification applied to the optical sensor signal, and indeed can be modeled as such. Accordingly, the signal path of a capture operation can be described by the example of Figure 3, where gain 201 represents the effect of changing the exposure parameters of sensor 201. Gain 205 represents a main gain common to all color channels, while gain 207 represents individual gains for each color channel.

[0103] Thus, for a given set of determined gains, these can be separated into a common gain and individual gains for each color channel. For example, the common gain can be applied to all optical sensor signals and set to, for example, a determined gain value for one of the color channels. The individual gain for that color channel can then be set to 1, and the individual gains for the other color channels can be set to a ratio between the determined gain for that color channel and the common gain.

[0104] The common gain can be achieved by controlling the exposure and setting the main gain. The exact approach can depend on the specific preferences and requirements of each individual embodiment. For example, in some embodiments, it may be desirable to set the exposure characteristics to maximize light capture (high numerical aperture, long exposure) to maximize the signal-to-noise ratio. The remaining common gain is achieved by the main gain. In other examples, it may be desirable to set the shutter speed and / or aperture to a desired value (or within a range) to provide a sharp image. The main gain can then be set again to provide the desired overall common gain.

[0105] As a specific example, the following approach can be used to implement the determined gain: 1. The minimum gain can be achieved using the exposure gain and the master gain, i.e. the common gain is set to the determined minimum color channel gain. 2. The gain controller 215 can first maximize the exposure, i.e., set the shutter speed and / or aperture to their maximum values ​​to improve the signal-to-noise ratio. The remaining contribution to the common gain is then realized by appropriately setting the master gain. In many embodiments, the master gain may only be set with a fairly coarse granularity, so the nearest quantization step can be set. To further fine-tune the capture operation, the following steps can be performed: 3. The gain controller 215 can determine a corresponding total gain for a given color channel for the determined exposure and master gain setting. This can be done specifically for a single color channel, typically the green color channel, because it has the highest contribution to luminance and is most important for blood volume pulse measurement. Because the master gain is generally coarse, the resulting multiplication factor generally does not match the desired multiplication factor for the green color channel.

[0106] 4. Fine-tuning of the exposure characteristics can be performed to fine-tune the gain of the green color channel, thus adjusting the shutter speed and / or aperture to more precisely achieve the overall desired gain for the green color channel.

[0107] 5. The individual color channel gains for the red and blue color channels can then be determined for the given exposure and master gain setting determined for the green color channel.

[0108] The above approach can be applied to all images in many embodiments. However, in many embodiments, the camera device can be configured to distinguish between images in which a skin region is detected or can be detected and images in which a skin region is not detected or cannot be detected. Specifically, in embodiments that rely on face detection to determine the first image region, the system can be configured to distinguish between images in which face detection is successful and images in which face detection is not successful.

[0109] In particular, in such embodiments, the gain processor 213 can be configured to determine color channel distributions for images in which skin regions are detected and store these for use in determining gains for images in which skin regions are not detected. The determined and stored color channel distributions are typically for regions that are different from, and specifically, are typically larger than, the first and second regions. In many embodiments, the gain processor 213 determines color channel distributions for the entire image in which skin regions are detected and stores the color channel distribution for the entire image.

[0110] When an image in which no skin regions are detected is being processed, the gain processor 213 may proceed to retrieve the stored color channel distributions and further determine the color channel distributions for the current image (also referred to as the no-skin image for brevity). The color channel distributions for the no-skin image (also referred to as the current color channel distributions for brevity) may be determined for the same image region for which the stored color channel distributions are determined, typically for the entire image, i.e., for the image region for which the current color channel distributions may be determined for the entire image region.

[0111] As a low-complexity example, if the comparison indicates that the stored and current color channel distributions are sufficiently similar (according to any suitable difference / similarity criterion), the gain processor 213 can proceed to directly reuse the gains determined for the image of the stored color channel distributions (and thus these gains can be stored along with the color channel distributions). If the distributions are not sufficiently similar, the gain processor 213 can, for example, proceed to apply nominal or default gain values ​​instead.

[0112] In many embodiments, the gain processor 213 can be configured to modify the gain determined for the stored color channel distribution to reflect the difference between this distribution and the current color channel distribution, the compensation being specifically such that the difference between the resulting color channel distributions is reduced.

[0113] For example, the gain processor 213 can determine a given (e.g., 90th) percentile intensity / pixel value for each color channel of both distributions. A compensation factor for each color channel can then be determined as the ratio between the 90th percentile value of the stored color channel distribution and the 90th percentile value of the current color channel distribution. The gain for each color channel can then be determined as the stored gain of the stored color channel distribution multiplied by the correction factor. This can result in gain settings that achieve the same 90th percentile value that are suitable for many practical scenarios where the changes in the capture environment and the time interval over which face detection occurs are relatively small.

[0114] As a specific example, in some embodiments, the gain processor 213 can be configured to determine a whole image percentile intensity value for the whole image, color channels, and percentile (0 to 100 percent for each color channel). These values ​​can be stored (i.e., the color channel distributions given by the percentile values ​​are stored). If a situation then arises in which no skin region is detected, the gain processor 213 can search the stored color channel distributions of the last image in which a skin region was detected. It can then determine the percentile for which all percentiles are equal to or less than the highest target percentile for the first image region. This percentile can then be used as the new target value for the channel with the highest intensity value for the given percentile.

[0115] Such an approach allows, for example, if a skin area is not detected, the capture operation is not altered unless there is a significant illumination change, in which case the gain processor 213 attempts to recreate the conditions immediately prior to the failure to detect the skin area.

[0116] Of particular concern in some embodiments may be the start-up process and how to set initial capture characteristics, particularly setting initial gain values ​​such that a good image is captured and, in many embodiments, accurate skin / face detection can be achieved.

[0117] In some embodiments, the gain processor 213 can be configured to determine the initial gain based on consideration of the characteristics of the entire image, without or requiring skin or face detection. The gain processor 213 can analyze the entire image to determine the intensity value / characteristic. For example, the gain processor 213 can determine a predetermined percentile intensity value for each color channel. For example, the 95th percentile pixel / intensity value of the entire image can be determined for each color channel. The determined pixel / intensity value can then be compared to a predetermined reference intensity value that is set to reflect a preferred intensity value, particularly for the 95th percentile. The predetermined reference intensity value can represent, for example, a 95th percentile value that, through appropriate testing and analysis, has been found to typically provide appropriate exposure settings that result in images that are particularly suitable for skin region and / or face detection. In some embodiments, a common initial gain for all color channels can be determined, particularly set to the ratio between the predetermined reference intensity value and the determined percentile intensity value.

[0118] In some embodiments, percentile intensity values ​​can be determined for only one of the color channels. For example, one color channel may be more important for face detection, and exposure can be optimized for that color channel. However, in many embodiments, values ​​are determined for all color channels, and the percentile intensity value that is compared to a predetermined reference intensity value can be selected as the highest value for each of the color channels. Thus, the initial gain can be adjusted so that the predetermined percentile value for the most intense color channel is set to a preferred level. This allows for capturing images suitable for skin / face detection in many scenarios.

[0119] As a specific example, the gain processor 213 can perform an initial image-wide exposure control to enable initial face detection and tracking. This initial camera control can control the maximum measured image-wide nth percentile towards a (configurable) target value.

[0120] 4 is a block diagram illustrating an exemplary processor 400 according to an embodiment of the present disclosure. Processor 400 can be used to implement one or more processors that implement devices or elements thereof as previously described. Processor 400 can be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) programmed to form a processor, a graphical processing unit (GPU), an application specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.

[0121] The processor 400 may include one or more cores 402. The cores 402 may include one or more arithmetic logic units (ALUs) 404. In some embodiments, the cores 402 may include a floating point logic unit (FPLU) 406 and / or a digital signal processing unit (DSPU) 408 in addition to or instead of the ALUs 404.

[0122] The processor 400 may include one or more registers 412 communicatively coupled to the cores 402. The registers 412 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 412 may be implemented using static memory. The registers may provide data, instructions, and addresses to the cores 402.

[0123] In some embodiments, processor 400 may include one or more levels of cache memory 410 communicatively coupled to cores 402. Cache memory 410 may provide computer-readable instructions to cores 402 for execution. Cache memory 410 may provide data for processing by cores 402. In some embodiments, computer-readable instructions may be provided to cache memory 410 by local memory, e.g., local memory connected to external bus 416. Cache memory 410 may be implemented with any suitable cache memory type, such as, for example, static random access memory (SRAM), metal-oxide-semiconductor (MOS) memory, such as dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0124] The processor 400 may include a controller 414 that may control input to the processor 400 from other processors and / or components included in the system and / or output from the processor 400 to other processors and / or components included in the system. The controller 414 may control data paths within the ALU 404, the FPLU 406, and / or the DSPU 408. The controller 414 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 414 may be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0125] The registers 412 and the cache 410 may communicate with the controller 414 and the cores 402 via internal connections 420A, 420B, 420C, and 420D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.

[0126] Input and output of processor 400 may be provided via bus 416, which may include one or more conductive lines. Bus 416 may be communicatively coupled to one or more components of processor 400, such as controller 414, cache 410, and / or registers 412. Bus 416 may be coupled to one or more components of the system.

[0127] The bus 416 may be coupled to one or more external memories. The external memory may include a read-only memory (ROM) 432. The ROM 432 may be masked ROM, an electrically programmable read-only memory (EPROM), or other suitable technology. The external memory may include a random access memory (RAM) 433. The RAM 433 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or other suitable technology. The external memory may include an electrically erasable programmable read-only memory (EEPROM) 435. The external memory may include a flash memory 434. The external memory may include a magnetic storage device such as a disk 436. In some embodiments, an external memory may be included in the system.

[0128] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units or as part of other functional units. As such, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits or processors.

[0129] In this application, any reference to any of the terms "responsive to," "based on," "dependent on," or "as a function of" should be considered a reference to the term "responsive to / based on / dependent on / as a function of." Any term should be considered to disclose any of the other terms, and use of only a single term should be considered shorthand for including other options / terms.

[0130] While the present invention has been described in connection with several embodiments, it is not intended that the present invention be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0131] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, they may be advantageously combined, and their inclusion in different claims does not imply that a combination of features is infeasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature may be applied to other claim categories as well, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must be performed, and in particular the order of individual steps in method claims does not imply that the steps must be performed in that order. Rather, steps may be performed in any suitable order. Furthermore, the singular does not exclude the plural. Thus, reference to "a," "an," "first," "second," etc. does not exclude a plural. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.

Claims

1. 1. An apparatus for estimating vital signs of a person from video images, comprising: a receiver configured to receive a video image including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; a first detector configured to determine a first image region in at least one of the video images, the first image region corresponding to a skin region of the person; and a second detector configured to determine a second image region in at least one of the video images, the second image region corresponding to a different region of the person than the first image region; and a determiner configured to determine a set of color channel distributions for pixels of the first image region; a processor configured to provide a skin tone index, the skin tone index indicative of a skin tone of the person; a reference processor configured to provide a reference color channel distribution characteristic for the skin tone; a gain processor configured to determine a gain for the optical sensor signal in dependence on the determined luminance characteristic for the second image region and in dependence on a comparison of the reference color channel distribution characteristic with characteristics of the set of color channel distributions; a gain controller configured to control a gain value applied to the optical sensor signal to match the determined gain; a vital sign determiner configured to estimate a vital sign characteristic of the person from the video image; A device having:

2. a gain circuit configured to apply the gain value to optical sensor signals from at least two color channel optical sensors; an image generator coupled to the gain circuit and configured to generate the video image from the optical sensor signal received from the gain circuit; 10. The apparatus of claim 1, comprising:

3. 3. The apparatus of claim 1, wherein the reference color channel distribution characteristic comprises a set of reference intensity values ​​for different color channels, and the characteristic of the color channel distribution comprises a set of determined intensity values ​​for the set of color channel distributions.

4. 4. The apparatus of claim 3, wherein the determined intensity values ​​are color channel pixel values ​​for a predetermined percentile of the set of color channel distributions.

5. 5. The apparatus of claim 3, wherein the gain processor is configured to determine a gain for a first color channel of the plurality of color channels in dependence on a difference between an intensity value determined for the first color channel and a reference intensity value for the first color channel.

6. 6. The apparatus of claim 3, wherein the gain processor is configured to set the gain so that a relationship between determined intensity values ​​for each color channel matches a relationship between reference intensity values ​​for the each color channel.

7. 6. The apparatus of claim 3, wherein the luminance characteristic is a luminance for the second image region resulting from applying the determined gain to the second image region, and the gain processor is configured to set the gain on the condition that the luminance characteristic indicates a luminance for the second image region that satisfies a criterion.

8. 8. The apparatus of claim 1, wherein the gain processor is configured to determine a common gain component depending on the luminance characteristic and to determine relative gain components depending on a comparison between the reference color channel distribution characteristic and a characteristic of the set of color channel distributions, the common gain component being a gain applied to all color channels of the plurality of color channels and the relative gain components being gain components applied to individual color channels of the plurality of color channels.

9. 9. The apparatus of claim 1, wherein the gain processor is configured to control exposure characteristics of at least one color channel optical sensor depending on at least one of the luminance characteristics and a comparison of the reference color channel distribution characteristics with characteristics of the set of color channel distributions.

10. 10. The apparatus of claim 1, wherein the gain processor is configured to determine and store a first color channel distribution for a third image region of an image in which a skin region is detected, and to determine a gain of the optical sensor signal for a first image in which a skin region is not detected depending on the first color channel distribution and a second color channel distribution determined for the third image region of the first image.

11. The apparatus of claim 10 , wherein the gain processor is configured to set the gain to reduce a difference between a characteristic of the first color channel distribution and a characteristic of the second color channel distribution.

12. 12. The apparatus of claim 10 or 11, wherein the third image area is the entire image area.

13. 14. The apparatus of claim 1, wherein the gain processor is configured to determine an initial gain for the optical sensor signal dependent on a comparison of a total image intensity value with a predetermined reference intensity value, the total image intensity value being a predetermined percentile intensity value of a color channel of the plurality of color channels.

14. 1. A method for estimating a person's vital signs from an image, comprising: receiving a video image including a plurality of color channels, each color channel representing a digitization of an optical sensor signal from a color channel optical sensor; determining a first image region in at least one of the video images, the first image region corresponding to a skin region of the person; determining a second image region in at least one of the video images, the second image region corresponding to a different region of the person than the first image region; determining a set of color channel distributions for pixels of the first image region; providing a skin tone index, the skin tone index indicative of the person's skin tone; providing a reference color channel distribution characteristic for the skin tone; determining a gain for the optical sensor signal in dependence on the determined luminance characteristic for the second image region and in dependence on a comparison of the reference color channel distribution characteristic with characteristics of the set of color channel distributions; controlling a gain value applied to the optical sensor signal to match the determined gain; estimating vital sign characteristics of the person from the video images; A method having the following.

15. A computer program product which, when run on a computer, causes the computer to carry out the method of claim 14.