Systems and methods for adjusting automatic image capture settings using a saliency-based region of interest

By generating saliency heat maps and tracking ROIs, the system addresses the challenge of exposing both shadows and highlights in images, enhancing image quality and reducing resource usage in image capture systems.

WO2026019420A1PCT designated stage Publication Date: 2026-01-22GOOGLE LLC

Patent Information

Application Number
PCT/US2024/038195
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing image capture systems struggle to properly expose both shadows and highlights in images, particularly when the object of interest is in a dark foreground against a bright background, as they often prioritize the background over the foreground in determining exposure settings.

Method used

The system generates saliency heat maps to detect a salient object and tracks a region of interest (ROI) associated with this object, reducing the frequency of heat map generation for subsequent frames and adjusting camera settings based on the tracked ROI to prioritize exposure on the object of interest.

Benefits of technology

This approach reduces computational resources and improves image quality by maintaining focus on the salient object, ensuring proper exposure and reducing power consumption, especially beneficial for resource-constrained devices like mobile cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038195_22012026_PF_FP_ABST
    Figure US2024038195_22012026_PF_FP_ABST
Patent Text Reader

Abstract

An example method includes generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The method also includes detecting, based on the one or more saliency heat maps, a salient object in the scene. The method additionally includes responsive to the detecting of the salient obj ect: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The method also includes adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR ADJUSTING AUTOMATIC IMAGE CAPTURE SETTINGS USING A SALIENCY-BASED REGION OF INTERESTBACKGROUND

[0001] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices, such as still and / or video cameras. The image capture devices can capture images, such as images that include people, animals, landscapes, and / or objects. Such objects may appear at different depths in the image.SUMMARY

[0002] This application generally relates to detecting salient regions of interest in videos, and automatically adjusting camera settings (e.g., for auto-focus, auto-exposure, brightness control, automatic white balance, etc.) based on such detection. Existing algorithms for automatic adjustment of camera settings generally prioritize a central portion of a frame. Also, for example, in high dynamic range (HDR) scenes, a foreground or a background is generally prioritized when determining an exposure setting for a camera. However, it may be challenging to properly expose both shadows and highlights in an image. For example, an object of interest (e.g., a person intended as the subject of a scene) may be in a dark foreground against a bright background. In the event such an object of interest is not prioritized, the background may dominate the exposure decision, resulting in an under-exposure of the object of interest. This may be especially problematic when the auto-focus algorithm focuses on the salient object (e.g., the object of interest), but auto-exposure does not meter based upon the salient object.

[0003] In one aspect, a computer-implemented method is provided. The method includes generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The method also includes detecting, based on the one or more saliency heat maps, a salient object in the scene. The method additionally includes responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The method also includes adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0004] In another aspect, a system is provided. The system may include one or more processors. The system may also include data storage, where the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the system to carry out operations. The operations may include generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The operations may also include detecting, based on the one or more saliency heat maps, a salient object in the scene. The operations may additionally include responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The operations may also include adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0005] In another aspect, a computing device is provided. The device includes a primary camera and a secondary camera that share a common field of view. The device also includes one or more processors and data storage that has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out operations. The operations may include generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The operations may also include detecting, based on the one or more saliency heat maps, a salient object in the scene. The operations may additionally include responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The operations may also include adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0006] In another aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having stored thereon program instructions that, upon execution by one or more processors of a computing device, cause the computing device to carry out operations. The operations may include generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The operations may also include detecting, based on the one or more saliency heat maps, a salient object in the scene. The operations may additionally include responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of thescene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The operations may also include adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0007] In another aspect, a program is provided. The program upon execution by one or more processors of a computing device, causes the computing device to carry out operations. The operations may include generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The operations may also include detecting, based on the one or more saliency heat maps, a salient object in the scene. The operations may additionally include responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The operations may also include adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device

[0008] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description and the accompanying drawings.BRIEF DESCRIPTION OF THE FIGURES

[0009] Figure 1 is an illustration of front, right-side, and rear views of a digital camera device, in accordance with example embodiments.

[0010] Figure 2 is an example illustration of automatic exposure (AE) settings without an ROI- based metering, in accordance with example embodiments.

[0011] Figure 3 is an example illustration of center-weighted and touch-weighted AE settings, in accordance with example embodiments.

[0012] Figure 4A is an example illustration of saliency heat map generation and AE settings, in accordance with example embodiments.

[0013] Figure 4B is an example illustration of AE settings based on the saliency heat map, in accordance with example embodiments.

[0014] Figure 5 is an example flowchart for ROI tracking, in accordance with example embodiments.

[0015] Figure 6 is another example flowchart for ROI tracking, in accordance with example embodiments.

[0016] Figure 7 is an example flowchart for luma determination based on ROI tracking, in accordance with example embodiments.

[0017] Figure 8 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments.

[0018] Figure 9 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments.

[0019] Figure 10 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments.

[0020] Figure 11 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments.

[0021] Figure 12 is a block diagram of an example computing device, in accordance with example embodiments.

[0022] Figure 13 is a flowchart of a method, in accordance with example embodiments.DETAILED DESCRIPTION

[0023] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as being an “example” or “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or features. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein.

[0024] Thus, the example embodiments described herein are not meant to be limiting. Aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are contemplated herein.

[0025] Further, unless context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be generally viewed as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.Overview

[0026] Heat maps are typically used to detect salient regions of interest. However, computing a heat map for each frame of a video may consume a considerable amount of compute resources. Also, for example, the heat map for a particular frame provides information about saliency in that frame, and does not carry over saliency based information from prior images.Accordingly, based on a frame-by-frame heat map approach, an object that is salient in one frame, may be determined to be non-salient in another frame, potentially resulting in a degraded image quality.

[0027] As described herein, a heat map may be used to determine when a region of interest (ROI) is significantly salient (e.g., where a saliency score based on the heat map exceeds a threshold saliency score, or a depth of an object in a scene is less than a depth threshold, etc.), the ROI may be tracked from frame-to-frame rather than re-computing the saliency heat map for each frame. Generally, due to their smaller size, tracking the ROI over successive frames involves fewer computational resources as compared to determining a heat map for the successive frames. The ROIs that are significantly salient in a frame may be indicated by a bounding contour, also referred to herein as a bounding box. For example, one or more feature points may be used to determine a contour of a bounding box for the ROI.

[0028] The term “region of interest” or “ROI” as used herein, generally refers to a bounding contour or a bounding box that indicates a portion of an image that includes an object. For example, the object may be a human face, and the ROI may be a circle around the human face. As another example, the object may be a flower pot and the ROI may be a bounding box that includes the flower pot. In some embodiments, an ROI may be indicated by a user. For example, the user may interact with a preview of a scene and select an object in a multimodal manner, such as by typing an identity of the object, giving voice instructions to the camera device, touching the object in the preview, hovering over the object in the preview, and so forth. In some embodiments, the object may be detected automatically (e.g., an object detection algorithm, a face detection algorithm, a segmentation algorithm, and so forth), and an ROI may be automatically generated for the detected object. In some implementations, a user-approved facial recognition algorithm may be applied to identify one or more individuals and generate ROIs for such detected individuals.

[0029] In addition to foreground objects, the techniques described herein may enable detection and / or tracking of various objects. When combined with HDR+ quality processing, objects such as rain drops, individual flower petals, grains of pollen, small living objects, such as plants, pets, insects, human eyes, animal eyes, feathers, mushrooms, and so forth. The techniques described herein also enable improved images for unique textures such as jeans, leather, cotton, any kind of fabric, stone, brick, rough surfaces, smooth surfaces, rust, paint, tissue fabric, mouse pads, tin foil, ice cubes, foam, and bubbles. Also, for example, images of natural subjects, such as fruits, vegetables, water droplets, trees, moss, grass, snowflakes, sea shells, seeds, and so forth can be enhanced. Also, for example, any object that may have adistinct appearance, and / or reveal new image information when viewed up close may be brought into sharp focus. Such objects may include, for example, coins, crayon tips, pencil tips, matches, needles, Q Tips, musical instruments, handwriting, paper, fingerprints, buttons, jewelry, floor tiles, and so forth.

[0030] Upon a successful detection of an ROI that is significantly salient, a tracking feature may be enabled that initiates tracking of the ROI from frame to frame. For example, a bounding box generated in a first frame (e.g., upon detecting an object), may be tracked over successive frames. In some embodiments, temporal feature matching and tracking may be performed. Generally speaking, the same landmark or ROI can appear in a plurality of frames (e.g., three or more frames), resulting in feature tracks. For example, temporal feature matching and tracking may be performed where a first feature in a first frame at time t — 1 is matched to a second feature in a second frame at time t. In some embodiments, tracking may involve intraframe tracking of features. In some embodiments, temporal feature tracking along with gyroscopic measurements may be used to triangulate and select the “up-to-scale” 3D points as good inlier landmarks. Subsequently, such inlier landmarks may be used to estimate per-frame scene depths.

[0031] Inter-frame tracking may involve several approaches, including detecting motion, optical flows, and so forth. Motion detection may be performed based on an analysis of an environmental brightness (e.g., based on data from an image sensor, data from an optical sensor, and so forth), an analysis of luminance across frames, by using motion vectors, and so forth. For example, a motion vector between a previous frame and a current frame may be determined to obtain an approximate model for frame by frame movement for various objects. Also, for example, various feature extraction and matching approaches may be used, such as, for example, optical flow estimation. For example, existing feature matching methods, such as ArCore features, or ILK features, may be used. The term “ILK” as used herein, generally refers to an inverse search version of the Lucas-Kanade algorithm for optical flow estimation.

[0032] Also, for example, a frequency at which a heat map engine is run to determine the heat map may be decreased. One or more auto adjustment settings of the camera (e.g. , auto-exposure (AE) algorithm, auto-focus (AF) algorithm, auto- white-balance (AWB) algorithm) may be configured to meter the frames based on the tracked ROI. For example, the one or more auto adjustment settings may be configured to associate a higher weight with a portion of the scene corresponding to the tracked ROI. This will enable the one or more auto adjustment settings of the camera to be correctly adjusted to prioritize the tracked ROI.

[0033] Generally, the heat map or saliency engine may continue to generate heat maps at a reduced frequency (e.g., every second, instead of every frame), and these heat maps may be analyzed for changes to the ROI, and / or an appearance of a new ROI. In some embodiments, the heat map engine may run at an initial rate of five (5) frames per second (FPS), and upon detection of a significant ROI, the ROI-tracker may be enabled, and the heat map engine may be configured to run at a reduced rate of one (1) FPS. Saliency can be an expensive algorithm. For example, generating saliency heat maps at 5 FPS may consume approximately 17 milli Watt (mW) of power.

[0034] Use of an ROI tracking based approach and a frame-by-frame heat map approach may be combined by using dynamic comparisons of saliency heat maps and ROI-based saliency. For example, a salient object may be asymmetrically shaped in a way that it appears significantly salient when viewed from one angle, but less salient when viewed from another angle. One example may be a rubber duck, which would appear salient when viewed sideways, but not so salient when it is rotated so that its rear is facing the camera. A frame to frame ROI tracking would maintain focus on the rubber duck even after it is rotated to a state where it no longer appears salient from the viewpoint of the camera. However, in the event a heat map were to be generated for each frame to detect saliency, such a heat map would not indicate saliency when the rubber duck is rotated so that a rear side faces the camera. Accordingly, a saliency detector would be able to detect the rubber duck as a salient object when viewed sideways, but not when viewed from the rear. In some embodiments, an object that is no longer salient based on the saliency heat map may continue to be an object of interest. In some embodiments, an object that is no longer salient based on the saliency heat map may no longer be an object of interest. As described herein, such determinations may be made by generating an ROI-based saliency and comparing it with a saliency heat map based saliency.

[0035] Another advantage of an ROI tracking based approach over a frame-by-frame heat map approach may be that a frequency of applying the heat map determination algorithm may be reduced, resulting in fewer computational resources (e.g., power, memory, processor time, etc.) being utilized. This can be especially significant for camera applications on mobile devices where resources are generally constrained and have to be distributed and prioritized over a large number of applications and services.Example Camera Systems

[0036] As image capture devices, such as cameras, become more popular, they may be employed as standalone hardware devices or integrated into various other types of devices. For instance, still and video cameras are now regularly included in wireless computing devices(e.g., mobile devices, such as mobile phones), tablet computers, laptop computers, video game interfaces, home automation devices, and even automobiles and other types of vehicles.

[0037] The physical components of a camera may include one or more apertures through which light enters, one or more recording surfaces for capturing the images represented by the light, and lenses positioned in front of each aperture to focus at least part of the image on the recording surface(s). The apertures may be of a fixed size or may be adjustable. In an analog camera, the recording surface may be a photographic film. In a digital camera, the recording surface may include an electronic image sensor (e.g., a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor) to transfer and / or store captured images in a data storage unit (e.g., memory).

[0038] One or more shutters may be coupled to, or positioned near, the lenses or the recording surfaces. Each shutter may either be in a closed position, in which it blocks light from reaching the recording surface, or an open position, in which light is allowed to reach the recording surface. The position of each shutter may be controlled by a shutter button. For instance, a shutter may be in the closed position by default. When the shutter button is triggered (e.g., pressed), the shutter may change from the closed position to the open position for a period of time, known as the shutter cycle. During the shutter cycle, an image may be captured on the recording surface. At the end of the shutter cycle, the shutter may change back to the closed position.

[0039] Alternatively, the shuttering process may be electronic. For example, before an electronic shutter of a CCD image sensor is “opened,” the sensor may be reset to remove any residual signal in its photodiodes. While the electronic shutter remains open, the photodiodes may accumulate charge. When or after the shutter closes, these charges may be transferred to longer-term data storage. Combinations of mechanical and electronic shuttering may also be possible.

[0040] Regardless of type, a shutter may be activated and / or controlled by something other than a shutter button. For instance, the shutter may be activated by a softkey, a timer, or some other trigger. Herein, the term “capture” may refer to any mechanical and / or electronic shuttering process that results in one or more images being recorded, regardless of how the shuttering process is triggered or controlled.

[0041] The exposure of a captured image may be determined by a combination of the size of the aperture, the brightness of the light entering the aperture, and the length of the shutter cycle (also referred to as the shutter length, the exposure length, or the exposure time). Additionally, a digital and / or analog gain (e.g., based on an ISO setting) may be applied to the image, therebyinfluencing the exposure. In some embodiments, the term “exposure length,” “exposure time,” or “exposure time interval” may refer to the shutter length multiplied by the gain for a particular aperture size. Thus, these terms may be used somewhat interchangeably, and should be interpreted as possibly being a shutter length, an exposure time, and / or any other metric that controls the amount of signal response that results from light reaching the recording surface.

[0042] In some implementations or modes of operation, a camera may capture one or more still images each time image capture is triggered. In other implementations or modes of operation, a camera may capture a video image by continuously capturing images at a particular rate (e.g., 24 frames per second) as long as image capture remains triggered (e.g., while the shutter button is held down). Some cameras, when operating in a mode to capture a still image, may open the shutter when the camera device or application is activated, and the shutter may remain in this position until the camera device or application is deactivated. While the shutter is open, the camera device or application may capture and display a representation of a scene on a viewfinder (sometimes referred to as displaying a “preview frame”). When image capture is triggered, one or more distinct payload images of the current scene may be captured.

[0043] Cameras, including digital and analog cameras, may include software to control one or more camera functions and / or settings, such as aperture size, exposure time, gain, and so on. Additionally, some cameras may include software that digitally processes images during or after image capture. While the description above refers to cameras in general, it may be particularly relevant to digital cameras. Digital cameras may be standalone devices (e.g., a DSLR camera) or may be integrated with other devices.

[0044] Either or both of a front-facing camera and a rear-facing camera may include or be associated with an ALS that may continuously or from time to time determine the ambient brightness of a scene that the camera can capture. In some devices, the ALS can be used to adjust the display brightness of a screen associated with the camera (e.g., a viewfinder). When the determined ambient brightness is high, the brightness level of the screen may be increased to make the screen easier to view. When the determined ambient brightness is low, the brightness level of the screen may be decreased, also to make the screen easier to view as well as to potentially save power. Additionally, the ambient light sensor’s input may be used to determine an exposure time of an associated camera, or to help in this determination.

[0045] Figure 1 is an illustration of front, right-side, and rear views of a digital camera device 100, in accordance with example embodiments. Digital camera device 100 may be, for example, a mobile device (e.g., a mobile phone), a tablet computer, or a wearable computing device. However, other embodiments are possible. Digital camera device 100 may includevarious elements, such as a body 102, a front-facing camera 104, a multi-element display 106, a shutter button 108, and other buttons 110. Digital camera device 100 could further include one or more rear-facing cameras 112, 114. Front-facing camera 104 may be positioned on a side of body 102 typically facing a user while in operation, or on the same side as multi-element display 106. Rear-facing cameras 112, 114 may be positioned on a side of body 102 opposite front-facing camera 104. Referring to the cameras as front-facing and rear-facing is arbitrary, and digital camera device 100 may include multiple cameras positioned on various sides of body 102.

[0046] Multi-element display 106 could represent a cathode ray tube (CRT) display, a lightemitting diode (LED) display, a liquid crystal display (LCD), a plasma display, or any other type of display known in the art. In some embodiments, multi-element display 106 may display a digital representation of the current image being captured by front-facing camera 104 and / or rear-facing cameras 112, 114, or an image that could be captured or was recently captured by either or both of these cameras. Thus, multi-element display 106 may serve as a viewfinder for either camera. Multi-element display 106 may also support touchscreen and / or presencesensitive functions that may be able to adjust the settings and / or configuration of any aspect of digital camera device 100.

[0047] Multi-element display 106 may include additional features related to a camera application. For example, multiple modes may be available for a user, including, a motion mode, portrait mode, video mode, video bokeh mode, and so forth. The camera application may be in camera mode and provide additional features, such as a reverse icon to activate reverse camera view, a trigger button to capture a previewed image, and a photo stream icon to access a database of captured images. Also for example, a magnification ratio slider may be displayed and a user can move a virtual object along the magnification ratio slider to select a magnification ratio. In some embodiments, a user may use the multi-element display 106, also referred to herein as the display screen, to adjust the magnification ratio (e.g., by moving two fingers on display screen in an outward motion away from each other), and magnification ratio slider may automatically display the magnification ratio.

[0048] Front-facing camera 104 may include an image sensor and associated optical elements such as lenses. Front-facing camera 104 may offer zoom capabilities or could have a fixed focal length. In other embodiments, interchangeable lenses could be used with front-facing camera 104. Front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. Front-facing camera 104 also could be configured to capture still images, video images, or both. Further, front-facing camera 104 could represent a monoscopic,stereoscopic, or multiscopic camera. Rear-facing cameras 112, 114 may be similarly or differently arranged. Additionally, front-facing camera 104, rear-facing cameras 112, 114, or both, may be an array of one or more cameras.

[0049] Either or both of front-facing camera 104 and rear-facing cameras 112, 114 may include or be associated with an illumination component that provides a light field to illuminate a target object. For instance, an illumination component could provide flash or constant illumination of the target object (e.g., using one or more LEDs). An illumination component could also be configured to provide a light field that includes one or more of structured light, polarized light, and light with specific spectral content. Other types of light fields known and used to recover three-dimensional (3D) models from an object are possible within the context of the embodiments herein.

[0050] In some digital camera devices 100, either or both of front-facing camera 104 and rearfacing cameras 112, 114 may include or be associated with an ambient light sensor that may continuously or from time to time determine the ambient brightness of a scene that the camera can capture. In some devices, the ambient light sensor can be used to adjust the display brightness of a screen associated with the camera (e.g., a viewfinder). When the determined ambient brightness is high, the brightness level of the screen may be increased to make the screen easier to view. When the determined ambient brightness is low, the brightness level of the screen may be decreased, also to make the screen easier to view as well as to potentially save power. Additionally, the ambient light sensor’s input may be used to determine an exposure time of an associated camera, or to help in this determination.

[0051] Digital camera device 100 could be configured to use multi-element display 106 and either front-facing camera 104 or rear-facing cameras 112, 114 to capture images of a target object (e.g., a subject within a scene). The captured images could be a plurality of still images or a video image (e.g., a series of still images captured in rapid succession with or without accompanying audio captured by a microphone). The image capture could be triggered by activating shutter button 108, pressing a softkey on multi-element display 106, or by some other mechanism. Depending upon the implementation, the images could be captured automatically at a specific time interval, for example, upon pressing shutter button 108, upon appropriate lighting conditions of the target object, upon moving digital camera device 100 a predetermined distance, or according to a predetermined capture schedule.

[0052] As noted above, the functions of digital camera device 100 (or another type of digital camera) may be integrated into a computing device, such as a wireless computing device, cell phone, tablet computer, laptop computer, and so on. For example, a camera controller may beintegrated with the digital camera device 100 to control one or more functions of the digital camera device 100.Saliency Heat Maps and ROI Tracking

[0053] Saliency heat maps may be used to generate regions of interest (ROI) in an image. Existing approaches use such identified regions of interest using a saliency heat may to adjust automatic exposure settings of a camera. For example, in the presence of bright salient objects in the image, the tone mapping strength may be lowered, and the overall image histogram may be shifted left. Adjusting auto-exposure settings and tone mappings involves trade-offs. For example, optimizing exposure on one ROI typically comes at the expense of quality for the rest of the image, particularly areas that have different luminance characteristics than the ROI. The heat map facilitates the identification of ROIs in the image, and information regarding whether such ROIs are of high enough significance to be worth emphasizing on at the expense of exposure with regards to the rest of the image.

[0054] In some embodiments, the generating of the one or more saliency heat maps is performed by a machine learning model. For example, a machine-learned technique ("Saliency Model") may be used to predict a saliency heat map for a given image. The Saliency Model may be implemented as one or more of a support vector machine (SVM), a recurrent neural network (RNN), a convolutional neural network (CNN), a dense neural network (DNN), other machine-learning techniques, and / or a combination thereof. The saliency heat map depicts the magnitude of the saliency probability on a scale (e.g., from black to white, where white indicates a high probability of saliency and black indicates a low probability of saliency). Each pixel within the saliency heat map may be assigned a saliency metric that represents how salient the region represented by the metric is based on a computing device applying a pre-trained machine learning model to the image frame, such that the pre-trained machine learning model outputs a saliency metric for each pixel within the saliency heat map.

[0055] In some embodiments, the Saliency Model may produce a bounding box enclosing the region with the greatest probability of saliency. The saliency heat map may be used to determine one or more objects of interest in an image, and generate one or more bounding boxes around the determined objects of interest. In some embodiments, anchor boxes of various aspect ratios and sizes located at various portions of the image may be used to determine an ROI. In some aspects, the ROI may be a portion of the image with a high or the highest average saliency value. For example, each pixel in the heat map may be associated with a saliency value, and an average of the saliency values of the pixels within the anchor bounding box may be used to determine the average saliency value.

[0056] In some embodiments, average saliency values of anchor bounding boxes of various sizes and aspect ratios may be assessed every few pixels. For example, the center of the anchor bounding boxes may be evenly spaced based on a stride value, and at each location, various sizes and aspect ratios of the anchor bounding boxes may be assessed to identify ROIs. For example, an anchor bounding box with a maximum average saliency value may be selected as a primary ROI. Other anchor bounding boxes that overlap with the selected anchor bounding box, and / or are less than a threshold distance away from the selected primary ROI are discarded. A secondary ROI may then be selected, and so forth.

[0057] For example, a primary ROI may be determined to be an ROI with a maximum average saliency. Next, based on the primary ROI, a secondary ROI may be determined to be an ROI with pixels of a lesser average saliency value when compared with the primary ROI. Also, for example, the secondary ROI may be required to be at least a threshold distance away from the primary ROI.

[0058] A more stable filtered ROI may be determined based on the primary ROI and the secondary ROI. The stable filtered ROI may then be potentially used as a tracking ROI. For example, a confidence value of a previous filtered ROI (associated with a previous frame) may be determined based on average saliency value in the previous filtered ROI with respect to a saliency heat map of a current frame. Likewise, the confidence value for the primary ROI may be determined based on the average saliency value in the primary ROI with respect to the saliency heat map of the current frame. In the event a saliency difference between the two confidence values exceeds a threshold, a determination may be made whether to maintain the previous filtered ROI as the stable filtered ROI for the current frame, or to update the previous filtered ROI to the primary ROI. For example, when the confidence value for the primary ROI exceeds the confidence value for the previous filtered ROI by the threshold value, the filtered ROI may be updated from the previous filtered ROI to the primary ROI. However, when the confidence value for the primary ROI does not exceed the confidence value for the previous filtered ROI by the threshold value, the previous filtered ROI may be maintained as the filtered ROI. The secondary ROI may also be used in a similar manner. Additional, and / or alternative techniques may be used to determine a filtered ROI.

[0059] In some embodiments, the computing device may select between the filtered ROI, the primary ROI, and the secondary ROI to determine an area in the image frame for automatic image capture settings. Saliency detection may be unstable over time. In order to address such potential instability, a Finite State Machine (FSM) may be used. For example, the FSM may require several consecutive frames with consistent saliency detections for auto-focus to committo a salient ROI and / or requires several consecutive frames with lack of consistent saliency detections for auto-focus to abandon the salient ROI. In this context, consistent detections refer to overlapping of detected bounding boxes between consecutive frames and / or high confidence in these detections.

[0060] An additional potential challenge is that saliency detection may report salient regions that are not in fact salient. This type of a false positive error may degrade the user experience with the camera. Imagine the camera trying to focus, for example, on a shiny object in the background of the scene. To address this challenge, the FSM helps because it takes care of instantaneous false positives: the FSM is less likely to detect consecutive false positives rather than a false positive in a single frame. In addition to this, in order to prevent false positives, the application of saliency auto-focus processes may be restricted based on global on-device signals. For example, certain requirements may be imposed (e.g., a minimum brightness value in the scene, a zoom ratio within a certain range, and / or lack of device motion) before allowing auto-focus. In addition, certain salient regions may be discarded when the estimated depth is out of range, when the estimated depth is too different from the estimated distance of a current ROI, and / or when the location of the bounding box is too far off-center.

[0061] Various techniques may be used to generate depth information for an image. In some cases, depth information may be generated for the entire image (e.g., for the entire image frame). In other cases, depth information may only be generated for a certain area or areas in an image. For instance, depth information may only be generated when image segmentation is used to identify one or more objects in an image. Depth information may be determined specifically for the identified object or objects.

[0062] In embodiments, stereo imaging may be utilized to generate a depth map. In such embodiments, a depth map may be obtained by correlating left and right stereoscopic images to match pixels between the stereoscopic images. The pixels may be matched by determining which pixels are the most similar between the left and right images. Pixels correlated between the left and right stereoscopic images may then be used to determine depth information. For example, a disparity between the location of the pixel in the left image and the location of the corresponding pixel in the right image may be used to calculate the depth information using binocular disparity techniques.

[0063] In some embodiments, the saliency heat map may be indirectly impacted by the depth of the scene. The Saliency Model is a multitask fission model that is trained to predict both a saliency heat map and a foreground / background probability heat map. Accordingly,foreground / background probability heat map influences the saliency heat map and the saliency score indirectly.

[0064] Depth information is also used later in the pipeline. For example, AF uses the depth of a saliency bounding box (along with the associated saliency score) as an additional requirement to consume and / or utilize the saliency result and initiate ROI tracking.

[0065] Figure 2 is an example illustration of automatic exposure (AE) settings without an ROI- based metering, in accordance with example embodiments. Image 200 illustrates a foreground object (located within box 205). Although the camera is focused on the foreground object, the AE settings fail to prioritize this region, resulting in underexposure of the foreground object.

[0066] Figure 3 is an example illustration of center-weighted and touch-weighted AE settings, in accordance with example embodiments. Images 305 and 315 were captured with center- weighted AE settings, and emphasize relatively uniform brightness for the entire region. Images 310 and 320 were captured with touch-weighted AE settings. This enables appropriate brightness levels for the candy jar in image 310, and the cat in image 320, when compared to the candy jar in image 305, or the cat in image 315. However, such brightness settings result in a trade-off where the rest of the image regions are not properly exposed. For example, in images 310, 320, the highlights in the background and the shadows underneath are overexposed. In image 310, for example, the candy jar being determined to be an object of interest may allow the AE settings to be configured to raise the overall exposure of all portions of image 310.

[0067] As used herein, the "shadows" or "shadow areas" of an image (or sequence of images) should be understood to include pixels or areas in an image (or across a sequence of images) that are the darkest. In practice, a pixel or area(s) in the image frame having a brightness level below a predetermined threshold may be identified as a shadow area. Of course, other methods for detecting a darker area that qualifies as a shadow area are also possible.

[0068] As used herein, the "highlights" or "highlights areas" of an image (or sequence of images) can include pixels or areas in an image (or across a sequence of images) that are the brightest. In practice, a pixel or area in the image frame having a brightness level above a predetermined threshold may be identified as a highlight area. Note that different thresholds may be utilized to classify shadow and highlight areas (such that areas can exist that are not classified as a shadow or highlight). Alternatively, the threshold for highlights and shadows may be the same (such that every area in the image is classified as a shadow or highlight). Further, the shadow and / or highlight threshold(s) might also be adaptive, and vary based on, e.g., the characteristics of the scene being captured. In some cases, the shadow and / or highlightthreshold(s) could vary spatially across the image frame. Other methods for detecting a bright area that qualifies as a highlight area are also possible.

[0069] Figure 4A is an example illustration of saliency heat map generation and AE settings, in accordance with example embodiments. Image 405 includes regions of varying brightness levels. Image 410 indicates that the mid-to-bright elements of image 405 are salient. This is indicated by a first ROI located inside a first bounding box 415 A, and a second ROI located inside a second bounding box 420A. Image 425 illustrates the corresponding portions of image 405. For example, first bounding box 415 A of image 410 corresponds to the third bounding box 415B of image 425, and second bounding box 420A of image 410 corresponds to the fourth bounding box 420B of image 425. Also, as indicated, third bounding box 415B of image 425 is associated with a saliency score of 0.21, and fourth bounding box 420B of image 425 is associated with a saliency score of 0.49.

[0070] Figure 4B is an example illustration of AE settings based on the saliency heat map, in accordance with example embodiments. For example, based on image 425 of Figure 4A, the AE settings may be configured so that the tone-mapping strength is reduced and the overall histogram of the resultant image is decreased. For example, image 435 illustrates a first histogram corresponding to image 430. A mid region of image 430 is indicated by the first portion 435A of the first histogram corresponding to image 430. As illustrated in image 445, reducing the tone-mapping strength and bringing the overall histogram of the image 430 down to generate image 440, results in isolating the mid-region in histogram space. For example, the first portion 435A of the first histogram corresponding to image 430 is further isolated in the second portion 445A of the second histogram corresponding to image 440. Quantitatively, image 435 represents a first histogram with a mean of 128.84, a standard deviation of 56.77, a median of 128, and 269,188 pixels. Image 445 represents a second histogram with a lower mean of 116.71, a higher standard deviation of 64.35, a lower median of 110, and 269,188 pixels. Accordingly, mid-regions are emphasized in image 440 at the expense of details of shadows, resulting in a perception of higher image contrast in image 440 as compared to image 430.

[0071] Figure 5 is an example flowchart 500 for ROI tracking, in accordance with example embodiments. Generally speaking, saliency heat map based ROI tracking may be integrated into existing auto-focus (AF) logic configurations. The system may check an ROI tracker to determine whether an ROI is being tracked by AF. In the event an ROI is being tracked, then it may be inferred that this is a saliency heat map based ROI, and the AE settings may be adjusted accordingly. In some embodiments, AF may be configured to use the saliency heatmap based ROI tracking described herein for salient regions in a rear-facing main camera and / or an ultrawide camera.

[0072] At step 505, a saliency engine may access data from saliency library 510 to detect a saliency heat map for an image. At step 515, a determination is made whether to use the saliency heat map.

[0073] Upon a determination that the saliency heat may not be used, the process proceeds to step 520. At step 520, existing AF logic continues to apply. Upon a determination that the saliency heat may be used, the process proceeds to step 525. At step 525, tracking is initiated.

[0074] At step 530, ROI tracking commences. For example, an ROI tracker may use tracking library 535 to track one or more ROIs in the image.

[0075] At step 540, a determination is made whether the ROI tracker may be able to continue tracking the one or more ROIs. For example, the saliency heat map may identify a region that has very low brightness (e.g., a dark region), and it may be challenging to identify a sufficient number of feature points in that region to identify a bounding box for purposes of ROI tracking. In the event the ROI tracker is unable to track, it may also be challenging for the AF logic to operate. For example, an attempt to focus on a dark region with no texture is likely to fail.

[0076] Upon a determination that the ROI tracker is no longer able to track the one or more ROIs, the process proceeds to step 545 to determine whether the one or more ROIs continue to be salient.

[0077] Some embodiments involve terminating the tracking of the ROI based on a determination that a saliency score associated with the salient object is below a saliency threshold. Such embodiments also involve increasing the heat map generation frequency for generation of saliency heat maps. For example, at step 545, a determination is made whether the one or more ROIs continue to be salient. Upon a determination that the one or more ROIs continue to be salient, the process returns to step 540 and determines whether the ROI tracker may be able to continue tracking the one or more ROIs. Upon a determination that the one or more ROIs are no longer salient, the process proceeds to step 520 where existing AF logic is applied.

[0078] Some embodiments involve adjusting, based on additional one or more saliency heat maps generated at the increased heat map generation frequency, another automatic image capture setting for the video. For example, the AF settings may be adjusted based on additional one or more saliency heat maps. Also, for example, AE settings may be adjusted based on additional one or more saliency heat maps. As another example, an auto white balance (AWB) settings may be adjusted based on additional one or more saliency heat maps.

[0079] At step 540, upon a determination that the ROI tracker is able to track the one or more ROIs, the process proceeds to step 550 to determine whether the one or more ROIs continue to be salient. The determination at step 550 is not based on the saliency heat map. Instead, the tracked ROI is overlaid over the saliency heat map saliency region, and a new saliency score may be determined. For example, one or more additional saliency heat maps may be generated at a reduced heat map generation frequency. The tracked ROI may be overlaid over saliency regions in the one or more additional saliency heat maps. Subsequently, a second saliency score associated with the tracked ROI may be determined. For example, a first tracked ROI may track a first object in a scene, and a second object may appear. Accordingly, a current heat map saliency of the scene may be different from an initial heat map saliency that triggered the first tracked ROI. The second saliency score may indicate whether the first tracked ROI continues to be relevant, or whether a second ROI associated with the second object is to be tracked. The determination, at step 550, whether the one or more ROIs continue to be salient can include a determination whether the second saliency score exceeds a second saliency threshold.

[0080] Generally, the transition from step 550 to step 545 may occur when the ROI tracker fails to initiate ROI tracking because there is not enough texture and / or features to track. This can happen, for example, if the user points the camera at a blank white wall. Another factor that can cause the transition from step 550 to step 545 is each time the ROI tracker runs to continue tracking (and update the ROI), the ROI tracker returns a confidence value along with the updated ROI. A sufficiently low confidence indicates that the tracked ROI is no longer useful for AF / AE determinations. For example, the sufficiently low confidence may indicate that the tracked ROI may not be tracking the object of interest. One example of when this can happen is when the object of interest being tracked leaves the camera field of view (FOV). In this example, the confidence is likely to be near zero (0) and the ROI tracking becomes unreliable.

[0081] Some embodiments involve, based upon a determination that the second saliency score exceeds the second saliency threshold, continuing the tracking of the tracked ROI. For example, the second saliency score for the first tracked ROI may be sufficiently high to continue tracking the first object. Some embodiments involve, based upon a determination that the second saliency score does not exceed the second saliency threshold, terminating the tracking of the tracked ROI. For example, the second saliency score for the first tracked ROI may not be sufficiently high to continue tracking the first object.

[0082] Some embodiments involve detecting, based on the one or more additional saliency heat maps, a second salient object in the scene. For example, the one or more additionalsaliency heat maps may indicate a second ROI associated with the second object. Such embodiments involve, responsive to the detecting of the second salient object, determining whether to initiate a second tracking of a second ROI associated with the second salient object. For example, it may be determined whether the second ROI associated with the second object is to be tracked. Such a determination may be based on the second saliency score associated with the tracked ROI.

[0083] Generally speaking, the saliency heat map and the ROI based on such heat map may change over time, and the actual ROI being tracked may be different. Accordingly, there is a need to review the current tracked ROI and the current saliency heat map to determine whether the tracked ROI is still salient. The configurations at step 550 are designed to achieve such continued consistency.

[0084] For example, the camera may initially capture a bush with a flower. The saliency heat map then indicates an ROI that includes a bounding box around the flower with a saliency score of 0.6. Upon a determination that the saliency score of 0.6 exceeds a threshold saliency, the ROI tracker may be triggered. As the camera continues to capture the scene with the bush and the flower, the saliency heat map and the tracked ROI are likely to be consistent, and there is likely to be a substantial overlap between the respective heat map region and the tracked ROI. As the ROI tracker is triggered, a frequency of the heat map generation may be reduced (e.g., from 5 FPS to 1 FPS).

[0085] For illustrative purposes, assume that the camera pans over the field of view and a dog enters the scene. At the next running of the saliency engine, the new saliency heat map is likely to capture the flower and the dog. However, the saliency score associated with the flower may be considerably smaller than the saliency score associated with the dog. However, the tracked ROI may continue to track the flower, and AF and AE adjustments may continue to be based on the tracked ROI for the flower. This may cause the portion of the image that includes the dog to be degraded as the dog is not the tracked ROI for which AE settings are being adjusted. Accordingly, a comparison of the tracked ROI and the new saliency heat map may indicate such discrepancy, and the AE logic may be updated.

[0086] As another example, consider a situation where the flower is no longer in the scene being captured. In this situation, at step 540, the tracker for the tracking ROI will indicate a low confidence and transition to step 545.

[0087] In some embodiments, the intended flow from 540 to 545 may fail due to an erroneous determination by the tracker. For example, at step 540, the tracker may provide a high confidence that the tracking succeeds, even though the flower is no longer in the scene.Accordingly, the process may transition from step 540 to step 550. In such an edge case, the saliency determination at step 550 can indicate that tracking ROI is no longer salient. One of the many advantages of the techniques described herein is that errors in edge cases such as the one illustrated here can be detected and corrected.

[0088] As another example, in addition to the flower no longer being present, there may be a dog in the scene. Therefore, a determination is made that the flower ROI is no longer tracked, and a further determination is made whether the tracked ROI is to be changed to track the dog. In the event the flower is no longer in the scene, a new saliency heat map is likely to indicate an absence of the flower. Also, for example, the new saliency heat map is likely to indicate a presence of the dog.

[0089] Similar to step 545, upon a determination that the one or more ROIs are no longer salient, the process proceeds to step 520 where existing AF logic is applied. In some embodiments, a new ROI tracking may be triggered.

[0090] Unlike step 545, upon a determination that the one or more ROIs continue to be salient, the process proceeds to step 555. At step 555, a determination is made to focus on the ROI being tracked and reduce a frequency of saliency heat map analysis.

[0091] In some embodiments, different thresholds may be applied to initiate ROI tracking and to continue ROI tracking. For example, a higher threshold may be used to initiate ROI tracking. Also, for example, a lower threshold may be used to continue focusing on the tracked ROI. One reason for two different values may be for hysteresis. For example, with one threshold, and with values close to the threshold, the logic may alternate frequently between continuing to focus on the tracked ROI and not tracking the ROI, resulting in a degraded user experience.

[0092] Also, for example, in a scene that initially tracks a flower as a tracked ROI, the camera may be panned to capture a dog. However, it may be desirable to continue to track the ROI associated with the flower. Accordingly, triggering a new ROI tracking based on the dog may have a higher threshold of saliency than a threshold for a continued ROI tracking based on the flower.

[0093] Figure 6 is another example flowchart 600 for ROI tracking, in accordance with example embodiments. Generally speaking, saliency heat map based ROI tracking may be associated with a lower priority when considering ROIs, with touch, face, and segmentation ROIs being associated with higher priorities. Flowchart 600 illustrates example priorities associated with various ROI policies for a camera.

[0094] At step 606, it is determined whether the ROI is a touch ROI. The term “touch ROI” as used herein generally refers to a region of interest identified by a user of the camera. Forexample, the user may touch a portion of the scene in a preview displayed and identify the ROI. Upon a determination that the ROI is a touch ROI, the process proceeds to step 610, where a camera processing logic (e.g., AF logic, AE settings, etc.) associated with touch ROI is initiated.

[0095] Upon a determination that the ROI is not a touch ROI, the process proceeds to step 615. At step 615, it is determined whether the ROI is a face ROI. The term “face ROI” as used herein generally refers to a region of interest that corresponds to a face (e.g., a human face, such as, for example, in a portrait).

[0096] Upon a determination that the ROI is a face ROI, the process proceeds to step 620. At step 620, it is determined whether a size of the face exceeds a threshold size. Upon a determination that the size of the face exceeds the threshold size, the process proceeds to step 625.

[0097] At step 625, it is determined whether the face pose is centered. Upon a determination that the face pose is centered, the process proceeds to step 630, where a camera processing logic (e.g., AF logic, AE settings, etc.) associated with face ROI is initiated.

[0098] Upon a determination that the ROI is not a face ROI, the process proceeds from step 615 to step 635. Similarly, upon a determination that the size of the face does not exceed the threshold size, the process proceeds to step 635. And likewise, upon a determination that the face pose is not centered, the process proceeds to step 635.

[0099] At step 635, it is determined whether a portion of the image includes segments corresponding to a person, and / or a skin of the person. Upon a determination that the portion of the image includes segments corresponding to a person, and / or a skin of the person, the process proceeds to step 640, where a camera processing logic (e.g., AF logic, AE settings, etc.) associated with segmentation ROI is initiated.

[0100] Upon a determination that the portion of the image does not include segments corresponding to a person, and / or a skin of the person, the process proceeds to step 645.

[0101] At step 645, it is determined whether the AF logic is based on a saliency heat map based ROI (e.g., as described with reference to Figure 5). Upon a determination that the AF logic is based on a saliency heat map based ROI, the process proceeds to step 650, where a camera processing logic (e.g., AF logic, AE settings, etc.) associated with saliency ROI is initiated (e.g., as described with reference to Figure 7).

[0102] Upon a determination that the AF logic is not based on a saliency heat map based ROI, the process proceeds to step 655. At step 655, it is determined that an ROI is not being used as part of the camera processing logic (e.g., AF logic, AE settings, etc.).

[0103] The saliency ROI may be used to determine a luma value referred to herein as ROIiuma. In turn, ROIiumamay be used to determine a Luma. The Luma may be a weighted average of a weighted average luma and the R0Iiuma. Generally speaking, the Luma may be a significant factor in determining long and short total exposure sensitivity (TES). Determining the Luma based on the saliency ROI enables the camera to be configured to assign a salient region a higher priority over the remainder of the scene when determining the AE settings. Subsequent to updating of the Luma, the Short tes and Long tes may be adjusted to bring the Luma towards Tar gettuma. The value for Targetiumamay be determined based on a scene. Scenes with different brightness levels may be associated with different values for T argetiuma. A similar approach may be applied to other ROIs to determine the AE settings. Such determinations are described with respect to Figure 7.

[0104] Figure 7 is an example flowchart 700 for luma determination based on ROI tracking, in accordance with example embodiments. Some embodiments involve generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. For example, the saliency node may operate at a first frequency of five (5) frames per second (FPS) to generate saliency heat maps.

[0105] Some embodiments involve determining a saliency score associated with the salient object. For example, at step 705, a saliency ROI box may be created with an associated saliency score (e.g., see 415b and 420B of image 425 in Figure 4A). The saliency node may generate saliency heat maps with an associated strength. In some embodiments, the determination of the saliency score may be based on a strength of the one or more saliency heat maps. In some embodiments, the determination of the saliency score is based on a depth of the salient object in the scene. For example, objects in a foreground may be associated with a higher saliency score, whereas objects in the background may be associated with a lower saliency score.

[0106] As part of the camera’s AF logic, at step 710, it may be determined whether the saliency score exceeds a threshold. In some embodiments, the threshold may be 80%. Some embodiments involve detecting, based on the one or more saliency heat maps, a salient object in the scene. For example, upon a determination that the saliency score exceeds a threshold, a salient object in the scene may be detected. In the event the saliency score does not exceed the threshold, the process proceeds to step 712. At step 712, the process does not track the saliency ROI. Generally, the camera’s AE logic can be based on determinations made by the AF logic.For example, in the event AF decides that the saliency score does not exceed the threshold, AE will not track the salient region, and AE will not have a saliency ROI box to use.

[0107] Some embodiments, involve, responsive to the detecting of the salient object, initiating a tracking of an ROI associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. In some embodiments, the initiating of the tracking of the ROI may be based on a determination that the saliency score exceeds a saliency threshold. For example, upon a determination that the saliency score exceeds a threshold, an ROI tracker logic may be triggered. For example, as part of the camera’s ROI tracker logic, at step 715, the saliency ROI box may be tracked.

[0108] Generally speaking, steps 710 and 715 are a summary of the process described with reference to Figure 5. Figure 5 is described in the context of the auto-focus (AF) mode. In some embodiments, the auto-exposure (AE) mode may consume the ROI tracking result from the AF mode instead of re-implementing the same logic..

[0109] Some embodiments involve adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device. For example, as part of a spatial ROI analysis, at step 720, the ROIiumamay be determined from the ROI box. The ROIiumais generally a center-weighted luminance determination. At step 725, a temporal filter may be applied to the ROIiuma(e.g., to maintain stability of luminance values). As indicated, the saliency heat map may not be stable during observation. During periods of high camera motion, based on gyroscope values, a lower temporal filter strength may be applied to enable faster and efficient adaptation to a changing scene. In the event of low gyro motion, it may be inferred that the scenery is stable, and a higher temporal filter strength may be applied to avoid fluctuations in AE tone mapping.

[0110] Also, for example, the ROIiumamay change from frame to frame, and this will likely affect the computation of Luma using Eqn. 3 below. This may cause larger differences in Luma and result in a degraded user experience. It may be more preferable to propagate changes gradually across frames. Accordingly, temporal filtering may be used to propagate a gradual change across frames. For example, starting with a first frame, the ROIiumafor a next frame may be determined as an average of the ROIiumaof the first frame and the ROIiumaof the next frame.

[0111] As part of a scene analysis, at step 730, an ROIweightmay be determined as:_ Touch_RO IweightU 'weight(Eqn. 1)

[0112] where the Touch_ROIweightis obtained after tuning. The factor of Yi used in Eqn. 1 may be another factor that indicates a confidence score in the Saliency ROI. The saliency score may be factored in to adjust the center-weighted luminance determination to a more ROI- weighted luminance determination. Accordingly, the saliency score may be obtained from step 705, and a modified ROI^eiglltmay be determined at step 735 as follows:R0Iweight =ROIweight* Saliency Score(Eqn. 2)

[0113] For example, when the SaliencyScore is 20%, then the ROIweightmay be multiplied by 0.2 to obtain the modified ROI^eigllt. Generally speaking, a higher saliency score results in a higher weighting. The SaliencyScore may be determined as a percentage, or a value between 0 andl.

[0114] Some embodiments involve determining an average luma value associated with the scene, wherein the average luma value causes an initial automatic image capture setting to prioritize a center of the scene. For example, Luma is a weighted average value of an entire image, without factoring in ROIs. Accordingly, in an absence of an ROI, Luma is a center- weighted luminance determination. For a given image, there may be a Targetiuma, and the Luma may be adjusted to be close to the desired T argetiuma.

[0115] Some embodiments involve determining an ROI luma value for the scene based on the tracked ROI, referred to herein as ROIiuma. For example, in the presence of an ROI, the Luma may be adjusted to factor in the ROIiuma. Some embodiments involve determining a modified luma value based on a weighted average of the average luma value and the ROI luma value. In some embodiments, as illustrated in Eqn. 3 below, the weight may correspond to ROIweight. For example, when the ROI is being tracked, the Luma may be determined at step 740 as follows:(Eqn. 3)

[0116] In such embodiments, the adjusting of the automatic image capture setting includes adjusting the initial automatic image capture setting based on the modified luma value to prioritize the tracked ROI in the scene.

[0117] In the event the AF logic does not commit to the ROI based on a saliency heat map, it may be preferable for the AE settings to factor in the ROI when determining the exposure. For example, when there is no ROI, and the saliency heat map indicates a salientregion in the scene, it may be preferable to prioritize that salient region. Accordingly, rather than using the tracked ROI from the AF logic, the camera may be configured to maintain and track its own saliency ROI.

[0118] Such a saliency ROI may be used in the same way as the ROI from the AF logic, however with lower weights. The weighting of the saliency ROI may be tuned accordingly. As previously described, it may be likely that the AF logic did not commit to the salient region because a confidence level was not sufficiently high. Accordingly, a scaled value of the saliency's confidence level may be used to weight the ROIiumawith respect to the Weighted J verageLumawhen determining the Luma in Eqn. 3. Generally, the saliency engine is likely to have been trained without considering focus and exposure, so exposing the salient region differently is unlikely to have an adverse effect on a confidence level of the saliency engine.

[0119] In some embodiments, the adjusting of the initial automatic image capture setting comprises maintaining the modified luma value within a threshold range of the average luma value. For example, a clamping factor may be added to constrain an effect of such Luma correction based on ROI tracking. For example, the clamping factor maintains consistency of the Luma with the luminance of the remainder of the image. At step 745, a clamp may be applied to Luma to generate a modified Luma' as follows:Luma' = clamp Luma, A * Weighted_Averageiuma, B * Weighted_Averageiuma)(Eqn. 4)

[0120] where A, B are constants. In some embodiments, the values may be chosen as .4 = 0.7 and # = 1.3. The purpose of the clamp is to avoid overcompensation for the presence of the ROI, and that the luminance stays within a bounded interval of an average luma value for the scene.

[0121] Subsequent to determining Luma based on Eqn. 4, the Short tes and Long tes may be adjusted to bring the Luma towards the T argetiuma.

[0122] In terms of interactions between various nodes of the image processing pipeline, a saliency node may provide, to an AE node for example, a saliency heat map and an ROI generated based on the saliency heat map. The AE node may provide to the saliency node, a running frequency for heat map generation and a flag to trigger saliency for a particular frame of the frames being captured. As between the AE node and the local tone mapping (LTM) node, the AE node may provide a short TES to the LTM node as brightness levels for highlight regions in the image. This may correspond to the brightness level at the sensor. The AE nodemay also provide a long TES to the LTM node as brightness levels for shadow regions in the image.

[0123] Figure 8 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments. Image 805 is a “before” image based on center-based AE setting, and image 810 is an “after” image using an ROI tracking. As illustrated, the overall exposure of the entire image in image 805 is higher than that of image 810. Subsequent to the ROI tracking based automatic image capture setting in image 810, the overall exposure has been reduced. Also, for example, in image 810, the region including the sign for “Electric Vehicle Parking Only” is brighter than the rest of the image.

[0124] Figure 9 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments. Image 905 is a “before” image based on center-based AE setting, and image 910 is an “after” image using an ROI tracking. As illustrated, the overall exposure in image 905 is higher than that of image 910. Subsequent to the ROI tracking based automatic image capture setting in image 910, the overall exposure is decreased. Also, for example, in image 910, the shadows and highlights in the region including the flower pot with the plant blend in with the remainder of the image. Also, for example, the carpet on the floor has a lower brightness level than the flower pot.

[0125] Figure 10 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments. Image 1005 is a “before” image based on center-based AE setting, and image 1010 is an “after” image using an ROI tracking. As illustrated, the brightness level of the entire image in image 805 is somewhat lower than that of image 1010, and the foreground is underexposed. Subsequent to the ROI tracking based automatic image capture setting in image 1010, the overall exposure level is increased, and the region including the planter with the plant is brighter and shows greater details of the design and texture of the planter.

[0126] Figure 11 is an example illustration of an image captured using an ROI tracking based automatic image capture setting, in accordance with example embodiments. Image 1105 is a “before” image based on center-based AE setting, and image 1110 is an “after” image using an ROI tracking. As illustrated, the brightness level of the entire image in image 1105 is higher than that of image 1110. Also, for example, in image 1110, the scene is overexposed. Subsequent to the ROI tracking based automatic image capture setting, the region including the sign for “Caution” is in sharper focus and the overall exposure of all portions of image 1110 has been lowered.Computing Device Architecture

[0127] Figure 12 is a block diagram of an example computing device 1200, in accordance with example embodiments. In particular, computing device 1200 shown in Figure 12 can be configured to perform at least one function described herein, including method 1300.

[0128] Computing device 1200 may include a user interface module 1201, a network communications module 1202, one or more processors 1203, data storage 1204, one or more cameras 1218, one or more sensors 1220, and power system 1222, all of which may be linked together via a system bus, network, or other connection mechanism 1205.

[0129] User interface module 1201 can be operable to send data to and / or receive data from external user input / output devices. For example, user interface module 1201 can be configured to send and / or receive data to and / or from user input devices such as a touch screen, a computer mouse, a keyboard, a keypad, a touch pad, a trackball, a joystick, a voice recognition module, and / or other similar devices. User interface module 1201 can also be configured to provide output to user display devices, such as one or more cathode ray tubes (CRT), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, either now known or later developed. User interface module 1201 can also be configured to generate audible outputs, with devices such as a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface module 1201 can further be configured with one or more haptic devices that can generate haptic outputs, such as vibrations and / or other outputs detectable by touch and / or physical contact with computing device 1200. In some examples, user interface module 1201 can be used to provide a graphical user interface (GUI) for utilizing computing device 1200.

[0130] Network communications module 1202 can include one or more devices that provide one or more wireless interfaces 1207 and / or one or more wireline interfaces 1208 that are configurable to communicate via a network. Wireless interface(s) 1207 can include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth™ transceiver, a Zigbee® transceiver, a Wi-Fi™ transceiver, a WiMAX™ transceiver, an LTE™ transceiver, and / or other type of wireless transceiver configurable to communicate via a wireless network. Wireline interface(s) 1208 can include one or more wireline transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate via a twisted pair wire, a coaxial cable, a fiberoptic link, or a similar physical connection to a wireline network.

[0131] In some examples, network communications module 1202 can be configured to provide reliable, secured, and / or authenticated communications. For each communication described herein, information for facilitating reliable communications (e.g., guaranteed message delivery) can be provided, perhaps as part of a message header and / or footer (e.g., packet / message sequencing information, encapsulation headers and / or footers, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity check values). Communications can be made secure (e.g., be encoded or encrypted) and / or decry pted / decoded using one or more cryptographic protocols and / or algorithms, such as, but not limited to, Data Encryption Standard (DES), Advanced Encryption Standard (AES), a Rivest-Shamir-Adelman (RSA) algorithm, a Diffie-Hellman algorithm, a secure sockets protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms can be used as well or in addition to those listed herein to secure (and then decry pt / decode) communications.

[0132] One or more processors 1203 can include one or more general purpose processors (e.g., central processing unit (CPU), etc.), and / or one or more special purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application specific integrated circuits, etc.). One or more processors 1203 can be configured to execute computer-readable instructions 1206 that are contained in data storage 1204 and / or other instructions as described herein.

[0133] Data storage 1204 can include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of one or more processors 1203. The one or more computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic or other memory or disc storage, which can be integrated in whole or in part with at least one of one or more processors 1203. In some examples, data storage 1204 can be implemented using a single physical device (e.g., one optical, magnetic, organic or other memory or disc storage unit), while in other examples, data storage 1204 can be implemented using two or more physical devices.

[0134] Data storage 1204 can include computer-readable instructions 1206 and perhaps additional data. In some examples, data storage 1204 can include storage required to perform at least part of the herein-described methods, scenarios, and techniques and / or at least part of the functionality of the herein-described devices and networks. In particular, computer- readable instructions 1206 can include instructions that, when executed by processor(s) 1203, enable computing device 1200 to provide for some or all of the functionality described herein.

[0135] In some embodiments, computer-readable instructions 1206 can include instructions that, when executed by processor(s) 1203, enable computing device 1200 to carry out operations. The operations may include generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device. The operations may also include detecting, based on the one or more saliency heat maps, a salient object in the scene. The operations may additionally include responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene. The operations may also include adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0136] In some examples, computing device 1200 can include tracked ROI module 1212. Tracked ROI module 1212 can be configured to track an ROI and cause adjustments to an automatic image capture setting of the image capturing device. For example, tracked ROI module 1212 can be configured to make decisions about initiating and / or terminating tracking of new and existing ROIs, compute various luma values, adjust a heat map generation frequency, and so forth.

[0137] In some examples, computing device 1200 can include one or more cameras 1218. Camera(s) 1218 can include one or more image capture devices, such as still and / or video cameras, equipped to capture light and record the captured light in one or more images; that is, camera(s) 1218 can generate image(s) of captured light. The one or more images can be one or more still images and / or one or more images utilized in video imagery. Camera(s) 1218 can capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as one or more other frequencies of light. Camera(s) 1218 can include a wide camera, a tele camera, an ultrawide camera, and so forth. Also, for example, camera(s) 1218 can be front-facing or rear-facing cameras with reference to computing device 1200. Camera(s) 1218 can include camera components such as, but are not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or an image sensor), lens, and / or shutter button. The camera components may be controlled at least in part by software executed by one or more processors 1203. For illustrative purposes, one camera of the one or more cameras 1218 may be a primary camera that is capturing a video of a scene, and a second camera of the one or more cameras 1218 may be a rear-facing main camera and / or an ultrawide camera used for purposes of saliency heat map based ROI tracking.

[0138] In some examples, computing device 1200 can include one or more sensors 1220. Sensors 1220 can be configured to measure conditions within computing device 1200 and / or conditions in an environment of computing device 1200 and provide data about these conditions. For example, sensors 1220 can include one or more of (i) sensors for obtaining data about computing device 1200, such as, but not limited to, a thermometer for measuring a temperature of computing device 1200, a battery sensor for measuring power of one or more batteries of power system 1222, and / or other sensors measuring conditions of computing device 1200; (ii) an identification sensor to identify other objects and / or devices, such as, but not limited to, a Radio Frequency Identification (RFID) reader, proximity sensor, one-dimensional barcode reader, two-dimensional barcode (e.g., Quick Response (QR) code) reader, and a laser tracker, where the identification sensors can be configured to read identifiers, such as RFID tags, barcodes, QR codes, and / or other devices and / or object configured to be read and provide at least identifying information; (iii) sensors to measure locations and / or movements of computing device 1200, such as, but not limited to, a tilt sensor, a gyroscope, an accelerometer, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser-displacement sensor, and a compass; (iv) an environmental sensor to obtain data indicative of an environment of computing device 1200, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor (e.g., an ambient light sensor), a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasound sensor and / or a smoke sensor; and / or (v) a force sensor to measure one or more forces (e.g., inertial forces and / or G-forces) acting about computing device 1200, such as, but not limited to one or more sensors that measure: forces in one or more dimensions, torque, ground force, friction, and / or a zero moment point (ZMP) sensor that identifies ZMPs and / or locations of the ZMPs. Many other examples of sensors 1220 are possible as well.

[0139] Power system 1222 can include one or more batteries 1224 and / or one or more external power interfaces 1226 for providing electrical power to computing device 1200. Each battery of the one or more batteries 1224 can, when electrically coupled to the computing device 1200, act as a source of stored electrical power for computing device 1200. One or more batteries 1224 of power system 1222 can be configured to be portable. Some or all of one or more batteries 1224 can be readily removable from computing device 1200. In other examples, some or all of one or more batteries 1224 can be internal to computing device 1200, and so may not be readily removable from computing device 1200. Some or all of one or more batteries 1224 can be rechargeable. For example, a rechargeable battery can be recharged via a wired connection between the battery and another power supply, such as by one or morepower supplies that are external to computing device 1200 and connected to computing device 1200 via the one or more external power interfaces. In other examples, some or all of one or more batteries 1224 can be non-rechargeable batteries.

[0140] One or more external power interfaces 1226 of power system 1222 can include one or more wired-power interfaces, such as a USB cable and / or a power cord, that enable wired electrical power connections to one or more power supplies that are external to computing device 1200. One or more external power interfaces 1226 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable wireless electrical power connections, such as via a Qi wireless charger, to one or more external power supplies. Once an electrical power connection is established to an external power source using one or more external power interfaces 1226, computing device 1200 can draw electrical power from the external power source the established electrical power connection. In some examples, power system 1222 can include related sensors, such as battery sensors associated with the one or more batteries or other types of electrical power sensors.

[0141] One or more external power interfaces 1226 of power system 1222 can include one or more wired-power interfaces, such as a USB cable and / or a power cord, that enable wired electrical power connections to one or more power supplies that are external to computing device 1200. One or more external power interfaces 1226 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable wireless electrical power connections, such as via a Qi wireless charger, to one or more external power supplies. Once an electrical power connection is established to an external power source using one or more external power interfaces 1226, computing device 1200 can draw electrical power from the external power source the established electrical power connection. In some examples, power system 1222 can include related sensors, such as battery sensors associated with the one or more batteries or other types of electrical power sensors.Example Methods of Operation

[0142] Figure 13 is a flowchart of a method, in accordance with example embodiments. Method 1300 may include various blocks or steps. The blocks or steps may be carried out individually or in combination. The blocks or steps may be carried out in any order and / or in series or in parallel. Further, blocks or steps may be omitted or added to method 1300.

[0143] The blocks of method 1300 may be carried out by various elements of computing device 1200 as illustrated and described in reference to Figure 12.

[0144] Block 1310 involves generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device.

[0145] Block 1320 involves detecting, based on the one or more saliency heat maps, a salient object in the scene.

[0146] Block 1330 involves responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene.

[0147] Block 1340 involves adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

[0148] Some embodiments involve determining a saliency score associated with the salient object. The initiating of the tracking of the ROI may be based on a determination that the saliency score exceeds a saliency threshold.

[0149] In some embodiments, the determining of the saliency score may be based on a strength of the one or more saliency heat maps.

[0150] In some embodiments, the generating of the one or more saliency heat maps is performed by a machine learning model.

[0151] Some embodiments involve terminating the tracking of the ROI based on a determination that a saliency score associated with the salient object is below a saliency threshold. Such embodiments involve increasing the heat map generation frequency for generation of saliency heat maps.

[0152] Some embodiments involve adjusting, based on additional one or more saliency heat maps generated at the increased heat map generation frequency, another automatic image capture setting for the video.

[0153] In some embodiments, the automatic image capture setting of the image capturing device may include one or more of an auto-exposure (AE) setting, an auto-focus (AF) setting, or an auto white balance (AWB) setting.

[0154] Some embodiments involve generating, at the reduced heat map generation frequency, one or more additional saliency heat maps. Such embodiments involve comparing the tracked ROI with the one or more additional saliency heat maps. Such embodiments also involve determining, based on the comparing, a second saliency score associated with the tracked ROI. Such embodiments further involve determining whether the second saliency score exceeds a second saliency threshold.

[0155] Some embodiments involve, based upon a determination that the second saliency score exceeds the second saliency threshold, continuing the tracking of the tracked ROI.

[0156] Some embodiments involve, based upon a determination that the second saliency score does not exceed the second saliency threshold, terminating the tracking of the tracked ROI.

[0157] Some embodiments involve detecting, based on the one or more additional saliency heat maps, a second salient object in the scene. Such embodiments involve, responsive to the detecting of the second salient object, determining whether to initiate a second tracking of a second ROI associated with the second salient object.

[0158] In some embodiments, the determining of whether to initiate the second tracking of the second ROI may be based on the second saliency score associated with the tracked ROI.

[0159] Some embodiments involve determining an average luma value associated with the scene, wherein the average luma value causes an initial automatic image capture setting to prioritize a center of the scene. Such embodiments involve determining an ROI luma value for the scene based on the tracked ROI. Such embodiments also involve determining a modified luma value based on a weighted average of the average luma value and the ROI luma value. The adjusting of the automatic image capture setting may involve adjusting the initial automatic image capture setting based on the modified luma value to prioritize the tracked ROI in the scene.

[0160] In some embodiments, the adjusting of the initial automatic image capture setting may involve maintaining the modified luma value within a threshold range of the average luma value.

[0161] In some embodiments, a first saliency threshold to continue tracking the ROI is lower than a second saliency threshold to initiate the tracking of the ROI.

[0162] Some embodiments involve adjusting, prior to the detecting of the salient object and based on the one or more saliency heat maps, another automatic image capture setting for the video.

[0163] Some embodiments involve capturing the image subsequent to the adjusting of the automatic image capture setting.

[0164] In some embodiments, the image capturing device may be a component of a computing device.

[0165] In some embodiments, the computing device may be a mobile device.

[0166] The particular arrangements shown in the Figures should not be viewed as limiting. It should be understood that other embodiments may include more or less of each element shown in a given Figure. Further, some of the illustrated elements may be combined or omitted. Yet further, an illustrative embodiment may include elements that are not illustrated in the Figures.

[0167] A step or block that represents a processing of information can correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a step or block that represents a processing of information can correspond to a module, a segment, or a portion of program code (including related data). The program code can include one or more instructions executable by a processor for implementing specific logical functions or actions in the method or technique. The program code and / or related data can be stored on any type of computer readable medium such as a storage device including a disk, hard drive, or other storage medium.

[0168] The computer readable medium can also include non-transitory computer readable media such as computer-readable media that store data for short periods of time like register memory, processor cache, and random access memory (RAM). The computer readable media can also include non-transitory computer readable media that store program code and / or data for longer periods. Thus, the computer readable media may include secondary or persistent long-term storage, like read only memory (ROM), optical or magnetic disks, compact disc read only memory (CD-ROM), for example. The computer readable media can also be any other volatile or non-volatile storage systems. A computer readable medium can be considered a computer readable storage medium, for example, or a tangible storage device.

[0169] While various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various disclosed examples and embodiments are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising: generating, at a heat map generation frequency, one or more saliency heat maps associated with a video of a scene being captured by an image capturing device; detecting, based on the one or more saliency heat maps, a salient object in the scene; responsive to the detecting of the salient object: initiating a tracking of a region of interest (ROI) associated with the salient object in subsequently captured video of the scene, and reducing the heat map generation frequency for generation of saliency heat maps for the subsequently captured video of the scene; and adjusting, based on the tracked ROI, an automatic image capture setting of the image capturing device.

2. The computer-implemented method of claim 1, further comprising: determining a saliency score associated with the salient object, and wherein the initiating of the tracking of the ROI is based on a determination that the saliency score exceeds a saliency threshold.

3. The computer-implemented method of claim 2, wherein the determining of the saliency score is based on a strength of the one or more saliency heat maps.

4. The computer-implemented method of any of claims 1-3, wherein the generating of the one or more saliency heat maps is performed by a machine learning model.

5. The computer-implemented method of any of claims 1-4, further comprising: terminating the tracking of the ROI based on a determination that a saliency score associated with the salient object is below a saliency threshold; and increasing the heat map generation frequency for generation of saliency heat maps.

6. The computer-implemented method of claim 5, further comprising:adjusting, based on additional one or more saliency heat maps generated at the increased heat map generation frequency, another automatic image capture setting for the video.

7. The computer-implemented method of any of claims 1-6, wherein the automatic image capture setting of the image capturing device comprises one or more of an auto-exposure (AE) setting, an auto-focus (AF) setting, or an auto white balance (AWB) setting.

8. The computer-implemented method of any of claims 1-7, further comprising: generating, at the reduced heat map generation frequency, one or more additional saliency heat maps; comparing the tracked ROI with the one or more additional saliency heat maps; determining, based on the comparing, a second saliency score associated with the tracked ROI; and determining whether the second saliency score exceeds a second saliency threshold.

9. The computer-implemented method of claim 8, further comprising: based upon a determination that the second saliency score exceeds the second saliency threshold, continuing the tracking of the tracked ROI.

10. The computer-implemented method of claim 8, further comprising: based upon a determination that the second saliency score does not exceed the second saliency threshold, terminating the tracking of the tracked ROI.

11. The computer-implemented method of claim 8, further comprising: detecting, based on the one or more additional saliency heat maps, a second salient object in the scene; and responsive to the detecting of the second salient object, determining whether to initiate a second tracking of a second ROI associated with the second salient object.

12. The computer-implemented method of claim 11, wherein the determining of whether to initiate the second tracking of the second ROI is based on the second saliency score associated with the tracked ROI.

13. The computer-implemented method of any of claims 1-12, further comprising: determining an average luma value associated with the scene, wherein the average luma value causes an initial automatic image capture setting to prioritize a center of the scene; determining an ROI luma value for the scene based on the tracked ROI; and determining a modified luma value based on a weighted average of the average luma value and the ROI luma value, and wherein the adjusting of the automatic image capture setting comprises adjusting the initial automatic image capture setting based on the modified luma value to prioritize the tracked ROI in the scene.

14. The computer-implemented method of claim 13, wherein the adjusting of the initial automatic image capture setting comprises maintaining the modified luma value within a threshold range of the average luma value.

15. The computer-implemented method of any of claims 1-14, wherein a first saliency threshold to continue tracking the ROI is lower than a second saliency threshold to initiate the tracking of the ROI.

16. The computer-implemented method of any of claims 1-15, further comprising: adjusting, prior to the detecting of the salient object and based on the one or more saliency heat maps, another automatic image capture setting for the video.

17. The computer-implemented method of any of claims 1-16, further comprising: capturing the image subsequent to the adjusting of the automatic image capture setting.

18. The computer-implemented method of any of claims 1-17, wherein the image capturing device is a component of a computing device.

19. The computer-implemented method of claim 18, wherein the computing device is a mobile device.

20. A computing device, comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executableinstructions that, when executed by the one or more processors, cause the computing device to carry out functions that comprise the computer-implemented method of any one of claims 1- 19.

21. An article of manufacture comprising one or more non-transitory computer readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to carry out functions that comprise the computer-implemented method of any one of claims 1-19.

22. A program that, when executed by one or more processors of a computing device, causes the computing device to carry out functions that comprise the computer- implemented method of any one of claims 1-19.

Citation Information

Patent Citations

  • Methods of compressing data and methods of assessing the same

    US20110255589A1

  • Hybrid object detector and tracker

    US20230021016A1

  • Methods for determining regions of interest for camera auto-focus

    WO2024076617A1

Cited By

  • System and method for performing object detection based on disparity image information

    US20260170847A1