Image Saliency-Based Smart Framing
Machine learning-based automatic zooming in image capture devices addresses the inefficiency of manual zoom adjustment by detecting salient regions and faces to achieve optimal image presentation.
Patent Information
- Application Number
- JP2025519883
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-10-05
- Publication Date
- 2025-10-09
AI Technical Summary
Current image capture devices require manual user input for adjusting zoom, which is time-consuming and inefficient, and often fail to provide an optimally zoomed presentation of the photographed scene.
Implementing machine learning techniques to detect visually salient regions and faces in an image, using bounding boxes to determine an appropriate zoom ratio automatically.
Facilitates automatic zooming to regions of interest, reducing manual input, enhancing image quality, and improving the functionality of image capture devices by providing faster and more accurate zoom and crop functions.
Smart Images

Figure 2025533881000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 378,444, filed October 5, 2022, the disclosure of which is incorporated herein by reference in its entirety. [Background technology]
[0002] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices, some of which are configured with telephoto capabilities. Summary of the Invention
[0003] In one embodiment, a method includes receiving an image captured by an image capture device. The method also includes determining a saliency bounding box based on the determined saliency metric for pixels of the image. The method further includes determining one or more face bounding boxes that enclose one or more faces identified in the image. The method further includes determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes. The method also includes determining a zoom ratio based on the determined zoom bounding box. The method further includes providing a zoomed image for display based on the determined zoom ratio.
[0004] In another embodiment, a system includes a processor and a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations. The operations include receiving an image captured by an image capture device. The operations also include determining a saliency bounding box based on a saliency metric determined for pixels of the image. The operations further include determining one or more face bounding boxes that enclose one or more faces identified in the image. The operations further include determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes. The operations also include determining a zoom ratio based on the determined zoom bounding box. The operations further include providing a zoomed image for display based on the determined zoom ratio.
[0005] In another embodiment, the image capture device comprises a camera and a control system. The control system is configured to receive an image captured by the image capture device. The control system is also configured to determine a saliency bounding box based on the saliency metric determined for the pixels of the image. The control system is further configured to determine one or more face bounding boxes that enclose one or more faces identified in the image. The control system is further configured to determine a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes. The control system is also configured to determine a zoom ratio based on the determined zoom bounding box. The control system is further configured to provide a zoomed image for display based on the determined zoom ratio.
[0006] In a further embodiment, a system is provided comprising: means for receiving an image captured by an image capture device; means for determining a saliency bounding box based on a saliency metric determined for pixels of the image; means for determining one or more face bounding boxes enclosing one or more faces identified in the image; means for determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes; means for determining a zoom ratio based on the determined zoom bounding box; and means for providing a zoomed image for display based on the determined zoom ratio.
[0007] The above summary is illustrative only and is not intended to be in any way limiting. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description, and accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] 1 illustrates an exemplary computing device, according to an exemplary embodiment. [Figure 2] FIG. 1 is a simplified block diagram illustrating some of the components of an exemplary computing system. [Figure 3] FIG. 1 illustrates the training and inference stages of one or more machine learning models, according to an example embodiment. [Figure 4a] 1 is an image according to an example embodiment. [Figure 4b] 1 is a heat map according to an example embodiment. [Figure 5] 1 illustrates a heatmap with bounding boxes according to an example embodiment. [Figure 6] 1 illustrates an image according to an exemplary embodiment. [Figure 7]1 illustrates an image with a bounding box on a phone, according to an example embodiment. [Figure 8] 1 illustrates an image on a phone, according to an exemplary embodiment. [Figure 9] 1 illustrates an image with bounding boxes and a heatmap according to an example embodiment. [Figure 10] 1 illustrates an image with bounding boxes and a heatmap according to an example embodiment. [Figure 11] 1 illustrates an image with bounding boxes and a heatmap according to an example embodiment. [Figure 12] 1 illustrates a heatmap with bounding boxes according to an example embodiment. [Figure 13] 1 is a flowchart of a method according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Exemplary methods, devices, and systems are described herein. It should be understood that the words "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or features, unless expressly stated. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.
[0010] Accordingly, the exemplary embodiments described herein are not intended to be limiting. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the Figures, could be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0011] Throughout this specification, the articles "a" or "an" are used to introduce elements of exemplary embodiments. Unless otherwise specified or clearly dictated otherwise by context, all references to "a" or "an" refer to "at least one" and all references to "the" refer to "at least one." The intention of using the conjunction "or" in a stated list of at least two terms is to refer to either of the listed terms or any combination of the listed terms.
[0012] The use of ordinal numbers such as "first," "second," "third," etc. is to distinguish between elements and does not indicate a particular order of these elements. For purposes of this description, the terms "multiple" and "a plurality of" refer to "two or more" or "more than one."
[0013] Furthermore, unless the context indicates otherwise, features shown in each of the figures may be used in combination with one another. Accordingly, the drawings should generally be viewed as aspects of components of one or more overall embodiments, with the understanding that not all of the illustrated features are required for each embodiment. In the drawings, like symbols typically identify like components unless the context dictates otherwise. Furthermore, unless otherwise noted, the drawings are not drawn to scale and are used for illustrative purposes only. Furthermore, the drawings are merely representative, and not all components are shown. For example, additional structural or constraint components may not be shown.
[0014] Furthermore, any recitation of elements, blocks, or steps in the specification or claims is for clarity purposes and, therefore, should not be construed as requiring or implying that these elements, blocks, or steps follow a particular arrangement or be performed in a particular order.
[0015] I. Overview To capture an image using an image capture device, including a digital camera, smartphone, laptop, etc., a user can power on the device and initiate a startup sequence for the image sensor (e.g., camera). The user can initiate the startup sequence by selecting an application or by simply powering on the device. The startup sequence typically includes an iterative optical and software setting adjustment process (e.g., autofocus, autoexposure, autowhite balance). After the startup sequence is complete, the image capture device can capture a high-quality image. However, the captured image rarely includes an optimally zoomed presentation of the photographed scene. For example, it may be preferable for the captured image to include a zoomed-in representation of the photographed scene, with an object of interest in focus.
[0016] Current techniques for adjusting the zoom of an image sensor involve the user providing manual input (e.g., adjusting a knob) or performing a gesture (e.g., a pinch gesture, a double-tap on the screen). Manually adjusting the zoom often requires time and effort from the user. Ideally, an image capture device would be able to automatically achieve an appropriate zoom for the captured scene.
[0017] Described herein are techniques for an image capture device to automatically zoom to regions of an image associated with a class(es) of interest in a captured scene, such as one or more faces. In some examples, by utilizing machine learning techniques, the image capture device may detect visually salient regions in an image, generate bounding boxes surrounding the visually salient regions, and calculate an appropriate zoom ratio. In doing so, the image capture device can implement the calculated zoom ratio to automatically achieve an appropriate zoom for the captured scene. The image capture device may also use a machine learning model to determine regions where faces or other classes of interest are predicted to be present, and the image capture device may use the determined regions in combination with the visually salient regions to determine regions to zoom into.
[0018] In particular, the image capture device may selectively determine which of the regions of interest defined by the bounding boxes combine with the visually salient regions defined by the saliency bounding boxes, which may improve the functionality of the camera and / or the functionality of a camera application on the image capture device. The image capture device may combine overlapping bounding boxes (e.g., the saliency bounding box and one or more bounding boxes determined from regions having classifications of interest) to determine a zoom bounding box, which may help facilitate the determination of a more appropriate zoom image that includes regions that a user may deem important while excluding other regions. Appropriately zooming an image may also facilitate automatic use of the appropriate camera when the image capture device has multiple cameras, reduced manual user input for zooming and / or cropping an image, faster and more accurate automatic zoom and / or crop functions, and higher quality images.
[0019] II. Exemplary Systems and Methods FIG. 1 illustrates an exemplary computing device 100. Computing device 100 is shown in a mobile phone form factor. However, computing device 100 may alternatively be implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. Computing device 100 may include various elements, such as a body 102, a display 106, and buttons 108 and 110. Computing device 100 may further include one or more cameras, such as a front-facing camera 104 and at least one rear-facing camera 112. In examples with multiple rear-facing cameras, such as shown in FIG. 1, each of the rear-facing cameras may have a different field of view. For example, the rear-facing cameras may include a wide-angle camera, a main camera, and a telephoto camera. The wide-angle camera may capture a wider portion of the environment compared to the main camera and the telephoto camera, and the telephoto camera may capture a more detailed image of a smaller portion of the environment compared to the main camera and the wide-angle camera.
[0020] The front camera 104 may be located on the side of the body 102 that normally faces the user during operation (e.g., the same side as the display 106). The rear camera 112 may be located on the side of the body 102 opposite the front camera 104. The designation of the cameras as front and rear is arbitrary, and the computing device 100 may include multiple cameras located on different sides of the body 102.
[0021] Display 106 may represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some examples, display 106 may display digital representations of a current image being captured by front camera 104 and / or rear camera 112, an image that may be captured by one or more of these cameras, an image that was recently captured by one or more of these cameras, and / or modified versions of one or more of these images. Thus, display 106 may function as a camera viewfinder. Display 106 may also support touchscreen functionality, which may allow settings and / or configurations of one or more aspects of computing device 100 to be adjusted.
[0022] The front-facing camera 104 may include an image sensor and associated optical elements, such as a lens. The front-facing camera 104 may provide zoom capabilities or have a fixed focal length. In other examples, interchangeable lenses may be used with the front-facing camera 104. The front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. The front-facing camera 104 may also be configured to capture still images, video images, or both. Furthermore, the front-facing camera 104 may represent, for example, a monocular camera, a stereoscopic camera, or a multi-lens camera. The rear-facing camera 112 may be similarly or differently positioned. Furthermore, one or more of the front-facing camera 104 and / or the rear-facing camera 112 may be an array of one or more cameras.
[0023] One or more of the front-facing camera 104 and / or the rear-facing camera 112 may include or be associated with an illumination component that provides a light field that illuminates the target object. For example, the illumination component may provide flash or constant illumination of the target object. The illumination component may also be configured to provide a light field that includes one or more of structured light, polarized light, and light with specific spectral content. Other types of light fields known and used to recover three-dimensional (3D) models from objects are possible within the context of the examples herein.
[0024] Computing device 100 may also include an ambient light sensor that may continuously or occasionally determine the ambient brightness of a scene that may be captured by cameras 104 and / or 112. In some implementations, the ambient light sensor may be used to adjust the display brightness of display 106. Additionally, the ambient light sensor may be used to determine or assist in determining the exposure length of one or more of cameras 104 or 112.
[0025] Computing device 100 can be configured to capture images of a target object using display 106 and front camera 104 and / or rear camera 112. The captured images can be multiple still images or a video stream. Image capture can be triggered by activating button 108, pressing a soft key on display 106, or some other mechanism. Depending on the implementation, images can be captured automatically at specific time intervals, for example, upon pressing button 108, in suitable lighting conditions for the target object, after moving computing device 100 a predetermined distance, or according to a predetermined capture schedule.
[0026] 2 is a simplified block diagram illustrating some of the components of an exemplary computing system 200. By way of example and not limitation, computing system 200 may be a cellular mobile phone (e.g., a smartphone), a computer (such as a desktop, notebook, tablet, server, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a game console, a robotic device, a vehicle, or some other type of device. Computing system 200 may represent, for example, an aspect of computing device 100.
[0027] 2, computing system 200 may include a communications interface 202, a user interface 204, a processor 206, data storage 208, and a camera component 224, all of which may be communicatively linked together by a system bus, network, or other connection mechanism 210. Computing system 200 may include at least some image capture and / or image processing functionality. It should be understood that computing system 200 may represent a physical image processing system, a specific physical hardware platform on which image sensing and / or processing applications run in software, or other combinations of hardware and software configured to perform image capture and / or image processing functions.
[0028] Communications interface 202 may enable computing system 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, communications interface 202 may facilitate circuit-switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet Protocol (IP) or other packet communications. For example, communications interface 202 may include a chipset and antenna arranged for wireless communication with a wireless access network or access point. Communications interface 202 may also take the form of or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) port, among other possibilities. Communications interface 202 may also take the form of or include a Wi-Fi, BLUETOOTH, Global Positioning System (GPS), or wide-area wireless interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)), among other possibilities. However, other forms of physical layer interfaces, as well as other types of standard or proprietary communication protocols, may be used via communication interface 202. Additionally, communication interface 202 may include multiple physical communication interfaces (e.g., a Wi-Fi interface, a BLUETOOTH interface, and a wide area wireless interface).
[0029] The user interface 204 may function to enable the computing system 200 to interact with a human or non-human user, such as by receiving input from the user and providing output to the user. Accordingly, the user interface 204 may include input components such as a keypad, keyboard, touch panel, computer mouse, trackball, joystick, microphone, etc. The user interface 204 may also include one or more output components, such as a display screen, which may be combined with a touch panel. The display screen may be based on CRT, LCD, LED, and / or OLED technology, or other technologies now known or later developed. The user interface 204 may also be configured to generate audible output(s) via speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. The user interface 204 may also be configured to receive and / or capture audible speech(es), noise(s), and / or signal(s) via a microphone and / or other similar devices.
[0030] In some examples, user interface 204 may include a display that functions as a viewfinder for still camera and / or video camera functions supported by computing system 200. Additionally, user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate configuring and focusing camera functions and capturing images. Some or all of these buttons, switches, knobs, and / or dials may be capable of being implemented by a touch panel.
[0031] Processor 206 may include one or more general-purpose processors, such as microprocessors, and / or one or more special-purpose processors, such as digital signal processors (DSPs), graphics processing units (GPUs), floating-point units (FPUs), network processors, or application-specific integrated circuits (ASICs). In some cases, the special-purpose processors may be capable of image processing, image alignment, and image combining, among other possibilities. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be wholly or partially integrated with processor 206. Data storage 208 may include removable and / or non-removable components.
[0032] Processor 206 may be capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in data storage 208 to perform various functions described herein. Thus, data storage 208 may include a non-transitory computer-readable medium having stored thereon program instructions that, when executed by computing system 200, cause computing system 200 to perform any of the methods, processes, or operations disclosed herein and / or in the accompanying drawings. Execution of program instructions 218 by processor 206 may result in processor 206 using data 212.
[0033] By way of example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device driver(s), and / or other modules) and one or more application programs 220 (e.g., camera functionality, address book, email, web browsing, social networking, audio-to-text functionality, text translation functionality, and / or gaming applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be primarily accessible to operating system 222, and application data 214 may be primarily accessible to one or more of application programs 220. Application data 214 may be located in a file system that is visible to or hidden from a user of computing system 200.
[0034] Application programs 220 may communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs may facilitate, for example, application programs 220 to read and / or write application data 214, send or receive information via communications interface 202, receive and / or display information on user interface 204, etc.
[0035] In some cases, application program 220 may be referred to as an "app" for short. Furthermore, application program 220 may be downloadable to computing system 200 through one or more online application stores or application markets. However, application programs may also be installed on computing system 200 in other ways, such as through a web browser or through a physical interface (e.g., a USB port) of computing system 200.
[0036] The camera component 224 may include, but is not limited to, an aperture, a shutter, a recording surface (e.g., photographic film and / or an image sensor), a lens, a shutter button, an infrared projector, and / or a visible light projector. The camera component 224 may include, among other possibilities, components configured to capture images in the visible light spectrum (e.g., electromagnetic radiation having wavelengths between 380 and 700 nanometers) and / or components configured to capture images in the infrared light spectrum (e.g., electromagnetic radiation having wavelengths between 701 nanometers and 1 millimeter). The camera component 224 may be controlled at least in part by software executed by the processor 206.
[0037] FIG. 3 is a diagram 300 illustrating the training stage 302 and inference stage 304 of trained machine learning model(s) 332, according to an example embodiment. Some machine learning techniques involve training one or more machine learning algorithms with an input set of training data to recognize patterns in the training data and provide output inferences and / or predictions regarding (the patterns in) the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, FIG. 3 illustrates the training stage 302 in which one or more machine learning algorithm(s) 320 are trained with training data 310 to become trained machine learning models 332. Creating the trained machine learning model(s) 332 during the training stage 302 may include determining one or more hyperparameters, such as one or more stride values, for one or more layers of the machine learning model, as described herein. Next, during the inference stage 304, the trained machine learning model(s) 332 may receive input data 330 and one or more inference / prediction requests 340 (perhaps as part of the input data 330) and, in response, provide as output one or more inferences and / or predictions 350. The one or more inferences and / or predictions 350 may be based in part on one or more learned hyperparameters, such as one or more learned stride values for one or more layers of the machine learning model as described herein.
[0038] Thus, the trained machine learning model(s) 332 may include one or more models of one or more machine learning algorithms 320. The machine learning algorithm(s) 320 may include, but are not limited to, artificial neural networks (e.g., convolutional neural networks described herein, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems). The machine learning algorithm(s) 120 may be supervised or unsupervised and may perform any suitable combination of online and offline learning.
[0039] In some examples, the machine learning algorithm(s) 320 and / or the trained machine learning model(s) 332 may be accelerated using on-device coprocessors, such as graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), and / or application-specific integrated circuits (ASICs). Such on-device coprocessors may be used to accelerate the machine learning algorithm(s) 320 and / or the trained machine learning model(s) 332. In some examples, the trained machine learning model(s) 332 may be trained to provide inferences on, reside on, execute on, and / or otherwise perform inferences for a particular computing device.
[0040] During the training phase 302, the machine learning algorithm(s) 320 can be trained by providing at least the training data 310 as training inputs using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing some (or all) of the training data 310 to the machine learning algorithm(s) 320, and the machine learning algorithm(s) 320 determining one or more output inferences based on the provided portion (or all) of the training data 310. Supervised learning involves providing some (or all) of the training data 310 to the machine learning algorithm(s) 320, and the machine learning algorithm(s) 320 determining one or more output inferences based on the provided portion (or all) of the training data 310, where the output inference(s) are either accepted or corrected based on the correct results associated with the training data 310. In some examples, the supervised learning of the machine learning algorithm(s) 320 can be governed by a set of rules and / or a set of labels for the training inputs, which may be used to correct the inferences of the machine learning algorithm(s) 320.
[0041] Semi-supervised learning involves having correct results for some, but not all, of the training data 310. During semi-supervised learning, supervised learning is used for the portions of the training data 310 that have correct results, and unsupervised learning is used for the portions of the training data 310 that do not have correct results.
[0042] Reinforcement learning involves machine learning algorithm(s) 320 receiving a reward signal related to a prior inference, where the reward signal may be a numerical value. During reinforcement learning, the machine learning algorithm(s) 320 may output an inference and receive a reward signal in response, where the machine learning algorithm(s) 320 are configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some examples, the machine learning algorithm(s) 320 and / or the trained machine learning model(s) 332 may be trained using other machine learning techniques, including, but not limited to, incremental learning and curriculum learning.
[0043] In some examples, the machine learning algorithm(s) 320 and / or the trained machine learning model(s) 332 can use transfer learning techniques. For example, transfer learning techniques may include pre-training the trained machine learning model(s) 332 on a set of data and further training them using the training data 310. More specifically, the machine learning algorithm(s) 320 can be pre-trained on data from one or more computing devices and the resulting trained machine learning models can be provided to the computing device CD1, which is intended to execute the trained machine learning models during the inference phase 304. Then, during the training phase 302, the pre-trained machine learning models can be further trained using the training data 310. This further training of the machine learning algorithm(s) 320 and / or the pre-trained machine learning models using the training data 310 of the data from CD1 can be performed using either supervised learning or unsupervised learning. Once the machine learning algorithm(s) 320 and / or the pre-trained machine learning models have been trained on at least the training data 310, the training phase 302 can be completed. The resulting trained machine learning model can be utilized as at least one of the trained machine learning model(s) 332.
[0044] In particular, once the training phase 302 is complete, the trained machine learning model(s) 332 may be provided to the computing device if not already on the computing device. The inference phase 304 can begin after the trained machine learning model(s) 332 are provided to the computing device CD1.
[0045] During the inference stage 304, the trained machine learning model(s) 332 can receive the input data 330 and generate and output one or more corresponding inferences and / or predictions 350 regarding the input data 330. Thus, the input data 330 can be used as input to the trained machine learning model(s) 332 to provide the corresponding inference(s) and / or prediction(s) 350. For example, the trained machine learning model(s) 332 can generate the inference(s) and / or prediction(s) 350 in response to one or more inference / prediction requests 340. In some examples, the trained machine learning model(s) 332 can be executed by another piece of software. For example, the trained machine learning model(s) 332 can be executed by an inference or prediction daemon so that they are readily available to provide inferences and / or predictions upon request. The input data 330 can include data from the computing device CD1 executing the trained machine learning model(s) 332 and / or input data from one or more computing devices other than CD1.
[0046] The exemplary image capture devices described herein may include, among other components, one or more cameras and sensors. The image capture device may be a smartphone, tablet, laptop, or digital camera, among other types of computing devices capable of performing the operations described herein.
[0047] As an example, a computing device may include one or more processors having logic for executing instructions, at least one built-in or peripheral image sensor (e.g., a camera), and an input / output device (e.g., a display panel) for displaying a user interface. The computing device may further include a computer-readable medium (CRM). The CRM may include any suitable memory or storage, such as random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), non-volatile RAM (NVRAM), read-only memory (ROM), or flash. The computing device stores device data (e.g., user data, multimedia data, applications, and / or the device's operating system) in the CRM. The device data may include executable instructions for an auto-zoom process. The auto-zoom process may be part of an operating system running on the image capture device or may be a separate component running within an application environment (e.g., a camera application) or "framework" provided by the operating system.
[0048] The computing device may implement machine-learned techniques ("visual saliency models"). The visual saliency models may be implemented as support vector machines (SVMs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), dense neural networks (DNNs), one or more heuristics, other machine learning techniques, combinations thereof, and the like. The visual saliency models may be iteratively trained off-device by exposure to training scenes, training sequences, and / or training events. For example, training may include exposing the visual saliency model to images (e.g., digital photographs) that include user-drawn bounding boxes that include visually salient regions (e.g., regions where one or more objects of particular interest to the user may reside). Exposure to images that include user-drawn bounding boxes may facilitate training the visual saliency model to identify visually salient regions within an image. As a result of the training, the visual saliency model may generate a visual saliency heat map for a given image and generate bounding boxes that enclose regions with the greatest probability of visual saliency. In this manner, the visual saliency model can predict visually salient regions within an image. After sufficient training, model compression using distillation can be performed on the visual saliency model, allowing for selection of the optimal model architecture based on model latency and power consumption. The visual saliency model can then be deployed as a separate module in the CRM of a computing device or implemented in the autozoom process.
[0049] The computing device may perform the auto-zoom process, perhaps automatically or in response to a received trigger signal including, for example, a user-performed gesture (e.g., a tap, a press) performed on an input / output device. The computing device may receive one or more captured images from an image sensor.
[0050] The computing device may utilize a visual saliency model to generate a visual saliency heatmap using one or more captured images.
[0051] For example, Figure 4a is an image 400 according to an example embodiment. Figure 4b is a heatmap 450 according to an example embodiment. A computing device may utilize a visual saliency model to generate a visual saliency heatmap of a captured image, as shown in Figures 4a and 4b. One or more processors calculate the visual saliency heatmap in the background operation of the device. In some examples, the image capture device does not display the visual saliency heatmap to a user. As shown, the visual saliency heatmap represents the magnitude of the probability of visual saliency on a scale from black to white, with white indicating a high probability of saliency and black indicating a low probability of saliency.
[0052] The visual saliency model may generate a bounding box that surrounds the region where the probability of visual saliency is greatest. Figure 5 shows a heatmap 500 with a bounding box 502, according to an example embodiment.
[0053] As shown in Figure 5, a visual saliency heatmap includes bounding boxes that enclose regions in an image that contain the greatest probability of visual saliency. If multiple objects of interest are present in a captured scene and the visual saliency model is tasked with identifying multiple salient regions in a captured image, the visual saliency model can be trained to generate bounding boxes that enclose the salient regions closest to the center of the captured image. This training technique assumes that the user is interested in the object that is most central in the image. Alternatively, the visual saliency model can be trained to generate bounding boxes that enclose all objects of interest in a captured image.
[0054] Using Equation 1, the computing device may calculate the target zoom ratio based on the dimensions of the bounding box. zmRatio=max(boundingBoxWidth / imageWidth,boundingBoxHeight / imageHeight) (1)
[0055] Equation 1 allows a computing device to calculate a zoom ratio (zmRatio) based on the bounding box width (boundingBoxWidth) and image width (imageWidth), as well as the bounding box height (boundingBoxHeight) and image height (imageHeight). The computing device may use the zoom ratio value to adjust the zoom setting of the image capture device.
[0056] Adjusting the zoom setting may include the computing device instructing one or more processors to adjust the placement of the optical lenses of the image sensor (i.e., optical zoom). In another aspect, the computing device utilizes a different image sensor to implement the calculated zoom ratio. In yet another aspect, the computing device digitally edits and enhances the image. For example, the computing device may not only crop and scale up the image, but also add pixels (i.e., digital zoom). A combination of these aspects (e.g., hybrid zoom) may also be utilized to achieve the appropriate zoom.
[0057] 6 shows an image 600 according to an example embodiment. Image 600 may be a zoomed-in image of image 400 of FIG. 4a based on bounding box 502 of FIG.
[0058] As shown in Figure 5, the visual saliency model identifies objects of interest in the image shown in Figure 4a, and the computing device implements the calculated zoom ratio. Figure 6 shows an automatically zoomed-in image of image 400 of Figure 4a.
[0059] In one aspect, the visual saliency model identifies visually salient regions and creates a bounding box only when the center of the visually salient region is within a predetermined distance from the center of the image. When this condition is met, the computing device zooms in to the center of the image. In another aspect, the computing device allows a user to zoom in to any region, where the zoom center is the center of the salient region rather than the center of the image.
[0060] Further to the above, when a computing device displays a captured image on an input / output device, the computing device may also display a bounding box surrounding the area in the image that contains the greatest probability of visual saliency. In this way, a user can visualize the proposed automatic zoom before the computing device implements the zoom ratio. Figure 7 shows an image with a bounding box on a phone, according to an exemplary embodiment.
[0061] As shown in Figure 7, the image capture device is a smartphone. The smartphone displays the captured image 700 to the user on an input / output device. A computing device running in the background of the device utilizes a visual saliency model to generate a visual saliency heat map, identify visual saliency regions, and create bounding boxes surrounding the visual saliency regions. As shown in Figure 7, the computing device presents a bounding box surrounding an object of interest in the captured image. The bounding box is presented to the user as a suggested automatic zoom.
[0062] In response to a user gesture, including, for example, selecting a bounding box or tapping an input / output device, the computing device implements the proposed automatic zoom. Figure 8 illustrates an image on a phone in accordance with an exemplary embodiment. Specifically, Figure 8 illustrates the smartphone of Figure 7 and the captured image with the proposed automatic zoom implemented.
[0063] As shown in Figure 8, the smartphone displays a zoomed-in version 800 of the captured image of Figure 7. In this manner, by utilizing a visual saliency model, the computing device can detect visually salient regions within the image, generate bounding boxes surrounding the visually salient regions, calculate an appropriate zoom ratio, and implement the calculated zoom ratio to automatically achieve a zoom appropriate for the captured scene.
[0064] One problem that can occur during an automatic zoom process is determining that a region of an image is not as salient as other regions of the image and cropping a portion of the image that includes important elements of the image. For example, in an image with a group of people, a computing device may determine that a person in the center of the image is the most salient region of the image and zoom to that region in the center of the image, which may result in the image not including all of the people in the group. Thus, in some examples, a computing device may be configured to zoom in on a particular region of an image frame based on the content of the image frame. For example, a computing device may determine a semantic classification of one or more objects in the image frame. Then, in addition to or as an alternative to a saliency heatmap, the computing device may determine the region to zoom in based on the semantic classification of the objects in the image frame.
[0065] 9 illustrates an image with bounding boxes and a heat map according to an example embodiment. FIG. 9 includes an image 900, from which a saliency heat map 910 may be generated. Based on the saliency heat map 910, a computing device may determine a saliency bounding box 914 by performing the saliency bounding box determination process as described above. Then, based on applying a machine learning model to predict regions of the image that have faces within the image 900, the computing device may determine face bounding boxes 912 and 916.
[0066] After determining one or more saliency bounding boxes and one or more face bounding boxes, the computing device may determine a zoom bounding box based on the determined saliency bounding box and face bounding box. In particular, the computing system may determine a zoom bounding box by combining the saliency bounding box with each of the one or more face bounding boxes that overlap the saliency bounding box, such that the zoom bounding box includes all regions within the saliency bounding box and all regions within each face bounding box that overlap the saliency bounding box. For example, based on the saliency bounding box 914 and the face bounding boxes 912 and 916, the computing device may determine a zoom bounding box 922.
[0067] 10 illustrates an image with bounding boxes and a heat map according to an example embodiment. In FIG. 10, image 1000 is an image including one or more people, and a computing device may determine a saliency heat map 1010 based on image 1000. Based on image 1000, the computing device may determine a face bounding box 1014, and based on the saliency heat map 1010, the computing device may determine a saliency bounding box 1012. Although face bounding box 1014 is outside the area of saliency bounding box 1012 and no overlap exists between face bounding box 1014 and saliency bounding box 1012, the computing device may determine that the area within face bounding box 1014 should be included in a zoom bounding box, e.g., zoom bounding box 1022.
[0068] In some examples, for an image in which a particular face bounding box is outside the saliency bounding box, the computing device may determine an average saliency of the saliency heatmap regions within the particular face bounding box, and the computing device may compare the determined average value with a threshold. If the determined average saliency exceeds the threshold, the computing device may expand the zoom bounding box to include the saliency bounding box. If the determined average saliency does not exceed the threshold, the computing device may ignore the particular face bounding box when determining the zoom bounding box. For example, for image 1000, the computing device may determine that the average saliency of the heatmap regions within face bounding box 1014 exceeds a threshold, and the computing device may determine a zoom bounding box 1022 that includes face bounding box 1014. If the average saliency of the heatmap regions within face bounding box 1014 does not exceed the threshold, the computing device may determine that the zoom bounding box includes the same region as saliency bounding box 1012 in the saliency heatmap 1010.
[0069] 11 shows an image with bounding boxes and a heat map according to an example embodiment. Based on image 1000, a computing device may determine at least a face bounding box 1114 and a heat map 1110. Then, based on the heat map 1110, the computing device may determine a saliency bounding box 1112, which may not overlap with the face bounding box 1114. The computing device may calculate an average saliency for regions within the face bounding box 1114 and determine that the average saliency does not exceed a threshold. Based on that determination, the computing device may determine a zoom bounding box 1122 to include only regions within the saliency bounding box 1112 and exclude regions within the face bounding box 1114. In some examples, the computing device may determine an additional face bounding box (e.g., a face bounding box at the position of the saliency bounding box 1112) that is included in the zoom bounding box 1122 such that such face bounding box may overlap with the saliency bounding box 1112.
[0070] In a further example, the computing device may determine whether one or more face bounding boxes that do not overlap with a saliency bounding box, for example, face bounding box 1114 that does not overlap with saliency bounding box 1112, are included in the zoom bounding box based on the respective sizes of the one or more face bounding boxes. For example, the percentage of pixels in face bounding box 1114 relative to the number of pixels in the entire image may be compared to a threshold pixel percentage. If the percentage of pixels in face bounding box 1114 exceeds the threshold pixel percentage, the computing device may include the face bounding box in the zoom bounding box. Conversely, if the percentage of pixels in face bounding box 1114 does not exceed the threshold pixel percentage, the computing device may exclude the face bounding box in the zoom bounding box.
[0071] After determining the zoom bounding box for the given image, the computing device may determine a zoom ratio based on Equation 1, as described above. At the determined zoom ratio, the computing device may provide a zoomed image for display, possibly by adjusting the field of view of the computing device and / or cropping the image to include only the region of the given image that is included in the zoom bounding box.
[0072] In some examples, the computing device may apply a predetermined amount of padding to the zoom bounding box, and based on the zoom bounding box with the predetermined amount of padding, the computing device may determine the zoom ratio. The predetermined amount of padding may facilitate displaying a zoomed image that is not overly cropped (e.g., when one or more objects in the image are on top of an edge of the cropped image).
[0073] As described above, to facilitate the process of determining the zoom box, a computing device may first determine a saliency bounding box. Figure 12 shows a heatmap with bounding boxes according to an example embodiment. Figure 12 includes within heatmap 1200 a candidate saliency bounding box 1202, a candidate saliency bounding box 1204, a candidate saliency bounding box 1206, and a candidate saliency bounding box 1208.
[0074] The heat map 1200 may include one or more pixels, with each pixel associated with a saliency metric. In the heat map 1200, more salient regions may be represented by brighter pixels and a larger saliency metric. To determine the saliency bounding boxes, the computing device may first normalize the saliency map by subtracting a predetermined value from the saliency metric associated with each pixel. The computing system may then determine the region within the normalized heat map (e.g., the region encompassed by the candidate saliency bounding box 1202) that maximizes the average saliency metric, perhaps by using a two-dimensional (2D) Kadane's algorithm or another iterative dynamic programming algorithm. In some examples, the computing system may also discover additional candidate saliency boxes by going through a similar process for the remainder of the heat map to identify candidate saliency bounding boxes 1204, 1206, and 1208. To determine which candidate saliency boxes to use when determining the zoom bounding boxes, the computing device may select the candidate bounding box that would generate the most conservative (e.g., smallest) zoom ratio. In the case of heatmap 1200, the computing device may determine to use candidate saliency bounding box 1202 to crop the image less. And because the computing device may crop the image less by using candidate saliency bounding box 1202, candidate saliency bounding box 1022 may be associated with the most conservative / smallest zoom ratio.
[0075] In some examples, one or more of the face bounding boxes included in the saliency bounding box and / or zoom bounding box may be located very close to an edge of an image (e.g., image 1000 of FIG. 10 ). Additionally and / or alternatively, one or more of the face bounding boxes included in the saliency bounding box and / or zoom bounding box may reside partially outside the edge of the image. If the computing device determines that one edge of the zoom bounding box is less than a threshold distance from the edge of the image and / or is outside the image's padding when padding is applied to the image's inner edge, the computing device may switch to using another camera on the computing device, for example, a camera with a wider field of view. The computing system may then capture another image with the camera with the wider field of view and determine a new zoom bounding box that may include areas not captured by the previous camera. Additionally and / or alternatively, if the computing device determines that the zoom bounding box exceeds a threshold distance from the edge of the image, the computing device may switch to a camera to use for zooming (e.g., a camera with a smaller field of view that provides more clarity for distant objects) and may determine a new image and a new zoom bounding box for that image.
[0076] Further, in some examples, the computing system need only zoom in on the center of the image. Thus, the computing system may determine the zoom ratio based on the two edges (e.g., left or right, top or bottom) of the zoom bounding box that are closest to the edge of the image, whereby additional regions of the image may be included on the other two edges of the zoom bounding box that are away from the edge of the image.
[0077] 13 is a flowchart of a method 1300 according to an example embodiment. Method 1300 may be performed by one or more computing systems (e.g., computing system 200 of FIG. 2) and / or one or more processors (e.g., processor 206 of FIG. 2). Method 1300 may be implemented in a computing device, such as computing device 100 of FIG. 1.
[0078] At block 1302, the method 1300 includes receiving an image captured by an image capture device.
[0079] At block 1304, the method 1300 includes determining a saliency bounding box based on the determined saliency metric for the pixels of the image.
[0080] At block 1306, the method 1300 includes determining one or more face bounding boxes that enclose one or more faces identified in the image.
[0081] At block 1308, the method 1300 includes determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes.
[0082] At block 1310, the method 1300 includes determining a zoom ratio based on the determined zoom bounding box.
[0083] At block 1312, the method 1300 includes providing a zoomed image for display based on the determined zoom ratio.
[0084] In some examples, determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes includes combining the saliency bounding box with each of the one or more face bounding boxes that overlap the saliency bounding box.
[0085] In some examples, determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes includes determining whether to combine the saliency bounding box with a given face bounding box of the one or more face bounding boxes based on a saliency metric determined for pixels of the given face bounding box.
[0086] In some examples, determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes includes determining the saliency bounding box to exclude a given face bounding box from the one or more face bounding boxes when the given face bounding box does not overlap with the saliency bounding box and when a saliency metric determined for pixels of the given face bounding box is less than a threshold.
[0087] In some examples, determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes includes determining whether to combine each of the one or more face bounding boxes with the saliency bounding box based on a respective size of each of the one or more face bounding boxes.
[0088] In some examples, determining the saliency bounding boxes includes applying a machine-learned saliency model to determine saliency metrics for pixels of the image.
[0089] In some examples, determining the one or more face bounding boxes includes applying a machine-learned face detection model to the image.
[0090] In some examples, the received image and the provided zoomed image are captured by different cameras on the image capture device, each of the different cameras having a different respective field of view.
[0091] In some examples, determining the saliency bounding boxes includes applying an iterative dynamic programming algorithm.
[0092] In some examples, determining the saliency bounding box includes determining a saliency heatmap of pixels of the image, normalizing the saliency heatmap by reducing each value in the saliency heatmap by a predetermined amount, and identifying a region in the normalized saliency heatmap that maximizes a mean saliency value in the normalized saliency heatmap.
[0093] In some examples, determining the saliency bounding box includes identifying a plurality of regions in the normalized saliency heat map that maximizes a mean saliency value in the normalized saliency heat map, and determining the saliency bounding box in a region of the plurality of regions that produces the most conservative zoom ratio.
[0094] In some examples, determining the zoom ratio includes applying a predetermined amount of padding to the zoom bounding box.
[0095] In some examples, providing the zoomed image is performed in response to a user input gesture received at the image capture device.
[0096] In some examples, the method 1300 is performed by an image capture device that includes a camera and a control system configured to perform the steps of the method 1300.
[0097] In such an example, the camera may be a first camera, the image capture device comprises a second camera, and the zoomed image is captured by the second camera.
[0098] III. Conclusion The present disclosure is not limited with respect to the specific embodiments described in this application, which are intended as examples of various aspects. Many modifications and variations are possible without departing from the scope thereof, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0099] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, like numerals generally identify like components unless the context dictates otherwise. The exemplary embodiments described in the specification and drawings are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0100] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the figures and as described herein, each step, block, and / or communication may represent the processing of information and / or the transmission of information according to the example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be executed in an order different from that shown or described, including substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or operations may be used in any of the message flow diagrams, scenarios, and flowcharts described herein, and these message flow diagrams, scenarios, and flowcharts may be combined with each other, either partially or in whole.
[0101] Steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to perform specific logical operations or actions in the method or technique. The program code and / or associated data may be stored in any type of computer-readable medium, such as a storage device, including a random access memory (RAM), a disk drive, a solid-state drive, or another storage medium.
[0102] Computer-readable media may also include non-transitory computer-readable media, such as computer-readable media that store data for the short term, such as register memory, processor cache, and RAM. Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for the long term. Thus, computer-readable media may include secondary storage or persistent long-term storage, such as, for example, read-only memory (ROM), optical or magnetic disks, solid-state drives, and compact disc read-only memories (CD-ROMs). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may also be considered computer-readable storage media, for example, tangible storage devices.
[0103] Additionally, steps or blocks representing one or more information transmissions may correspond to information transmissions between software and / or hardware modules within the same physical device, however, other information transmissions may be between software and / or hardware modules of different physical devices.
[0104] The particular arrangement shown in the drawings should not be considered limiting. It should be understood that other embodiments may include more or less of each element shown in a given drawing. Furthermore, some of the illustrated elements may be combined or omitted. Furthermore, example embodiments may include elements that are not shown.
[0105] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, the true scope being indicated by the following claims.
Claims
1. receiving an image captured by an image capture device; determining a saliency bounding box based on the determined saliency metric for pixels of the image; determining one or more face bounding boxes surrounding one or more faces identified in the image; determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes; determining a zoom ratio based on the determined zoom bounding box; providing a zoomed image for display based on the determined zoom ratio; and A method comprising:
2. 2. The method of claim 1 , wherein determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes comprises combining the saliency bounding box with each of the one or more face bounding boxes that overlap the saliency bounding box.
3. 2. The method of claim 1 , wherein determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes comprises determining whether to merge the saliency bounding box with a given face bounding box of the one or more face bounding boxes based on the saliency metric determined for pixels of the given face bounding box.
4. 2. The method of claim 1 , wherein determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes comprises determining the saliency bounding box to exclude a given face bounding box of the one or more face bounding boxes when the given face bounding box of the one or more face bounding boxes does not overlap the saliency bounding box and when the saliency metric determined for pixels of the given face bounding box is less than a threshold.
5. 2. The method of claim 1 , wherein determining the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes comprises determining whether to merge each of the one or more face bounding boxes with the saliency bounding box based on a respective size of each of the one or more face bounding boxes.
6. The method of claim 1 , wherein determining the saliency bounding boxes comprises applying a machine-learned saliency model to determine the saliency metrics for the pixels of the image.
7. The method of claim 1 , wherein determining the one or more face bounding boxes comprises applying a machine-learned face detection model to the image.
8. The method of claim 1 , wherein the received image and the provided zoomed image are captured by different cameras on the image capture device, each of the different cameras having a different respective field of view.
9. The method of claim 1 , wherein determining the salient bounding boxes comprises applying an iterative dynamic programming algorithm.
10. Determining the saliency bounding box comprises: determining a saliency heatmap of the pixels of the image; normalizing the saliency heatmap by decreasing each value in the saliency heatmap by a predetermined amount; identifying a region in the normalized saliency heatmap that maximizes a mean saliency value in the normalized saliency heatmap; The method of claim 1 , comprising:
11. Determining the saliency bounding box comprises: identifying a plurality of regions in the normalized saliency heatmap that maximize a mean saliency value in the normalized saliency heatmap; determining the saliency bounding box to be the region among the plurality of regions that produces the most conservative zoom ratio; The method of claim 10, comprising:
12. The method of claim 1 , wherein determining the zoom ratio comprises applying a predetermined amount of padding to the zoom bounding box.
13. The method of claim 1 , wherein providing the zoomed image is performed in response to a user input gesture received at the image capture device.
14. 1. An image capture device, comprising: A camera and a control system, the control system comprising: receiving an image captured by the capture; determining a saliency bounding box based on the determined saliency metric for pixels of the image; determining one or more face bounding boxes surrounding one or more faces identified in the image; determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes; determining a zoom ratio based on the determined zoom bounding box; providing a zoomed image for display based on the determined zoom ratio; and an image capture device configured to:
15. The image capture device of claim 14 , wherein the camera is a first camera, the image capture device comprises a second camera, and the zoomed image is captured by the second camera.
16. 15. The image capture device of claim 14, wherein the control system is configured to determine the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes by combining the saliency bounding box with each of the one or more face bounding boxes that overlap the saliency bounding box.
17. 15. The image capture device of claim 14, wherein the control system is configured to determine the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes by determining whether to combine the saliency bounding box with a given face bounding box of the one or more face bounding boxes based on the saliency metric determined for pixels of the given face bounding box.
18. 15. The image capture device of claim 14, wherein the control system configured to determine the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes comprises determining the saliency bounding box by excluding a given face bounding box of the one or more face bounding boxes when the given face bounding box does not overlap the saliency bounding box and when the saliency metric determined for pixels of the given face bounding box is less than a threshold.
19. 15. The image capture device of claim 14, wherein the control system is configured to determine the zoom bounding box based on the saliency bounding box and the one or more face bounding boxes by determining whether to combine each of the one or more face bounding boxes with the saliency bounding box based on a total number of the one or more face bounding boxes and based on a respective size of each of the one or more face bounding boxes.
20. A non-transitory computer-readable medium storing program instructions executable by one or more processors to cause the one or more processors to perform operations, the operations comprising: receiving an image captured by an image capture device; determining a saliency bounding box based on the determined saliency metric for pixels of the image; determining one or more face bounding boxes surrounding one or more faces identified in the image; determining a zoom bounding box based on the saliency bounding box and the one or more face bounding boxes; determining a zoom ratio based on the determined zoom bounding box; providing a zoomed image for display based on the determined zoom ratio; and 1. A non-transitory computer-readable medium comprising:
Citation Information
Patent Citations
Method for receiving multimedia signals comprising audio frames and video frames
JP2009508386A
Image processor and method
JP2011035634A
Systems and methods utilizing deep learning for selectively storing audiovisual content
JP2021507550A
Video decoding method and video coding method
JP2022531759A