A method for determining the region of interest for camera autofocus.
The image capture device uses machine-learned techniques to detect and stabilize autofocus on visually salient regions, addressing instability and misjudgment in focusing by employing a primary and secondary ROI, enhancing autofocus stability and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2023-10-04
- Publication Date
- 2026-05-21
AI Technical Summary
Determining which area of an environment to focus on in an image capture device is complicated by the presence of various objects, and objects of interest may move or be outside the frame, causing instability and misjudgment in autofocus processes.
The image capture device utilizes machine-learned techniques to detect visually salient regions, generate bounding boxes, and apply autofocus processes to these regions, using a primary and secondary region of interest (ROI) to stabilize focus and switch between areas efficiently.
This approach enhances autofocus stability and efficiency by focusing on high-saliency areas, reducing instability and misjudgment, and allowing seamless transitions between focus areas.
Smart Images

Figure 0007863685000001 
Figure 0007863685000002 
Figure 0007863685000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the priority of U.S. Provisional Patent Application No. 63 / 378,648, filed on October 6, 2022, the entire disclosure of which is incorporated herein by reference.
Background Art
[0002] Many modern computing devices, including mobile phones, personal computers, and tablets, are equipped with image capture devices. Some image capture devices are configured with a telephoto function.
Summary of the Invention
[0003] In an embodiment, the method includes receiving an image frame captured by an image capture device. The method also includes determining a saliency heatmap representing the saliency of pixels in the image frame. The method further includes determining a primary region of interest (ROI) and a secondary ROI of the image frame based on the saliency heatmap. The method further includes determining a filtered ROI of the image frame, where the filtered ROI is updated from a previous filtered ROI to the primary ROI based on the saliency difference between the previous filtered ROI and the primary ROI exceeding a first threshold. The method also includes applying one or more autofocus processes based on at least one of the filtered ROI, the primary ROI, or the secondary ROI.
[0004] In other embodiments, the system includes a processor and a non-temporary computer-readable medium storing instructions, which, when executed by the processor, cause the processor to perform an operation. The operation includes receiving an image frame captured by an image capture device. The operation also includes determining a splendor heatmap representing the splendor of pixels in the image frame. The operation further includes determining a primary region of interest (ROI) and a secondary ROI of the image frame based on the splendor heatmap. The operation further includes determining a filtered ROI of the image frame, which is updated from a previous filtered ROI to a primary ROI based on the splendor difference between the previous filtered ROI and the primary ROI exceeding a first threshold. The operation also includes applying one or more autofocus processes based on at least one of the filtered ROI, primary ROI, or secondary ROI.
[0005] In this embodiment, the image capture device includes a camera and a control system. The control system is configured to receive image frames captured by the image capture device. The control system is also configured to determine a splendor heatmap representing the splendor of pixels within the image frame. The control system is further configured to determine a primary region of interest (ROI) and a secondary ROI of the image frame based on the splendor heatmap. The control system is further configured to determine a filtered ROI of the image frame, which is updated from a previous filtered ROI to a primary ROI based on the splendor difference between the previous filtered ROI and the primary ROI exceeding a first threshold. The control system is also configured to apply one or more autofocus processes based on at least one of the filtered ROI, the primary ROI, or the secondary ROI.
[0006] In other embodiments, a system is provided that includes means for receiving an image frame captured by an image capture device. The system also includes means for determining a splendor heatmap representing the splendor of pixels within the image frame. The system further includes means for determining a primary region of interest (ROI) and a secondary ROI of the image frame based on the splendor heatmap. The system further includes means for determining a filtered ROI of the image frame, which is updated from a previous filtered ROI to a primary ROI based on the splendor difference between the previous filtered ROI and the primary ROI exceeding a first threshold. The system also includes means for applying one or more autofocus processes based on at least one of the filtered ROI, primary ROI, or secondary ROI.
[0007] The above summary is illustrative and not intended to be limiting. Further embodiments, features, and characteristics beyond those described above will become apparent from the figures and the following detailed description and accompanying drawings. [Brief explanation of the drawing]
[0008] [Figure 1] An exemplary computing device is shown according to an exemplary embodiment. [Figure 2] This is a simplified block diagram showing some of the components of an exemplary computing system. [Figure 3] This figure shows the training and inference phases of one or more trained machine learning models according to an exemplary embodiment. [Figure 4a] This is an image illustrating an exemplary embodiment. [Figure 4b] This is a heatmap based on an exemplary embodiment. [Figure 5] A heatmap with a bounding box is shown according to an exemplary embodiment. [Figure 6A]An anchor boundary box according to an exemplary embodiment is shown. [Figure 6B] The anchor boundary box position is shown according to an exemplary embodiment. [Figure 7] A representative embodiment shows a prominent region of interest (ROI). [Figure 8] An image with an ROI is shown according to an exemplary embodiment. [Figure 9] A finite state machine is shown by an exemplary embodiment. [Figure 10] An exemplary embodiment of a finite-state machine manager is shown. [Figure 11] This is a flowchart of the method according to an exemplary embodiment. [Modes for carrying out the invention]
[0009] This specification describes exemplary methods, devices, and systems. The terms “example” and “exemplary” are used herein to mean “serving as an example, case, or illustration.” Any embodiment or feature described herein as “example” or “exemplary” should not be construed as necessarily preferable or advantageous to other embodiments or features unless otherwise indicated. Other embodiments may be used, or other modifications may be made, without departing from the scope of the subject matter presented herein.
[0010] Therefore, the exemplary embodiments described herein are not intended to be limiting. It will be readily apparent that the aspects of this disclosure described herein and shown in the drawings may be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0011] Throughout this description, the articles “a” or “an” are used to introduce elements of exemplary embodiments. Any reference to “a” or “an” refers to “at least one,” and any reference to “the” refers to “that at least one,” unless otherwise specified or clearly indicated by the context. The use of the conjunction “or” in lists containing at least two terms is intended to indicate any of the listed terms, or any combination thereof.
[0012] The use of ordinal numbers such as "first," "second," and "third" is to distinguish each element, not to indicate a specific order of these elements. For the purposes of this explanation, the terms "multiple" and "a plurality of" refer to "two or more" or "more than one."
[0013] Furthermore, unless otherwise indicated by the context, the features illustrated in each figure may be used in combination with each other. Therefore, the drawings should be viewed generally as representations of components of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment. In the figures, similar symbols generally identify similar components unless otherwise indicated by the context. Furthermore, unless otherwise noted, the figures are not drawn to scale and are for illustrative purposes only. Additionally, the figures are representative and do not show all components; for example, additional structural or restrictive components may not be shown.
[0014] Furthermore, any enumeration of elements, blocks, or steps in this specification or in the claims is for clarification purposes only. Therefore, such enumeration should not be construed as requiring or implying that these elements, blocks, or steps must be performed in a particular arrangement or order.
[0015] I. Overview To capture an image using an image capture device, including digital cameras, smartphones, and laptops, a user may power on the device and initiate a startup sequence for the image sensor (e.g., a camera). The user may initiate the startup sequence by selecting an application or simply by powering on the device. The startup sequence typically involves iterative optical and software setting adjustment processes (e.g., autofocus, auto exposure, auto white balance). After the startup sequence is complete, the image capture device may capture an image. Ideally, the image capture device may have the ability to precisely focus on the image and / or apply various autofocus processes to the image so that the captured image and / or preview of the captured image include a focused view of any object of interest within the image.
[0016] However, determining which area of the environment to focus on (for example, which area of the environment contains the object of interest) can be complicated by the various objects present in the environment, many of which may be focusable areas. Furthermore, the object that the image capture device is focusing on may move to a different location and / or be moved outside the frame of the image capture device. Therefore, it can be important for the image capture device to be able to quickly switch to other focus areas. Moreover, the presence of various objects in the environment may cause the image capture device to associate similar levels of interest with different areas, and the image capture device may fluctuate between focusing on one area and focusing on other areas with similar levels of interest.
[0017] This specification describes techniques for an image capture device to autofocus on areas of an image frame associated with high saliency while reducing instability and misjudgment of salient areas. In some examples, by utilizing machine-learned techniques, the image capture device can detect visually salient regions within an image frame, generate one or more bounding boxes surrounding one or more visually salient regions, determine the visually salient regions to focus on, and apply one or more autofill processes to those visually salient regions within the image frame.
[0018] In some embodiments, the image capture device may determine a primary region of interest (ROI) and a more stable filtered ROI based on the primary ROI. For each image frame, the image capture device may first set the filtered ROI to be the same as the previous filtered ROI and update the confidence value of the filtered ROI to the average of the spleness values at the pixels of the updated heatmap. Based on the amount of overlap and / or relative confidence between the primary ROI and the previous filtered ROI, the image capture device may decide whether to match the filtered ROI to the primary ROI or to leave the filtered ROI the same as the previous filtered ROI. In some embodiments, the image capture device may determine whether the amount of overlap between the primary ROI and the previous filtered ROI exceeds a threshold, and based on that determination, the image capture device may update the filtered ROI to the primary ROI. If the image capture device determines that the amount of overlap exceeds the threshold, the image capture device may maintain the filtered ROI. In this case, the image capture device may set the filtered ROI to be the same as the previous filtered ROI and update the confidence value of the filtered ROI to the average of the splendor values at the pixels of the updated heatmap. Additionally and / or alternatively, the computing system may identify the splendor difference between the previous filtered ROI and the primary ROI, and based on whether the splendor difference exceeds a threshold, the computing system may update the previous filtered ROI to the primary ROI. If the computing system determines that the splendor difference does not exceed a threshold, the computing system may maintain the filtered ROI in the state of the previous ROI.
[0019] When updating the filtered ROI based on the primary ROI and the previously filtered ROI, the filtered ROI can remain the previously filtered ROI with updated confidence values based on the updated heatmap until the region becomes insignificant. Thus, when the previous filtered ROI becomes insignificant, the filtered ROI can be updated to the primary ROI. By allowing the filtered ROI to remain the previous filtered ROI, stability can be promoted when the image frame changes slightly from frame to frame and when the frame has multiple objects of similar saliency.
[0020] In a further embodiment, the image capture device can determine both a primary region of interest (ROI) and a secondary ROI. The secondary ROI can be determined such that it does not overlap with the primary ROI or a prohibited region around the primary ROI. Since the computing device uses at least two ROIs of the current image frame to determine where to focus, the computing device can consider two salient objects and / or regions simultaneously, potentially making the focus switch from one region to another faster and more seamless. Additionally, the computing device can potentially apply a low-pass filter to the primary ROI by applying a threshold to the saliency difference and / or amount of overlap before switching from the filtered ROI to the primary ROI, thereby potentially making the filtered ROI more stable and power efficient.
[0021] In further embodiments, a computing device may select between filtered ROIs, primary ROIs, and secondary ROIs to determine an area within an image frame to focus on in one or more autofocus processes. To address the potential problem that splendor detection may exhibit instability over time, a finite state machine (FSM) may be used. This FSM may require a number of consecutive frames in which splendor detection is consistent for autofocus to commit to a splendid ROI, and / or require a number of consecutive frames in which splendor detection is inconsistent for autofocus to abandon a splendid ROI. In this context, consistent detection refers to overlapping bounding boxes detected across consecutive frames, and / or high reliability of these detections.
[0022] A further potential challenge is that splendor detection often reports splendor areas that are not actually splendor. This is known as a false positive, and this type of error can negatively impact the user experience, particularly with cameras. Imagine a camera trying to focus on a glowing object in the background of a scene, for example. While FSMs are useful in addressing this challenge because they handle momentary false positives, FSMs are less likely to detect consecutive false positives than a single false positive. In addition, to prevent false positives, the application of the splendor autofocus process may be restricted based on global on-device signals. For example, certain requirements (e.g., minimum brightness value of the scene, zoom ratio within a specific range, and / or no device movement) may be imposed before autofocus can be enabled. Furthermore, certain splendor areas may be discarded when the estimated depth is out of range, when the estimated depth differs too much from the estimated distance of the current ROI, and / or when the bounding box is too far from the center.
[0023] II. Exemplary Systems and Methods Figure 1 shows an exemplary computing device 100. The computing device 100 is shown in the form factor of a mobile phone. However, the computing device 100 may alternatively be implemented in a laptop computer, tablet computer, and / or wearable computing device, among many other possibilities. The computing device 100 may include various elements such as a body 102, a display 106, and buttons 108 and 110. The computing device 100 may further include one or more cameras, such as a front camera 104 and one or more rear cameras 112. Each of the rear cameras may have a different field of view. For example, the rear cameras may include a wide-angle camera, a main camera, and a telephoto camera. The wide-angle camera may capture a wider portion of the environment compared to the main camera and the telephoto camera, and the telephoto camera may capture a more detailed image of a smaller portion of the environment compared to the main camera and the wide-angle camera.
[0024] The front camera 104 may be located on the side of the main unit 102 that normally faces the user during operation (for example, the same side as the display 106). The rear camera 112 may be located on the side of the main unit 102 opposite to the front camera 104. It is optional to refer to the cameras as front and rear, and the computing device 100 may include multiple cameras located on various sides of the main unit 102.
[0025] The display 106 may represent a cathode ray tube (CRT) display, a light-emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light-emitting diode (OLED) display, or any other type of display known in the art. In some embodiments, the display 106 may display a digital representation of the current image captured by the front camera 104 and / or the rear camera 112, an image that may be captured by one or more of these cameras, a recently captured image by one or more of these cameras, and / or a modified version of one or more of these images. Thus, the display 106 may function as a camera viewfinder. The display 106 may also support touchscreen functionality, which may allow adjustment of settings and / or configurations of one or more aspects of the computing device 100.
[0026] The front camera 104 may include an image sensor and associated optical elements such as a lens. The front camera 104 may provide a zoom function or may have a fixed focal length. In other embodiments, interchangeable lenses may be used with the front camera 104. The front camera 104 may have a variable mechanical aperture and a mechanical shutter and / or an electronic shutter. The front camera 104 may also be configured to capture still images, video images, or both. Furthermore, the front camera 104 may represent, for example, a monocular camera, a stereoscopic camera, or a multi-lens camera. The rear camera 112 may be arranged similarly or differently. Furthermore, one or more of the front camera 104 and / or the rear camera 112 may be an array of one or more cameras.
[0027] One or more of the front camera 104 and / or rear camera 112 may include, or be associated with, an illumination component that provides a light field for illuminating the object of interest. For example, the illumination component may provide flash or constant illumination of the object of interest. The illumination component may also be configured to provide a light field that includes one or more of structured light, polarized light, and light having specific spectral components. In the context of the embodiments herein, it is known to reconstruct three-dimensional (3D) models from objects, and other types of light fields used in this manner are possible.
[0028] The computing device 100 may also include an ambient light sensor that can continuously or intermittently determine the ambient brightness of the scene that cameras 104 and / or 112 can capture. In some embodiments, the ambient light sensor may be used to adjust the display brightness of the display 106. Furthermore, the ambient light sensor may be used to determine, or facilitate, the exposure length of one or more of the cameras 104 or 112.
[0029] The computing device 100 may be configured to capture images of a target object using the display 106 and the front camera 104 and / or rear camera 112. The captured images may be multiple still images or a video stream. Image capture may be triggered by activating button 108, pressing a soft key on the display 106, or by some other mechanism. Depending on the embodiment, for example, when button 108 is pressed, images may be captured automatically at specific time intervals according to a predetermined capture schedule, or by moving the computing device 100 a predetermined distance under favorable lighting conditions for the target object.
[0030] Figure 2 is a simplified block diagram showing some of the components of an exemplary computing system 200. Examples, but not limited to, of the computing system 200 may include a cellular mobile phone (e.g., a smartphone), a computer (desktop, notebook, tablet, server, or handheld computer, etc.), home automation components, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a game console, a robotic device, a vehicle, or any other type of device. The computing system 200 may, for example, represent an embodiment of computing device 100.
[0031] As shown in Figure 2, the computing system 200 may include a communication interface 202, a user interface 204, a processor 206, data storage 208, and a camera component 224, all of which may be linked together communicatively by a system bus, network, or other connection mechanism 210. The computing system 200 may be equipped with at least some image capture and / or image processing functions. It should be understood that the computing system 200 may represent a physical image processing system, or a specific physical hardware platform on which an image sensing and / or processing application operates in software, or any other combination of hardware and software configured to perform image capture and / or processing functions.
[0032] The communication interface 202 may enable the computing system 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, the communication interface 202 may facilitate conventional telephone service (POTS) communication and / or circuit-switched and / or packet-switched communications such as Internet Protocol (IP), or other packetized communications. For example, the communication interface 202 may include a chipset and antenna positioned for wireless communication with a wireless access network or access point. The communication interface 202 may also take the form of a wired interface, such as an Ethernet®, Universal Serial Bus (USB), or High Resolution Multimedia Interface (HDMI®) port, or may include a wired interface. The communication interface 202 may also take the form of a wireless interface, such as Wi-Fi, Bluetooth®, Global Positioning System (GPS), or Wide Area Wireless Interface (e.g., WiMAX® or 3GPP® Long-Term Evolution (LTE)), or may include a wireless interface. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used via the communication interface 202. Furthermore, the communication interface 202 may include multiple physical communication interfaces (e.g., Wi-Fi interface, Bluetooth® interface, and wide-area wireless interface).
[0033] The user interface 204 may function to enable the computing system 200 to interact with a human or non-human user, such as by receiving input from the user and providing output to the user. Thus, the user interface 204 may include input components such as a keypad, keyboard, touch-sensitive panel, computer mouse, trackball, joystick, and microphone. The user interface 204 may also include one or more output components, such as a display screen which may be combined with a touch-sensitive panel. The display screen may be based on CRT, LCD, LED, and / or OLED technology, or other currently known or future technologies. The user interface 204 may also be configured to generate audible output(s) via speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. The user interface 204 may also be configured to receive and / or capture audible speech(s), noise(s), and / or signals(s) via microphones and / or other similar devices.
[0034] In some embodiments, the user interface 204 may include a display that functions as a viewfinder for still camera and / or video camera functions supported by the computing system 200. Furthermore, the user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of camera functions, as well as image capture. Some or all of these buttons, switches, knobs, and / or dials may be implemented by touch-sensitive panels.
[0035] The processor 206 may comprise one or more general-purpose processors, such as microprocessors, and / or one or more dedicated processors, such as digital signal processors (DSPs), graphics processing units (GPUs), floating-point units (FPUs), network processors, or application-specific integrated circuits (ASICs). In some cases, the dedicated processors may be capable of image processing, image alignment, and image merging, among many other possibilities. The data storage 208 may comprise one or more volatile and / or non-volatile storage components, such as magnetic storage, optical storage, flash storage, or organic storage, and may be integrated with the processor 206 in whole or in part. The data storage 208 may comprise removable and / or non-removable components.
[0036] The processor 206 may be capable of executing program instructions 218 (e.g., compiled or uncompiled program logic and / or machine code) stored in the data storage 208 to perform various functions described herein. Thus, the data storage 208 may include a non-temporary computer-readable medium on which the program instructions are stored, and when executed by the computing system 200, the program instructions cause the computing system 200 to perform any of the methods, processes, or operations disclosed herein and / or in the accompanying drawings. By executing the program instructions 218, the processor 206 may use the data 212.
[0037] For example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more application programs 220 (e.g., camera functions, address book, email, web browsing, social networking, audio-to-text conversion functions, text translation functions, and / or game applications) installed on the computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. The operating system data 216 may be primarily accessible by the operating system 222, and the application data 214 may be primarily accessible by one or more of the application programs 220. The application data 214 may be located in a file system visible to the user of the computing system 200, or in a file system hidden from the user.
[0038] The application program 220 may communicate with the operating system 222 via one or more application programming interfaces (APIs). These APIs may facilitate, for example, the application program 220 reading and / or writing application data 214, sending or receiving information via the communication interface 202, and receiving and / or displaying information on the user interface 204.
[0039] In some cases, the application program 220 may be referred to simply as "the app." Furthermore, the application program 220 may be downloadable to the computing system 200 via one or more online application stores or application marketplaces. However, the application program can also be installed on the computing system 200 by other means, such as via a web browser or through the physical interface of the computing system 200 (e.g., a USB port).
[0040] The camera component 224 may include, but is not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or image sensor), lens, shutter button, infrared projector, and / or visible light projector. Among the many possibilities, the camera component 224 may include a component configured to capture images of the visible light spectrum (e.g., electromagnetic radiation with wavelengths of 380 to 700 nanometers) and / or a component configured to capture images of the infrared light spectrum (e.g., electromagnetic radiation with wavelengths of 701 nanometers to 1 millimeter). The camera component 224 may be controlled, at least partially, by software executed by the processor 206.
[0041] Figure 3 shows Figure 300 illustrating the training phase 302 and inference phase 304 of a trained machine learning model(s) 332 according to an exemplary embodiment. Some machine learning techniques involve training one or more machine learning algorithms on an input set of training data in order to recognize patterns in the training data and provide output inferences and / or predictions about the patterns in the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, Figure 3 shows a training phase 302 in which one or more machine learning algorithms 320 are trained on training data 310 to become a trained machine learning model 332. Generating a trained machine learning model(s) 332 during the training phase 302 may include determining one or more hyperparameters, such as one or more stride values for one or more layers of the machine learning model, as described herein. Next, during the inference phase 304, the trained machine learning model 332 receives input data 330 and one or more inference / prediction requests 340 (possibly as part of the input data 330) and, in response, may provide one or more inferences and / or predictions 350 as outputs. One or more inferences and / or predictions 350 may be based in part on one or more learned hyperparameters, such as one or more learned stride values for one or more layers of the machine learning model, as described herein.
[0042] Therefore, a trained machine learning model(s) 332 may include one or more models of one or more machine learning algorithms 320. The machine learning algorithm(s) 320 may include, but are not limited to, artificial neural networks (e.g., convolutional neural networks, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, appropriate statistical machine learning algorithms, and / or heuristic machine learning systems as described herein). The machine learning algorithm(s) 320 may implement supervised or unsupervised learning, and may implement any appropriate combination of online and offline learning.
[0043] In some embodiments, machine learning algorithms 320 and / or trained machine learning models 332 may be accelerated using on-device coprocessors such as graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), and / or application-specific integrated circuits (ASICs). Using such on-device coprocessors, machine learning algorithms 320 and / or trained machine learning models 332 may be accelerated. In some embodiments, trained machine learning models 332 may be trained, exist, and run to provide inference on a particular computing device, and / or may perform inference in other ways for a particular computing device.
[0044] During training phase 302, machine learning algorithms 320 may be trained by providing at least training data 310 as training input, using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning includes providing some (or all) of the training data 310 to machine learning algorithms 320, and machine learning algorithms 320 determining one or more output inferences based on the provided portion (or all) of the training data 310. Supervised learning includes providing some (or all) of the training data 310 to machine learning algorithms 320, and machine learning algorithms 320 determining one or more output inferences based on the provided portion of the training data 310, where the output inferences are either accepted or corrected based on the correct results related to the training data 310. In some embodiments, supervised learning of a machine learning algorithm(s) 320 may be managed by a set of rules and / or labels about the training input, and the inference of the machine learning algorithm(s) 320 may be corrected using the set of rules and / or labels.
[0045] Semi-supervised learning involves having correct results for some, but not all, of the training data 310. During semi-supervised learning, supervised learning is used for the portion of the training data 310 for which correct results are obtained, and unsupervised learning is used for the portion of the training data 310 for which correct results are not obtained.
[0046] Reinforcement learning involves a machine learning algorithm(s) 320 receiving a reward signal for a previous inference, which may be numerical. During reinforcement learning, the machine learning algorithm(s) 320 may output an inference and receive a reward signal in response, and the machine learning algorithm(s) 320 is configured to attempt to maximize the numerical value of the reward signal. In some embodiments, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some embodiments, the machine learning algorithm(s) 320 and / or the trained machine learning model(s) 332 may be trained using other machine learning techniques, including but not limited to incremental learning and curriculum learning.
[0047] In some embodiments, the machine learning algorithm(s) 320 and / or the pre-trained machine learning model(s) 332 may use transfer learning techniques. For example, the transfer learning technique may include the pre-trained machine learning model(s) 332 being pre-trained on a dataset and then further trained using training data 310. More specifically, the machine learning algorithm(s) 320 may be pre-trained on data from one or more computing devices, and the resulting pre-trained machine learning model may be provided to computing device CD1, which is intended to run the pre-trained machine learning model during inference phase 304. Then, during training phase 302, the pre-trained machine learning model may be further trained using training data 310. This further training of the machine learning algorithm(s) 320 and / or the pre-trained machine learning model using training data 310 on CD1 may be performed using either supervised or unsupervised learning. Once the machine learning algorithm(s) 320 and / or the pre-trained machine learning model have been trained on at least the training data 310, training phase 302 may be completed. The resulting machine learning model can be used as at least one of the 332 trained machine learning models.
[0048] Specifically, once training phase 302 is complete, the trained machine learning model(s) 332 may be provided to the computing device if they do not already exist on the computing device. After the trained machine learning model(s) 332 have been provided to computing device CD1, inference phase 304 may begin.
[0049] During the inference phase 304, the trained machine learning model(s) 332 may receive input data 330 and generate and output one or more corresponding inferences and / or predictions 350 with respect to the input data 330. Thus, the input data 330 may be used as input to the trained machine learning model(s) 332 in order to provide corresponding inferences and / or predictions 350. For example, the trained machine learning model(s) 332 may generate inferences and / or predictions 350 in response to one or more inference / prediction requests 340. In some embodiments, the trained machine learning model(s) 332 may be run by part of other software. For example, the trained machine learning model(s) 332 may be run by an inference or prediction daemon so that it is readily available to provide inferences and / or predictions on demand. The input data 330 may include data from computing device CD1 running a trained machine learning model(s) 332, and / or input data from one or more computing devices other than CD1.
[0050] The exemplary image capture devices described herein may include, among many components, one or more cameras and sensors. The image capture device may be a smartphone, tablet, laptop, or digital camera, among many types of computing devices capable of performing the operations described herein.
[0051] As an embodiment, a computing device may include one or more processors having logic for executing instructions, at least one built-in or peripheral image sensor (e.g., a camera), and an input / output device (e.g., a display panel) for displaying a user interface. The computer device may further include computer-readable media (CRM). The CRM may include any suitable memory or storage device, such as random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), non-volatile RAM (NVRAM), read-only memory (ROM), or flash. The computing device stores device data (e.g., user data, multimedia data, applications, and / or the device's operating system) in the CRM. The device data may include executable instructions for an auto-zoom process. The auto-zoom process may be part of a process that the operating system performs on the image capture device, or it may be a process performed by a separate component within an application environment (e.g., a camera application) or within a “framework” provided by the operating system.
[0052] Computing devices may implement machine learning techniques ("visual splendor models"). Visual splendor models may be implemented using one or more of the following: support vector machines (SVMs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), dense neural networks (DNNs), one or more heuristics, other machine learning techniques, and combinations thereof. Visual splendor models may be trained iteratively outside the device, given training scenes, training sequences, and / or training events. For example, training may involve providing the visual splendor model with images (e.g., digital photographs) that include bounding boxes of user drawings surrounding visual splendor regions (e.g., regions where one or more objects of particular interest to the user may exist). In some embodiments, these images may include bounding boxes or heatmaps, which are generated by tracking the annotator's eyes while the annotator views the image to determine which areas of the image are most splendid. Furthermore, in some embodiments, the visual splendor model may be trained with heatmaps that would have been generated by tracking the positions of the image the annotator was viewing. For example, if five annotators look at a first area and ten look at a second area, the first area in the image may be determined to have half the splendor of the second area. Providing an image with user-drawn bounding boxes can facilitate training of the visual splendor model to identify visually splendor regions within the image. As a result of training, the visual splendor model may generate a visual splendor heatmap of a given image. The computing device can then generate various bounding boxes, including those enclosing the region with the highest visual splendor score, based on the visual splendor heatmap. In this way, the visual splendor model can predict visually splendor regions within an image. After sufficient training, model compression using distillation may be performed on the visual splendor model, which allows for the selection of an optimal model architecture based on the model's latency and power consumption.Subsequently, the visual spleness model can be deployed as an independent module in the CRM of a computing device, or implemented in an automated zoom process.
[0053] The computing device may perform an automatic zoom process, possibly automatically or in response to a received trigger signal, such as a user gesture (e.g., tap, press) made on an input / output device. The computing device may receive one or more captured images from an image sensor.
[0054] Computing devices may utilize visual splendor models to generate a visual splendor heatmap using one or more captured images.
[0055] For example, Figure 4a is an image 400 according to an exemplary embodiment. Figure 4b is a heatmap 450 according to an exemplary embodiment. As shown in Figures 4a and 4b, a computing device may utilize a visual splendor model to generate a visual splendor heatmap of a captured image. One or more processors may compute the visual splendor heatmap in the background operation of the device. In some embodiments, the image capture device may not display the visual splendor heatmap to the user. As illustrated, the visual splendor heatmap represents the magnitude of visual splendor probability on a scale from black to white, where white indicates a high splendor probability and black indicates a low splendor probability. Each pixel in the visual splendor heatmap may be assigned a splendor metric, which represents the degree of splendor of the area represented by the metric. This is based on the computing device applying a pre-trained machine learning model to the image frame, so that the pre-trained machine learning model outputs a splendor metric for each pixel in the visual splendor heatmap.
[0056] A visual spleness model can generate a bounding box that encloses the region where the visual spleness probability is highest. Figure 5 shows a heatmap 500 with a bounding box 502 according to an exemplary embodiment.
[0057] As shown in Figure 5, the visual splendor heatmap includes bounding boxes 502 surrounding the areas in the image where the visual splendor probability is highest. A visual splendor model may be trained to output a heatmap, and a computing device may use the heatmap to determine one or more objects of interest in a captured image and to generate one or more bounding boxes around these objects of interest. These generated bounding boxes may contain one or more areas predicted by the computing device to have splendor.
[0058] In some embodiments, the computing device may experiment with anchor boxes of various sizes and positions, as described in Faster R-CNN by Ren et al. in 2016. Specifically, as described herein, anchor boxes may be considered to determine regions of interest, e.g., areas with high or best mean splendor values. For example, Figure 6A shows anchor bounding boxes of sizes 602, 604, and 606 according to an exemplary embodiment. As shown in Figure 6A, the computing device may evaluate anchor bounding boxes of different aspect ratios (e.g., 1:2, 1:1, and 2:1), and the computing device may evaluate anchor bounding boxes of various sizes for each of the different aspect ratios. The computing device may determine the mean splendor value for each aspect ratio and size of anchor bounding box with respect to various anchor bounding box positions. Specifically, each pixel in the heatmap may be associated with a splendor value, and the computing device may determine the average splendor value of the pixels in the anchor bounding box to determine the mean splendor value.
[0059] Figure 6B shows the anchor bounding box locations according to an exemplary embodiment. As shown in Figure 6B, a computing device can evaluate the average splendor of anchor bounding boxes of various sizes and aspect ratios every few pixels. The centers of the anchor bounding boxes may be spaced equally based on stride values, and at each location, the computing device can evaluate various sizes and aspect ratios of anchor bounding boxes. For example, among several locations, the computing device can determine the average splendor of nine anchor bounding boxes of sizes 602, 604, and 606 with aspect ratios of 1:2, 1:1, and 2:1 at locations 650, 652, and 654, separated by a stride of 3 pixels. Based on the average splendor of the anchor bounding boxes, the computing device can determine one or more regions of interest (ROIs).
[0060] Figure 7 shows a prominent ROI according to an exemplary embodiment. In the exemplary process, the computing device may determine the splendor heatmap 700 according to the process described above, possibly using a machine learning model or other algorithms that predict the splendor metric at each pixel in the image frame. Based on the splendor heatmap 700, the computing device may determine the primary ROI with the highest average splendor by calculating the average splendor value of pixels in various anchor bounding boxes, possibly as described above in relation to Figures 6A-6B.
[0061] Next, the computing device may determine a secondary ROI based on the primary ROI, where the secondary ROI may have pixels with a lower mean splendor compared to the primary ROI. The secondary ROI may need to be at least a threshold distance away from the primary ROI. As shown in Figure 704, the computing device may determine a secondary ROI based on a restricted zone 714. The restricted zone 714 may be the area around the primary ROI 712, within which the secondary ROI cannot be located. The computing device may determine a secondary ROI 716 such that it does not overlap with the primary ROI 712 or the restricted zone 714. In some embodiments, the computing device may apply the aforementioned anchor bounding box method to pixels outside the primary ROI 712 and the restricted zone 714 to determine the secondary ROI 716. Specifically, the primary ROI 712 may be the region with the highest mean splendor, and the secondary ROI 716 may be the region with the second-highest mean splendor, subject to the aforementioned constraints.
[0062] The computing system may then determine a more stable filtered ROI based on the primary ROI 712 and secondary ROI 716, which may then be used in one or more autofocus processes. For example, Figure 8 shows an image 800 having ROIs 802, 804, and 806 according to an exemplary embodiment. Image 800 may include the primary ROI 806, the secondary ROI 802, and the previously filtered ROI 804.
[0063] The primary ROI 806 may include the most prominent area within the image frame, which the computing device may determine using the aforementioned anchor bounding box method. The secondary ROI 802 may include an area within the image frame that is less prominent than the primary ROI 806. The previous filtered ROI 804 may consist of areas within previous image frames that the computing device previously determined based on the previous primary ROI.
[0064] Based on the primary ROI 806 and the previous filtered ROI 804, the computing device may determine a new filtered ROI for an area within the image frame. The filtered ROI may then be used to apply one or more autofocus processes. For example, the computing device may determine the confidence values of the previous filtered ROI 804 and the primary ROI 806. Specifically, the computing device may determine the confidence value of the previous filtered ROI 804 based on the average spleness value of the previous filtered ROI calculated from the spleness heatmap of the image frame, even though the previous filtered ROI 804 is associated with the previous image frame. The confidence value of the primary ROI may be based on the average spleness value for each ROI calculated from the spleness heatmap generated from the image frame. The computing device may determine that the spleness difference between the previous filtered ROI 804 and the primary ROI 806 exceeds a threshold (for example, the confidence value of the primary ROI 806 exceeds the confidence value of the previous filtered ROI 804 by a threshold). Based on this determination, the computing device may update the filtered ROI from the previous filtered ROI 804 to the primary ROI 806. Additionally and / or alternatively, if the computing device determines that the difference in significance between the previous filtered ROI 804 and the primary ROI 806 does not exceed a threshold, the computing device may retain the previous filtered ROI.
[0065] In some alternative embodiments, the computing device may also consider the secondary ROI 802 and determine that the significance difference between the previous filtered ROI 804 and the secondary ROI 802 exceeds a threshold (for example, the confidence value of the secondary ROI 802 exceeds the confidence value of the previous filtered ROI 804 by a threshold), and the computing device may update the filtered ROI from the previous filtered ROI 804 to the secondary ROI 802. Other considerations may also be taken into account, such as the stability of the primary ROI, secondary ROI, and / or the previous filtered ROI.
[0066] In some embodiments, a computing device may determine a filtered ROI based on the amount of overlap between the primary ROI 806 and / or secondary ROI 802 with the previously filtered ROI 804. Specifically, the computing system may determine a first overlap amount of the previously filtered ROI 804 with respect to the primary ROI 806, and a second overlap amount of the secondary ROI 802 with respect to the previously filtered ROI 804. If both the first and second overlap amounts do not exceed a threshold, the computing device may update the filtered ROI to either the primary ROI 806 or the secondary ROI 802 based on which of the primary ROI 806 and the secondary ROI 802 is associated with a larger mean stellar value. It may be shown that if the overlap is small (e.g., the overlap does not exceed a threshold), the impact of updating the filtered ROI may be large, while if the overlap is large (e.g., the overlap exceeds a threshold), the impact of updating the filtered ROI may be small. When a filtered ROI has a large overlap with another location, updating the filtered ROI to that other location can degrade the user experience because the area of focus may change rapidly over time. In further embodiments, both the difference in significance and the amount of overlap may be considered when deciding whether to update the filtered ROI.
[0067] In some embodiments, a computing system may decide whether to update a filtered ROI from a primary ROI to a secondary ROI based on whether the previous filtered ROI overlaps with a primary ROI or a secondary ROI. For example, a computing device may update a filtered ROI to a secondary ROI if the difference in significance between the primary and secondary ROIs does not exceed a threshold, and the secondary ROI overlaps with the previous ROI, or overlaps by at least a threshold amount. Alternatively, if the difference in significance between the primary and secondary ROIs exceeds a threshold, the computing device may decide to update the filtered ROI to a primary ROI.
[0068] After determining the filtered ROI, the computing device may apply one or more autofocus processes to the image frame. In some embodiments, applying one or more autofocus processes may include adjusting the camera lens so that the lens focuses on the filtered ROI. In further embodiments, the computing device may apply blur to areas of the image frame outside the filtered ROI, possibly by artificially blurring the background of the image frame, so that the focus of the image frame becomes an area within the filtered ROI.
[0069] As described above, in some embodiments, splendor detection can become unstable over time and may report splendor regions that are not actually splendor. For example, a computing system may alternate between two regions with nearly equivalent splendor, which can result in periodic and / or random out-of-focus image frames. Out-of-focus image frames can make it difficult to run further algorithms (e.g., a classifier that detects objects in an image to make the image easier to search) and can degrade the user experience.
[0070] To facilitate the determination of the image frame to focus on, the computing device may decide whether to apply one or more autofocus processes based on the filtered ROI being associated with a particular state of a finite state machine. Figure 9 shows a finite state machine 900 according to an exemplary embodiment. The finite state machine 900 may help to avoid instantaneous misjudgments.
[0071] As shown in Figure 9, the finite state machine 900 includes a committed state 902, a pending state 904, a standby state 906, and a probation state 908. The committed state 902 may indicate that the filtered ROI is available; the pending state 904 may indicate that the filtered ROI is awaiting stability verification; the probation state 908 may indicate that the filtered ROI is available but awaiting unsuccessful stability verification; and the standby state 906 indicates that the filtered ROI is unavailable. In the standby state 906, the computing device may choose not to proceed with applying one or more autofocus processes based on the ROI, whereas in the committed state 902 or probation state 908, the computing device may choose to proceed with applying one or more autofocus processes. In the pending state 904, the computing device may choose not to proceed with applying one or more autofocus processes until it is verified that the ROI is stable or unstable.
[0072] For an image frame containing primary ROIs, secondary ROIs, and filtered ROIs, the computing device may assign a state to each ROI, and this state may be updated each time a new primary ROI, secondary ROI, and / or filtered ROI is determined. For example, if a primary ROI, secondary ROI, or filtered ROI is associated with a reliability measurement (e.g., mean splendor) that does not exceed a threshold, the computing device may update the state of each ROI from pending state 904 to standby state 906, from committed state 902 to probation state 908, and / or from probation state 908 to standby state 906.
[0073] Furthermore, if a primary ROI, secondary ROI, or filtered ROI is inconsistent and / or inconsistent with any of the previous primary ROI, previous secondary ROI, and / or previous filtered ROI, the computing device may update the state of each ROI from pending state 904 to standby state 906, from committed state 902 to probation state 908, and / or from probation state 908 to standby state 906. In some embodiments, consistency of a primary ROI (or secondary ROI, or filtered ROI) may be defined as having an overlap in threshold amounts with the previous primary ROI and not having too abrupt changes in depth.
[0074] Furthermore, if the confidence value (e.g., the determined mean spleness value of pixels within the ROI) is above the threshold and the consistency is also adequate (e.g., the amount of overlap, perhaps among several factors, is above the threshold), the computing device may update the state of each ROI from pending state 904 to committed state 902, or from probation state 908 to committed state 902.
[0075] In some embodiments, filtered ROIs, primary ROIs, and secondary ROIs may function as object detectors, with each ROI representing one or more objects. However, filtered ROIs, primary ROIs, and / or secondary ROIs do not necessarily detect objects; rather, they may detect the most prominent location within the image frame. For example, filtered ROIs, primary ROIs, and / or secondary ROIs may also track off-center objects next to a textured wall, or groups of people in the background, as long as these subjects are most prominent. Thus, a computing system performing the method described herein may output an off-center focus area over a textured wall or a group of people in the background, and the focus area is moving, but not necessarily following an object.
[0076] Furthermore, a computing system performing the method described herein may be able to rapidly switch focus from one area represented by one ROI to another area represented by another ROI, since the computing system can identify multiple ROIs. Thus, if an object that was located in the filtered ROI no longer exists and the filtered ROI is no longer very prominent, the computing system may rapidly switch focus to the primary or secondary ROI. Moreover, if an image frame contains two objects of similar prominence in the primary and secondary ROIs, a computing device performing the method described herein may maintain stability on the primary ROI with respect to a certain number of image frames, perhaps before switching to the secondary ROI, rather than continuously switching between the primary and secondary ROIs.
[0077] Figure 10 shows a finite state machine manager 1000 according to an exemplary embodiment. The finite state manager 1000 includes a splendor ROI state machine 1002 and a splendor ROI state machine 1004. The finite state manager 1000 may prepare candidate ROIs (e.g., primary ROIs, secondary ROIs, previously filtered ROIs, or filtered ROIs) to input into either the splendor ROI state machine 1002 or the splendor ROI state machine 1004 by verifying several factors. For example, if a candidate ROI is already associated with a splendor ROI state machine, the candidate ROI cannot be input into the other splendor ROI state machine. A candidate ROI may be discarded if the objects inside the candidate ROI are not within a valid distance range. A candidate ROI may also be discarded if it is not within a valid window close to the center of the image.
[0078] Having two splendor ROI state machines can help facilitate switching between applying one or more autofocus algorithms to one area of an image frame and applying them to other areas of the image frame. For example, if an object that is in focus within an image frame disappears from the frame for a short time and that object was associated with splendor ROI state machine 1002, the computing device can check whether the state of the candidate ROI region associated with splendor ROI state machine 1004 is acceptable. If the state is acceptable, the computing device can quickly switch to focusing on the candidate ROI.
[0079] Figure 11 is a flowchart of Method 1100 according to an exemplary embodiment. Method 1100 may be performed by one or more computing systems (e.g., computing system 200 in Figure 2) and / or one or more processors (e.g., processor 206 in Figure 2). Method 1100 may be performed by a computing device such as computing device 100 in Figure 1.
[0080] Method 1100 includes receiving an image frame captured by an image capture device in block 1102.
[0081] Method 1100 includes determining a spleness heatmap in block 1104 that represents the spleness of pixels within an image frame.
[0082] Method 1100 includes determining the primary and secondary ROIs of an image frame based on a spleness heatmap in block 1106.
[0083] Method 1100 includes determining the filtered ROI of the image frame in block 1108. The filtered ROI is updated from the previous filtered ROI to the primary ROI based on the sampling difference between the previous filtered ROI and the primary ROI exceeding a first threshold.
[0084] Method 1100 includes applying one or more autofocus processes in block 1110 based on at least one of a filtered ROI, a primary ROI, or a secondary ROI.
[0085] In some embodiments, determining the primary and secondary ROIs is based on the assumption that the primary ROI has a greater mean significance than the secondary ROI.
[0086] In some embodiments, the filtered ROI is updated from the previous filtered ROI to the primary ROI based on the fact that both the first overlap amount of the previous filtered ROI with respect to the primary ROI and the second overlap amount of the previous filtered ROI do not exceed a second threshold.
[0087] In some embodiments, once a filtered ROI is set to a previous filtered ROI, the filtered ROI is associated with an updated mean spiciness value based on the spiciness heatmap.
[0088] In some embodiments, determining primary and secondary ROIs based on a splendor heatmap involves determining a number of candidate anchor boxes distributed on the splendor heatmap, each of which is associated with a splendor measurement, and determining primary and secondary ROIs based on the respective splendor measurements of each candidate anchor box.
[0089] In some embodiments, the multiple candidate anchor boxes include multiple anchor boxes having multiple different aspect ratios at a given position within the image frame.
[0090] In some embodiments, candidate anchor boxes are evenly distributed on the splendor heatmap.
[0091] In some embodiments, the multiple candidate anchor boxes include multiple anchor boxes having multiple sizes at a given position within the image frame.
[0092] In some embodiments, based on the sampling heatmap, the primary ROI is associated with the primary reliability measurement, the previous filtered ROI is associated with the filtered reliability measurement, and the sampling difference is based on the primary reliability measurement and the filtered reliability measurement.
[0093] In some embodiments, the primary reliability measurement is based on the average of one or more splendor values at one or more pixels within the primary ROI, and the filtered reliability measurement is based on the average of one or more splendor values at one or more pixels within the previous filtered ROI.
[0094] In some embodiments, determining the primary and secondary ROIs is based on the assumption that the primary ROI is at least a threshold distance away from the secondary ROI.
[0095] In some embodiments, determining primary and secondary ROIs includes determining the primary ROI based on a saturation heatmap, determining a restricted area around the primary ROI, and determining the secondary ROI based on the restricted area around the primary ROI and the saturation heatmap such that the secondary ROI is not located within the primary ROI or within the restricted area around the primary ROI.
[0096] In some embodiments, the previous filtered ROI is based on the previous image frame captured before the current image frame.
[0097] In some embodiments, determining a spleness heatmap representing the spleness of each pixel in an image frame involves applying a pre-trained machine learning model to the image frame to determine the spleness heatmap.
[0098] In some embodiments, applying one or more autofocus processes includes causing the camera lens to adjust its focus to the filtered ROI.
[0099] In some embodiments, applying one or more autofocus processes includes applying blur to areas of the image frame outside the filtered ROI.
[0100] In some embodiments, Method 1100 further includes applying a finite state machine to the filtered ROI, and applying one or more autofocus processes is based on the fact that the filtered ROI is associated with a particular state of the finite state machine.
[0101] In some embodiments, the finite state machine includes a committed state indicating that the filtered ROI is available, a pending state indicating that the filtered ROI is awaiting stability verification, a probation state indicating that the filtered ROI is available but awaiting failure of stability verification, and a standby state indicating that the filtered ROI is unavailable.
[0102] In some embodiments, Method 1100 further includes updating the state associated with the filtered ROI, which includes updating the state from pending to standby, from committed to probation, or from probation to standby, based on the determination that the confidence measurement associated with the filtered ROI does not exceed a second threshold.
[0103] In some embodiments, method 1100 further includes updating the state associated with the filtered ROI, which includes updating the state from pending to standby, from committed to probation, or from probation to standby, based on the determination that the filtered ROI does not overlap with a previous filtered ROI.
[0104] In some embodiments, a particular state of a finite state machine is a committed state.
[0105] In some embodiments, Method 1100 further includes applying a finite state machine to each of the filtered ROI, primary ROI, and secondary ROI, and applying one or more autofocus processes based on the respective states of the finite state machines associated with each of the filtered ROI, primary ROI, and secondary ROI.
[0106] In some embodiments, Method 1100 is performed by an image capture device including a camera and a control system configured to perform the steps of Method 1100.
[0107] In such embodiments, the image capture device is a mobile device, and the image frame is captured by a camera.
[0108] In some embodiments, the method may include receiving an image frame captured by an image capture device. The method may also include determining a splendor heatmap representing the splendor of pixels within the image frame. The method may further include determining a primary ROI based on the splendor heatmap. The method may further include determining a forbidden area around the primary ROI. The method may also include determining a secondary ROI based on the splendor heatmap such that the secondary ROI does not overlap with the primary ROI and does not overlap with the forbidden area. The method may further include controlling one or more autofocus processes based on the primary and secondary ROIs.
[0109] III. Conclusion This disclosure is not limited to the specific embodiments described in this application, and these embodiments are intended to be illustrative of various aspects. As will be apparent to those skilled in the art, many modifications and variations may be made without departing from the scope. In addition to the methods and apparatus described herein, functionally equivalent methods and apparatus included within the scope of this disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims.
[0110] The above detailed description, with reference to the accompanying drawings, illustrates various features and operations of the disclosed systems, devices, and methods. In the drawings, unless otherwise indicated by context, similar reference numerals generally identify similar components. The exemplary embodiments described herein and in the drawings are not intended to be limiting. Other embodiments may be used and other modifications may be made without departing from the scope of the subject matter presented herein. It will be readily apparent that the aspects of this disclosure described herein and shown in the drawings may be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0111] With respect to any or all of the message flow diagrams, scenarios, and flowcharts shown in the drawings and discussed herein, each step, each block, and / or each communication may represent the processing and / or transmission of information according to exemplary embodiments. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, actions described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be performed in an order different from that shown or discussed, including substantially simultaneous or reversed orders, depending on the functions involved. Furthermore, more or fewer blocks and / or actions may be used in any of the message flow diagrams, scenarios, and flowcharts discussed herein, and these message flow diagrams, scenarios, and flowcharts may be combined with each other in part or as a whole.
[0112] A step or block representing the processing of information may correspond to a circuit that can be configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a block representing the processing of information may correspond to a module, segment, or portion of program code (including associated data). The program code may include one or more instructions that can be executed by a processor to perform a specific logical operation or action in the method or technique. The program code and / or associated data may be stored in any type of computer-readable medium, such as a storage device including random access memory (RAM), disk drives, solid-state drives, or other storage media.
[0113] Computer-readable media may also include non-temporary computer-readable media such as register memory, processor cache, and RAM, which store data for short periods. Computer-readable media may also include non-temporary computer-readable media that store program code and / or data for long periods. Therefore, computer-readable media may include secondary storage or persistent long-term storage such as read-only memory (ROM), optical or magnetic disks, solid-state drives, and compact disk read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage systems. Computer-readable media may be considered, for example, computer-readable storage media or tangible storage devices.
[0114] Furthermore, a step or block representing one or more information transmissions may correspond to information transmissions between software modules and / or hardware modules within the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.
[0115] The specific arrangements shown in the drawings should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in the given drawings. Furthermore, some of the exemplary elements may be combined or omitted. Moreover, exemplary embodiments may include elements not shown in the drawings.
[0116] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are illustrative and not intended to limit, and the true scope is shown by the following claims.
Claims
1. It is a method, Receiving image frames captured by an image capture device, Determining a saturation heatmap that represents the saturation of pixels within the aforementioned image frame, Based on the aforementioned sampling heatmap, the primary region of interest (ROI) and secondary ROI of the image frame are determined. The method includes determining the filtered ROI of the image frame, wherein the filtered ROI is updated from the previous filtered ROI to the primary ROI based on the significance difference between the previous filtered ROI and the primary ROI exceeding a first threshold, and the method further includes Applying one or more autofocus processes based on at least one of the filtered ROI, the primary ROI, or the secondary ROI, Methods that include...
2. The method according to claim 1, wherein determining the primary ROI and the secondary ROI is based on the fact that the primary ROI has a greater mean significantity than the secondary ROI.
3. The method according to claim 1, wherein the filtered ROI is updated from the previous filtered ROI to the primary ROI based on the fact that the amount of overlap of the previous filtered ROI with respect to the primary ROI does not exceed a second threshold.
4. The method according to claim 1, wherein when the filtered ROI is set to the previous filtered ROI, the filtered ROI is associated with the updated mean splendor value based on the splendor heatmap.
5. Determining the primary ROI and the secondary ROI based on the sampling heatmap includes determining a plurality of candidate anchor boxes distributed on the sampling heatmap. Each of the candidate anchor boxes is associated with a significantness measurement, The method according to claim 1, wherein determining the primary ROI and the secondary ROI is based on the respective significantity measurements of the candidate anchor boxes.
6. The method according to claim 5, wherein the plurality of candidate anchor boxes include a plurality of anchor boxes having a plurality of different aspect ratios at a given position within the image frame.
7. The method according to claim 5, wherein the candidate anchor boxes are evenly distributed on the saturation heatmap.
8. The method according to claim 5, wherein the plurality of candidate anchor boxes include a plurality of anchor boxes having a plurality of sizes at a given position within the image frame.
9. Based on the aforementioned sampling heatmap, the primary ROI is associated with the primary reliability measurement, and the previous filtered ROI is associated with the filtered reliability measurement. The method according to claim 1, wherein the significant difference is based on the primary reliability measurement value and the filtered reliability measurement value.
10. The method according to claim 9, wherein the primary reliability measurement is based on the average of one or more significantity values in one or more pixels in the primary ROI, and the filtered reliability measurement is based on the average of one or more significantity values in one or more pixels in the previous filtered ROI.
11. The method according to claim 1, wherein determining the primary ROI and the secondary ROI is based on the fact that the primary ROI is at least a threshold distance away from the secondary ROI.
12. Determining the primary ROI and the secondary ROI is Based on the aforementioned sampling heatmap, the primary ROI is determined, Determining the restricted area around the aforementioned primary ROI, Based on the prohibited area around the primary ROI and the significance heatmap, the secondary ROI is determined such that it does not exist within the primary ROI or within the prohibited area around the primary ROI. The method according to claim 1, including the method described in claim 1.
13. The method according to claim 1, wherein the previously filtered ROI is based on a previous image frame captured before the image frame.
14. The method according to claim 1, wherein determining a spleness heatmap representing the spleness of each pixel in the image frame comprises applying a pre-trained machine learning model to the image frame to determine the spleness heatmap.
15. The method according to claim 1, wherein applying one or more autofocus processes includes causing a camera lens to focus on the filtered ROI.
16. The method according to claim 1, wherein applying one or more autofocus processes includes applying blur to the region of the image frame outside the filtered ROI.
17. The method further includes applying a finite state machine to the filtered ROI, The method according to claim 1, wherein applying one or more autofocus processes is based on the fact that the filtered ROI is associated with a specific state of the finite state machine.
18. The method according to claim 17, wherein the finite state machine includes a committed state indicating that the filtered ROI is available; a pending state indicating that the filtered ROI is awaiting stability verification; a probation state indicating that the filtered ROI is available but awaiting failure of stability verification; and a standby state indicating that the filtered ROI is unavailable.
19. The method further includes updating the status associated with the filtered ROI, The method according to claim 18, wherein updating the state associated with the filtered ROI includes updating the state from the pending state to the standby state, from the committed state to the probation state, or from the probation state to the standby state, based on the determination that the reliability measurement associated with the filtered ROI does not exceed a second threshold.
20. The method further includes updating the status associated with the filtered ROI, The method according to claim 18, wherein updating the state associated with the filtered ROI includes updating the state from the pending state to the standby state, from the committed state to the probation state, or from the probation state to the standby state, based on the determination that the filtered ROI does not overlap with the previous filtered ROI.
21. The method according to claim 17, wherein the particular state of the finite state machine is a committed state.
22. The method further includes applying a finite state machine to each of the filtered ROI, the primary ROI, and the secondary ROI, The method according to any one of claims 1 to 21, wherein applying one or more autofocus processes is based on the respective states of the finite state machine associated with the filtered ROI, the primary ROI, and the secondary ROI, respectively.
23. It is an image capture device, Camera and, A control system is provided, and the control system is Receiving an image frame captured by the aforementioned image capture device, Determining a saturation heatmap that represents the saturation of pixels within the aforementioned image frame, Based on the aforementioned sampling heatmap, the primary region of interest (ROI) and secondary ROI of the image frame are determined. The control system is configured to determine the filtered ROI of the image frame, and the filtered ROI is updated from the previous filtered ROI to the primary ROI based on the fact that the difference in sampling between the previous filtered ROI and the primary ROI exceeds a first threshold, and the control system further performs Applying one or more autofocus processes based on at least one of the filtered ROI, the primary ROI, or the secondary ROI, An image capture device configured to perform the following actions.
24. The image capture device according to claim 23, wherein the image capture device is a mobile device, and the image frame is captured by the camera.
25. A program that stores instructions executable by one or more processors, wherein the instructions cause the one or more processors to perform an operation, and the operation is Receiving image frames captured by an image capture device, Determining a saturation heatmap that represents the saturation of pixels within the aforementioned image frame, Based on the aforementioned sampling heatmap, the primary region of interest (ROI) and secondary ROI of the image frame are determined. The operation includes determining the filtered ROI of the image frame, wherein the filtered ROI is updated from the previous filtered ROI to the primary ROI based on the significance difference between the previous filtered ROI and the primary ROI exceeding a first threshold, and the operation further includes Applying one or more autofocus processes based on at least one of the filtered ROI, the primary ROI, or the secondary ROI, A program that includes this.