Bokeh effect in variable aperture (VA) camera systems

JP2025513696A5Pending Publication Date: 2026-03-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

The prior art is difficult to realize simulation effects of different aperture sizes in the variable aperture optical system, which limits the flexibility of background blur processing on images.

Method used

By receiving image data of different aperture sizes, corresponding depth maps and focus maps are generated, and combined with the technology of simulating aperture size, an output image with simulated aperture effect is generated.

Benefits of technology

The effect of simulating different aperture sizes in the variable aperture optical system is realized, which enhances the flexibility and quality of image processing, especially in portrait photography, which can simulate the background blur effect of large apertures more naturally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides a system, method, and device for image processing that supports enhancement of image effects, such as bokeh effects, applied in image processing. In a first aspect, a method of image processing includes determining a depth map corresponding to a first scene based on first image data and second image data captured with different aperture sizes, determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, and determining an output image frame based on the focus map, the first image data, and the second image data. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Aspects of the present disclosure relate generally to image processing, and more particularly to applying effects to captured image data, such as background blurring to recreate a bokeh effect. Several features can enable and provide improved image processing, including improved photographic appearance of subjects and objects in scenes having objects at multiple depths.

[0002] introduction An image capture device is a device capable of capturing one or more digital images, whether a still image for a photograph or a sequence of images for a video. Capture devices can be incorporated into a wide variety of devices. By way of example, image capture devices may include stand-alone digital cameras or digital video camcorders; camera-equipped wireless communication device handsets, such as mobile phones, cellular or satellite radio phones, personal digital assistants (PDAs), panels or tablets, gaming devices, etc.; computing devices, such as webcams, video surveillance cameras, etc.; or other devices with digital imaging or video capabilities.

[0003] In a particular scene, a photographer may want the viewer to focus on one part of the scene. For example, in a portrait photo of a person, the photographer may want the viewer to focus on the person and not the rest of the scenery. The photographer may choose a low aperture lens for such a photo, because a low aperture results in objects at different depths than the person being significantly blurred. A smaller aperture lens produces more blur than a larger aperture lens. However, a smaller aperture lens is generally larger in size and made from more costly materials. Summary of the Invention

[0004] The following summarizes some aspects of the present disclosure in order to provide a basic understanding of the discussed technology. This summary is not an extensive overview of all of the contemplated features of the present disclosure, and is not intended to identify key or critical elements of all aspects of the present disclosure or to delineate the scope of any or all aspects of the present disclosure. Its sole purpose is to present some concepts of one or more aspects of the present disclosure in summary form as a prelude to the more detailed description that is presented later.

[0005] Computational photography can be used to process image data to apply effects to a scene, such as applying a bokeh effect that mimics the blurring that traditionally occurs with lenses with small apertures. In a variable aperture (VA) camera system, blurring at aperture sizes different from those available in the VA camera system may be desired. Image processing can be applied to obtain aperture sizes different from those available in the VA camera system or aperture sizes different from those used to capture the image data. Visual effects on a photograph, such as blurring to replicate a bokeh effect, can be applied through image processing based on image data captured at multiple aperture sizes by controlling the aperture of the VA camera system.

[0006] In one aspect of the disclosure, a method for image processing includes receiving first image data captured with a first aperture size and second image data captured with a second aperture size, where each of the first image data and the second image data represents a first scene; determining a depth map corresponding to the first scene based on the first image data and the second image data, where the depth map includes a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, where the focus map includes a second value indicative of an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and an object in the first scene; and determining an output image frame based on the focus map, the first image data, and the second image data.

[0007] In an additional aspect of the disclosure, an apparatus includes at least one processor and a memory coupled to the at least one processor, the at least one processor configured to perform operations including receiving first image data captured with a first aperture size and second image data captured with a second aperture size, each of the first image data and the second image data representing a first scene, determining a depth map corresponding to the first scene based on the first image data and the second image data, the depth map including a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene, determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, the focus map including a second value indicative of an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and an object in the first scene, and determining an output image frame based on the focus map, the first image data, and the second image data.

[0008] In an additional aspect of the disclosure, an apparatus includes means for receiving first image data captured with a first aperture size and second image data captured with a second aperture size, each of the first image data and the second image data representing a first scene; means for determining a depth map corresponding to the first scene based on the first image data and the second image data, the depth map including a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; means for determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, the focus map including a second value indicative of an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and an object in the first scene; and means for determining an output image frame based on the focus map, the first image data, and the second image data.

[0009] In an additional aspect of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform an operation. These operations include receiving first image data captured with a first aperture size and second image data captured with a second aperture size, each of the first image data and the second image data representing a first scene; determining a depth map corresponding to the first scene based on the first image data and the second image data, the depth map including a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, the focus map including a second value indicative of an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and an object in the first scene; and determining an output image frame based on the focus map, the first image data, and the second image data.

[0010] An image capture device is a device capable of capturing one or more digital images, whether still image photographs or a sequence of images for a video, and may be incorporated within a wide variety of devices. By way of example, an image capture device may include a stand-alone digital camera or digital video camcorder; a camera-equipped wireless communication device handset, such as a mobile phone, cellular phone or satellite radio phone, personal digital assistants (PDAs), panels or tablets, gaming devices, etc.; a computing device, such as a webcam, video surveillance camera, etc.; or other device with digital imaging or video capabilities.

[0011] Generally, this disclosure describes image processing techniques for digital cameras having image sensors and image signal processors (ISPs). The ISP can be configured to control the capture of image frames from one or more image sensors and to process one or more image frames from the one or more image sensors to generate a view of a scene in a corrected image frame. The corrected image frame can be part of a sequence of image frames forming a video sequence. The video sequence can include other image frames received from the image sensor or other image sensors and / or other corrected image frames based on input from the image sensor or other image sensors. In some embodiments, the processing of one or more image frames can be performed in the image sensor, such as in a binning module. The image processing techniques described in the embodiments disclosed herein can be performed by circuitry in the image sensor, such as a binning module, circuitry in an image signal processor (ISP), circuitry in an application processor (AP), or a combination of these components, or two or all of these components.

[0012] In one embodiment, the image signal processor may receive instructions to capture a sequence of image frames in response to loading of software, such as a camera application, to generate a preview display from an image capture device. The image signal processor may be configured to generate a single flow of output frames based on image frames received from one or more image sensors. This single flow of output frames may include raw image data from the image sensor, binned image data from the image sensor, or corrected image frames that have been processed by one or more algorithms, such as in a binning module, within the image signal processor. For example, image frames acquired from an image sensor, where some processing may have been performed on the data before being output to the image signal processor, may be processed within the image signal processor by processing the image frames through an image post-processing engine (IPE) and / or other image processing circuitry to perform one or more of tone mapping, portrait lighting, contrast enhancement, gamma correction, etc.

[0013] After an output frame representative of a scene is determined by the image signal processor using image corrections such as binning, as described in various embodiments herein, the output frame may be displayed on a device display as a single still image and / or as part of a video sequence, may be saved to a storage device as a picture or video sequence, may be transmitted over a network, and / or may be printed on an output medium. For example, the image signal processor may be configured to obtain input frames of image data (e.g., pixel values) from various image sensors and then generate corresponding output frames of image data (e.g., preview display frames, still image capture, frames for video, frames for object tracking, etc.). In other examples, the image signal processor may output frames of image data to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., auto focus (AF), auto white balance (AWB), and auto exposure control (AEC)), generating a video file via the output frames, composing frames for display, composing frames for storage, transmitting frames over a network connection, etc. That is, the image signal processor can acquire incoming frames from one or more image sensors, each coupled to one or more camera lenses, and then generate a flow of output frames to be output to various destinations.

[0014] In some aspects, the corrected image frame can be generated by combining aspects of the image correction of the present disclosure with other computational photography techniques, such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). For HDR photography, the first and second image frames are captured using different exposure times, different apertures, different lenses, and / or other different characteristics, which may result in an improved dynamic range of the fused image when the two image frames are combined. In some aspects, the method can be performed with respect to MFNR photography, where the first and second image frames are captured using the same or different exposure times and are fused to generate a corrected first image frame having reduced noise compared to the captured first image frame.

[0015] In some aspects, the device may include an image signal processor or processor (e.g., an application processor) that includes certain functions related to camera control and / or processing, such as enabling or disabling a binning module or otherwise controlling aspects of image correction. The methods and techniques described herein may be performed entirely by the image signal processor or processor, or various operations may be divided between the image signal processor and a processor, or in some aspects across additional processors.

[0016] The apparatus may include one, two, or more image sensors, such as including a first image sensor. If there are multiple image sensors, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have a different sensitivity or different dynamic range than the second image sensor. In one embodiment, the first image sensor may be a wide-angle image sensor and the second image sensor may be a telephoto image sensor. In another embodiment, the first sensor is configured to acquire images through a first lens having a first optical axis, and the second sensor is configured to acquire images through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification and the second lens may have a second magnification different from the first magnification. This configuration may occur using a lens cluster on the mobile device, such as when multiple image sensors and associated lenses are located at offset locations on the front or back of the mobile device. Additional image sensors with larger, smaller, or the same field of view may also be included. The image correction techniques described herein may be applied to image frames captured from any of the image sensors in a multi-sensor device.

[0017] In an additional aspect of the present disclosure, a device configured for image processing and / or image capture is disclosed. The device includes a means for capturing an image frame. The device further includes one or more means for capturing data representative of a scene, such as image sensors (including charge-coupled devices (CCDs), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal oxide semiconductor (CMOS) sensors), time-of-flight detectors, etc. The device may further include one or more means for integrating and / or focusing light rays into the one or more image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components can be controlled to capture a first image frame and / or a second image frame, which are input to the image processing techniques described herein.

[0018] Other aspects, features, and implementations will become apparent to those skilled in the art upon reviewing the following description of certain exemplary aspects in conjunction with the accompanying figures. Although features may be discussed in conjunction with certain aspects and figures below, various aspects may include one or more of the advantageous features discussed herein. In other words, although one or more aspects may be discussed as having certain advantageous features, one or more of such features may also be used in accordance with various aspects. Similarly, although exemplary aspects may be discussed below as device, system, or method aspects, the exemplary aspects may be implemented in a variety of devices, systems, and methods.

[0019] The method can be embedded in a computer-readable medium as computer program code including instructions that cause a processor to perform the steps of the method. In some embodiments, the processor can be part of a mobile device including a first network adapter configured to transmit data, such as an image or video, as recorded data or streaming data over a first network connection of a plurality of network connections, a processor coupled to the first network adapter, and a memory. The processor can cause transmission of the corrected image frames described herein over a wireless communication network, such as a 5G NR communication network.

[0020] The foregoing has outlined rather broadly the features and technical advantages of the embodiments according to the present disclosure in order that the following Detailed Description may be better understood. Additional features and advantages are described hereinafter. The concepts and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The concepts disclosed herein, both their organization and method of operation, characteristic of the concepts disclosed herein, together with associated advantages, will be better understood in consideration of the following description in conjunction with the accompanying figures. Each of the figures is provided for the purpose of illustration and description, and not as a definition of the limits of the claims.

[0021] Although aspects and implementations are described in this application by way of example for some embodiments, those skilled in the art will appreciate that additional implementations and use cases may occur in many different configurations and scenarios. The innovations described herein may be implemented across many different platform types, devices, systems, shapes, sizes, and packaging configurations. For example, aspects and / or applications may occur via integrated chip implementations and other non-modular component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, artificial intelligence (AI)-enabled devices, etc.). Some embodiments may or may not be specifically targeted to a use case or application, but a wide variety of combination applicability of the described innovations may occur. Implementations may range from chip-level or modular components to non-modular, non-chip-level implementations, and even aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical settings, devices incorporating the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, transmitting and receiving wireless signals necessarily includes a number of components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, summers / analog summers, etc.). It is contemplated that the innovations described herein may be practiced in a wide variety of devices, chip-level components, systems, distributed configurations, end-user devices, etc., of various sizes, shapes, and configurations.

[0022] A further understanding of the nature and advantages of the present disclosure can be realized by referring to the following drawings. In the accompanying figures, similar components or features may have the same reference label. Furthermore, various components of the same type can be distinguished by following the reference label with a dash and a second label that distinguishes the similar components. When only a first reference label is used in this specification, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. [Brief description of the drawings]

[0023] [Figure 1] 1 illustrates a block diagram of an example device for performing image capture from one or more image sensors in accordance with one or more aspects of the present disclosure. [Diagram 2] 1 illustrates a block diagram of an exemplary processing configuration for applying a bokeh effect, in accordance with one or more aspects of the present disclosure. [Diagram 3] 1 illustrates a flowchart of an exemplary method for applying a bokeh effect with a variable aperture (VA) camera system, in accordance with one or more aspects of the present disclosure. [Figure 4] 1 illustrates a flowchart of an exemplary method for applying a bokeh effect using a sharpness value, in accordance with one or more aspects of the present disclosure. [Diagram 5] 1 illustrates a block diagram of an exemplary processing configuration for applying a bokeh effect using a sharpness value, in accordance with one or more aspects of the present disclosure. [Figure 6] 1 illustrates a flowchart of an example method for applying a bokeh effect with a variable aperture (VA) camera system using machine learning, in accordance with one or more aspects of the present disclosure. [Figure 7] 1 illustrates a block diagram of an example processing configuration for applying a bokeh effect using machine learning, in accordance with one or more aspects of the present disclosure.

[0024] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] The detailed description of the present invention, set forth below in conjunction with the accompanying drawings, is intended as an illustration of various configurations and is not intended to limit the scope of the present disclosure. Rather, the detailed description of the present invention includes specific details intended to provide a thorough understanding of the subject matter of the present invention. Those skilled in the art will appreciate that these specific details are not required in every instance and that in some instances, well-known structures and components are shown in block diagram form for clarity of presentation.

[0026] The present disclosure provides systems, apparatus, methods, and computer-readable media that support improved image processing for applying effects to image data captured by a camera system, such as a variable aperture (VA) camera system. The improved image processing can apply blur to a scene captured by image data from multiple aperture sizes by using image data captured with two (or more) different aperture sizes to generate a focus map that can be combined with a predetermined relationship for the aperture sizes.

[0027] Particular implementations of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages or benefits: In some aspects, the present disclosure provides techniques for improving the appearance of images by creating more natural effects, such as bokeh effects, that provide an appearance more similar to natural effects obtained through optical systems. This technique may enable camera systems with smaller available aperture sizes to reproduce photographic effects available in camera systems with larger available aperture sizes, which are generally more expensive and less portable.

[0028] An exemplary device for capturing image frames using one or more image sensors, such as a smartphone, may include a configuration of two, three, four, or more cameras on the back (e.g., opposite the user display) or front (e.g., on the same side as the user display) of the device. A device with multiple image sensors includes one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing images captured by the image sensors. The one or more image signal processors can provide the processed image frames to a memory and / or processor (such as an application processor, an image front end (IFE), an image processing engine (IPE), or other suitable processing circuitry) for further processing, such as for encoding, storage, transmission, or other manipulation.

[0029] As used herein, an image sensor may refer to the image sensor itself as well as certain other components coupled to the image sensor that are used to generate an image frame for processing by an image signal processor or other logic circuitry, or for storage in memory, whether a short-term buffer or longer-term non-volatile memory. For example, an image sensor may include other components of a camera, including shutters, buffers, or other readout circuitry for accessing individual pixels of the image sensor. An image sensor may also refer to analog front-end or other circuitry for converting analog signals into a digital representation of the image frame that is provided to digital circuitry coupled to the image sensor.

[0030] In the following description, numerous specific details are set forth, such as examples of specific components, circuits, and processes, to provide a thorough understanding of the present disclosure. The term "coupled" as used herein means directly connected or connected through one or more intervening components or circuits. In addition, in the following description, for the purpose of explanation, specific terminology is set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that these specific details may not be required to practice the teachings disclosed herein. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the teachings of the present disclosure.

[0031] Some portions of the following Detailed Description are presented in terms of procedures, logic blocks, processes, and other symbolic representations of operations on data bits within a computer memory. In this disclosure, a procedure, logic block, process, etc., is conceived to be a self-consistent sequence of steps or instructions leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system.

[0032] In the figures, a single block may be described as performing one function or multiple functions. The function or functions performed by the block may be performed in a single component or across multiple components, and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various example components, blocks, modules, circuits, and steps are generally described below in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each specific application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure. Also, the example device may include components other than those shown, including well-known components such as a processor, memory, etc.

[0033] Aspects of the present disclosure are applicable to any electronic device that includes or is coupled to two or more image sensors capable of capturing image frames (or "frames"). Moreover, aspects of the present disclosure can be implemented in devices having or coupled to image sensors of the same or different capabilities and characteristics (resolution, shutter speed, sensor type, etc.). Furthermore, aspects of the present disclosure can be implemented in devices for processing image frames, such as processing devices capable of retrieving stored images for processing, including processing devices present in cloud computing systems, regardless of whether the device includes or is coupled to an image sensor.

[0034] Unless otherwise indicated, and as will be apparent from the discussion that follows, discussions utilizing terms such as "accessing," "receiving," "sending," "using," "selecting," "determining," "normalizing," "multiplying," "averaging," "monitoring," "comparing," "applying," "updating," "measuring," "deriving," "solving," "generating," and the like throughout this application will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities in the computer system's registers and memory, and converts such data to other data that is similarly represented as physical quantities in the computer system's registers, memory, or other such information storage, transmission, or display devices.

[0035] The terms "device" and "apparatus" are not limited to one physical object (such as one smartphone, one camera controller, one processing system, etc.) or a particular number of physical objects. As used herein, a device can be any electronic device having one or more parts capable of implementing at least some parts of the present disclosure. Although the following description and examples use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. As used herein, an apparatus may include a device for performing the described operations, or a portion of such a device.

[0036] 1 illustrates a block diagram of an exemplary device 100 for performing image capture from one or more image sensors. The device 100 may include or be otherwise coupled to an image signal processor 112 for processing image frames from one or more image sensors, such as a first image sensor 101, a second image sensor 102, and a depth sensor 140. In some implementations, the device 100 also includes or is coupled to a processor 104 and a memory 106 that stores instructions 108. The device 100 may also include or be coupled to a display 114 and input / output (I / O) components 116. The I / O components 116, such as a touch screen interface and / or physical buttons, may be used to interact with a user. The I / O components 116 may also include network interfaces for communicating with other devices, including a wide area network (WAN) adapter 152, a local area network (LAN) adapter 153, and / or a personal area network (PAN) adapter 154. An exemplary WAN adapter is a 4G LTE or 5G NR wireless network adapter. An exemplary LAN adapter 153 is an IEEE 802.11 WiFi wireless network adapter. An exemplary PAN adapter 154 is a Bluetooth wireless network adapter. Each of the adapters 152, 153, and / or 154 may be coupled to an antenna, including multiple antennas configured for primary and diversity reception and / or configured to receive a particular frequency band. The device 100 may further include or be coupled to a power source 118 for the device 100, such as a battery or components for coupling the device 100 to an energy source. The device 100 may also include or be coupled to additional features or components not shown in FIG. 1 .In one embodiment, a wireless interface, which may include a number of transceivers and a baseband processor, may be coupled to or included within WAN adapter 152 for a wireless communication device. In a further embodiment, an analog front end (AFE) for converting analog image frame data to digital image frame data may be coupled between image sensor 101 and image sensor 102 and image signal processor 112.

[0037] The device may include or be coupled to a sensor hub 150 for interfacing with sensors for receiving data regarding the movement of the device 100, data regarding the environment surrounding the device 100, and / or other non-camera sensor data. One exemplary non-camera sensor is a gyroscope, which is a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another exemplary non-camera sensor is an accelerometer, which is a device configured to measure acceleration, which can also be used to determine the speed and distance of movement by appropriately integrating the measured acceleration, and one or more of the acceleration, speed, and / or distance can be included in the generated motion data. In some aspects, a gyroscope in an electronic image stabilization system (EIS) can be coupled to the sensor hub or directly to the image signal processor 112. In another example, the non-camera sensor can be a global positioning system (GPS) receiver.

[0038] The image signal processor 112 can receive image data, such as that used to form an image frame. In one embodiment, a local bus connection couples the image signal processor 112 to the image sensor 101 of the first camera and the image sensor 102 of the second camera. In another embodiment, a wired interface couples the image signal processor 112 to an external image sensor. In a further embodiment, a wireless interface couples the image signal processor 112 to the image sensors 101, 102.

[0039] The first camera may include a first image sensor 101 and a corresponding first lens 131. The second camera may include a second image sensor 102 and a corresponding second lens 132. Each of the lenses 131 and 132 may be controlled by an associated auto-focus (AF) algorithm 133 running in the ISP 112, which adjusts the lenses 131 and 132 to focus on a particular focal plane at a particular scene depth from the image sensors 101 and 102. The AF algorithm 133 may be assisted by a depth sensor 140.

[0040] The first image sensor 101 and the second image sensor 102 are configured to capture one or more image frames. The lenses 131 and 132 focus light onto the image sensors 101 and 102, respectively, through one or more apertures for receiving light, one or more shutters for blocking light when outside an exposure window, one or more color filter arrays (CFAs) for filtering light other than a certain frequency range, one or more analog front ends for converting analog measurements to digital information, and / or other suitable components for imaging. The first lens 131 and the second lens 132 may have different fields of view for capturing different representations of a scene. For example, the first lens 131 may be an ultra-wide angle (UW) lens and the second lens 132 may be a wide angle (W) lens. The multiple image sensors may include a combination of ultra-wide angle (high field of view (FOV)), wide angle, telephoto, and super telephoto (low FOV) sensors. That is, each image sensor can be configured via hardware configuration and / or software settings to provide different but overlapping fields of view. In one configuration, the image sensors are configured using different lenses with different magnifications resulting in different fields of view. The sensors can be configured such that the UW sensor has a larger FOV than the W sensor, which has a larger FOV than the T sensor, which has a larger FOV than the UT sensor. For example, a sensor configured for a wide FOV can capture a field of view ranging from 64 to 84 degrees, a sensor configured for an ultra-wide FOV can capture a field of view ranging from 100 to 140 degrees, a sensor configured for a telephoto FOV can capture a field of view ranging from 10 to 30 degrees, and a sensor configured for a super telephoto FOV can capture a field of view ranging from 1 to 8 degrees.

[0041] Image signal processor 112 processes image frames captured by image sensor 101 and image sensor 102. Although FIG. 1 illustrates device 100 as including two image sensors 101 and 102 coupled to image signal processor 112, any number of image sensors (e.g., one, two, three, four, five, six, etc.) may be coupled to image signal processor 112. In some aspects, a depth sensor, such as depth sensor 140, may be coupled to image signal processor 112, and output from the depth sensor may be processed in a similar manner as the output of image sensor 101 and image sensor 102. Furthermore, any number of additional image sensors or image signal processors may be present for device 100.

[0042] In some embodiments, the image signal processor 112 can execute instructions from a memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included within the image signal processor 112, or instructions provided by the processor 104. Additionally or alternatively, the image signal processor 112 can include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, the image signal processor 112 can include one or more image front ends (IFEs) 135, one or more image post-processing engines 136 (IPEs), and / or one or more automatic exposure correction (AEC) 134 engines. The AF 133, AEC 134, AFE 135, and APE 136 can each include application specific circuitry and can be embodied as software code executed by the ISP 112 and / or as a combination of hardware in the ISP 112 and software code executed on the ISP 112. The ISP 112 may further implement an automatic white balancing (AWB) engine for performing white balancing operations. The AWB engine may be implemented in the ISP 112 or other dedicated or general-purpose processing circuitry within the image capture device 100, such as in the image front ends (IFEs) 135 or on a digital signal processor (DSP).

[0043] In some implementations, memory 106 may include a non-transient or non-transitory computer-readable medium having stored thereon computer-executable instructions 108 for performing all or a portion of one or more operations described in this disclosure. In some implementations, instructions 108 include a camera application (or other suitable application) to be executed by device 100 to generate images or videos. Instructions 108 may also include other applications or programs executed by device 100, such as an operating system and specific applications other than for image or video generation. Execution of the camera application, such as by processor 104, may cause device 100 to generate images using image sensors 101 and 102 and image signal processor 112. Memory 106 may also be accessed by image signal processor 112 to store processed frames or by processor 104 to obtain processed frames. In some embodiments, device 100 does not include memory 106. For example, device 100 can be a circuit that includes image signal processor 112, and the memory can be external to device 100. Device 100 can be coupled to external memory and configured to access the memory to write output frames for display or long-term storage. In some embodiments, device 100 is a system-on-chip (SoC) that incorporates image signal processor 112, processor 104, sensor hub 150, memory 106, and input / output components 116 in a single package.

[0044] In some embodiments, at least one of the image signal processor 112 or the processor 104 executes instructions to perform various operations described herein, including noise reduction operations. For example, execution of the instructions may instruct the image signal processor 112 to begin or end capturing an image frame or a sequence of image frames, including noise reduction as described in the embodiments herein. In some embodiments, the processor 104 may include one or more general-purpose processor cores 104A capable of executing scripts or instructions of one or more software programs, such as instructions 108 stored in the memory 106. For example, the processor 104 may include one or more application processors configured to execute a camera application (or other suitable application for generating images or videos) stored in the memory 106.

[0045] When executing the camera application, the processor 104 can be configured to instruct the image signal processor 112 to perform one or more operations related to the image sensor 101 or image sensor 102. For example, the camera application can receive a command to initiate a video preview display in which a video including a sequence of image frames from one or more image sensors 101 or image sensor 102 is captured and processed. Image correction, such as by cascaded IPE, can be applied to one or more image frames in the sequence. Execution of instructions 108 outside of the camera application by the processor 104 can also cause the device 100 to perform any number of functions or operations. In some embodiments, the processor 104 can include ICs or other hardware (e.g., an artificial intelligence (AI) engine 124) in addition to the ability to execute software to cause the device 100 to perform any number of functions or operations, such as those described herein. In some other embodiments, the device 100 does not include the processor 104, such as when all of the described functions are configured in the image signal processor 112.

[0046] In some embodiments, display 114 may include one or more suitable displays or screens that allow for user interaction and / or allow for presenting items to a user, such as previews of image frames being captured by image sensor 101 and image sensor 102. In some embodiments, display 114 is a touch-sensitive display. I / O components 116 may be or include any suitable mechanisms, interfaces, or devices for receiving input (e.g., commands) from a user and for providing output to a user via display 114. For example, I / O components 116 may include (without limitation) a graphical user interface (GUI), a keyboard, a mouse, a microphone, a speaker, a compressible bezel, one or more buttons (e.g., a power button), sliders, switches, etc.

[0047] Although shown coupled together via the processor 104, the components (such as the processor 104, memory 106, image signal processor 112, display 114, and I / O components 116) can be coupled together in various other configurations, such as via one or more local buses, not shown for simplicity. Although the image signal processor 112 is shown as separate from the processor 104, the image signal processor 112 can also be a core of the processor 104, an application processor unit (APU), included in a system on chip (SoC) or otherwise included with the processor 104. Although the device 100 is referred to in the examples herein for carrying out aspects of the present disclosure, some device components may not be shown in FIG. 1 to avoid obscuring aspects of the present disclosure. Moreover, other components, multiple components, or combinations of components can be included in a device suitable for carrying out aspects of the present disclosure. Thus, the present disclosure is not limited to any particular device or configuration of components, including the device 100.

[0048] 2 illustrates a block diagram of an exemplary processing configuration for applying a bokeh effect according to one or more aspects of the present disclosure. The processor 104 of the system 200 can communicate with an image signal processor (ISP) 112 via a bidirectional bus and / or separate control and data lines. The processor 104 can control the camera 202 via a camera control block 212 of the ISP 112 to acquire first image data at a first aperture size and second image data at a second aperture size. For example, if the camera 202 is a variable aperture (VA) camera system, the processor 104 can execute a camera application to instruct the camera 202 to configure a first aperture size to acquire the first image data from the camera 202 and to instruct the camera 202 to configure a second aperture size to acquire the second image data from the camera 202. This aperture reconfiguration and acquisition of the first and second image data can be performed with little or no change present in the scene captured at the first and second aperture sizes. Exemplary aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture numbers correspond to smaller aperture sizes and smaller aperture numbers correspond to larger aperture sizes, i.e., f / 2.0 is a larger aperture size than f / 8.0.

[0049] Image data received from the camera 202 may be processed in one or more blocks 214 of the ISP 112 and provided to the processor 104. The processor 104 may further process the image data to apply effects based on image data captured at multiple aperture sizes. In some embodiments, applying a bokeh effect using image data captured at multiple aperture sizes may provide a more realistic bokeh effect that more closely resembles the bokeh effect obtained by a camera lens having a desired aperture size. For example, first image data captured at a first aperture size of f / 2.0 and second image data captured at a second aperture size of f / 2.4 may be used to apply a bokeh effect with a simulated aperture size of f / 1.4 that is not obtainable with image data from a single aperture size.

[0050] The processor 104 may include a focus map generator 222 and a blur applicator 224. The focus map generator 222 may determine a desired amount of focus for portions of a scene, which is determined to match the blur effect for a desired aperture size. The focus map may be based on a depth map that describes the distance of objects in the scene from the camera 202. The depth map may be determined from image data acquired at two or more aperture sizes. The focus map generator 222 provides the focus map to the blur applicator 224, which also receives the image data collected at two or more aperture sizes. The blur applicator 224 determines output image frames 204 based on the focus map and the image data acquired at the various aperture sizes. The output image frames may be stored in the memory 106 and then recorded in another memory, such as a non-volatile memory, transmitted via a network adapter, and / or output to a display device. Although not shown, other logic and memory circuits may be present in the system 200. For example, a buffer may exist between the processor 104 and the ISP 112, or between the ISP 112 and the camera 202. As another example, additional processing circuitry may exist within the processor 104 to apply additional effects to the image data (e.g., lighting, color casting, high dynamic range (HDR) merging). In some embodiments, the processing circuitry providing the functionality of the blur applicator 224 may be reconfigured to provide the functionality of the additional effects. In some embodiments, the processing circuitry providing the functionality of the focus map generator 222 and / or the blur applicator 224 may be incorporated within a different component, such as the ISP 112, a DSP, an ASIC, or other custom logic.

[0051] The system 200 of FIG. 2 may be configured to perform the operations described with reference to FIG. 3 to determine an output image frame. FIG. 3 illustrates a flowchart of an example method for applying a bokeh effect with a variable aperture (VA) camera system according to one or more aspects of the present disclosure. The method 300 includes, at block 302, receiving first image data captured with a first aperture size and second image data captured with a second aperture size. The first image data and the second image data may represent the same scene, such as by being captured closely in time, such as consecutively, from the same or similar viewpoint, such that objects in the scene are present in the same or similar location in the first image data and the second image data.

[0052] The first image data and the second image data may be captured by a variable aperture (VA) camera system. In a variable aperture (VA) camera system, block 302 may include configuring the camera 202 with a first aperture size to acquire the first image data, and then subsequently configuring the camera 202 with a second aperture size to acquire the second image data. In other camera systems with multiple fixed aperture sizes, block 302 may include capturing image data from multiple cameras, each with a different aperture size. Capturing image data from a single camera system with variable aperture (VA) capabilities or other techniques for adjusting the amount of light reaching the image sensor may be more desirable than multiple camera systems. For example, multiple camera systems may have different fields of view, different alignments, and different light and color sensitivities. These differences may need to be compensated for before merging image data from multiple camera systems. If the image data is acquired from a single camera system, these alignment and other challenges may be mitigated or avoided.

[0053] In block 304, a depth map is determined. A depth map may be a set of values ​​arranged in an array or arrays (e.g., a matrix or a table), each value representing a portion of a scene and one or more pixels in that portion of the scene. Thus, image data organized in an image frame of N×M pixels may have a corresponding depth map of K×L values, where N is an integer multiple of K and M is an integer multiple of L. These values ​​indicate the distance between an image sensor that records the image data and an object in the first scene. The depth map may be determined from the first image data and the second image data received in block 302, as described with reference to FIG. 4, FIG. 5, FIG. 6, or FIG. 7. For example, the depth map may be a blur map indicating the amount of blur for a corresponding portion of the scene. A depth map as described with respect to block 304 may include any representation of depth, regardless of whether the units of the values ​​in the depth map are distance or not.

[0054] At block 306, a focus map can be determined based on the depth map and the simulated aperture size. The simulated aperture size can be a desired aperture size for representing the scene and can be input by a user to a camera application running on the processor 104. The focus map can use predefined information about the camera system at the simulated aperture size that reflects the sharpness of the image at various depths. By combining values ​​in the depth map with the predefined information, a value of the focus map can be determined that indicates the amount of blur at the simulated aperture size for a corresponding distance in the depth map between the image sensor and an object in the scene.

[0055] At block 308, an output image frame is determined from the focus map and the first image data. For example, the output image frame can be determined by applying a blurring algorithm to the first image data with a blurring intensity based on a corresponding portion of the focus map. In some embodiments, the output image frame is also based on the second image data. For example, the output image frame can be determined by blending the first image data with corresponding data of the second image data based on the focus map, where weights assigned to the first image data and the second image data during blending are based on the focus map. The resulting output image frame includes portions that are blurred based on the focus map to obtain a photograph corresponding to a simulated aperture size. The blurred portions can provide the photograph with an appearance that reflects the blurring produced by a lens having an aperture size different from either the first aperture size or the second aperture size of block 402. The larger aperture size can be larger than the aperture size available with the camera system that acquired the first image data and the second image data, allowing for image characteristics similar to those of more expensive, larger camera systems to be obtained without the additional cost and size of a larger aperture camera system.

[0056] Some embodiments of the image processing described in Fig. 3 are based on an image-based contrast calculation as described with reference to Fig. 4. Fig. 4 shows a flow chart of an exemplary method for applying a bokeh effect using a sharpness value according to one or more aspects of the present disclosure. A block diagram illustrating the image data processing of Fig. 4 is shown in Fig. 5. Fig. 5 shows a block diagram of an exemplary processing configuration for applying a bokeh effect using a sharpness value according to one or more aspects of the present disclosure.

[0057] Method 400 includes, at block 402, receiving first image data 502 captured with a first aperture size and second image data 504 captured with a second aperture size. The first image data and second image data may be received after capture, such as described with reference to block 302 of FIG. 3. The operations of method 400 may also be performed in real time with the capture of the first image data and the second image data, for example, to generate a preview image on a display. The operations of method 400 may also, or alternatively, be performed on the image capture device and / or by a remote device after capture of the first image data and the second image data.

[0058] In block 404, a blurriness map is determined based on the first sharpness of the first image data and the second sharpness of the second image data. The blurriness map may indicate the depth of the object from the camera. The first sharpness may be a first sharpness map 512 determined from the first image data 502, and the second sharpness may be a second sharpness map 514 determined from the second image data 504. Similar to the depth map described above, image data organized in an image frame of N×M pixels may have a corresponding sharpness map of K×L values, where N is an integer multiple of K and M is an integer multiple of L. The values ​​in the sharpness map may indicate the distance between the image sensor recording the image data and the object in the first scene, since parts of the scene closer to the focus of the camera are sharper. The sharpness may measure the level of clarity of image details. In some embodiments, sharpness may be determined as a function of its Laplacian, normalized by the local average luminance at the surrounding pixels according to the following formula:

[0059]

number

[0060] In block 406, a focus map may be determined based on the blurriness map and the blurriness lookup table. A focus map generator 224 may receive the blurriness map 516 and a blurriness lookup table (LUT) 518 (or other predetermined relationship between aperture size and depth). The LUT 518 may provide a relationship 518A and a relationship 518B that map blurriness to depth for various aperture sizes. The relationship 518A may correspond to a smaller aperture size than the relationship 518B based on its lower blurriness at higher depths. The LUT 518 may include one or more relationships that may be stored as a table, although alternatively or additionally, the predetermined relationship may be stored as a representative formula. The focus map generator 224 determines a focus map 520 by mapping the blurriness map to a focus map using the relationships in the LUT 518 that correspond to the simulated aperture sizes. In some embodiments, the relationship for the simulated aperture size may not be present in the LUT 518, and the relationship may instead be interpolated or otherwise estimated from other relationships present in the LUT 518.

[0061] At block 408, the first image data 502 and the second image data 504 may be blended by the blur applicator 222 based on the focus map 520 to determine an output image frame. The output image frame 204 may be a photograph or a video sequence with background (e.g., bokeh) blur corresponding to the simulated aperture size. The blending process may apply a heavier weight to the large aperture image data for the background portion based on a larger focus map score, and the blending process may also apply a heavier weight to the small aperture image data for the foreground (or object) portion based on a smaller focus map score. Although the processing of two image data at two aperture sizes to generate each output image frame is described, the process may also involve three or more image data at additional aperture sizes.

[0062] Some embodiments of the image processing described in Figure 3 are based on machine learning based calculations, as described with reference to Figure 6. Figure 6 shows a flow chart of an example method for applying a bokeh effect with a variable aperture (VA) camera system using machine learning, according to one or more aspects of the present disclosure. A block diagram illustrating the image data processing of Figure 6 is shown in Figure 7. Figure 7 shows a block diagram of an example processing configuration for applying a bokeh effect using machine learning, according to one or more aspects of the present disclosure.

[0063] Method 600 includes receiving, at block 602, first image data 502 captured with a first aperture size and second image data 504 captured with a second aperture size. The first image data and second image data may be received as streaming data during capture from a camera system, such as described with reference to block 302 of FIG. 3. The operations of method 600 may also be performed in real time with the capture of the first image data and the second image data, for example, to generate a preview image on a display. The operations of method 600 may also or alternatively be performed on the image capture device and / or by a remote device after capture of the first image data and the second image data.

[0064] In block 604, a blur kernel size map is determined based on machine learning analysis in a pre-trained artificial intelligence (AI) model 710 of the first image data 502 and the second image data 504. The blur kernel size map may indicate the depth of the object from the camera. The blur kernel may be a small matrix of values ​​representing a blurring process applied to a portion of the image data. The AI ​​model may be trained using images captured with various aperture sizes to prepare an offline model. The model may be applied to the image data 502 and the image data 504 to determine representative blur kernel sizes for various portions of the image data. Similar to the depth map described above, image data organized in an image frame of N×M pixels may have a corresponding blur kernel size map of K×L values, where N is an integer multiple of K and M is an integer multiple of L. The values ​​in the blur kernel size map may indicate the distance between the image sensor recording the image data and the object in the first scene, since portions of the scene closer to the focal point of the camera have smaller blur kernel sizes. In some embodiments, the AI ​​model 710 is a Resnet-34 model.

[0065] At block 606, a focus map may be determined based on the blur kernel size map and the blur kernel size lookup table. A focus map generator 224 may receive the blur kernel size map 716 and a blur kernel size lookup table (LUT) 718 (or other predefined relationship between aperture size and depth). The LUT 718 may provide a relationship 718A and a relationship 718B that map blur kernel size to depth for various aperture sizes. The relationship 718A may correspond to a smaller aperture size than the relationship 718B based on the smaller blur kernel size at higher depths. The LUT 718 may include one or more relationships that may be stored as a table, although alternatively or additionally, the predefined relationship may be stored as a representative formula. The focus map generator 224 determines the focus map 520 by mapping the blur kernel size map 716 to the focus map 520 using a relationship in the LUT 718 that corresponds to the simulated aperture size. In some embodiments, the relationship for the simulated aperture size may not be present in the LUT 718, and the relationship may instead be interpolated or otherwise estimated from other relationships present in the LUT 718. For example, predefined data may be available for aperture sizes of f / 2.0 and f / 2.4, and if a simulated aperture size of f / 1.4 is selected by the user, the relationship for f / 1.4 may be obtained by interpolating the predefined data for f / 2.0 and f / 2.4.

[0066] At block 608, the first image data 502 and the second image data 504 may be blended by the blur applicator 222 based on the focus map 520 to determine an output image frame. The output image frame 204 may be a photograph or a video sequence with background (e.g., bokeh) blur corresponding to the simulated aperture size. The blending process may apply a heavier weight to the large aperture image data for the background portion based on a larger focus map score, and the blending process may also apply a heavier weight to the small aperture image data for the foreground (or object) portion based on a smaller focus map score. Although the processing of two image data at two aperture sizes to generate each output image frame is described, the process may also involve three or more image data at additional aperture sizes.

[0067] The bokeh effect applied according to the above-described embodiment using the focus map 520 and predefined relationships such as lookup table 518 and lookup table 718 provides a more natural looking bokeh. The bokeh appears closer to the bokeh naturally produced by an optical lens with a large aperture, based in part on the predefined relationships. Conventional bokeh effects may assume a linear relationship between blur and depth from the camera's focus. However, optical lenses generally have a non-linear behavior, as shown in relationships 518A, 518B, 718A, and 718B. Therefore, the predefined data provides a more realistic and natural bokeh effect. Furthermore, the focus map generated from image data captured at two different aperture sizes provides a more accurate representation of the depth of objects in the scene, thereby allowing for a more realistic bokeh effect to be applied.

[0068] In one or more aspects, the techniques for supporting image processing may include additional aspects, such as any single aspect or any combination of aspects, described below or in conjunction with one or more other processes or devices described elsewhere herein. In a first aspect, supporting image processing may include an apparatus configured to perform operations including receiving first image data captured with a first aperture size and second image data captured with a second aperture size, where each of the first image data and the second image data represents a first scene; determining a depth map corresponding to the first scene based on the first image data and the second image data, where the depth map includes a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, where the focus map includes a second value indicative of an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and an object in the first scene; and determining an output image frame based on the focus map, the first image data, and the second image data. Moreover, the apparatus may perform or operate according to one or more aspects as described below. In some implementations, the apparatus includes a wireless device, such as a UE or a BS. In some implementations, the apparatus may include at least one processor and a memory coupled to the processor. The processor may be configured to perform the operations described herein with respect to the apparatus. In some other implementations, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, the program code being executable by a computer to cause the computer to perform the operations described herein with respect to the apparatus.In some implementations, the apparatus may include one or more means configured to perform the operations described herein. In some implementations, a method of wireless communication may include one or more operations described herein with respect to the apparatus.

[0069] In a second aspect, in combination with the first aspect, determining the depth map includes determining a blur map corresponding to the first scene based on the first image data and the second image data.

[0070] In a third aspect, in combination with one or more of the first or second aspect, determining the blur map includes determining a first sharpness value for the first image data and determining a second sharpness value for the second image data, and the blur map includes a blur value based on the first sharpness value and the second sharpness value.

[0071] In a fourth aspect, in combination with one or more of the first to third aspects, determining the blur map includes determining, based on supervised machine learning, a blur kernel size value corresponding to a distance between an image sensor recording the first image data and an object in the first scene.

[0072] In a fifth aspect, in combination with one or more of the first to fourth aspects, determining the focus map includes determining a second value based on the depth map and a predetermined relationship between depth and blur.

[0073] In a sixth aspect, in combination with one or more of the first to fifth aspects, the image sensor is configured as a variable aperture (VA) camera system.

[0074] In a seventh aspect, in combination with one or more of the first to sixth aspects, the image sensor captures the first image data and the second image data.

[0075] In an eighth aspect, in combination with one or more of the first to seventh aspects, determining the depth map includes determining a blurriness map including determining a first sharpness value for the first image data and determining a second sharpness value for the second image data, wherein the blur map includes a blur value based on the first sharpness value and the second sharpness value.

[0076] In a ninth aspect, in combination with one or more of the first to eighth aspects, determining the focus map includes determining the focus map based on a blurriness map and a predetermined relationship between aperture size and depth.

[0077] In a tenth aspect, in combination with one or more of the first to ninth aspects, determining the output image frame includes blending the first image data with corresponding data of the second image data based on a focus map.

[0078] In an eleventh aspect, in combination with one or more of the first to tenth aspects, determining an output image frame includes determining an output image frame that includes a portion that has been blurred based on the focus map to obtain a photograph that corresponds to the simulated aperture size.

[0079] In a twelfth aspect, in combination with one or more of the first to eleventh aspects, determining the depth map includes determining a blur kernel size value based on supervised machine learning.

[0080] In a thirteenth aspect, in combination with one or more of the first to twelfth aspects, determining the focus map includes determining the focus map based on a value of a blur kernel size and a predetermined relationship between the blur kernel size and depth.

[0081] In a fourteenth aspect, in combination with one or more of the first to thirteenth aspects, determining the output image frame includes blending the first image data with corresponding data of the second image data based on the focus map.

[0082] In a fifteenth aspect, in combination with one or more of the first to fourteenth aspects, determining an output image frame includes determining an output image frame that includes a portion that has been blurred based on a focus map to obtain a photograph that corresponds to a simulated aperture size.

[0083] In a sixteenth aspect, in combination with one or more of the first to fifteenth aspects, the operation is performed by an image capture device comprising a variable aperture (VA) camera system configured to capture first image data at a first aperture size and second image data at a second aperture size.

[0084] Those skilled in the art will appreciate that information and signals may be represented using any of a wide variety of technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0085] The components, functional blocks, and modules described herein with respect to Figures 1-5 include, among many examples, processors, electronic devices, hardware devices, electronic components, logic circuits, memories, software code, firmware code, or any combination thereof. Software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, and / or functions, among many examples, whether referred to as software, firmware, middleware, microcode, hardware description language, or the like. Furthermore, features discussed herein may be implemented via dedicated processor circuitry, via executable instructions, or combinations thereof.

[0086] Those skilled in the art will further appreciate that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein are merely examples, and that the components, methods, or interactions of various aspects of the disclosure can be combined or performed in ways other than those shown and described herein.

[0087] The various example logic, logic blocks, modules, circuits, and algorithmic processes described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. Interoperability between hardware and software has been described generally in terms of functionality and is illustrated in the various example components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented in hardware or software depends on the particular application and design constraints imposed on the overall system.

[0088] The hardware and data processing devices used to implement the various example logic, logic blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using general purpose single-chip or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. A general purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some implementations, a processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. In some implementations, certain processes and methods may be performed by circuits that are specialized for a given function.

[0089] In one or more aspects, the functions described can be implemented in hardware, digital electronic circuitry, computer software, firmware, or any combination thereof, including the structures disclosed herein and their structural equivalents. Implementations of the subject matter described herein can also be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a computer storage medium for execution by or for controlling the operation of a data processing apparatus.

[0090] If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code. The processes of the methods or algorithms disclosed herein may be implemented in processor-executable software modules that may reside on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that can be enabled to transfer a computer program from one place to another. A storage medium may be any available medium that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection may be properly termed a computer-readable medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc ("Blu-ray" is a registered trademark), where a disk typically reproduces data magnetically, while a disc reproduces data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.Furthermore, the operations of a method or algorithm may reside as code and instructions, one or any combination or set, on a machine-readable medium and a computer-readable medium, which may be embodied in a computer program product.

[0091] Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to several other implementations without departing from the spirit or scope of the present disclosure. Thus, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with the present disclosure, the principles and novel features disclosed herein.

[0092] Moreover, those skilled in the art will readily appreciate that the terms "upper" and "lower" may be used to facilitate description of the figures and are intended to indicate relative positions corresponding to the orientation of the figures on a properly oriented page, and may not reflect the proper orientation of any device in which they may be implemented.

[0093] Certain features that are described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable subcombination. Furthermore, although features may be described above as acting in a particular combination, and may even be initially claimed as such, one or more features from a claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0094] Similarly, although operations are shown in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order shown, or in sequential order, or that all of the operations shown be performed in order to achieve a desired result. Furthermore, the figures may generally depict one or more exemplary processes in the form of a flow diagram. However, other operations not shown may be incorporated into the exemplary process depicted in the schematic. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the depicted operations. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged in multiple software products. Furthermore, several other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired results.

[0095] As used herein, including the claims, the term "or", when used in a list of two or more items, means that any one of the listed items can be employed alone, or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing components A, B, or C, the composition may contain only A, only B, only C, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C. Also, as used herein, including the claims, "or" as used in a list of items ending with "at least one of" indicates a disjunctive list, such as, for example, a list of "at least one of A, B, or C" means any of A, B, or C, or AB, or AC, or BC, or ABC (i.e., A and B and C), or any of them in any combination thereof. The term "substantially" is defined as a majority, but not necessarily the entirety, of what is specified, as will be understood by one of ordinary skill in the art (and is inclusive of what is specified, e.g., substantially 90 degrees includes 90 degrees and substantially parallel includes parallel). In any of the disclosed implementations, the term "substantially" can be replaced with "within a [percentage] of" what is specified, where percentage includes 0.1, 1, 5, or 10 percent.

[0096] The above description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the embodiments and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. 1. A method comprising: receiving first image data captured with a first aperture size and second image data captured with a second aperture size, each of the first image data and the second image data representing a first scene; determining a depth map corresponding to the first scene based on the first image data and the second image data, the depth map including a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, the focus map including a second value indicating an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and the object in the first scene; determining an output image frame based on the focus map, the first image data, and the second image data.

2. The method of claim 1 , wherein determining the depth map comprises determining a blurriness map corresponding to the first scene based on the first image data and the second image data.

3. 3. The method of claim 2, wherein determining the blurriness map comprises determining a first sharpness value for the first image data and determining a second sharpness value for the second image data, and the blurriness map comprises a blur value based on the first sharpness value and the second sharpness value.

4. 3. The method of claim 2, wherein determining the blurriness map comprises determining, based on supervised machine learning, a blur kernel size value corresponding to the distance between the image sensor recording the first image data and the object in the first scene.

5. The method of claim 2 , wherein determining the focus map comprises determining the second value based on the depth map and a predetermined relationship between depth and blur.

6. The method of claim 1 , wherein the image sensor captures the first image data and the second image data, and the image sensor is configured as a variable aperture (VA) camera system.

7. determining a blurriness map, wherein determining the depth map includes determining a first sharpness value for the first image data and determining a second sharpness value for the second image data, the blurriness map including a blur value based on the first sharpness value and the second sharpness value; determining the focus map includes determining the focus map based on the blurriness map and a predetermined relationship between aperture size and depth; determining the output image frame includes blending the first image data with corresponding data of the second image data based on the focus map; The method of claim 6.

8. 8. The method of claim 7, wherein determining the output image frame comprises determining the output image frame including a blurred portion based on the focus map to obtain a photograph corresponding to the simulated aperture size.

9. determining the depth map includes determining a blur kernel size value based on supervised machine learning; determining the focus map includes determining the focus map based on the blur kernel size value and a predetermined relationship between blur kernel size and depth; determining the output image frame includes blending the first image data with corresponding data of the second image data based on the focus map; The method of claim 6.

10. 10. The method of claim 9, wherein determining the output image frame comprises determining the output image frame including a blurred portion based on the focus map to obtain a photograph corresponding to the simulated aperture size.

11. 1. An apparatus comprising: a memory storing processor-readable code; at least one processor coupled to the memory; Equipped with The at least one processor executes the processor-readable code to cause the at least one processor to: receiving first image data captured with a first aperture size and second image data captured with a second aperture size, each of the first image data and the second image data representing a first scene; determining a depth map corresponding to the first scene based on the first image data and the second image data, the depth map including a first value indicative of a distance between an image sensor recording the first image data and an object in the first scene; determining a focus map based on the depth map and a simulated aperture size different from the first aperture size and the second aperture size, the focus map including a second value indicating an amount of blur at the simulated aperture size for a corresponding distance between the image sensor recording the first image data and the object in the first scene; determining an output image frame based on the focus map, the first image data, and the second image data; 10. An apparatus configured to cause an operation to be performed, comprising:

12. The apparatus described in claim 11, wherein the at least one processor is configured to execute the processor-readable code and cause the at least one processor to further perform a method described in any one of claims 2 to 10.

13. 11. A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 10.