Merging separate pixel data to achieve greater depth of field

By generating multiple sub-image data from a pixel-separated camera and utilizing the weight differences of high-frequency content, the problem of image blurring outside the depth of field is solved, thereby improving image sharpness and expanding the depth of field.

CN115244570BActive Publication Date: 2025-11-14GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180006421.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-11-14
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

Because the corresponding part of the scene is outside the depth of field of the camera device that captures the image, part of the image may become blurred, and existing technologies have difficulty in effectively correcting or adjusting this blurring phenomenon.

Method used

By using a split-pixel camera to generate split-pixel image data comprising multiple sub-images, and by utilizing the differences in the frequency content of the sub-images, especially by giving greater weight to high-frequency content, the sharpness of the image is enhanced, resulting in an enhanced image with extended depth of field.

Benefits of technology

It improves image sharpness, expands the depth of field, and effectively corrects blurring caused by scene features outside the depth of field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115244570B_ABST
    Figure CN115244570B_ABST
Patent Text Reader

Abstract

A method includes obtaining separated pixel image data comprising a first sub-image and a second sub-image. The method further includes: for each corresponding pixel in the separated pixel image data, determining a corresponding position of a scene feature represented by the corresponding pixel relative to a depth of field, and identifying an out-of-focus pixel based on the corresponding position. The method additionally includes: for each corresponding out-of-focus pixel, determining a corresponding pixel value based on the corresponding position, the position of the corresponding out-of-focus pixel within the separated pixel image data, and at least one of a first value of a corresponding first pixel in the first sub-image or a second value of a corresponding second pixel in the second sub-image. The method further includes generating an enhanced image with extended depth of field based on the corresponding pixel values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image data processing, and more specifically to image data processing of separated pixels. Background Technology

[0002] Because a corresponding part of the scene lies outside the depth of field of the camera device capturing the image, a portion of the image may become blurred. The degree of blurring can depend on the position of the corresponding part of the scene relative to the depth of field, with the amount of blur increasing as the corresponding part of the scene moves away from the depth of field in a direction toward or away from the camera. In some cases, image blurring is undesirable and can be adjusted or corrected using various image processing techniques, models, and / or algorithms. Summary of the Invention

[0003] A split-pixel camera can be configured to generate split-pixel image data comprising multiple sub-images. For a defocused pixel in the split-pixel image data, the frequency content of the sub-image can vary as a function of (i) the position of the defocused pixel within the image data and (ii) the position of the scene features represented by the defocused pixel relative to the depth of field (i.e., pixel depth) of the split-pixel camera. Specifically, for a given defocused pixel, one of the sub-images may appear sharper than the others, depending on the position and depth associated with the given defocused pixel. Accordingly, the relationship between pixel position, pixel depth, and sub-image frequency content can be characterized for the split-pixel camera and used to improve the sharpness of portions of the split-pixel image data. In particular, instead of summing or equally weighting the pixels of the sub-images, sub-image pixels containing both low-frequency and high-frequency content can be given greater weight than sub-image pixels containing only low-frequency content, thereby increasing the apparent sharpness of the resulting image.

[0004] In a first example embodiment, a method may include obtaining split-pixel image data captured by a split-pixel camera. The split-pixel image data may include a first sub-image and a second sub-image. The method may further include: for each corresponding pixel among a plurality of pixels in the split-pixel image data, determining a corresponding position of a scene feature represented by the corresponding pixel relative to the depth of field of the split-pixel camera. The method may further include: based on the corresponding position of the scene feature represented by each corresponding pixel among the plurality of pixels, identifying one or more out-of-focus pixels among the plurality of pixels, wherein the one or more out-of-focus pixels are located outside the depth of field. The method may further include: for each corresponding out-of-focus pixel among the one or more out-of-focus pixels, determining a corresponding pixel value based on: (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field, (ii) the position of the corresponding out-of-focus pixel within the split-pixel image data, and (iii) at least one of a first value of a corresponding first pixel in the first sub-image or a second value of a corresponding second pixel in the second sub-image. The method may further include generating an enhanced image with extended depth of field based on the corresponding pixel value determined for each corresponding out-of-focus pixel.

[0005] In a second example embodiment, a system may include a processor and a non-transitory computer-readable medium having instructions stored thereon, which, when executed by the processor, cause the processor to perform operations. The operations may include acquiring split-pixel image data captured by a split-pixel camera. The split-pixel image data may include a first sub-image and a second sub-image. The operations may further include: for each corresponding pixel among a plurality of pixels in the split-pixel image data, determining a corresponding position of a scene feature represented by the corresponding pixel relative to the depth of field of the split-pixel camera. The operations may further include: based on the corresponding position of the scene feature represented by each corresponding pixel among the plurality of pixels, identifying one or more out-of-focus pixels among the plurality of pixels, wherein the one or more out-of-focus pixels are located outside the depth of field. The operations may further include: for each corresponding out-of-focus pixel among the one or more out-of-focus pixels, determining a corresponding pixel value based on: (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field, (ii) the position of the corresponding out-of-focus pixel within the split-pixel image data, and (iii) at least one of a first value of a first pixel in the first sub-image or a second value of a second pixel in the second sub-image. The operation may also include generating an enhanced image with extended depth of field based on the corresponding pixel value determined for each respective out-of-focus pixel.

[0006] In a third example embodiment, a non-transitory computer-readable medium may store instructions thereon that, when executed by a computing device, cause the computing device to perform operations. The operations may include obtaining split-pixel image data captured by a split-pixel camera. The split-pixel image data may include a first sub-image and a second sub-image. The operations may further include: for each corresponding pixel among a plurality of pixels in the split-pixel image data, determining the corresponding position of a scene feature represented by the corresponding pixel relative to the depth of field of the split-pixel camera. The operations may further include: based on the corresponding position of the scene feature represented by each corresponding pixel among the plurality of pixels, identifying one or more out-of-focus pixels among the plurality of pixels, wherein the one or more out-of-focus pixels are located outside the depth of field. The operations may further include: for each corresponding out-of-focus pixel among the one or more out-of-focus pixels, determining a corresponding pixel value based on: (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field, (ii) the position of the corresponding out-of-focus pixel within the split-pixel image data, and (iii) at least one of a first value of a first pixel in the first sub-image or a second value of a second pixel in the second sub-image. The operation may also include generating an enhanced image with extended depth of field based on the corresponding pixel value determined for each respective out-of-focus pixel.

[0007] In a fourth example embodiment, a system may include components for obtaining split-pixel image data captured by a split-pixel camera. The split-pixel image data may include a first sub-image and a second sub-image. The system may also include components for determining, for each corresponding pixel of the plurality of pixels in the split-pixel image data, the corresponding position of a scene feature represented by the corresponding pixel relative to the depth of field of the split-pixel camera. The system may also include components for identifying one or more off-focus pixels among the plurality of pixels based on the corresponding position of the scene feature represented by each corresponding pixel of the plurality of pixels, wherein the one or more off-focus pixels are located outside the depth of field. The system may also include components for determining a corresponding pixel value for each of the one or more off-focus pixels based on: (i) the corresponding position of the scene feature represented by the corresponding off-focus pixel relative to the depth of field, (ii) the position of the corresponding off-focus pixel within the split-pixel image data, and (iii) at least one of a first value of a corresponding first pixel in the first sub-image or a second value of a corresponding second pixel in the second sub-image. The system may also include components for generating an enhanced image with extended depth of field based on the corresponding pixel value determined for each corresponding off-focus pixel.

[0008] These and other embodiments, aspects, advantages, and alternatives will become apparent to those skilled in the art upon reading the following detailed description with appropriate reference to the accompanying drawings. Furthermore, the summary of the invention provided herein, along with the other descriptions and drawings, are intended to illustrate embodiments by way of example only, and thus many variations are possible. For example, structural elements and process steps may be rearranged, combined, distributed, eliminated, or otherwise altered while remaining within the scope of the claimed embodiments. Attached Figure Description

[0009] Figure 1 A computing device is shown as an example according to the description herein.

[0010] Figure 2 A computing system based on an example described herein is shown.

[0011] Figure 3 A dual-pixel image sensor, as described herein, is shown as an example.

[0012] Figure 4 The point spread function associated with split-pixel image data, as described in this article, is illustrated.

[0013] Figure 5 A system based on an example described in this article is shown.

[0014] Figure 6A , Figure 6B , Figure 6C , Figure 6D , Figure 6E and Figure 6F The examples described herein show pixel sources corresponding to different combinations of pixel depth and pixel position in dual-pixel image data.

[0015] Figure 6G This is a summary based on the examples described in this article. Figure 6A , Figure 6B , Figure 6C , Figure 6D , Figure 6E and Figure 6F The table showing the relationships is shown.

[0016] Figure 7A , Figure 7B , Figure 7C , Figure 7D , Figure 7E and Figure 7F The pixel sources corresponding to different combinations of pixel depth and pixel position of four-pixel image data are shown according to the examples described herein.

[0017] Figure 7G This is a summary based on the examples described in this article. Figure 7A , Figure 7B , Figure 7C , Figure 7D , Figure 7E and Figure 7F The table showing the relationships is shown.

[0018] Figure 8 This is a flowchart based on the example described in this article. Detailed Implementation

[0019] This document describes exemplary methods, apparatus, and systems. It should be understood that the terms “exemplary” and “illustrative” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as “exemplary,” “illustrative,” and / or “illustrative” is not necessarily to be construed as preferred or superior to other embodiments or features unless stated otherwise. Therefore, other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0020] Accordingly, the exemplary embodiments described herein are not intended to be limiting. It will be readily understood that, as generally described herein and shown in the accompanying drawings, aspects of this disclosure can be arranged, replaced, combined, separated, and designed in a variety of different configurations.

[0021] Furthermore, unless the context otherwise suggests, the features shown in each of the figures can be used in combination with each other. Therefore, the figures should generally be considered as aspects of one or more overall embodiments, where it should be understood that not all features shown are necessary for every embodiment.

[0022] Furthermore, any enumeration of elements, blocks, or steps in this specification or claims is for clarity purposes. Therefore, such enumeration should not be construed as requiring or implying that these elements, blocks, or steps follow a particular arrangement or are performed in a particular order. Unless otherwise stated, the drawings are not to scale.

[0023] I. Overview

[0024] Separate pixel cameras can be used to generate separated pixel image data comprising two or more sub-images, each sub-image generated by a corresponding subset of the photosites of the separate pixel camera. For example, a separate pixel camera may include a dual-pixel image sensor, where each pixel of the dual-pixel image sensor consists of two photosites. Accordingly, the separated pixel image data may be dual-pixel image data, which includes a first sub-image generated by the photosite to the left of the pixel and a second sub-image generated by the photosite to the right of the pixel. In another example, a separate pixel camera may include a quad-pixel image sensor, where each pixel of the quad-pixel image sensor consists of four photosites. Accordingly, the separated pixel image data may be quad-pixel image data, which includes a first sub-image generated by the photosite to the upper left of the pixel, a second sub-image generated by the photosite to the upper right of the pixel, a third sub-image generated by the photosite to the lower left of the pixel, and a fourth sub-image generated by the photosite to the lower right of the pixel.

[0025] When a split-pixel camera is used to capture images of scene features (e.g., an object or a portion thereof) that are out of focus (i.e., outside the depth of field of the split-pixel camera), the frequency content of the sub-images may differ. Specifically, for each out-of-focus pixel, at least one split-pixel sub-image of the split-pixel image data (referred to as a high-frequency pixel source) may represent a spatial frequency above a threshold frequency that is not represented by other split-pixel sub-images of the split-pixel image data (referred to as (a plurality of) low-frequency pixel sources). The difference in the frequency content of the sub-images may be a function of (i) the position of the scene feature (e.g., an object or a portion thereof) relative to the depth of field (and thus also the position of the image representing the scene feature relative to the camera's depth of focus) and (ii) the position of (a plurality of) out-of-focus pixels representing the scene feature relative to the image sensor region (e.g., represented by coordinates within the image).

[0026] When these two sub-images are summed to generate a separate pixel image, some sharpness of the high-frequency pixel sources may be lost because spatial frequencies above a threshold frequency are represented only in one of the sub-images. Accordingly, by giving the high-frequency pixel sources a heavier weight than the low-frequency pixel sources (instead of simply summing the sub-images), the sharpness of the overall separate pixel image generated by combining the sub-images can be enhanced, resulting in an increase in spatial frequencies above the threshold frequency.

[0027] In some cases, high-frequency and low-frequency pixel sources can be merged in the image space. In one example, each out-of-focus pixel can be a weighted sum of spatially corresponding pixels in the sub-image, where spatially corresponding high-frequency pixel sources are given greater weight than their spatially corresponding low-frequency pixel sources. In another example, a corresponding source pixel can be selected from the spatially corresponding pixels of the sub-image for each out-of-focus pixel. Specifically, spatially corresponding high-frequency pixel sources can be selected as source pixels, and spatially corresponding low-frequency pixel sources can be discarded.

[0028] In other cases, high-frequency and low-frequency pixel sources can be merged in the frequency space. Specifically, high-frequency and low-frequency pixel sources can each be assigned frequency-specific weights. For example, frequencies above a threshold frequency that exist in the high-frequency pixel source but not in the low-frequency pixel source, or are not adequately represented in the low-frequency pixel source, can be boosted to increase the sharpness of the resulting enhanced image. Frequencies below a threshold frequency and / or frequencies present in both the high-frequency and low-frequency pixel sources can be weighted equally to preserve the content of both sub-images at these frequencies.

[0029] Differences in the frequency content of sub-images may be a result of optical defects present in the optical path of a split-pixel camera device. Therefore, in some cases, the relationship between frequency content, pixel depth, and pixel position can be determined on a per-camera and / or per-camera-model basis, and can then be used in combination with the corresponding camera instance and / or camera model to enhance the sharpness of the thus captured split-pixel image.

[0030] II. Example Computing Devices and Systems

[0031] Figure 1 An example computing device 100 is shown. The computing device 100 is shown in the form factor of a mobile phone. However, the computing device 100 may alternatively be implemented as a laptop computer, tablet computer, and / or wearable computing device, among other possibilities. The computing device 100 may include various components such as a body 102, a display 106, and buttons 108 and 110. The computing device 100 may also include one or more cameras, such as a front-facing camera 104 and a rear-facing camera 112, one or more of which may be configured to generate dual-pixel image data.

[0032] The front-facing camera 104 may be located on the side of the body 102 that is typically facing the user during operation (e.g., on the same side as the display 106). The rear-facing camera 112 may be located on the side of the body 102 opposite to the front-facing camera 104. Referring to the cameras as front and rear is arbitrary, and the computing device 100 may include multiple cameras placed on various sides of the body 102.

[0033] Display 106 may represent a cathode ray tube (CRT) display, a light-emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light-emitting diode (OLED) display, or any other type of display known in the art. In some examples, display 106 may display a digital representation of the current image captured by front camera 104 and / or rear camera 112, images that may be captured by one or more of these cameras, recently captured images by one or more of these cameras, and / or modified versions of one or more of these images. Thus, display 106 may act as a viewfinder for the camera. Display 106 may also support touchscreen functionality that allows adjustment of settings and / or configurations of one or more aspects of computing device 100.

[0034] The front-facing camera 104 may include an image sensor and associated optical elements, such as a lens. The front-facing camera 104 may provide zoom capability or may have a fixed focal length. In other examples, interchangeable lenses may be used with the front-facing camera 104. The front-facing camera 104 may have a variable mechanical aperture and a mechanical shutter and / or an electronic shutter. The front-facing camera 104 may also be configured to capture still images, video images, or both. Furthermore, the front-facing camera 104 may represent, for example, a single-field-of-view camera. The rear-facing camera 112 may be arranged similarly or differently. Additionally, one or more front-facing cameras 104 and / or rear-facing cameras 112 may be an array of one or more cameras.

[0035] One or more of the front camera 104 and / or the rear camera 112 may include or be associated with an illumination component that provides a light field to illuminate a target object. For example, the illumination component may provide flash or constant illumination of the target object. The illumination component may also be configured to provide a light field that includes one or more of structured light, polarized light, and light with specific spectral content. Other types of light fields known and possible for recovering a three-dimensional (3D) model from an object are possible within the context of the examples herein.

[0036] The computing device 100 may also include an ambient light sensor that can continuously or intermittently determine the ambient brightness of the scene that the cameras 104 and / or 112 can capture. In some embodiments, the ambient light sensor may be used to adjust the display brightness of the display 106. Furthermore, the ambient light sensor may be used to determine the exposure length of one or more of the cameras 104 or 112, or to assist in that determination.

[0037] The computing device 100 can be configured to capture images of a target object using a display 106 and a front-facing camera 104 and / or a rear-facing camera 112. The captured images can be multiple still images or a video stream. Image capture can be triggered by activating button 108, pressing a softkey on the display 106, or through some other mechanism. Depending on the implementation, images can be captured automatically at specific time intervals, such as when button 108 is pressed, under appropriate lighting conditions on the target object, when the digital camera device 100 is moved a predetermined distance, or according to a predetermined capture schedule.

[0038] Figure 2 This is a simplified block diagram illustrating some components of an example computing system 200. As an example and not a limitation, computing system 200 can be a cellular mobile phone (e.g., a smartphone), a computer (such as a desktop, laptop, tablet, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a game console, a robotic device, a vehicle, or some other type of device. Computing system 200 can represent, for example, aspects of computing device 100.

[0039] like Figure 2 As shown, the computing system 200 may include a communication interface 202, a user interface 204, a processor 206, a data storage 208, and a camera assembly 224, all of which may be communicatively linked together via a system bus, network, or other connection mechanism 210. The computing system 200 may be equipped with at least some image capture and / or image processing capabilities. It should be understood that the computing system 200 may represent a physical image processing system, a specific physical hardware platform on which image sensing and / or processing applications operate in software, or other combinations of hardware and software configured to perform image capture and / or processing functions.

[0040] Communication interface 202 allows computing system 200 to communicate with other devices, access networks, and / or transmission networks using analog or digital modulation. Therefore, communication interface 202 can facilitate circuit-switched and / or packet-switched communications, such as Common Old Telephone Service (POTS) communications and / or Internet Protocol (IP) or other packet-switched communications. For example, communication interface 202 may include a chipset and antenna arranged for wireless communication with a radio access network or access point. Furthermore, communication interface 202 may take the form of a wired interface or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI) port. Communication interface 202 may also take the form of a wireless interface or include a wireless interface, such as Wi-Fi. Global Positioning System (GPS) or wide-area wireless interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols can be used on communication interface 202. Furthermore, communication interface 202 may include multiple physical communication interfaces (e.g., Wi-Fi interfaces, ...). Interface and wide area wireless interface).

[0041] User interface 204 can be used to allow computing system 200 to interact with human or non-human users, such as receiving input from the user and providing output to the user. Therefore, user interface 204 may include input components such as a keypad, keyboard, touch panel, computer mouse, trackball, joystick, microphone, etc. User interface 204 may also include one or more output components, such as a display screen, for example, a display screen may be combined with a touch panel. The display screen may be based on CRT, LCD, and / or LED technology, or other technologies now known or developed in the future. User interface 204 may also be configured to generate audible output(s) via speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. User interface 204 may also be configured to receive and / or capture audible speech, noise, and / or signals via microphones and / or other similar devices.

[0042] In some examples, the user interface 204 may include a display that serves as a viewfinder for still camera functions and / or video camera functions supported by the computing system 200. Furthermore, the user interface 204 may include one or more buttons, switches, knobs, and / or dials to facilitate the configuration and focusing of camera functions and image capture. Some or all of these buttons, switches, knobs, and / or dials may be implemented via a touch-sensitive panel.

[0043] Processor 206 may include one or more general-purpose processors (e.g., microprocessors) and / or one or more special-purpose processors (e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating-point units (FPUs), network processors, or application-specific integrated circuits (ASICs)). In some cases, the special-purpose processor may be capable of image processing, image alignment, and image merging, among other possibilities. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated wholly or partially with processor 206. Data storage 208 may include removable and / or non-removable components.

[0044] Processor 206 may be able to execute program instructions 218 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 208 to perform the various functions described herein. Therefore, data storage 208 may include a non-transitory computer-readable medium thereon storing program instructions that, when executed by computing system 200, cause computing system 200 to perform any methods, processes, or operations disclosed in this specification and / or the accompanying drawings. Execution of program instructions 218 by processor 206 may cause processor 206 to use data 212.

[0045] As an example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device drivers (multiple) and other modules) and one or more applications 220 (e.g., camera functionality, address book, email, web browsing, social networking, audio-to-text functionality, text translation functionality, and / or game applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be primarily accessible by operating system 222, and application data 214 may be primarily accessible by one or more of the applications 220. Application data 214 may be located in a file system that is visible or hidden from the user of computing system 200.

[0046] Application 220 can communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs can facilitate, for example, application 220 reading and / or writing application data 214, sending or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on.

[0047] In some cases, application 220 may be simply referred to as "app". Furthermore, application 220 may be downloadable to computing system 200 through one or more online application stores or app markets. However, applications may also be installed on computing system 200 in other ways, such as via a web browser or through a physical interface on computing system 200 (e.g., a USB port).

[0048] Camera assembly 224 may include, but is not limited to, aperture, shutter, recording surface (e.g., photographic film and / or image sensor), lens, shutter button, infrared projector, and / or visible light projector. Camera assembly 224 may include components configured to capture images in the visible spectrum (e.g., electromagnetic radiation with wavelengths of 380-700 nm) and components configured to capture images in the infrared spectrum (e.g., electromagnetic radiation with wavelengths of 701 nm-1 mm). Camera assembly 224 may be controlled at least in part by software executed by processor 206.

[0049] III. Example of a dual-pixel image sensor

[0050] Figure 3 A split-pixel image sensor 300 configured to generate split-pixel image data is shown. Specifically, the split-pixel image sensor 300 is shown as a dual-pixel image sensor comprising a plurality of pixels arranged in a grid comprising columns 302, 304, 306, and 308 to 310 (i.e., columns 302-310) and rows 312, 314, 316, and 318 to 320 (i.e., rows 312-320). Each pixel is shown as divided into a first (e.g., left) photosensitive point represented by a corresponding shaded area and a second (e.g., right) photosensitive point represented by a corresponding white-filled area. Thus, the right half of the pixel located in column 302, row 312 is labeled “R” to indicate the right photosensitive point, and the left half of the pixel is labeled “L” to indicate the left photosensitive point.

[0051] Although the photosensitive point of each pixel is shown as dividing each pixel into two equal vertical halves, the photosensitive point can alternatively divide each pixel in other ways. For example, each pixel can be divided into an upper photosensitive point and a lower photosensitive point. The areas of the photosensitive points may not be equal. Furthermore, although the split-pixel image sensor 300 is shown as a dual-pixel image sensor in which each pixel includes two photosensitive points, the split-pixel image sensor 300 can alternatively be implemented in which each pixel is divided into a different number of photosensitive points. For example, the split-pixel image sensor 300 can be implemented as a quad-pixel image sensor in which each corresponding pixel of the quad-pixel image sensor is divided into four photosensitive points defining four quadrants of the corresponding pixel (e.g., (first) upper left quadrant, (second) upper right quadrant, (third) lower left quadrant, and (fourth) lower right quadrant).

[0052] Each photosensitive point of a given pixel may include a corresponding photodiode, the output signal of which can be read independently of other photodiodes. Furthermore, each pixel of the split-pixel image sensor 300 may be associated with a corresponding color filter (e.g., red, green, or blue). A demosaic algorithm can be applied to the output of the split-pixel image sensor 300 to generate a color image. In some cases, fewer pixels than all pixels of the split-pixel image sensor 300 may be divided into multiple photosensitive points. For example, each pixel associated with a green color filter may be divided into two independent photosensitive points, while each pixel associated with a red or blue color filter may include a single photosensitive point. In some cases, the split-pixel image sensor 300 may be used to implement a front-facing camera 104 and / or a rear-facing camera 112, and may form part of a camera assembly 224.

[0053] The split-pixel image sensor 300 can be configured to generate split-pixel image data. In one example, the split-pixel image data can be dual-pixel image data, which includes a first sub-image generated by a first set of photosensitive points (e.g., only left-side photosensitive points) and a second sub-image generated by a second set of photosensitive points (e.g., only right-side photosensitive points). In another example, the split-pixel image data can be quad-pixel image data, which includes a first sub-image generated by a first set of photosensitive points (e.g., only upper-left photosensitive points), a second sub-image generated by a second set of photosensitive points (e.g., only upper-right photosensitive points), a third sub-image generated by a third set of photosensitive points (e.g., only lower-left photosensitive points), and a fourth sub-image generated by a fourth set of photosensitive points (e.g., only lower-right photosensitive points).

[0054] Sub-images can be generated as part of a single exposure. For example, sub-images can be captured substantially simultaneously, with the capture time of one sub-image falling within a threshold time of the capture time of another sub-image. The signals generated by each photosensitive point of a given pixel can be combined into a single output signal to generate conventional (e.g., RGB) image data.

[0055] When an imaged scene feature (such as a foreground object, background object, environment, and / or portions thereof) is in focus (i.e., the scene feature is within the camera's depth of field, and / or the light reflected from it is focused within the camera's depth of focus), the corresponding signals generated by each photosensitive point of a given pixel can be substantially the same (e.g., the signals of separate pixels can be within each other's thresholds). When an imaged scene feature is out of focus (i.e., the scene feature is before or after the camera's depth of field, and / or the light reflected from it is focused before or after the camera's depth of focus), the corresponding signal generated by the first photosensitive point of a given pixel can be different from the corresponding signals generated by the other photosensitive points of the given pixel. The degree of this difference can be proportional to the degree of defocus and can indicate the position of the scene feature relative to the depth of field (and the position of the light reflected from it relative to the depth of focus). Accordingly, separate pixel image data can be used to determine whether the imaged scene feature is within, before, and / or after the depth of field of the camera device.

[0056] IV. Example Point Spread Function for Separating Pixel Image Data

[0057] Figure 4 An example spatial variation of the point spread function (PSF) of a dual-pixel image sensor (e.g., a split-pixel image sensor 300) associated with imaging a defocused plane is shown. Specifically, Figure 4Regions 400, 402, and 404 are shown, each corresponding to a region of the dual-pixel image sensor. The PSF in region 402 shows the spatial variation of the PSF associated with the left sub-image, the PSF in region 404 shows the spatial variation of the PSF associated with the right sub-image, and the PSF in region 400 represents the spatial variation of the PSF associated with the overall dual-pixel image, each PSF simultaneously imaging the defocus plane. Figure 4 As shown, the PSF of the overall dual-pixel image is equal to the sum of the PSFs of the left and right sub-images. For clarity, a PSF size ratio related to the size of the dual-pixel image sensor has been chosen, and this ratio can vary in various implementations.

[0058] Each of regions 400, 402, and 404 comprises 16 PSFs arranged in rows 410, 412, 414, and 416 and columns 420, 422, 424, and 426. Furthermore, corresponding dashed lines indicate the vertical midline of each region, thus dividing the region into two equal halves: a left half and a right half. The left half of region 402 includes PSFs that allow capturing a wider range of spatial high-frequency information than the PSFs in the right half of region 402, as indicated by differences in the shading patterns of these PSFs. Specifically, the PSFs in columns 420 and 422 of region 402 have higher cutoff spatial frequencies than the PSFs in columns 424 and 426 of region 402. Similarly, the right half of region 404 includes PSFs that allow capturing a wider range of spatial high-frequency information than the PSFs in the left half of region 404, as indicated by differences in the shading patterns of these PSFs. Specifically, the PSFs in columns 424 and 426 of region 404 have higher cutoff spatial frequencies than the PSFs in columns 420 and 422 of region 404.

[0059] Accordingly, when imaging an out-of-focus area of ​​the scene, the left half of the first sub-image corresponding to region 402 may appear sharper than (i) the right half of the first sub-image and (ii) the left half of the second sub-image corresponding to region 404. Similarly, when imaging an out-of-focus area of ​​the scene, the right half of the second sub-image corresponding to region 404 may appear sharper than (i) the left half of the second sub-image and (ii) the right half of the first sub-image corresponding to region 402.

[0060] This spatial variability in the frequency content across split-pixel sub-images may be a result of various real-world imperfections in the optical path of dual-pixel camera devices and may not be apparent from idealized optical models. In some cases, spatial variability can be empirically characterized on a per-camera basis and subsequently used to generate an enhanced version of the image captured by that camera model.

[0061] When the first and second sub-images are added together, the resulting overall dual-pixel image corresponds to region 400. That is, the resulting dual-pixel image appears to be generated using a dual-pixel image sensor associated with the PSF of region 400. Accordingly, the relatively sharper content of the left half of the first sub-image (corresponding to region 402) is combined with the content of the left half of the second sub-image (corresponding to region 404), thus blurring it. Similarly, the relatively sharper content of the right half of the second sub-image (corresponding to region 404) is combined with the content of the right half of the first sub-image (corresponding to region 402), thus blurring it.

[0062] Specifically, frequencies up to the first cutoff frequency of the PSF in the right half of region 402 and / or the left half of region 404 are represented in both halves of regions 402 and 404. However, frequencies between (i) the first cutoff frequency and (ii) the second cutoff frequency of the PSF in the left half of region 402 and / or the right half of region 404 are represented in the left half of region 402 and the right half of region 404, but not in the right half of region 402 and the left half of region 404. Accordingly, when the PSFs of regions 402 and 404 are summed to form the PSF of region 400, the frequencies between the first and second cutoff frequencies are not adequately represented compared to frequencies below the first cutoff frequency (e.g., their relative power is lower). Therefore, summing the pixel values ​​of the sub-image does not utilize the differences in spatial frequency content present in different parts of the sub-image.

[0063] Figure 4 The PSF in the diagram corresponds to a defocus plane located on a first side of the focal plane and / or depth of field (e.g., between (i) the camera device and (ii) the focal plane and / or depth of field). When the defocus plane is located on a second side of the focal plane (e.g., outside the focal plane), Figure 4 The pattern of the PSF shown can be different. For example, the PSF pattern can be flipped, and it can be approximated by swapping the PSF positions of regions 402 and 404. Corresponding PSF changes can be observed additionally or alternatively when each separate pixel is instead divided into upper and lower photosensitive points, and / or into four photosensitive points that divide the separate pixel into four quadrants, and other possibilities. The relationship between the PSF cutoff frequency across the image sensor region and the scene feature position relative to the depth of field can vary between camera models, and can therefore be determined empirically on a per-camera basis.

[0064] V. Example System for Generating Enhanced Images

[0065] The presence of high-frequency spatial information in different parts of the separated pixel sub-image can be used to enhance the separated pixel image by increasing its sharpness, thereby effectively extending the corresponding depth of field. Specifically, Figure 5 An example system for generating enhanced images by utilizing high-frequency information present in certain portions of separated pixel sub-images is shown. Figure 5 A system 500 configured to generate an enhanced image 528 based on separated pixel image data 502 is shown. The system 500 may include a pixel depth calculator 508, a pixel depth classifier 512, a pixel frequency classifier 520, and a pixel value merger 526. The components of the system 500 may be implemented as hardware, software, or a combination thereof.

[0066] The split-pixel image data 502 may include sub-images 504 to 506 (i.e., sub-images 504-506). The split-pixel image data 502 may be captured by the split-pixel image sensor 300. Each sub-image in sub-images 504-506 may have the same resolution as the split-pixel image data 502 and may be captured as part of a single exposure. Accordingly, each corresponding pixel of the split-pixel image data 502 may be associated with a corresponding pixel in each sub-image in sub-images 504-506. In one example, the split-pixel image data 502 may include two sub-images, thus it may be referred to as dual-pixel image data. In another example, the split-pixel image data 502 may include four sub-images, thus it may be referred to as quad-pixel image data.

[0067] When the separated pixel image data 502 represents scene features located outside the depth of field of the separated pixel camera (causing the corresponding light to focus outside the depth of field of the separated pixel camera), some of the sub-images 504-506 may include high-frequency spatial information that may not be present in other sub-images 504-506. The term "high frequency" and / or variations thereof, as used herein, refers to frequencies above a threshold frequency, where the first separated pixel sub-image contains frequency content above the threshold frequency, while the corresponding second separated pixel sub-image does not. Conversely, the term "low frequency" and / or variations thereof, as used herein, refers to frequencies below and including the threshold frequency. The threshold frequency may vary depending on the separated pixel camera and / or the scene being photographed, as well as other factors.

[0068] The pixel depth calculator 508 can be configured to determine multiple pixel depths 510 of multiple pixels of the separated pixel image data 502. For example, the multiple pixel depths 510 may correspond to all pixels of the separated pixel image data 502, or fewer than all pixels of the separated pixel image data 502. The multiple pixel depths 510 may indicate the depth of multiple corresponding scene features (e.g., objects and / or portions thereof) relative to the depth of field of the separated pixel camera, and / or the depth of the corresponding image of the multiple scene features relative to the depth of focus of the separated pixel camera. In some embodiments, the multiple pixel depths 510 may include, for example, a binary representation of the depth associated with the corresponding scene feature to indicate whether the corresponding scene feature is located behind (e.g., on a first side) or before (e.g., on a second side) the depth of field of the separated pixel camera. In some cases, the multiple pixel depths may include a ternary representation that is also configured to indicate, for example, that the corresponding scene feature is located and / or focused within the depth of field of the separated pixel camera.

[0069] In other implementations, the pixel depth 510 may take more than three values, thereby indicating, for example, how far in front of and / or behind the depth of field the corresponding scene feature is located. It should be understood that when a scene feature is located outside the depth of field of the split-pixel camera (i.e., the area in front of the lens where the scene feature would produce an image that appears sufficiently in focus), the corresponding image (i.e., the light representing the corresponding scene feature) is also located (i.e., focused) outside the depth of field of the split-pixel camera (i.e., the area behind the lens where the image appears sufficiently in focus).

[0070] The pixel depth calculator 508 can be configured to determine the depth value of a corresponding pixel based on the signal disparity between (i) a first pixel in a first sub-image of sub-images 504-506 and (ii) a second pixel in a second sub-image of sub-images 504-506. Specifically, the signal disparity can be positive when the scene feature is located on a first side of the depth of field, and negative when the scene feature is located on a second side of the depth of field. Therefore, the sign of the disparity can indicate the direction of the depth value of the corresponding pixel relative to the depth of field, while the magnitude of the disparity can indicate the magnitude of the depth value. In the case of four-pixel image data, the depth value can be additionally or alternatively based on the third pixel in a third sub-image of sub-images 504-506 and the fourth pixel in a fourth sub-image of sub-images 504-506.

[0071] Pixel depth classifier 512 can be configured to identify (i) multiple focused pixels 514 of the separated pixel image data 502 and (ii) multiple defocused pixels 516 of the separated pixel image data 502 based on pixel depth(s) 510. The multiple focused pixels 514 may represent scene features located within the depth of field (e.g., within a threshold distance on either side of the focal plane) and thus may not undergo depth-of-field and / or sharpness enhancement. The multiple defocused pixels 516 may represent scene features located outside the depth of field (e.g., beyond a threshold distance on either side of the focal plane) and thus may undergo depth-of-field and / or sharpness enhancement. Each corresponding defocused pixel among the multiple defocused pixels 516 can be associated with a corresponding pixel position in the multiple pixel positions 518, where the corresponding pixel position indicates, for example, the coordinates of the corresponding defocused pixel within the separated pixel image data 502. Each corresponding defocused pixel can also be associated with a corresponding pixel depth in the multiple pixel depths 510.

[0072] Pixel frequency classifier 520 can be configured to identify multiple high-frequency pixel sources 522 and multiple low-frequency pixel sources 524 for each corresponding defocused pixel in the multiple defocused pixels 516. Specifically, pixel frequency classifier 520 can be configured to identify multiple high-frequency pixel sources 522 and multiple low-frequency pixel sources 524 based on the position of the corresponding pixel within the separated pixel image data 502 and the depth value associated with the corresponding pixel. The multiple high-frequency pixel sources 522 of the corresponding defocused pixel may include pixels corresponding to positions of multiple locations in a first subset of sub-images 504-506, while the multiple low-frequency pixel sources 524 of the corresponding defocused pixel may include pixels corresponding to positions of multiple locations in a second subset of sub-images 504-506. The first subset and the second subset determined for the corresponding pixel can be mutually exclusive.

[0073] In the case of dual-pixel image data, multiple high-frequency pixel sources 522 can indicate, for example, that sub-image 504 contains more vivid (i.e., higher frequency) image content of the corresponding pixels, and multiple low-frequency pixel sources 524 can indicate that sub-image 506 contains less vivid (i.e., lower frequency) image content of the corresponding pixels. In the case of quad-pixel image data, multiple high-frequency pixel sources 522 can indicate, for example, that sub-image 506 contains more vivid image content of the corresponding pixels, and multiple low-frequency pixel sources 524 can indicate that all other sub-images (including sub-image 504) contain less vivid image content of the corresponding pixels. Regarding Figures 6A-7G The selection of pixel sources is explained and discussed in more detail.

[0074] Pixel value merger 526 can be configured to generate an enhanced image 528 based on (i) (a plurality of) focused pixels 514 and (ii) (a plurality of) high-frequency pixel sources 522 and (a plurality of) low-frequency pixel sources 524 determined for each corresponding defocused pixel in (a plurality of) defocused pixels 516. Specifically, pixel value merger 526 can be configured to generate the corresponding pixel values ​​of (a plurality of) focused pixels 514 by summing the spatially corresponding pixel values ​​of sub-images 504-506. For (a plurality of) defocused pixels 516, pixel value merger 526 can be configured to generate the corresponding pixel values ​​by giving (a plurality of) high-frequency pixel sources 522 a greater weight than (a plurality of) low-frequency pixel sources 524 (at least relative to some frequencies) to increase the apparent sharpness and / or depth of field of the corresponding portion of the separated pixel image data 502.

[0075] VI. Example relationships between pixel depth, pixel position, and sub-image frequency content

[0076] Figures 6A-6F An example mapping is shown between the position of a corresponding pixel within two-pixel image data, the depth value associated with that pixel, and a sub-image containing high-frequency image content. Specifically, Figures 6A-6F Each of the images shows a two-pixel image region divided into four quadrants labeled with corresponding depths on the left, and a corresponding sub-image quadrant on the right, which provides high-frequency content for each quadrant of the two-pixel image given the corresponding depth.

[0077] A quadrant of an image can be a rectangular region spanning one-quarter of an image frame and is generated by bisecting the image frame horizontally and vertically. Therefore, four quadrants can span the entire image frame and divide the image frame into four equal sub-parts. Similarly, half of an image can be the union of two adjacent quadrants. For example, the union of two horizontally adjacent quadrants can define the upper or lower half, and the union of two vertically adjacent quadrants can define the left or right half. In other words, the upper and lower halves of an image can be defined by bisecting the image horizontally into two equal rectangular regions, while the left and right halves can be defined by bisecting the image vertically into two equal rectangular regions. The overall separated pixel image (e.g., image data 502), sub-images (e.g., sub-images 504-506), and / or enhanced images (e.g., enhanced image 528) can each be divided into corresponding halves and / or quarters.

[0078] Figure 6AIt is shown that when quadrants 600, 602, 604, and 606 of the dual-pixel image sensor are each used to image a plane located on the first side of the depth of field (e.g., after) (as shown in "Depth: -1") (i.e., a scene with a constant depth relative to the split-pixel camera), then: (i) quadrant 610 of the first dual-pixel sub-image contains higher frequency content of quadrant 600 than quadrant 620 of the correspondingly located quadrant of the second dual-pixel sub-image; (ii) quadrant 622 of the second sub-image contains higher frequency content of quadrant 602 than quadrant 612 of the correspondingly located quadrant of the first sub-image; (iii) quadrant 614 of the first sub-image contains higher frequency content of quadrant 604 than quadrant 624 of the correspondingly located quadrant of the second sub-image; and (iv) quadrant 626 of the second sub-image contains higher frequency content of quadrant 606 than quadrant 616 of the correspondingly located quadrant of the first sub-image.

[0079] Figure 6B It is shown that when quadrants 600, 602, 604, and 606 of a dual-pixel image sensor are each used to image a plane located on the second side of the depth of field (e.g., before) (as shown in "Depth: +1"), then: (i) quadrant 620 of the second sub-image contains higher frequency content of quadrant 600 than quadrant 610 of the corresponding location of the first sub-image; (ii) quadrant 612 of the first sub-image contains higher frequency content of quadrant 602 than quadrant 622 of the corresponding location of the second sub-image; (iii) quadrant 624 of the second sub-image contains higher frequency content of quadrant 604 than quadrant 614 of the corresponding location of the first sub-image; and (iv) quadrant 616 of the first sub-image contains higher frequency content of quadrant 606 than quadrant 626 of the corresponding location of the second sub-image.

[0080] Figure 6C It is shown that when quadrants 600 and 604 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), and quadrants 602 and 606 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), then quadrants 610, 612, 614, and 616 of the first sub-image contain higher frequency content of quadrants 600, 602, 604, and 606, respectively, than quadrants 620, 622, 624, and 626 of the second sub-image.

[0081] Figure 6DThis illustrates that when quadrants 600 and 604 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), and quadrants 602 and 606 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), then quadrants 620, 622, 624, and 626 of the second sub-image contain higher frequency content of quadrants 600, 602, 604, and 606, respectively, than quadrants 610, 612, 614, and 616 of the first sub-image.

[0082] Figure 6E It is shown that when quadrants 600 and 602 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), and quadrants 604 and 606 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), then: (i) quadrants 612 and 614 of the first sub-image contain higher frequency content of quadrants 602 and 604 than quadrants 622 and 624 of the second sub-image, respectively, and (ii) quadrants 620 and 626 of the second sub-image contain higher frequency content of quadrants 600 and 606 than quadrants 610 and 616 of the first sub-image, respectively.

[0083] Figure 6F It is shown that when quadrants 600 and 602 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), and quadrants 604 and 606 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), then: (i) quadrants 610 and 616 of the first sub-image contain higher frequency content of quadrants 600 and 606 than quadrants 620 and 626 of the second sub-image, respectively, and (ii) quadrants 622 and 624 of the second sub-image contain higher frequency content of quadrants 602 and 604 than quadrants 612 and 614 of the first sub-image, respectively.

[0084] Figure 6G Table 630 is shown, which summarizes the results of... Figures 6A-6F The example shows the relationship between the combination of pixel location and pixel depth. Specifically, when scene features are within the depth of field (i.e., scene depth = 0), the frequency content of the first sub-image is substantially and / or approximately the same as the frequency content of the second sub-image (e.g., the signal power per frequency differs by no more than a threshold amount). Therefore, the pixel value of the focused pixel can be obtained by adding the values ​​of the corresponding pixels in the first and second sub-images without applying unequal weights to these values ​​to improve sharpness.

[0085] When the scene feature is located on the first side of the depth of field (i.e., scene depth = -1), the first (e.g., left) sub-image provides higher frequency content for out-of-focus pixels located in the first (e.g., left) half of the dual-pixel image (e.g., quadrants 600 and 604), and the second (e.g., right) sub-image provides higher frequency content for out-of-focus pixels located in the second (e.g., right) half of the dual-pixel image (e.g., quadrants 602 and 606). When the scene feature is located on the second side of the depth of field (i.e., scene depth = +1), the second sub-image provides higher frequency content for out-of-focus pixels located in the first half of the dual-pixel image, and the first sub-image provides higher frequency content for out-of-focus pixels located in the second half of the dual-pixel image.

[0086] Figures 7A-7F An example mapping is shown between the position of a corresponding pixel within four-pixel image data, the depth value associated with that pixel, and a sub-image containing high-frequency image content. Specifically, Figures 7A-7F Each of the images shows a four-pixel image region divided into quadrants labeled with corresponding depths on the left, and sub-image quadrants on the right, which provide high-frequency content for each quadrant of the four-pixel image given the corresponding depth.

[0087] Figure 7A It is shown that when quadrants 700, 702, 704, and 706 of a four-pixel image sensor are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), then: (i) quadrant 710 of the first four-pixel sub-image contains higher frequency content of quadrant 700 than the corresponding quadrants of the other three four-pixel sub-images (e.g., quadrant 740 of the correspondingly located fourth four-pixel sub-image); (ii) quadrant 722 of the second four-pixel sub-image contains higher frequency content of quadrant 700 than the corresponding quadrants of the other three four-pixel sub-images (e.g., quadrant 740 of the corresponding fourth four-pixel sub-image). (iii) The quadrant 732 of the corresponding sub-image contains higher frequency content of quadrant 702, (iv) the quadrant 734 of the third four-pixel sub-image contains higher frequency content of quadrant 704 than the corresponding quadrants of the other three four-pixel sub-images (e.g., the quadrant 724 of the corresponding second four-pixel sub-image), and (iv) the quadrant 746 of the fourth four-pixel sub-image contains higher frequency content of quadrant 706 than the corresponding quadrants of the other three four-pixel sub-images (e.g., the quadrant 716 of the corresponding first four-pixel sub-image).

[0088] Figure 7BIt is shown that when quadrants 700, 702, 704, and 706 of a four-pixel image sensor are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), then: (i) quadrant 740 of the fourth sub-image contains higher frequency content of quadrant 700 than the corresponding quadrants of the other three sub-images (e.g., quadrant 710 of the correspondingly located first sub-image); (ii) quadrant 732 of the third sub-image contains higher frequency content of quadrant 700 than the corresponding quadrants of the other three sub-images (e.g., quadrant 710 of the corresponding first sub-image). (iii) The corresponding quadrant 722 of the sub-image contains higher frequency content of quadrant 702, (iv) the quadrant 724 of the second sub-image contains higher frequency content of quadrant 704 than the corresponding quadrants of the other three sub-images (e.g., the corresponding quadrant 734 of the third sub-image), and (iv) the quadrant 716 of the first sub-image contains higher frequency content of quadrant 706 than the corresponding quadrants of the other three sub-images (e.g., the corresponding quadrant 746 of the fourth sub-image).

[0089] Figure 7C It is shown that when quadrants 700 and 704 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), and quadrants 702 and 706 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), then: (i) quadrants 710 and 716 of the first sub-image contain higher frequency content of quadrants 700 and 706 than the corresponding quadrants of the other three sub-images (e.g., quadrants 740 and 746 of the fourth sub-image, respectively), and (ii) quadrants 732 and 734 of the third sub-image contain higher frequency content of quadrants 702 and 704 than the corresponding quadrants of the other three sub-images (e.g., quadrants 722 and 724 of the second sub-image, respectively).

[0090] Figure 7D It is shown that when quadrants 700 and 704 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), and quadrants 702 and 706 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), then: (i) quadrants 740 and 746 of the fourth sub-image contain higher frequency content of quadrants 700 and 706 than the corresponding quadrants of the other three sub-images (e.g., quadrants 710 and 716 of the first sub-image, respectively), and (ii) quadrants 722 and 724 of the second sub-image contain higher frequency content of quadrants 702 and 704 than the corresponding quadrants of the other three sub-images (e.g., quadrants 732 and 734 of the third sub-image, respectively).

[0091] Figure 7EIt is shown that when quadrants 700 and 702 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), and quadrants 704 and 706 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), then: (i) quadrants 740 and 746 of the fourth sub-image contain higher frequency content of quadrants 700 and 706 than the corresponding quadrants of the other three sub-images (e.g., quadrants 710 and 716 of the first sub-image, respectively), and (ii) quadrants 732 and 734 of the third sub-image contain higher frequency content of quadrants 702 and 704 than the corresponding quadrants of the other three sub-images (e.g., quadrants 722 and 724 of the second sub-image, respectively).

[0092] Figure 7F It is shown that when quadrants 700 and 702 are each used to image a plane located on the first side of the depth of field (as shown in "Depth: -1"), and quadrants 704 and 706 are each used to image a plane located on the second side of the depth of field (as shown in "Depth: +1"), then: (i) quadrants 710 and 716 of the first sub-image contain higher frequency content of quadrants 700 and 706 than the corresponding quadrants of the other three sub-images (e.g., quadrants 740 and 746 of the fourth sub-image, respectively), and (ii) quadrants 722 and 724 of the second sub-image contain higher frequency content of quadrants 702 and 704 than the corresponding quadrants of the other three sub-images (e.g., quadrants 732 and 734 of the third sub-image, respectively).

[0093] Figure 7G Table 750 is shown, which summarizes the results of... Figures 7A-7F The example shows the relationship between the combination of pixel position and pixel depth. Specifically, when the scene features are within the depth of field (i.e., scene depth = 0), the frequency contents of the first, second, third, and fourth sub-images are substantially and / or approximately the same. Therefore, the pixel value of the focused pixel can be obtained by summing the values ​​of corresponding pixels in the first to fourth sub-images without applying unequal weights to these values ​​to improve sharpness.

[0094] When the scene features are located on the first side of the depth of field (i.e., scene depth = -1), the first sub-image provides higher frequency content for out-of-focus pixels located in the first quadrant (e.g., quadrant 700) of the four-pixel image, the second sub-image provides higher frequency content for out-of-focus pixels located in the second quadrant (e.g., quadrant 702) of the four-pixel image, the third sub-image provides higher frequency content for out-of-focus pixels located in the third quadrant (e.g., quadrant 704) of the four-pixel image, and the fourth sub-image provides higher frequency content for out-of-focus pixels located in the fourth quadrant (e.g., quadrant 706) of the four-pixel image.

[0095] When scene features are located on the second side of the depth of field (i.e., scene depth = +1), the fourth sub-image provides higher frequency content for out-of-focus pixels in the first quadrant (e.g., quadrant 700) of the four-pixel image, the third sub-image provides higher frequency content for out-of-focus pixels in the second quadrant (e.g., quadrant 702) of the four-pixel image, the second sub-image provides higher frequency content for out-of-focus pixels in the third quadrant (e.g., quadrant 704) of the four-pixel image, and the first sub-image provides higher frequency content for out-of-focus pixels in the fourth quadrant (e.g., quadrant 706) of the four-pixel image.

[0096] Although Figure 6A – Figure 6G and Figures 7A-7G The relationship between pixel location, pixel depth, and high-frequency data source is shown at the level of each quadrant (for clarity), but in practice, the determination of high-frequency pixel sources can be performed at the level of each pixel (based on each pixel location and each pixel depth). Therefore, based on (i) the corresponding depth of a particular pixel and (ii) the half of the separated pixel image in which the particular pixel is located (in the case of a two-pixel image) or quadrant (in the case of a four-pixel image), it can be used... Figures 6A-6G and / or Figures 7A-7G The relationships shown are used to determine the high-frequency sources of a specific pixel. Accordingly, although... Figures 6A-6F and Figures 7A-7F For the purpose of explanation Figure 6G and Figure 7G The purpose of the relationships shown is to illustrate a specific arrangement of pixel depth, but in practice, different distributions and / or combinations of pixel depth can be observed. Figure 6G and Figure 7G The relationship shown can be used to identify corresponding sub-images containing high-frequency information that may not exist in other separated pixel images for each pixel in the separated pixel image.

[0097] VII. Example Pixel Value Merging

[0098] Turn back Figure 5The pixel value merger 526 can be configured to combine corresponding pixel values ​​of (multiple) high-frequency pixel sources 522 and (multiple) low-frequency pixel sources 524 in various ways. Specifically, for (multiple) out-of-focus pixels 516, the pixel value merger 526 can be configured to favor (multiple) high-frequency pixel sources 522, or to give (multiple) high-frequency pixel sources 522 a greater weight than (multiple) low-frequency pixel sources 524. In some cases, (multiple) high-frequency pixel sources 522 can be given a greater weight than (multiple) low-frequency pixel sources 524 relative to all frequencies. In other cases, (multiple) high-frequency pixel sources 522 can be given a greater weight than (multiple) low-frequency pixel sources 524 relative to frequencies above a threshold frequency (e.g., a cutoff frequency that separates the frequency content of different sub-images), while frequencies below or equal to the threshold frequency can be given equal weight.

[0099] In one example, the combination of pixel values ​​can be performed in the spatial domain. A given pixel in the enhanced image 528 corresponding to the out-of-focus pixel of the separated pixel image data 502 (i.e., where...) ) can be represented as Where P i,j P represents the pixel value at pixel position (i.e., coordinates) i, j in a given image. ENHANCED This indicates an enhanced image of 528, P. 1 Sub-image 504, P N Sub-image 506, D i,j Let w1(i,j,D) represent the depth value associated with the pixel at pixel position i,j, and let w1(i,j,D) represent the depth value associated with the pixel at pixel position i,j. i,j ) to w N (i, j, D) i,j ) is a weighting function configured to favor the corresponding high-frequency pixel source 522 relative to the low-frequency pixel source 524. In some cases, for each corresponding pixel position (i, j), the sum of weights can be normalized to a predetermined value, such as N (e.g., ). A given pixel in the enhanced image 528 corresponding to the focused pixel of the separated pixel image data 502 (i.e., where D i,j ∈DOF) can be represented as Where w1(i,j,D) i,j )=...=w N (i, j, D) i,j ) = 1.

[0100] In one example, in the case of two-pixel image data, when P 1 It is a high-frequency pixel source at pixel positions i and j (and P) 2Therefore, when it is a low-frequency pixel source, In other words, for a given out-of-focus pixel in the separated pixel image data 502, pixel values ​​from high-frequency pixel sources can be weighted more heavily (i.e., assigned a larger weight) than pixel values ​​from low-frequency pixel sources in order to sharpen the given out-of-focus pixel. In another example, in the case of dual-pixel image data, when P 1 When P2 is a high-frequency pixel source at pixel positions i and j (and P2 is therefore a low-frequency pixel source), In other words, for a given out-of-focus pixel in the separated pixel image data 502, a high-frequency pixel source can be selected as the exclusive source of the pixel value, and a low-frequency pixel source can be discarded.

[0101] Weighting function w1(i,j,D) i,j ) to w N (i, j, D) i,j ) can be D i,j Discrete or continuous functions. Weighted function w1(i, j, D) i,j ) to w N (i, j, D) i,j ) can be D i,j A linear function, or it could be D i,j The exponential function, and other possibilities. The weights assigned to a particular high-frequency pixel source by its corresponding weighting function can cause the high-frequency pixel source to form 50% of the signal of the corresponding pixel in (i) enhanced image 528 (e.g., when D i,j (ii) Enhance 100% of the signal of the corresponding pixel in image 528 when |D ∈ DOF (e.g., when |D ∈ DOF) and (ii) enhance 100% of the signal of the corresponding pixel in image 528 (e.g., when |D ∈ DOF) i,j |>D THRESHOLD Between (time).

[0102] In another example, the combination of pixel values ​​can be performed in the frequency domain. For example, a given pixel in the enhanced image 528 corresponding to the out-of-focus pixel of the separated pixel image data 502 (i.e., where...). ) can be represented as Where ω=(ω x ω y F represents the horizontal and vertical spatial frequencies that may exist in the separated pixel image data 502, respectively. ENHANCED Represents the enhanced image 528, F in the frequency domain 1 (ω) represents the sub-image 504 in the frequency domain, F N (ω) represents the sub-image 506 in the frequency domain, D i,jRepresents the depth value associated with the pixel at coordinates i, j. IFT() represents the inverse frequency transform (e.g., inverse Fourier transform, inverse cosine transform, etc.), and v1(ω, i, j, D) i,j ) to v N (ω, i, j, D) i,j ) is a frequency-specific weighting function configured to boost the high frequencies present in the corresponding high-frequency pixel source(s) 522 relative to the low frequencies present in both the corresponding high-frequency pixel source(s) 522 and the corresponding low-frequency pixel source(s) 524. In some cases, for each corresponding spatial frequency ω of a given pixel, the sum of weights can be normalized to a predetermined value, such as N (e.g., ). A given pixel in the enhanced image 528 corresponding to the focused pixel of the separated pixel image data 502 (i.e., where D i,j ∈DOF) can be represented as Where v1(ω, i, j, D) i,j )=...=v N (ω, i, j, D) i,j ) = 1.

[0103] In some implementations, the weighting function v1(ω, i, j, D) i,j ) to v N (ω, i, j, D) i,j This can additionally or alternatively be a function of the difference in frequency content between sub-images. For example, for spatial frequencies above a threshold frequency (i.e., When the difference between a high-frequency pixel source and a low-frequency pixel source exceeds a threshold difference, the high-frequency pixel source can be weighted more heavily than the low-frequency pixel source. Conversely, if the difference between these two pixel sources is less than or equal to the threshold difference, they can be weighted equally. For spatial frequencies below or equal to the threshold frequency (i.e.,...), High-frequency pixel sources and low-frequency pixel sources can be weighted equally.

[0104] Therefore, in the case of two-pixel image data, When F 1 (ω)-F 2 (ω)>F THRESHOLD Time (where F) 1 (ω) is a high-frequency pixel source), W1(ω, i, j, D) i,j )>W2(ω,i,j,D i,j When F 2 (ω)-F 1 (ω)>F THRESHOLD Time (where F) 2 (ω) is a high-frequency pixel source), W2(ω, i, j, D) i,j)>W1(ω,i,j,D i,j ), and when |F 1 (ω)-F 2 (ω)|≤F THRESHOLD At that time, W1(ω, i, j, D) i,j )=W2(ω,i,j,D i,j ). W1(ω, i, j, D) i,j )=W2(ω,i,j,D i,j ).

[0105] The threshold frequency can be based on (e.g., equal to) the cutoff frequency of the low-frequency pixel source. The threshold difference can be based on the expected noise level present in the sub-image at different frequencies (e.g., greater than twice the average or peak of that noise level). Therefore, for out-of-focus pixels, the high-frequency content of (multiple) high-frequency pixel sources 522 (which is not present in the low-frequency pixel sources) can be boosted to sharpen the separated pixel image, while the low-frequency content present in both pixel sources 522 and 524 can be equally weighted to preserve the content of both sources at these frequencies.

[0106] Weighting function v1(ω, i, j, D) i,j ) to v N (ω, i, j, D) i,j ) can be D i,j Discrete or continuous functions, and / or may be D i,j The linear or exponential function, and other possibilities. The weights assigned by its corresponding weighting function to the specific frequencies present in the high-frequency pixel sources can result in the high-frequency pixel sources forming 50% of the frequency-specific signal of the corresponding pixels in (i) the enhanced image 528 (e.g., when D i,j When ∈DOF and / or when ω≤ω THRESHOLD (i) 100% of the frequency of a specific signal of the corresponding pixel in image 528 (e.g., when |D i,j |>D THRESHOLD When and / or when ω > ω THRESHOLD Between (time).

[0107] In another example, one or more algorithms configured to merge an image focus stack can be used to perform the combination of pixel values. Specifically, the image focus stack can include multiple images, each captured at a corresponding different focus point (resulting in different depths of focus and / or depth-of-field locations). Thus, the different images in the image focus stack can include different in-focus and out-of-focus portions. Algorithms configured to merge the image focus stack can involve, for example: (i) calculating a weight per pixel based on pixel contrast and using the weight per pixel to combine the images in the focus stack, (ii) determining a depth map for each image in the focus stack and using the depth map to identify the sharpest pixel value in the focus stack, and / or (iii) using a pyramid-based method to identify the sharpest pixel, and other possibilities.

[0108] In the context of non-separated pixel image data, the image focus stack may include motion blur because different images in the image focus stack are captured at different times. Therefore, for scenes involving motion, reconstructing an image with enhanced sharpness can be difficult.

[0109] In the context of split-pixel image data, multiple sub-images of the split-pixel image data can be used to form an image focus stack. Variations in the spatial frequency content of sub-images in the out-of-focus regions of the split-pixel image data can approximate different focus levels (and thus also different depths of focus and depth-of-field positions). Because variations in the spatial frequency content of the sub-images can be achieved without explicitly adjusting the focal length of the split-pixel camera, the split-pixel sub-images can be captured as part of a single exposure, thus potentially excluding motion blur (or at least including less motion blur than a comparable non-split-pixel image focus stack). Accordingly, one or more focus stacking algorithms can be used to merge the split-pixel sub-images to generate enhanced images of static and / or dynamic scenes.

[0110] VIII. Additional Example Operations

[0111] Figure 8 A flowchart is shown relating to operations related to generating an image with enhanced sharpness and / or depth of field. These operations can be performed by computing device 100, computing system 200, and / or system 500, among other possibilities. Figure 8 The embodiments can be simplified by removing any one or more features shown therein. Furthermore, these embodiments can be combined with any features, aspects, and / or implementations of any of the previously drawn figures or otherwise described herein.

[0112] Box 800 may involve obtaining split-pixel image data captured by a split-pixel camera, wherein the split-pixel image data includes a first sub-image and a second sub-image.

[0113] Box 802 may involve: for each corresponding pixel among multiple pixels of the separated pixel image data, determining the corresponding position of the scene feature represented by the corresponding pixel relative to the depth of field of the separated pixel camera.

[0114] Box 804 may involve: identifying one or more out-of-focus pixels among a plurality of pixels, based on the corresponding position of scene features represented by each corresponding pixel among a plurality of pixels, wherein the one or more out-of-focus pixels are located outside the depth of field.

[0115] Box 806 may involve determining a corresponding pixel value for each of one or more defocus pixels based on: (i) the corresponding position of the scene feature represented by the corresponding defocus pixel relative to the depth of field, (ii) the position of the corresponding defocus pixel within the separated pixel image data, and (iii) at least one of a first value of the corresponding first pixel in the first sub-image or a second value of the corresponding second pixel in the second sub-image.

[0116] Box 808 may involve generating an enhanced image with extended depth of field based on the corresponding pixel value determined for each respective out-of-focus pixel.

[0117] In some embodiments, determining the corresponding pixel value may include: for each corresponding out-of-focus pixel, selecting either a corresponding first pixel or a corresponding second pixel as the source pixel of the corresponding out-of-focus pixel based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data. The corresponding pixel value may be determined for each corresponding out-of-focus pixel based on the value of the source pixel.

[0118] In some embodiments, for off-focus pixels located in the first half of the separated pixel image data: when the scene feature represented by the corresponding off-focus pixel is behind the depth of field, the corresponding first pixel can be selected as the source pixel, and when the scene feature represented by the corresponding off-focus pixel is in front of the depth of field, the corresponding second pixel can be selected as the source pixel. For off-focus pixels located in the second half of the separated pixel image data: when the scene feature represented by the corresponding off-focus pixel is behind the depth of field, the corresponding second pixel can be selected as the source pixel, and when the scene feature represented by the corresponding off-focus pixel is in front of the depth of field, the corresponding first pixel can be selected as the source pixel.

[0119] In some embodiments, selecting a source pixel for a corresponding off-focus pixel may include: for an off-focus pixel located in a first half of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the first half of the separated pixel image data, selecting a corresponding first pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding second pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0120] Selecting a source pixel for a corresponding off-focus pixel may additionally include: for an off-focus pixel located in the second half of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may also include: for an off-focus pixel located in the second half of the separated pixel image data, selecting a corresponding second pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding first pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0121] In some embodiments, determining the corresponding pixel value may include: for each corresponding out-of-focus pixel, determining a first weight for the corresponding first pixel and a second weight for the corresponding second pixel based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data. The corresponding pixel value may be determined for each corresponding out-of-focus pixel based on (i) the first product of the first weight and the first value of the corresponding first pixel and (ii) the second product of the second weight and the second value of the corresponding second pixel.

[0122] In some embodiments, for defocused pixels located in the first half of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the first weight of the corresponding first pixel can be greater than the second weight of the corresponding second pixel; and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the second weight of the corresponding second pixel can be greater than the first weight of the corresponding first pixel. For defocused pixels located in the second half of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the second weight of the corresponding second pixel can be greater than the first weight of the corresponding first pixel; and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the first weight of the corresponding first pixel can be greater than the second weight of the corresponding second pixel.

[0123] In some embodiments, determining the first weight and the second weight may include: for a defocused pixel located in a first half of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first weight and the second weight may further include: for a defocused pixel located in the first half of the separated pixel image data, determining a first weight greater than the second weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a second weight greater than the first weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0124] Determining the first and second weights may additionally include: for a defocused pixel located in the second half of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first and second weights may also include: for a defocused pixel located in the second half of the separated pixel image data, determining a second weight greater than the first weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a first weight greater than the second weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0125] In some embodiments, determining the first weight and the second weight may include: for each of a plurality of spatial frequencies present in the separated pixel image data, determining (i) a first amplitude of the corresponding spatial frequency in a first sub-image, and (ii) a second amplitude of the corresponding spatial frequency in a second sub-image. A difference between the first amplitude and the second amplitude may be determined for each corresponding spatial frequency. A first weight for the first amplitude and a second weight for the second amplitude may be determined for each corresponding spatial frequency. For each corresponding spatial frequency that is above a threshold frequency and associated with a difference exceeding the threshold, the first weight may differ from the second weight, and the first weight and the second weight may be based on the corresponding position of the scene feature represented by the corresponding defocus pixel relative to the depth of field. For each corresponding spatial frequency that is (i) below a threshold frequency or (ii) above a threshold frequency and associated with a difference not exceeding the threshold, the first weight may be equal to the second weight. A corresponding pixel value may be determined for each corresponding defocus pixel based on the sum of (i) a first plurality of products of the first weight and the first amplitude of each corresponding spatial frequency represented by a first value of the corresponding first pixel and (ii) a second plurality of products of the second weight and the second amplitude of each corresponding spatial frequency represented by a second value of the corresponding second pixel.

[0126] In some embodiments, the separated pixel image data may include a first sub-image, a second sub-image, a third sub-image, and a fourth sub-image. Determining the corresponding pixel value may include: for each corresponding out-of-focus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data, selecting one of the corresponding first pixel in the first sub-image, the corresponding second pixel in the second sub-image, the corresponding third pixel in the third sub-image, or the corresponding fourth pixel in the fourth sub-image as the source pixel of the corresponding out-of-focus pixel. The corresponding pixel value may be determined for each corresponding out-of-focus pixel based on the value of the source pixel.

[0127] In some embodiments, for defocused pixels located in the first quadrant of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the corresponding first pixel can be selected as the source pixel, and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the corresponding fourth pixel can be selected as the source pixel. For defocused pixels located in the second quadrant of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the corresponding second pixel can be selected as the source pixel, and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the corresponding third pixel can be selected as the source pixel. For defocused pixels located in the third quadrant of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the corresponding third pixel can be selected as the source pixel, and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the corresponding second pixel can be selected as the source pixel. For out-of-focus pixels located in the fourth quadrant of the separated pixel image data: when the scene feature represented by the corresponding out-of-focus pixel is behind the depth of field, the corresponding fourth pixel can be selected as the source pixel, and when the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the corresponding first pixel can be selected as the source pixel.

[0128] In some embodiments, selecting a source pixel for a corresponding off-focus pixel may include: for an off-focus pixel located in a first quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the first quadrant of the separated pixel image data, selecting a corresponding first pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding fourth pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0129] Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the second quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the second quadrant of the separated pixel image data, selecting a corresponding second pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding third pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0130] Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the third quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may further include: for an off-focus pixel located in the third quadrant of the separated pixel image data, selecting a corresponding third pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding second pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0131] Selecting a source pixel for a corresponding off-focus pixel may additionally include: for an off-focus pixel located in the fourth quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding off-focus pixel is located behind or before the depth of field. Selecting a source pixel for a corresponding off-focus pixel may also include: for an off-focus pixel located in the fourth quadrant of the separated pixel image data, selecting a corresponding fourth pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is behind the depth of field, or selecting a corresponding first pixel as a source pixel based on and / or in response to determining that the scene feature represented by the corresponding off-focus pixel is before the depth of field.

[0132] In some embodiments, the separated pixel image data may include a first sub-image, a second sub-image, a third sub-image, and a fourth sub-image. Determining the corresponding pixel value may include: for each corresponding defocus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding defocus pixel relative to the depth of field and (ii) the position of the corresponding defocus pixel within the separated pixel image data, determining a first weight for the corresponding first pixel, a second weight for the corresponding second pixel, a third weight for the corresponding third pixel in the third sub-image, and a fourth weight for the corresponding fourth pixel in the fourth sub-image. The corresponding pixel value may be determined for each corresponding defocus pixel based on (i) the first product of the first weight and the first value of the corresponding first pixel, (ii) the second product of the second weight and the second value of the corresponding second pixel, (iii) the third product of the third weight and the third value of the corresponding third pixel, and (ii) the fourth product of the fourth weight and the fourth value of the corresponding fourth pixel.

[0133] In some embodiments, for defocused pixels located in the first quadrant of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the first weight of the corresponding first pixel can be greater than the second weight of the corresponding second pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel; and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the fourth weight of the corresponding fourth pixel can be greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the third weight of the corresponding third pixel. For defocused pixels located in the second quadrant of the separated pixel image data: when the scene feature represented by the corresponding defocused pixel is behind the depth of field, the second weight of the corresponding second pixel can be greater than the first weight of the corresponding first pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel; and when the scene feature represented by the corresponding defocused pixel is in front of the depth of field, the third weight of the corresponding third pixel can be greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the fourth weight of the corresponding fourth pixel. For out-of-focus pixels located in the third quadrant of the separated pixel image data: when the scene feature represented by the corresponding out-of-focus pixel is behind the depth of field, the third weight of the corresponding third pixel can be greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the fourth weight of the corresponding fourth pixel. Furthermore, when the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the second weight of the corresponding second pixel can be greater than the first weight of the corresponding first pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel. For out-of-focus pixels located in the fourth quadrant of the separated pixel image data: when the scene feature represented by the corresponding out-of-focus pixel is behind the depth of field, the fourth weight of the corresponding fourth pixel can be greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the third weight of the corresponding third pixel. Furthermore, when the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the first weight of the corresponding first pixel can be greater than the second weight of the corresponding second pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel.

[0134] In some embodiments, determining the first weight, second weight, third weight, and fourth weight may include: for a defocused pixel located in a first quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first weight, second weight, third weight, and fourth weight may further include: for a defocused pixel located in the first quadrant of the separated pixel image data, determining a first weight greater than the second weight, third weight, and fourth weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a fourth weight greater than the first weight, second weight, and third weight based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0135] Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the second quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the second quadrant of the separated pixel image data, determining a second weight greater than the first, third, and fourth weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a third weight greater than the first, second, and fourth weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0136] Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the third quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the third quadrant of the separated pixel image data, determining a third weight greater than the first, second, and fourth weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a second weight greater than the first, third, and fourth weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0137] Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the fourth quadrant of the separated pixel image data, determining whether the scene feature represented by the corresponding defocused pixel is located behind or before the depth of field. Determining the first, second, third, and fourth weights may further include: for a defocused pixel located in the fourth quadrant of the separated pixel image data, determining a fourth weight greater than the first, second, and third weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is behind the depth of field, or determining a first weight greater than the second, third, and fourth weights based on and / or in response to determining that the scene feature represented by the corresponding defocused pixel is before the depth of field.

[0138] In some embodiments, the corresponding sub-image of the split-pixel image data may have already been captured as part of a single exposure by the corresponding photosensitive point of the split-pixel camera.

[0139] In some embodiments, one or more focal pixels can be identified based on the corresponding location of scene features represented by each corresponding pixel among a plurality of pixels. The scene features represented by the one or more focal pixels may be located within the depth of field. For each corresponding focal pixel among the one or more focal pixels, the corresponding pixel value can be determined by adding a first value of the corresponding first pixel in the first sub-image and a second value of the corresponding second pixel in the second sub-image. An enhanced image can be further generated based on the corresponding pixel values ​​determined for each corresponding focal pixel.

[0140] In some embodiments, determining the corresponding position of a scene feature represented by a corresponding pixel relative to the depth of field of a split-pixel camera may include: for each corresponding pixel in a plurality of pixels of the split-pixel image data, determining the difference between a corresponding first pixel in a first sub-image and a corresponding second pixel in a second sub-image, and for each corresponding pixel in a plurality of pixels of the split-pixel image data, determining the corresponding position based on the difference.

[0141] IX. Conclusion

[0142] This disclosure is not intended to limit the scope of the specific embodiments described herein, which are intended to illustrate various aspects. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from its scope. In addition to the methods and apparatus described herein, functionally equivalent methods and apparatuses within the scope of this disclosure will be apparent to those skilled in the art based on the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0143] The above detailed description, with reference to the accompanying drawings, illustrates various features and operations of the disclosed systems, devices, and methods. In the drawings, similar symbols generally identify similar components unless the context otherwise indicates. The exemplary embodiments described herein and in the drawings are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein. It will be readily understood that, as generally described herein and shown in the accompanying drawings, aspects of this disclosure can be arranged, replaced, combined, separated, and designed in a variety of different configurations.

[0144] Regarding any or all message flow diagrams, scenarios, and flowcharts in the accompanying drawings, and as discussed herein, each step, block, and / or communication may represent the processing and / or transmission of information according to exemplary embodiments. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages may not be performed in the order shown or discussed, including performing them substantially concurrently or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or operations may be associated with any message flow diagrams, scenarios, and flowcharts discussed herein. Figure 1 These message flow diagrams, scenarios, and flowcharts can be used together, and they can be combined with each other, either partially or entirely.

[0145] The steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, the blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). Program code may include one or more instructions executable by a processor to implement specific logical operations or actions in the method or technique. Program code and / or associated data may be stored on any type of computer-readable medium, such as storage devices including random access memory (RAM), disk drives, solid-state drives, or other storage media.

[0146] Computer-readable media can also include non-transitory computer-readable media, such as short-term data storage media like register memory, processor cache, and RAM. Computer-readable media can also include long-term storage media for program code and / or data. Therefore, computer-readable media can include secondary or permanent long-term storage, such as read-only memory (ROM), optical discs or disks, solid-state drives, and optical disc read-only memory (CD-ROM). Computer-readable media can also be any other volatile or non-volatile storage system. For example, a computer-readable medium can be considered a computer-readable storage medium or a tangible storage device.

[0147] Furthermore, a step or block representing one or more information transfers may correspond to information transfers between software and / or hardware modules within the same physical device. However, other information transfers may occur between software and / or hardware modules in different physical devices.

[0148] The specific arrangements shown in the accompanying drawings should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in the given drawings. Furthermore, some of the elements shown may be combined or omitted. Additionally, exemplary embodiments may include elements not shown in the drawings.

[0149] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and not for limitation, and the true scope is indicated by the appended claims.

Claims

1. A computer-implemented method for processing image data, the method comprising: Obtain split-pixel image data captured by a split-pixel camera including an image sensor, wherein each pixel in the image sensor includes a first photosensitive point and a second photosensitive point, and wherein the split-pixel image data includes a first sub-image corresponding to image data generated by the first photosensitive point and a second sub-image corresponding to image data generated by the second photosensitive point; For each corresponding pixel among multiple pixels of the separated pixel image data, determine the corresponding position of the scene feature represented by the corresponding pixel relative to the depth of field of the separated pixel camera; Based on the corresponding position of the scene feature represented by each of the plurality of pixels, one or more off-focus pixels are identified among the plurality of pixels, wherein the scene feature represented by the one or more off-focus pixels is located outside the depth of field; For each of the one or more defocused pixels, a corresponding pixel value is determined based on (i) the corresponding position of the scene feature represented by the corresponding defocused pixel relative to the depth of field, (ii) the position of the corresponding defocused pixel within the separated pixel image data, and (iii) at least one of a first value of the corresponding first pixel in the first sub-image or a second value of the corresponding second pixel in the second sub-image. as well as An enhanced image with extended depth of field is generated based on the corresponding pixel value determined for each defocused pixel.

2. The computer-implemented method according to claim 1, wherein, Determining the corresponding pixel value includes: For each corresponding out-of-focus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data, one of the corresponding first pixel or the corresponding second pixel is selected as the source pixel of the corresponding out-of-focus pixel; and For each corresponding out-of-focus pixel, the corresponding pixel value is determined based on the value of the source pixel.

3. The computer-implemented method according to claim 2, wherein: For the off-focus pixels located in the first half of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding first pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding second pixel is selected as the source pixel; as well as For the off-focus pixels located in the second half of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding second pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding first pixel is selected as the source pixel.

4. The computer-implemented method according to claim 1, wherein, Determining the corresponding pixel value includes: For each corresponding out-of-focus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data, a first weight for the corresponding first pixel and a second weight for the corresponding second pixel are determined; and For each corresponding out-of-focus pixel, the corresponding pixel value is determined based on (i) the first product of the first weight and the first value of the corresponding first pixel and (ii) the second product of the second weight and the second value of the corresponding second pixel.

5. The computer-implemented method according to claim 4, wherein: For the off-focus pixels located in the first half of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the first weight of the corresponding first pixel is greater than the second weight of the corresponding second pixel; as well as When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the second weight of the corresponding second pixel is greater than the first weight of the corresponding first pixel; And for the off-focus pixels located in the second half of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the second weight of the corresponding second pixel is greater than the first weight of the corresponding first pixel; as well as When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the first weight of the corresponding first pixel is greater than the second weight of the corresponding second pixel.

6. The computer-implemented method according to any one of claims 4-5, wherein, Determining the first weight and the second weight includes: For each of the multiple spatial frequencies present in the separated pixel image data, determine (i) the first amplitude of the corresponding spatial frequency in the first sub-image and (ii) the second amplitude of the corresponding spatial frequency in the second sub-image; For each corresponding spatial frequency, determine the difference between the first amplitude and the second amplitude; For each corresponding spatial frequency, a first weight for the first amplitude and a second weight for the second amplitude are determined, wherein: For each corresponding spatial frequency above a threshold frequency and associated with a difference exceeding the threshold, the first weight differs from the second weight, and the first weight and the second weight are based on the corresponding position of the scene features represented by the corresponding out-of-focus pixels relative to the depth of field. For each corresponding spatial frequency that is (i) below the threshold frequency or (ii) above the threshold frequency and associated with a difference not exceeding the threshold, the first weight is equal to the second weight; and For each corresponding out-of-focus pixel, the corresponding pixel value is determined based on (i) a first plurality of products of a first weight and a first amplitude of each corresponding spatial frequency represented by a first value of the corresponding first pixel and (ii) a second plurality of products of a second weight and a second amplitude of each corresponding spatial frequency represented by a second value of the corresponding second pixel.

7. The computer-implemented method according to claim 1, wherein, The separated pixel image data includes a first sub-image, a second sub-image, a third sub-image, and a fourth sub-image, wherein determining the corresponding pixel value includes: For each corresponding out-of-focus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data, one of the corresponding first pixel in the first sub-image, the corresponding second pixel in the second sub-image, the corresponding third pixel in the third sub-image, or the corresponding fourth pixel in the fourth sub-image is selected as the source pixel of the corresponding out-of-focus pixel; and For each corresponding out-of-focus pixel, the corresponding pixel value is determined based on the value of the source pixel.

8. The computer-implemented method according to claim 7, wherein: For defocused pixels located in the first quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding first pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding fourth pixel is selected as the source pixel; For defocused pixels located in the second quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding second pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding third pixel is selected as the source pixel; For defocused pixels located in the third quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding third pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding second pixel is selected as the source pixel; as well as For the off-focus pixels located in the fourth quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located beyond the depth of field, the corresponding fourth pixel is selected as the source pixel; and When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the corresponding first pixel is selected as the source pixel.

9. The computer-implemented method according to claim 1, wherein, The separated pixel image data includes a first sub-image, a second sub-image, a third sub-image, and a fourth sub-image, wherein determining the corresponding pixel value includes: For each corresponding out-of-focus pixel, based on (i) the corresponding position of the scene feature represented by the corresponding out-of-focus pixel relative to the depth of field and (ii) the position of the corresponding out-of-focus pixel within the separated pixel image data, determine the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, the third weight of the corresponding third pixel in the third sub-image, and the fourth weight of the corresponding fourth pixel in the fourth sub-image; and For each corresponding out-of-focus pixel, the corresponding pixel value is determined based on (i) the first product of the first weight and the first value of the corresponding first pixel, (ii) the second product of the second weight and the second value of the corresponding second pixel, (iii) the third product of the third weight and the third value of the corresponding third pixel, and (ii) the fourth product of the fourth weight and the fourth value of the corresponding fourth pixel.

10. The computer-implemented method according to claim 9, wherein: For defocused pixels located in the first quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the first weight of the corresponding first pixel is greater than the second weight of the corresponding second pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel. as well as When the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the fourth weight of the corresponding fourth pixel is greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the third weight of the corresponding third pixel. For defocused pixels located in the second quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the second weight of the corresponding second pixel is greater than the first weight of the corresponding first pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel. as well as When the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the third weight of the corresponding third pixel is greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the fourth weight of the corresponding fourth pixel. For defocused pixels located in the third quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the third weight of the corresponding third pixel is greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the fourth weight of the corresponding fourth pixel. as well as When the scene feature represented by the corresponding out-of-focus pixel is located in front of the depth of field, the second weight of the corresponding second pixel is greater than the first weight of the corresponding first pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel; And for the off-focus pixels located in the fourth quadrant of the separated pixel image data: When the scene feature represented by the corresponding out-of-focus pixel is located behind the depth of field, the fourth weight of the corresponding fourth pixel is greater than the first weight of the corresponding first pixel, the second weight of the corresponding second pixel, and the third weight of the corresponding third pixel. as well as When the scene feature represented by the corresponding out-of-focus pixel is in front of the depth of field, the first weight of the corresponding first pixel is greater than the second weight of the corresponding second pixel, the third weight of the corresponding third pixel, and the fourth weight of the corresponding fourth pixel.

11. The computer-implemented method according to any one of claims 1-5 and 7-10, wherein, The corresponding sub-images of the separated pixel image data are captured by using the corresponding photosensitive points of the separated pixel camera as part of a single exposure.

12. The computer-implemented method according to any one of claims 1-5 and 7-10, further comprising: Based on the corresponding position of the scene feature represented by each of the plurality of pixels, one or more focused pixels among the plurality of pixels are identified, wherein the scene feature represented by the one or more focused pixels is located within the depth of field; as well as For each of the one or more focused pixels, a corresponding pixel value is determined by adding a first value of the corresponding first pixel in the first sub-image and a second value of the corresponding second pixel in the second sub-image, wherein the enhanced image is further generated based on the corresponding pixel value determined for each corresponding focused pixel.

13. The computer-implemented method according to any one of claims 1-5 and 7-10, wherein, Determining the corresponding position of the scene feature represented by the corresponding pixel relative to the depth of field of the separate pixel camera includes: For each corresponding pixel among the multiple pixels of the separated pixel image data, determine the difference between the corresponding first pixel in the first sub-image and the corresponding second pixel in the second sub-image; and For each corresponding pixel among the multiple pixels of the separated pixel image data, the corresponding position is determined based on the difference.

14. An image data processing system, comprising: processor; as well as A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a processor, cause the processor to perform a computer-implemented method according to any one of claims 1-13.

15. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a computing device, cause the computing device to perform a computer-implemented method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Method and device of imaging control of image with shallow depth of field effect as well as imaging equipment

    CN104159038A

  • Image processing method and terminal

    CN105120154A