Method and image processor unit for processing image data

The image processing method for OCL sensors enhances autofocus speed and accuracy by combining multiple views to account for variations in sensitivity and exposure, resulting in high-quality image processing and depth estimation.

JP2026501477APending Publication Date: 2026-01-16DREAM CHIP TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024531382
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing autofocus systems in digital imagers, such as those in smartphones and cameras, are slow and require mechanical lens movement, and phase detection autofocus (PDAF) systems can be improved for faster and more accurate focus estimation.

Method used

An image processing method that utilizes multiple views from an on-chip lens (OCL) sensor to capture and process sub-pixel values, accounting for variations in sensitivity, brightness, exposure, and crosstalk, and combines these views to enhance autofocus accuracy and speed.

Benefits of technology

The method allows for improved autofocus speed and accuracy by leveraging the properties of OCL sensors, increasing brightness and capturing image characteristics with longer exposure times, enabling high-quality image processing and depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501477000001_ABST
    Figure 2026501477000001_ABST
Patent Text Reader

Abstract

The disclosed invention relates to a method for generating image data (IMG) from an image sensor (4) comprising a matrix of photosensitive elements and a plurality of lens elements (7) and / or filter elements (9) arranged in a pixel matrix in front of a sub-pixel matrix of finer photosensitive elements. RAW The present invention relates to a method for processing an image of a subject, the method comprising: a group of photosensitive elements being placed behind a common lens element (7) and / or a common filter element (9) to provide subpixel values ​​for respective positions in the pixel matrix; an image sensor (4) being adapted to capture image data for a plurality of views, each view comprising a matrix of selected subpixel values ​​of the pixel matrix captured by the matrix of photosensitive elements; according to the invention, for each captured view of the image, a first variation of the subpixel values ​​in the respective view is determined separately from other views of the same image; a second variation of the subpixel values ​​associated with the same positions in the pixel matrix behind the respective lens element (7) and / or filter element (9) is determined among a set of views of the same image; the image data for the image is processed using the determined first and second variations.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for processing image data of an image sensor, the image sensor comprising a matrix of photosensitive elements and a plurality of lens elements and / or filter elements arranged in a pixel matrix in front of sub-pixel matrices of finer photosensitive elements, groups of photosensitive elements being placed behind a common lens element and / or common filter element to provide sub-pixel values ​​for their respective positions in the pixel matrix, the image sensor (4) being adapted to capture image data for a plurality of views, each view comprising a matrix of groups of selected sub-pixel values ​​of the pixel matrix captured by the matrix of photosensitive elements.

[0002] The invention further relates to an image processor unit for processing image data provided by an image sensor, the image sensor comprising a matrix of photosensitive elements and a plurality of lens elements and / or filter elements arranged in a pixel matrix in front of sub-pixel matrices of finer photosensitive elements, groups of photosensitive elements being behind a common lens element and / or common filter element to provide sub-pixel values ​​for each position in the pixel matrix, the raw image data comprising a matrix of pixel values ​​for the pixel matrix captured by the matrix of photosensitive elements, the matrix of pixel values ​​being divided into a set of views, each view comprising a matrix of sub-pixel values ​​comprising a view-specific sub-pixel of the sub-pixel matrix for each position in the matrix associated with a respective lens element and / or filter element.

[0003] The invention further relates to a computer program comprising instructions which, when the program is executed by a processing unit, cause the processing unit to carry out the steps of the above-mentioned method. [Background technology]

[0004] Digital imagers are widely used in everyday products such as smartphones, tablets, notebooks, cameras, cars, and wearables. Many of the imaging systems in these products have an autofocus function to produce clear images or sharp videos.

[0005] Traditionally, autofocus is controlled based on contrast detection, where the lens is mechanically moved to the position where the scene contrast is highest. This control process is generally slow and requires a lens mechanism to mechanically change the lens position.

[0006] The speed and accuracy of autofocusing can be improved with the technique of phase detection autofocus (PDAF) using a PDAF sensor or focus sensor with so-called phase detection (PD) pixels that are discretely and regularly located throughout the image sensor area. Phase information (parallax) estimated from the phase detection pixels is used to determine the lens position to achieve optimal focus in a predefined region of interest (ROI) in the scene. This process is typically very fast if the parallax is sufficiently accurate.

[0007] For example, omnidirectional focus sensors based on on-chip-lens sensor (OCL) technology make it possible to use all image pixels for phase detection, improving the accuracy and speed of autofocusing.

[0008] "Learning Single Camera Depth Estimation Using Dual Pixels" by R. Garg, N. Wadhwa, S. Ansari, and Y.T. Barron in Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, 2019, pp. 1556-1565, describes a learning-based method for working with dual-pixel images to estimate depth due to the inherent ambiguity in estimated depth from the use of dual pixels.

[0009] "Synthetic Depth-of-Field with a Single-Camera Mobile Phone" by N. Wadhwa, R. Garg, D.E. Jacobs, B.E. Feldman, N. Kanazawa, R. Carroll, Y. Movshovitz-Attias, J.T. Barron, Y. Pritch, and M. Levoy, in ACM Trans. Graph., vol. 37, no. 4, art. 64, pp. 64:1-64:13, describes a system for computing synthetic shallow depth-of-field images on a mobile phone by using a dual-pixel sensor per lens element. A neural network trained to segment people and accessories is combined with a sensor with dual-pixel (DP) autofocus hardware, which provides a two-sample light field with a narrow baseline to extract a dense depth map.

[0010] "Du" by Y. Zhang, N. Wadhwa, RSS Orts-Escolano, C. Haene, S. Fanello, R. Garg in Computer Vision, ECCV 2020, eds.: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm, Springer International Publishing, Cham 2020, pp. 582-598. 2 "Net: Learning Depth Estimation from Dual Cameras and Dual-Pixels" describes the use of neural networks for depth estimation, combining stereo from dual cameras with stereo from dual-pixel sensors to provide a dense depth map with sharp edges.

[0011] "Defocus Deblurring Using Dual-Pixel Data" by A. Abuolaim and M.S. Brown, presented at the European Conference on Computer Vision 2020, pp. 110-126, addresses the problem of defocus blurring in images captured with a shallow depth of field due to the use of a wide aperture. The proposed method utilizes a dual-pixel sensor that captures two sub-aperture images of a scene in a single image shot. The two sub-aperture images are used to calculate the appropriate lens position to focus on a specific scene region and are then discarded. A duplicate neural network architecture uses the discarded sub-aperture images to reduce the defocus blur.

[0012] "Ghost Free Deep High-Dynamic Range Imaging Using Focus Pixels for Complex Motion Scenes" by S. Woo, YH Ryu, and Y.O. Kim in IEEE Transactions on Image Processing, vol. 30, pp. 5001-5016, describes deep learning for seamless fusion of multi-exposure low-dynamic range images using a focal pixel sensor that provides left and right luminance images simultaneously with full-resolution RGB images.

[0013] WO2020 / 237366A1 discloses a system and method for reflection removal of an image from a dual pixel sensor by determining a first light point in a left view and a second light point in a right view and determining the disparity between the gradients. A confidence value and a weighted gradient map are determined, and a background layer and a reflection layer are achieved after iteratively minimizing a loss function.

[0014] This method is described in more detail in "Reflection Removal Using a Dual-Pixel Sensor" by A. Punnappuarath and M.S. Brown in Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition 2019, pp. 7628-7637.

[0015] "Super-Resolution Imaging Using a Focus Pixel Sensor" by SM Woo, YW Ha, and YO Kim at the IEEE Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA-ASC), Japan 2021, describes the injection of multiple low-resolution focus pixel images into a normal image based on a repetitive channel and spatial tension layer structure.

[0016] "Joint Electromagnetic and Ray-Tracing Simulations for Quad-Pixel Sensor and Computational Imaging" by G. Shataignier, B. Vandame, and J. Vaillant in Optics Express, vol. 27, no. 21, 14-10-2019, pp. 3046-3051, describes a design for a quad-pixel sensor in which microlenses cover 2x2 subpixels, and a method for blending wave optics simulations with ray-tracing simulations to generate physically accurate composite images. Summary of the Invention

[0017] Based on such known omnidirectional focus sensors comprising multiple photosensitive elements per pixel, i.e. per common lens element and / or filter element, it is an object of the present invention to provide an improved method for processing image data of such image sensors, as well as an improved processor unit and computer program.

[0018] This object is achieved by a method comprising the features of claim 1 and by an image processor unit comprising the features of claim 13. Preferred embodiments are disclosed in the dependent claims.

[0019] Multiple views are captured for each image, each view comprising a matrix of selected subpixel values ​​of the pixel matrix captured for the image by a matrix of photosensitive elements. From this, for each pixel location associated with a common lens element and / or filter element, multiple finer subpixel values ​​are assigned to and distributed across the set of views. Each view represents at least one of these finer subpixel values, but not all of the subpixel values ​​for the pixel location, and the combined views together contain all of the subpixel value information for each pixel location. These multiple views of images allow for the determination of multiple effects, such as different variations caused by crosstalk and shading effects resulting from a common main lens, and variations in sensitivity, brightness, and exposure time associated with the common lens element and / or filter element for the pixel location of the associated set of subpixels.

[0020] According to the present invention, the method comprises: - for each captured image view, determining a first variation of sub-pixel values ​​in the respective view separately from other views of the same image; - determining a second variation of sub-pixel values ​​associated with the same position in the pixel matrix behind each lens element (7) and / or filter element (9) in the set of views of the same image; - processing image data for the image using the determined first and second variations; Equipped with.

[0021] Variations in subpixel values ​​can be specified through the use of operators to determine different types of effects when capturing an image with subpixel values. Subpixels at common pixel locations associated with common lens elements (e.g., on-chip lenses) and / or filter elements do not behave similarly. They may have variations in sensitivity and brightness. Brightness variations may also occur due to shading from a common main lens. Additionally, crosstalk may occur between photosensitive elements and their associated subpixel values. Exposure time also has an effect on subpixel values. These effects can be assigned to a set of operators, each associated with either a first variation in subpixel values ​​in a respective view distinct from other views of the same image, or a second variation in subpixel values ​​associated with the same location in a pixel matrix among a set of views of the same image.

[0022] Operations on the captured raw image data can be performed automatically by using a general image formation model that represents the nth view image. This general image formation model can be established by the following relationship:

[0023]

number

[0024] where Y n is the nth view image pixel matrix, and F n is the color filter array subsampling operator associated with the nth view image pixel matrix, and S n is a spatially-varying and color-dependent lens shading operator for the nth view image pixel matrix, X denotes a region of interest (ROI) of the unknown scene with focal depth (D),

[0025]

number

[0026] denotes the two-dimensional convolution operation, and H n (D) is the point spread function associated with the nth view image pixel matrix at a constant focal depth (D), and η n is the signal-dependent noise for the nth view image pixel matrix, and E n,m is the color-dependent exposure variation operator between the nth view image pixel matrix and the mth view image pixel matrix, and C n,m is the wavelength and color dependent crosstalk operator between the nth view image pixel matrix and the mth view image pixel matrix, and C n,m captures wavelength and color dependent crosstalk between co-located sub-pixels in a pixel matrix that share the same lens and / or filter elements.

[0027] That is, processing image data captured by an on-chip lens all-focus sensor through the use of the image formation model described above is very effective and allows for consideration of multiple characteristics of the image sensor by correlating different matrices, operators, and factors.

[0028] The image data from the image sensor is processed as follows: - combining a set of views of the image by combining, for each location in the pixel matrix, a plurality of sub-pixel values ​​associated with a common location in the pixel matrix, and pre-processing the combined sub-pixel values ​​of the captured pixel matrix for the image; - preprocessing the views separately by preprocessing the captured pixel matrix with a plurality of subpixel values ​​of each view by evaluating the variation between subpixel values ​​associated with the same position in the pixel matrix; - preprocessing both the preprocessed combined sub-pixel values ​​of the captured pixel matrix in the set of views and the preprocessed pixel matrix of the plurality of sub-pixel values ​​for each position in the pixel matrix for each view by using the results of the variability assessment; can have:

[0029] Combining multiple subpixel values ​​results in binning of the image data, which makes it possible to determine the variance between these binned image data for an image and the image data in at least one view of the image.

[0030] Processing the views of the image by using the determined variations described by the operators of the image formation model allows for separate pre-processing of both the combined view of the image, i.e., the pixel matrix with the combined pixel values, and the separate views of the image, i.e., the pixel matrix with the separate sub-pixel values, which provides an improved image processing flow that takes advantage of the properties of on-chip lens all-focus sensors.

[0031] The method does not discard the results after calculating the disparity, but instead takes advantage of all the image information captured by, for example, an OCL all-focus sensor.

[0032] Multiple subpixel values ​​associated with a common pixel location in a set of views can be combined, for example, by automatically processing the sum of the subpixel values ​​associated with the same location in the matrix to achieve a combined view comprising a sum-binned pixel matrix. Thus, for example, subpixel values ​​captured by photosensitive elements positioned under a common lens element and / or common filter element are combined into one combined pixel value in the combined view by processing the sum of these pixel values.

[0033] Multiple sub-pixel values ​​associated with a common location in the set of views can be combined to achieve an average pixel matrix by automatically averaging the sub-pixel values ​​associated with the same location in the set of views, whereby the combined pixel value is processed for, for example, a pixel location in the pixel matrix associated with a common lens element and / or a common filter element by calculating the average of the sub-pixel values ​​in the set of views captured by a light sensitive element at a location below the common lens element and / or common filter element.

[0034] From this, the sub-pixel values ​​associated with a common pixel location in the pixel matrix are combined into a common resulting sum and / or average pixel value for that pixel location in the pixel matrix.

[0035] Combining pixel values ​​in multiple views of an image has the effect of increasing the brightness compared to the brightness of a single view, i.e., a sub-pixel value captured by only one photosensitive element. The combined pixel matrix reflects an image captured with a longer exposure time compared to a pixel matrix comprising only one sub-pixel value captured by one photosensitive element.

[0036] A captured view (i.e., captured pixel matrix of an image) comprising multiple subpixel values ​​per position in the view can be preprocessed using a set of pixel matrices, with each view in the set being a subpixel matrix comprising subpixel values ​​of an associated set of photosensitive elements. Each of these photosensitive elements of the associated set has the same relative position in the group of photosensitive elements with respect to a common position in the view associated with a common lens element and / or common filter element. Thus, the group of subpixel values ​​forms a pixel matrix for the view, i.e., image. The subpixel values ​​are captured by the photosensitive elements, i.e., associated lens elements and / or filter elements, each having the same relative position with respect to their associated pixel position in the pixel matrix. For example, when four photosensitive elements, i.e., lens elements and / or filter elements, are positioned below a common pixel position, the subpixels at the upper left position form a first pixel matrix, the subpixels at the upper right position form a second pixel matrix, the subpixels at the lower left position form a third pixel matrix, and the subpixels at the lower right position form a fourth pixel matrix.

[0037] The variation in exposure characteristics can be assessed by correlating the combined view of the image with at least one view of the image, where the combined view, each comprising a combined sub-pixel value processed from a combination of associated sub-pixel values ​​for each position in the matrix, represents a longer exposure time compared to the separate pixel matrices, each comprising a matrix of sub-pixel values, each captured by one respective photosensitive element.

[0038] From this, one of the views of the set can be used to evaluate image characteristics, for example, related to exposure time, by comparing the combined pixel value of the combined view with its associated subpixel value. For example, a summed binned pixel matrix provides a set of pixel values ​​captured by at least two of the photosensitive elements at a pixel location in the pixel matrix, with a brighter, i.e., a summed increased pixel value. A pixel matrix with one subpixel value for a pixel location provides an image with a characteristic, such as a shorter exposure time, i.e., a smaller single pixel value, than a longer summed binned pixel value. The difference in pixel values ​​associated with common pixel locations in the pixel matrix is ​​valuable information.

[0039] View blur can be assessed by correlating related subpixel values ​​of at least two views of the image, where a first subpixel matrix of a first group of subpixel values ​​can be correlated with a second subpixel matrix of a second group of subpixel values.

[0040] Because the scene is optionally blurred based on the distance from the focal plane, there is a shift in the disparity between the number N of sub-aperture (focus) views. This disparity depends on the depth in the blur shape. The disparity between the N number of views can be modeled as an explicit shift in the image content with sub-pixel accuracy. If sub-pixels sharing one common pixel location in the pixel matrix, i.e., a common on-chip lens and / or filter element, have the same exposure, shift estimation can be simplified. Due to the split pixel arrangement, the disparity is directional and has a specific direction for a particular view. This allows for a simpler estimation of the shift since the surge direction is known.

[0041] The subpixel matrices can be correlated with each other by automatically treating the difference value for each location in the pixel matrix as the difference between a subpixel value of one subpixel matrix in the set of subpixel matrices and a subpixel value of another subpixel matrix in the same set of subpixel matrices.

[0042] From this, for each pixel location in the view, i.e., associated pixel matrix, at least one difference value is processed from the sub-pixel values ​​associated with the same pixel location, i.e., common lens element and / or filter element. The resulting at least one pixel matrix with the resulting different values ​​for each associated pixel location can be stored and used for further processing of the image data to achieve an improved pre-processed image of the sequence of images.

[0043] The disparity between views of an image captured by an image sensor can be automatically estimated, and the disparity of sub-pixel values ​​associated with the same pixel location determined.

[0044] From this, sub-pixel values ​​associated with different relative positions among common pixel positions of the pixel matrix, i.e., among different relative positions associated with common lens elements and / or filter elements, are correlated with each other to find the disparity.

[0045] A depth value of the captured image, representing the depth of focus of the captured image relative to the focal plane of the image sensor, can be automatically estimated.

[0046] The depth value can be achieved from the difference in sub-pixel values ​​for one common pixel position in the pixel matrix, i.e., the difference between sets of sub-pixels associated with a common lens element and / or filter element due to the relative displacement of the photosensitive elements below the common lens element and / or filter element.

[0047] The method can be implemented as a routine in the image processor unit, i.e. by respectively programming the processor, or by providing the image processor unit with appropriate hardware logic to perform the steps of the method described above. To this end, the image processor unit can be implemented as a programmed microcontroller, a microprocessor, or as dedicated hard and software logic, e.g. in the form of an ASIC, or as hardware logic, e.g. an FPGA.

[0048] The invention will now be further explained, by way of example, with the aid of the enclosed drawings in which: [Brief explanation of the drawings]

[0049] [Figure 1] 1 is a block diagram of an image camera comprising an image sensor and an image processor unit; [Figure 2] FIG. 1 is a flow diagram of an embodiment of a method for processing image data. [Figure 3] FIG. 10 is a flow diagram of an illustrative second portion of a method for processing image data. [Figure 4] A rough pixel matrix captured by the image sensor and an associated set of pixel matrix processes from the captured pixel matrix. [Figure 5] FIG. 10 is a diagram of the focus effect using an on-chip lens sensor depending on the focal plane. [Figure 6] FIG. 1 is a flow diagram of an illustrative method for processing image data. [Figure 7] FIG. 1 is a flow diagram of an illustrative method for constrained block matching. DETAILED DESCRIPTION OF THE INVENTION

[0050] FIG. 1 shows a camera 2 and raw image data IMG provided by an image sensor 4 of the camera 2. RAW1 presents an illustrative schematic block diagram of an image processor unit 1 comprising an image processor unit 3 for processing an image of interest.

[0051] The image sensor 4 receives the raw image IMG RAW is a dataset provided by a matrix of pixels per image, P x,y , an array, i.e., a pixel matrix.

[0052] The camera 2 comprises an opto-mechanical lens system 5, for example a fixed uncontrolled lens 15 or a variable controllable lens system.

[0053] Groups of photosensitive elements of image sensor 4 are placed behind a common lens element 7. This means that light passing through one lens element 7 of lens matrix 6 reaches at least two photosensitive elements associated with that lens element 7. For example, the groups of photosensitive elements may be a group of two elements placed horizontally adjacent to each other, or a group of four photosensitive elements placed in two rows and two columns behind the common lens element 7.

[0054] To capture colors in an image, a color filter array 8 (CFA) comprising a matrix of filter elements 9 with alternating filter characteristics (e.g., Bayer RGGB, RGBE, RYYB, CYYM, CYGM, RGBW, RGBW#1...3, X-Trans, Quad Bayer, RYYB Quad Bayer, Nonacell, RCCC, RCCB, etc.) can (optionally) be provided in the optical path in front of the image sensor 4. Preferably, the color filter array 8 is located between the lens matrix 6 and the photosensitive elements of the image sensor 4. The color filter array 8 can also be located in front of the lens matrix 6, i.e., between the lens system 5 and the lens matrix 6.

[0055] A group of photosensitive elements of the image sensor 4 assigned to a common lens element 7 may share a common filter element 9, i.e. the same color. In another embodiment, each photosensitive element of the image sensor 4 may also have its particular filter element 9 having the same color as the filter elements 9 of the group assigned to the same lens element 7, or at least partially having a different color compared to the other photosensitive elements of the group.

[0056] The image processor unit 1 may be integrated into a handheld device such as a smartphone, a tablet, a wearable, a photo or video camera, or the like.

[0057] The image processor unit 1 receives an image IMG from an image sensor 4 that captures an image. RAW , and frames are also considered to be images in the sense of the present invention.

[0058] The image processor unit 3 may be integrated into the image camera 2 or may be provided as a separate unit connected to the image camera 2 by wired or wireless means.

[0059] The image processor unit 3 outputs the final image IMG FIN To achieve raw image data IMG RAW This also includes the option of receiving a final image sequence of frames.

[0060] FIG. 2 shows raw image data IMG captured by the image sensor 4. RAW 1 is an exemplary flow diagram of a method for processing an image of a scene using sub-aperture views. An image sensor 4, for example, a multiple on-chip lens sensor, for example, a 2x2 OCL sensor, captures different information about the scene using sub-aperture views.

[0061] The camera can be controlled with 3A for auto exposure, auto balancing, and auto focusing control.

[0062] In the first sequence PRE_1, a plurality of composition-related sub-pixel values ​​IMG in the pixel matrix are TL,TR,BL,BR are combined in step a1) to achieve a summed binned raw image. The sums of the pixel values ​​of the groups of photosensitive elements associated with a common lens element 7 are summed to achieve a resulting summed pixel value for the pixel location in the pixel matrix associated with the common lens element 7.

[0063] In the second pre-processing flow PRE_2, a number n of pixel values ​​IMG captured by a group of photosensitive elements associated with the common lens element 7 are TL,TR,BL,BR are preprocessed separately from each other. Hence, a set of N pixel matrices, each comprising a subpixel value for a particular position in the pixel matrix, is preprocessed separately from the combined total binned image of step a1) in sequence PRE_1. The captured pixel matrix comprising multiple subpixel values ​​for each position in the pixel matrix is ​​processed using the set of pixel matrices in preprocessing sequence PRE_2, starting from step a2). Each view captured by the image sensor of the set of image views comprises a pixel matrix with a subpixel matrix for each pixel position comprising subpixel values ​​captured from an associated set of photosensitive elements. Each photosensitive element of the associated set having the same relative position in the group of photosensitive elements to a common position in the pixel matrix is ​​associated with a common lens element 7 or a common filter element 9.

[0064] Both separate sequences PRE_1 and PRE_2 may provide, in steps b1) and b2), for example, black level correction of the summed binned raw image (step b1)) or black level correction of the N focused images (step b2)).

[0065] Furthermore, the pre-processing PRE_1 and PRE_2 flows may each comprise a step c1) of crosstalk correction of the summed binned raw image (step c1)) and a step c2) of crosstalk correction of N separately focused images (step c2)).

[0066] Further, separate flows of pre-processing PRE_1 and PRE_2 may comprise a step d1) of lens shading correction LSC for the sum binned raw image (step d1) and a step d2) of lens shading correction LSC separately for the N focused images (step d2)).

[0067] As a result, statistical control values ​​can optionally be obtained from the results of both separate sequences of pre-processing (3A statistics). The statistics can be used for 3A control.

[0068] N focused images IMG TL,TR,BL,BR The pre-processing flow PRE_2 can be followed by a step e) of disparity estimation. The result of the disparity estimation can be used for the 3A control.

[0069] Furthermore, the results can be used for an additional step f) of depth estimation.

[0070] The results of both the separate pre-processing PRE_1 and PRE_2 of the sum binned raw image and the N focused images can be used as input for further pre-processing steps g) of white balancing, h) of raw image denoising, i) of defective pixel correction, and j) of the computational photography engine CPE, which can include at least one of a set of image processing routines based on the results of the separate and combined pre-processing flows and the results of step e) of disparity estimation DispEst and depth estimation DepthEst.

[0071] White balancing can be further controlled by the 3A control.

[0072] In the 3A control, the sensor can be exposed for the focused image of the sum binned image according to the characteristics. The exposure variation shift, e.g., log2(1 / N), can be controlled, for example, between the view and the sum binned image. Sub-pixel exposure variation is possible. For the white balancing step g), the same white balancing can be applied to the focused image and the sum binned image, i.e., to the results of both separate pre-processing flows.

[0073] Autofocus statistics can be achieved from unidirectional and multidirectional parallax.

[0074] FIG. 3 shows an illustrative flow diagram that follows up on the flow shown in FIG.

[0075] Step j) of the computational photography engine CPE may include at least one of the following set of processes: Raw image demosaicing, raw image alignment, raw image fusion, synthetic blur, haze removal, chroma noise removal, color image alignment, color image fusion, defocus blur removal, veiling glare removal, and the like.

[0076] Bokeh is the aesthetic quality of the blur produced in out-of-focus parts of an image. Differences in lens aberrations and aperture shapes cause very different blur effects.

[0077] Step j) of the computational photography engine may be followed by combined post-processing of the image or sequence of images processed in step j). These post-processing steps may include a step k) of color correction, a step l) of local tone mapping LTM, a step m) of global tone mapping, a step n) of sharpening, and the like. These post-processing steps may be provided in the form of additional separate processing routines of the processing unit. These post-processing steps may also be part of the computational photography engine CPE.

[0078] These pre-processing and post-processing steps of the CPE are commonly known and are currently performed based on the combination of sub-pixel values ​​of the captured pixel matrix.

[0079] The method for processing image data can be based on a computational photography engine that has the preprocessed N-focus image and the preprocessed sum-binned image as input pixel matrices separately from each other. Furthermore, the computational photography engine can also include estimated disparity and estimated depth as inputs, which provides increased image information related to the characteristics of the on-chip lens image sensor.

[0080] The computational photography engine is capable of delivering excellent quality, high dynamic range, high resolution, full color images at output that are already denoised and sharpened.

[0081] 4 presents an exemplary pixel matrix comprising a 4x4 matrix of pixel locations in the matrix. The pixel matrix represents a view of a raw image captured by the image sensor 4, or a portion of a larger view of the image, as illustrated by the small section. Each pixel location is associated with one common lens element 7. From this, a group of 2x2 photosensitive elements is placed behind the common lens element 7. For each pixel location, there is a top row T, a bottom row B, a left column L, and a right column R. The four photosensitive elements provide top left (TL), top right (TR), bottom left (BL), and bottom right (BR) pixel values ​​for each sub-pixel location behind the common lens element 7 at the pixel location of the pixel matrix.

[0082] The image sensor 4 captures these sub-pixel values ​​TL, TR, BL, BR for each image and provides a pixel matrix 10 comprising a set of sub-pixel values.

[0083] In step a1), the composition-related sub-pixel values ​​in the pixel matrix are combined to achieve a resulting combined pixel value for each pixel position in the pixel matrix. For example, the summed binned pixel matrix 11 is automatically processed by summing the pixel values ​​for each pixel position and associated color.

[0084] For example, the summed pixel values ​​S(red), S(green), and S(blue) are processed from the subpixel values ​​TL(RGB), TR(RGB), BL(RGB), and BR(RGB) according to the following formula: TL(R)+TR(R)+BL(R)+BR(R)=S(R) TL(G)+TR(G)+BL(G)+BR(G)=S(G) TL(B)+TR(B)+BL(B)+BR(B)=S(B) Here, R=red, G=green, and B=blue.

[0085] Optionally, additional combined pixel values ​​S(color) can be achieved for other colors or white depending on the color filter array.

[0086] Pixel matrix 11 represents the combined pixel values ​​of captured pixel matrix 10 .

[0087] Additionally, multiple sub-pixel matrices 12a, 12b, 12c, 12d can be achieved in separate pre-processing streams.

[0088] Each of the N subpixel matrices (N focused images) comprises one specific subpixel value for each pixel position of the pixel matrix 10 captured by a respective photosensitive element. Each pixel matrix 12a, 12b, 12c, 12d of the set comprises subpixel values ​​TL (color), TR (color), BR (color), and BL (color) of an associated set of photosensitive elements having the same relative position in the group of photosensitive elements relative to a common position in the pixel matrix associated with the common lens element. Thus, the first subpixel matrix 12a comprises a subpixel value TL (color) in the upper left position of the 2x2 subpixel matrix for each pixel position. The second subpixel matrix 12b comprises a subpixel value TR (color) in the upper right position, the third subpixel matrix 12c comprises a subpixel value BL (color) in the lower left position, and the fourth subpixel matrix 12d comprises a subpixel value BR (color) in the lower right position. The color of each of the sub-pixel values ​​depends on the filter elements of the associated lens element or, optionally, the filter elements of the associated sub-pixel photosensitive element.

[0089] The resulting set of sub-pixel matrices 12a, 12b, 12c, 12d is pre-processed in a stream starting from step a2) of FIG. 2, independently from the combined pixel matrix 11 starting from step a1).

[0090] These different views of the same image contain different information. Performing such separate and combined pre-processing allows for better image quality of the resulting final image.

[0091] FIG. 5 presents a flow diagram of a method for processing image data of the image sensor 4 showing the pre-processing of the different views in more detail.

[0092] In step a1), the reference view V REF are processed with a specific color channel interpolation, e.g., green channel interpolation GCI. N views of an image V 1,2,...,N-1 One of the views V is selected as the reference view. REFmay be, for example, a TR view. In this example, the color channel that has the highest sampling or provides more detail than the other channels is chosen for the reference view. In step b1), smoothing is performed, for example, by low-pass filtering LPF.

[0093] Alternate View V ALT_1,2,...,N-1 In step a2), the reference view V REF A color channel interpolation CCI, e.g., a green channel interpolation GCI, uses the associated pixel matrix to interpolate the respective alternative view V ALT_1,2,...,N-1 ,=V1;V2;...;V N-1 , for example, processed for TL, BL, and BR views.

[0094] Steps a1) and a2) may comprise smoothing, for example by low-pass filtering. One purpose of smoothing is to make the estimate robust to noise.

[0095] Due to the positions of the TR, TL, BL, and BR views in the full-resolution raw data, the shift direction for alternate views is known, which allows for simplification and speed-up in the next step b) of constrained block matching, as well as constraints on the search direction.

[0096] This step b) is followed by a step c) of sub-pixel refinement using the results of step d) of determining a foreground and / or background map MV-R that provides estimated motion model parameters. The motion model is translational.

[0097] Finally, the displacement(s) estimates are processed in step e).

[0098] In the following, we describe the properties of the OCL all-focus sensor and, based on those properties, describe in more detail the proposed image formation model, which facilitates, for example, disparity estimation, autofocusing, demosaicing, multiview fusion, and many other applications.

[0099] The on-chip lens (OCL) all-focus sensor has a unique design. Directly beneath the same shared microlens 7, several (N) sub-pixels with their own photodiodes 14a, 14b are housed, each capable of detecting light and recording a corresponding signal independently. The incident light through the OCL is split into N directions, resulting in N sub-aperture views of the scene captured simultaneously (see Figure 5). In a sense, the focus sensor functions as a coarse N-view light-field camera.

[0100] In particular, in a dual pixel OCL sensor where the OCL is shared by two sub-pixels, e.g., left and right sub-pixels, the incident light is separated in the left and right directions, so the result is two views of the scene, i.e., left and right views.

[0101] Similarly, in a quad-pixel OCL sensor where the OCL is shared by a group of 2x2 sub-pixels, the result is those four views of the scene as the light is separated into four directions: top left (TL), top right (TR), bottom left (BL), and bottom right (BR), and so on. This is illustrated in Figure 4, which shows a 4x4 pixel matrix with 16 sub-pixels TL, TR, BL, BR of a 2x2 OCL all-focus sensor.

[0102] An OCL all-focus sensor can operate in different modes, such as native full-resolution mode, re-mosaic full-resolution mode, or binned versions of either mode. As an example, the pixel matrix in Figure 4 illustrates the total binned focus image arrangement for a 2x2 OCL sensor with a Quad-Bayer color filter array (QB-CFA).

[0103] Particular operating modes are selected for optimal performance under certain conditions or to enable certain functions, for example, in sum-binned mode, sub-pixels are merged into larger pixels with higher sensitivity, allowing for better low-light performance.

[0104] The disparity between the N sub-aperture views increases in either the forward or backward direction (in front of or behind the focal plane) as the scene moves away from the focal plane, as shown in Figure 6. Therefore, signed unidirectional / multidirectional disparities are derived from multiple views (focused images) and then used to calculate the optimal lens position for high-precision and fast autofocus. Typically, the focused images (or images derived from them) are discarded by the camera hardware after calculating the disparity. The type of data passed through the image signal processor is determined based on the operating mode chosen for the intended features or use case.

[0105] Regardless of the specific sensor design with respect to OCL, sub-pixel color filters and their placement, and exposure settings, the important features that an OCL all-focus sensor has / can provide are: characteristics Indeed, these characteristics may encourage sensor vendors to adopt advanced sensor design techniques that will enable new modes of operation, which may enable differentiating features for the consumer market. Below, these characteristics are discussed to set the stage for the proposed image formation model and the proposed image processing pipeline.

[0106] 1. Brightness variations between views Because the N views are captured by the same sensor, they should be synchronized in time and space. However, subpixels sharing a common lens element 7 (OCL) do not need to have the same exposure. They can have varying exposures per pixel, for example, a Quad-Bayer filter array (QB-CFA) with two exposures: a short exposure and a long exposure. This is a technique commonly known as digital overlap HDR (high density resolution) or DOL-HDR (digital overlap HDR). Due to the camera's main lens shading effect and possibly per-subpixel exposure settings, the N views will experience brightness variations between views that need to be taken into account (or even exploited) in image processing and fusion operations.

[0107] 2. Variable exposure in a single shot Due to the split subpixel arrangement directly below the OCL, incident light is split into multiple (N) directions. When subpixels sharing an OCL have the same exposure, the exposure of the focused image is approximately 1 / N of the exposure of the sum-binned image. This approximation is provided here because the splitting is not perfect in an actual sensor. The result is an exposure value for the focused image, EV, of approximately log2(1 / N). The focused image effectively has less exposure than the sum-binned image, and saturated bright areas in the sum-binned image will not be saturated in the focused image. The set of sum-binned and focused images can therefore be viewed as images captured at two different exposures with an exposure value of approximately log2(1 / N), enabling HDR imaging from two exposures in a single shot. When M shots with M exposure times are captured by an OCL-focused sensor, the effective number of different EVs is approximately 2M, enabling HDR fusion from 2M exposures.

[0108] The above can be generalized to the case where the N subpixels have different exposures, such as those highlighted above for long and short exposures. Thus, effectively, with a single capture, an OCL all-focus sensor can provide multiple differently exposed images, where each image carries different information about the scene being photographed. Fusion of single-shot images (or images from multiple shots) can enable various computational photography (CP) functions, such as HDR, joint demosaicing and HDR, joint HDR and SR (super-resolution), to name a few possibilities.

[0109] 3. Parallax as Blur Shift or Pixel Shift Because the scene is optically blurred based on its distance from the focal plane, there is a shift or disparity between the N sub-aperture (focal) views, and the disparity depends on the depth and shape of the blur kernel. The disparity between the N views can be modeled as an explicit shift in the image content with sub-pixel accuracy. Shift estimation is simplified when sub-pixels sharing an OCL have the same exposure (and white balancing). In addition, due to the split pixel arrangement, the disparity is directional and has a specific direction for a particular view. This fact also simplifies shift estimation because the search direction is known. The search size for possible disparities depends on the scene depth and the shape of the blur kernel, but is relatively small, typically on the order of a few pixels.

[0110] For scene regions that are out of focus (far from the focal plane, either in the front or back direction), the effective (defocus) blur varies greatly between views, and modeling disparity as an explicit pixel shift will fail and lead to erroneous disparity estimates. Note that it is actually the difference in the point spread functions (PSFs) of the sub-aperture views that generates the disparity, not an explicit shift in the image content.

[0111] Disparity is the result of PSF difference Without loss of generality, consider the case of a 2x2 OCL all-focus sensor, such as one with a QB-CFA configuration (QB = Quad Bayer). At a given depth D, H TL , H TR , H BL , and H BR Let X denote the point spread function PSF for the TL, TR, BL, and BR views, respectively. X denotes the unknown scene ROI with depth D. The four views can be expressed as:

[0112]

number

[0113] η noise is the signal dependent noise. Here we assume that crosstalk and sensitivity and lens / color shading corrections are performed. The inter-view horizontal and vertical disparities can be expressed as:

[0114]

number

[0115] From Equations 7-10, it can be shown that it is the difference in the inter-view PSFs of the sub-aperture views that generates the disparity, not an explicit shift in the image content.

[0116] When the N sub-pixels have different exposures, such as those highlighted above for long and short exposures, the disparity estimation will not only be expressed as a pixel shift estimate and / or a blur difference estimate, but will also include a photometric shift estimate due to the varying exposures for the N sub-pixels, unless brightness variations are considered prior to / in the disparity estimation algorithm. Whether exposure is constant or varying for sub-pixels that share an OCL, primary lens shading will result in brightness variations between views, i.e., photometric shifts between views. This variation also needs to be considered prior to / in the disparity estimation solution.

[0117] The constrained displacement model as described above is an example of a disparity estimation block in the pipeline. Based on an understanding of the blur characteristics, the displacement model can be easily generalized to other types of sensors, for example, OCL sensors that provide more than four views (N>4).

[0118] 4. View-to-view blur symmetry When the scene is far from the focal plane, there will be defocus blur. Due to the split subpixel arrangement, the defocus blur is split between the views, resulting in different point spread functions for the different views. The difference between the sets of N point spread functions is what actually results in parallax, as previously mentioned.

[0119] The point spread functions for the N views vary in a complex manner and depend on several factors, such as the main lens focal length, the aperture diameter, the pixel angular response, the scene depth, the amount of defocus, and the focal distance from the camera, etc. That said, the point spread functions for those N views will be approximately symmetrical due to the OCL and subpixel division design d.

[0120] For example, for a dual-pixel OCL all-focus sensor, the point spread function of the left view will be approximately equal to the point spread function of the right view flipped about the vertical axis. And for a quad-pixel OCL all-focus sensor, the point spread function PSF of the TL view will be approximately equal to the point spread function PSF of the BL view flipped about the horizontal axis, and so on. Due to imperfections in OCL positioning, crosstalk between subpixels, and other manufacturing limitations, the symmetry is only approximate. The shape of the focus blur and the point spread functions of the N views can be parametrically modeled as a translating disk, a damped Gaussian function, or any other shape that can be derived from factory or laboratory calibration. Regardless of the defocus blur model, its division into mutually symmetric (mirror-image) point spread functions for the N views is a very useful cue that can enable various computational photography features, such as reflection removal, veiling glare removal, defocus blur removal, SR, and HDR, to name just a few.

[0121] Disparity Estimation Block In the following, the disparity estimation block e) of FIG. 5 is explained in more detail with the use of FIG.

[0122] 6 presents a schematic diagram comprising a side view of the image sensor 4. A row of photosensitive elements is shown, with a pair of photosensitive elements 14a, 14b per row being placed adjacent to each other in a common row below a common lens element 7. A pixel matrix comprises a number of lens elements 7 placed adjacent to each other in rows as can be seen, and correspondingly in columns as shown in FIG.

[0123] Each lens element 7 is a microlens integrally formed on the sensor chip of the image sensor 4 .

[0124] The imaging camera 2 further comprises an optical lens system 15 comprising at least one lens (ie an objective lens).

[0125] Lens 15 has an associated focal plane FP. The position of a light beam initiated from a point source P through lens 15 depends on the position of the point source P relative to the focal plane FP.

[0126] In the left part a) of Figure 6, a point light source P lies on the focal plane FP. This results in a light beam being focused onto one microlens 7 and the associated photosensitive elements 14a, 14b behind this common lens element 7. The distance between the objective lens 15 and the focal plane FP corresponds to the distance between the objective lens 15 and the plane of the lens elements 7 of the image sensor.

[0127] The central part b) of Fig. 6 presents a situation in which the point light source P is at a greater distance from the focal plane FP. This results in a focus of the light beam starting from the point of light P. For this purpose, an objective lens 15 is placed in front of the lens element 7, between the plane of the lens element 7 and the objective lens 15. From this, the light beam is then spread out and traverses several lens elements 7 and the associated photosensitive elements 14a, 14b. It can be seen that the light beam is not distributed equally on all photosensitive elements 14a, 14b behind the common lens element. This depends on the angle of incidence of the light beam through the lens element 7.

[0128] The right part c) of Figure 6 shows the situation where the point light source P is behind the focal plane closer to the objective lens 15. The focal point of the objective lens 15 is therefore behind the image sensor so that the light beam passes through multiple lens elements 7. The photosensitive elements 14a, 14b behind the common lens element are affected differently depending on the angle of incidence of the light beam through each lens element 7.

[0129] Depending on the position of the light source point P relative to the focus or focal plane of the image sensor 4, a displacement of the image, ie a sub-pixel value, occurs.

[0130] For example, when each full resolution frame can be separated into four views, namely, top right (TR), top left (TL), bottom left (BL), and bottom right (BR), the pixel values ​​TR, TL, BL, BR are different from each other according to the displacement. One of the views, for example, the TR view, can be set as a reference frame.

[0131] When a point light source P is in front of the focal plane, the following sub-pixel level shifts S occur. Below we list the shifts from the reference view, TR view, to the alternative views TL, BL, and BR:

[0132] [Table 1]

[0133] If the point source P is behind the focal plane, the sub-pixel level shift s from the reference view TR to the alternate view is:

[0134] [Table 2]

[0135] The blur kernel acts in the opposite direction when the scene is in front of or behind the focal plane. The direction of the blur itself, either positive or negative, is determined by the position of the view in the N-subpixel grid and the position of the scene relative to the focal plane. Also, in the ideal case, assuming perfect OCL symmetric positioning, no aberrations, and no other sensor imperfections, the magnitude s is constant for the TL, BL, and BR views. This assumption is OK to calculate the disparity parameter s with a reasonable accuracy sufficient to move the lens for autofocusing, multiview fusion, re-mosaicing, or any other process requiring knowledge of the disparity between views.

[0136] Basically, the direction of view displacement is known. For a 4-view OCL sensor (N=4), they are listed in a table. Therefore, the search direction for s is known and can be constrained. This speeds up the calculation. The described model can be generalized to any number and is not limited to the illustratively illustrated N=4 view OCL sensor. The method can also be applied accordingly to views N>4.

[0137] In the following, the proposed image formation model is explained.

[0138] Proposed image formation model Now that some important features of the OCL all-focus sensor have been discussed, a general image formation model is described below. The purpose of the model is to serve as a basis for the order of operations in an image processing pipeline and, by extension, for image processing / fusion algorithms, in particular to determine the requirements for processing an image by a particular algorithm.

[0139] The following notation is used: -Y n : Raw image of the nth view (focus) -Y sum : total binned images for N views -E n,m : color-dependent exposure variation operator between the nth and mth views -C n,m : wavelength and color dependent crosstalk operator between the nth and mth views -S n : a spatially varying and color-dependent lens shading operator for the nth view -H n (D): PSF associated with the nth view image at a given depth D -F n : the color filter array subsampling operator associated with the nth view image -η n : signal-dependent noise for the nth view

[0140] Sum

[0141]

number

[0142] corresponds to the defocus blur kernel at a constant depth D. The PSH system,

[0143]

number

[0144] is reasonably subject to certain conditions: 1.Non-negative;H n ≧0 2. View-to-view symmetry 3. Equal inter-view contribution; if defocus blur is isotropic,

[0145]

number

[0146] is.

[0147] Let X denote the unknown scene ROI with depth D. Following the above notation, at a given exposure time, the nth view image can be expressed as:

[0148]

number

[0149] where:

[0150]

number

[0151] denotes a two-dimensional convolution operation. When there is exposure variation for subpixels that share an OCL, the N views are related to each other by an exposure variation operation, which is: If the sensor is operating in the linear range, this can be simplified to a linear operator. When all views have the same exposure, this operator is 1. C n,m captures the wavelength and color dependent crosstalk between sub-pixels that share the same OCL.

[0152] The proposed image formation model does not make any assumptions regarding the color filter placement underneath the OCL—subpixels sharing the same OCL can have similar or different color filters—and the model does not restrict the operating mode of the OCL all-focus sensor, e.g., with respect to subpixel exposure settings or type of binning.

[0153] Constrained Multiview Block Matching By using the incoming multiview or view region of interest (ROI), constrained multiview block matching can be processed not only for precise registration but also for improved performance (see step b) in Figure 5). The output is a translation transformation from the reference view to the alternate view with constant shift and symmetric sign with (sub)pixel accuracy. A schematic diagram for constrained multiview block matching is illustrated in Figure 7, and the following is an illustrative pseudocode:

[0154] A) Illustrative pseudocode for constrained block matching: for each patch_TR in reference_frame do for each step_s in range from -max_S to +max_S do [patch_TL / BL / BR] = get_patch_position(step_s, patch_TR) [cost_TL / BL / BR] = compute_patch_cost(patch_TR, patch_TL / BL / BR) total cost = (cost_TL + cost_BL + cost_BR) [min_total_cost, best_step_value, compensated_patches_min_idx, best_step_counter] = calculate_best_step_s(total_cost, step_s) end for best_int_step(patch_TR) = best_step_value min_patch_cost(patch_TR) = min_total_cost [patch_subpixel_shift(idx_patch) = subpixel_refine(compensated_patches_min_idx) end for

[0155] B) Pseudocode for translational estimation: ~ , best_step_index] = max(best_step_sounter); Global_int_translation = (best_step_index - 1) - search_size; for every_patch in reference_view do if global_int_translation == best_int_step(patch_TR) best_subpixel_step(patch_TR) = best_int_step(patch_TR) + patch_supixel_shift(idx_patch) end end for global_subpixel_translation = mean(best_subpixel_step);

[0156] ​Constrained block matching is based on the idea of ​​block matching algorithms and can be performed faster on multi-view datasets because it only requires a one-way search of blocks. This constrained block matching allows for simultaneous alignment of three views.

[0157] Before performing constrained block matching, the reference view TR and three other alternate views, namely, TL, BL, and BR, are interpolated through their main channels. First, for each patch in the reference view-TR, sliding patches for TL, BL, and BR in the alternate views are calculated. A one-dimensional search is performed, searching only horizontally for TL, only diagonally for BL, and only vertically for BR.

[0158] For this symmetric search, only one loop is used for step_S from -max_S to max_S to calculate the coordinates of patches TL, BL, and BR. The calculation of each alternative patch position is written as follows: patch_TL_X = patch_TR_X - step_s patch_TL_Y = patch_TR_Y patch_BR_X = patch_TR_X patch_BR_Y = patch_TR_Y + step_s patch_BL_X = patch_TR_X - step_s patch_BL_Y = patch_TR_Y + step_s The sign of step_s is hard-coded in these formulas, and later the same sign of step_s is applied to form a translation matrix: in the Y direction, the sign of step_s is applied directly, and in the X direction, the inverse sign of step_s is applied.

[0159] To calculate best_step_s, the modified patch is compensated for brightness changes using the TR patch. Then, the cost for each compensated modified patch TL / BL / BR for each step is calculated by the mean of absolute differences (MAD) between the reference patch TR and the modified patch TL / BL / BR, and the patch costs are summarized to form a total cost for each step. Finally, best_step_s and min_cost_value for each patch are calculated with respect to the minimum cost over all steps. For each patch from the reference frame TR, the number of minimum steps, best_step_counter for each patch, is counted, and the subpixel shift for each patch with best_step_s is calculated.

[0160] Sub-pixel shift sub-block To calculate the sub-pixel shift for each change patch, a Taylor approximation algorithm is applied to calculate the sub-pixel refinement between TR and TL / BL / BR. The sign of the sub-pixel shift is taken for the larger absolute value of the vertical sub-pixel shift. Finally, the sub-pixel shift is calculated by averaging the x-shift of TL, the y-shift of BR, and the x-shift and y-shift of BL.

[0161] Infer integer global transformations The step s index, Step_s_index, is derived from the counter array of best_step_s and then converted to a globally converted integer.

[0162] Estimate the best sub-pixel step For each patch in the reference frame TR, the best sub-pixel step will be applied only if the patch has a global integer transformation; sub-pixel shifts larger than 1 will not be considered.

[0163] Calculate the global transformation for the sub-pixel step The sub-pixel global translation is calculated by taking the average of all the best sub-pixel steps.

[0164] Computational Complexity Analysis Complexity is presented here in terms of big O notation (O notation) to provide an approximate estimate of algorithm execution time as a function of the number of pixels to be processed.

[0165] [Table 3]

[0166] Complexity Estimation for Constrained and Unconstrained Multiview Block Matching

[0167] [Table 4]

[0168] From the above table, it can be seen that constrained block matching can be performed faster than conventional global motion estimation using a block matching algorithm.

[0169] Proposed image processing pipeline In the following, we describe a proposed image processing pipeline that follows an image formation model (based on the characteristics of the OCL all-focus sensor), which can be used for appropriate processing of image data to maximize image quality, SW / HW efficiency, and enable various functions.

[0170] Based on the introduced image formation model, the sequence of operations in the image processing pipeline for processing OCL all-focus data for CP development is illustrated below. It is worth mentioning that not all sensor operation modes are captured in this pipeline. The order of processing related to CP is the main concern. Also, the proposed data flow is intended to maximize data sharing in order to reduce CPU cycles and minimize memory requirements. Some important operations are discussed in some detail below.

[0171] A. Sensor output data In addition to the obvious advantages for autofocus, the OCL all-focus sensor delivers rich information content that can enable various computational photography functions. Typically, phase images (or their difference) are used for disparity estimation and then discarded by the camera hardware. However, the in-focus image enjoys unique characteristics over non-phase sum-binned image(s), and exploiting them will enable various computational photography functions. Thus, in the proposed image processing pipeline, the two data outputs from the all-focus sensor are: 1. N focal (phase) images 2. Sum binned image(s) (higher sensitivity, no phase information)

[0172] B.3A control The 3A (auto-exposure, auto-white balancing, and auto-focus) algorithms are key in most modern consumer cameras. In the proposed image processing pipeline, the 3A statistics are calculated from the sensor raw data. In particular, the AE (auto-exposure) and AWB (auto-white balancing) algorithms use statistics calculated from either the sum-binned image(s) or the focused image. This approach will create an EV shift of approximately log2(1 / N) between the focused and sum-binned image(s). The same white balancing gain is applied to the sum-binned image(s) and the focused image. The AF (auto-focus) algorithm will utilize statistics calculated from the unidirectional / multidirectional disparity estimation algorithm to determine the optimal lens position for the focusing ROI.

[0173] If the sub-pixels are allowed to have different exposures, the AE / AWB statistics can be calculated from the sum binned image(s) of the same exposure (e.g., long-exposure or short-exposure sum binned images) or one of the in-focus images of the same exposure. Whatever the approach, an EV shift will be observed between the sum binned image(s) and the corresponding in-focus views.

[0174] C. Pre- / Post-processing of pixel data In addition to black level correction, the sum binned image and the focused image may need to be corrected for crosstalk because the separation between subpixels under the OCL is not perfect and light passing through one subpixel may leak into adjacent subpixels. Additionally, due to primary lens shading, not only will the sum binned image experience color lens shading, but the N views will also experience brightness variations between views due to the combined effects of lens shading and directional integration of light in different views (as discussed above). Lens shading correction is performed to compensate for brightness variations in the sum binned image(s) and the focused image. Image denoising (and defective pixel correction) in the raw domain is also performed. The image output from the CP engine is post-processed for color correction, local / global tone mapping, sharpening, and other image enhancement and color rendering operations.

[0175] D. Disparity and Depth Estimation After the effect of lens shading is corrected for the N views (which manifests as color-dependent brightness variations across the image), the unidirectional / multidirectional disparities between the views are calculated and used to provide statistics for the AF algorithm to drive lens movement for optimal focus of the scene focus ROI. From the estimated disparities, a depth map is constructed by the depth estimation block. The depth map is then utilized in the CP engine to enable several features such as defocus blur removal and synthetic blur, to name just a few.

[0176] E. Computational Photography Engine In the proposed pipeline, the computational photography engine encompasses a collection of algorithms that operate on one or more of the following inputs: -Estimated scene disparity -Estimated scene depth - Preprocessed sum binned image(s) -Preprocessed focused image

[0177] The engine output then goes through post-processing operations of color correction, sharpening, contrast enhancement, and other image enhancement operations. The CP engine may include (but is not limited to) the blocks illustrated below. The details of those blocks are highly dependent on the underlying algorithm design and are left for another separate disclosure.

[0178] The table below summarizes the possible input data for each of the CP blocks. Naturally, there are dependencies between blocks; for example, image alignment is a key step for the success of some multi-image fusion functions such as HDR and SR. Some features can also be performed jointly. To avoid complexity in presentation, the dependencies and possible joint operations are not shown in the pipeline. However, the data flow through the CP sub-engines should flexibly enable dependent / joint functions in a seamless manner. Also, data / memory sharing should always be performed to minimize CPU cycles and hardware / memory costs.

[0179] From this table, it is worth noting that the estimated disparity and depth as well as the focused image are valuable data that can enable various CP functions. Traditionally, the focused image is discarded by the camera hardware after disparity estimation, but in the proposed image processing pipeline, the focused image is fully used.

[0180] An illustrative list of computational photography functions and their input data

[0181] [Table 5]

[0182] The processing pipeline also implies certain sensor modes to be enabled that are not present in the current OCL all-focus sensor.

Claims

1. Image data (IMG) from the image sensor (4) RAW 1. A method for processing an image of a subject, the method comprising: an image sensor (4) comprising a matrix of photosensitive elements and a plurality of lens elements (7) and / or filter elements (9) arranged in a pixel matrix in front of sub-pixel matrices of finer photosensitive elements, groups of photosensitive elements being placed behind a common lens element (7) and / or common filter element (9) to provide sub-pixel values ​​for their respective positions in the pixel matrix; and the image sensor (4) being adapted to capture image data for a plurality of views, each view comprising a matrix of groups of selected sub-pixel values ​​of the pixel matrix captured by the matrix of photosensitive elements, determining, for each captured image view, a first variation in subpixel values ​​in each view separately from other views of the same image; determining a second variation of sub-pixel values ​​associated with the same position in the pixel matrix behind each lens element (7) and / or filter element (9) in a set of views of the same image; processing the image data for the image using the determined first and second variations; and A method characterized by:

2. 2. The method of claim 1, wherein the operation on the captured raw image data is automatically performed by using a general image formation model that represents the Nth view image by the following relationship: [Equation 1] Here, Y n is the nth view image pixel matrix, and F n is a color filter array subsampling operator associated with the nth view image pixel matrix, and S n is a spatially-varying and color-dependent lens shading operator for the nth view image pixel matrix determined as the first variation, X denotes a region of interest (ROI) of an unknown scene having a depth of focus (D), [Equation 2] denotes a two-dimensional convolution operation, and H n (D) is the point spread function associated with the nth view image pixel matrix at a constant focal depth (D), and η n is the signal-dependent noise for the n-th view image pixel matrix, and E n,m is a color-dependent exposure variation operator between the n-th view image pixel matrix and the m-th view image pixel matrix determined as the second variation, and C n,m is a wavelength and color dependent crosstalk operator between the n-th view image pixel matrix and the m-th view image pixel matrix determined as the second variation, and C n,m captures the wavelength and color dependent crosstalk between co-located sub-pixels in the pixel matrix that share the same lens element (7) and / or filter element (9).

3. combining a set of views of an image by, for each location in the pixel matrix, combining a plurality of the sub-pixel values ​​associated with a common location in the pixel matrix, and pre-processing the combined sub-pixel values ​​of the captured pixel matrix for an image; preprocessing the views separately by preprocessing the captured pixel matrix with a plurality of the subpixel values ​​of each view by evaluating the variation between the subpixel values ​​associated with the same location in the pixel matrix; using the results of the evaluation of the variability to preprocess both the preprocessed combined sub-pixel values ​​of the captured pixel matrix in the set of views and the preprocessed pixel matrix of a plurality of the sub-pixel values ​​for each location in the pixel matrix for each view; 3. The method according to claim 1 or 2, characterized in that

4. 4. The method of claim 3, characterized in that the sub-pixel values ​​associated with a common position in the pixel matrix, i.e. a common lens element (7) and / or filter element (9), are combined to achieve a summed binned pixel matrix by automatically processing the sum of the sub-pixel values ​​associated with the same position in the pixel matrix.

5. 5. The method according to claim 1, characterized in that the sub-pixel values ​​associated with a common position in the pixel matrix, i.e. a common lens element (7) and / or filter element (9), are combined to achieve an average pixel matrix by automatically processing the average of the sub-pixel values ​​associated with the same position in the pixel matrix.

6. 6. The method according to claim 1, characterized in that the captured pixel matrix comprises a plurality of sub-pixel values ​​per position in the pixel matrix by using a set of pixel matrices, wherein each pixel matrix of the set is a sub-pixel matrix comprising sub-pixel values ​​of an associated set of photosensitive elements, each photosensitive element of the associated set having the same relative position in the group of photosensitive elements with respect to a common position in the pixel matrix associated with a common lens element (7) and / or a common filter element (9).

7. A method according to any one of claims 1 to 6, characterized in that the brightness variation of sub-pixel values ​​associated with the same pixel location in said set of image views is evaluated.

8. 8. The method of claim 1, further comprising assessing variability in exposure characteristics by correlating the combined views of an image with at least one view of the image, wherein the combined views comprising the combined sub-pixel values ​​each processed from the combination of associated sub-pixel values ​​for a respective position in the pixel matrix represent a longer exposure time compared to separate pixel matrices each comprising a matrix of sub-pixel values, each sub-pixel value being captured by one respective photosensitive element.

9. 9. The method according to claim 1, further comprising: assessing view blur by correlating related sub-pixel values ​​of at least two views of the image, wherein a first sub-pixel matrix of a first group of sub-pixel values ​​is correlated with a second sub-pixel matrix of a second group of sub-pixel values.

10. 10. The method of claim 9, wherein the subpixel matrices are correlated with each other by automatically treating a difference value for each of the locations in the pixel matrix as the difference between the subpixel value of one subpixel matrix in the set of subpixel matrices and the subpixel value of another subpixel matrix in the set of subpixel matrices.

11. 11. A method according to any one of claims 1 to 10, characterized in that the disparity between the views of an image is automatically estimated, wherein the disparity of sub-pixel values ​​associated with the same pixel location is determined.

12. 12. The method according to any one of claims 1 to 11, characterized in that the depth value of the captured image is related to the depth of focus of the captured image related to the focal plane of the image sensor (4).

13. The raw image data (IMG) provided by the image sensor (4) RAW 1. An image processor unit (3) for processing an image sensor (4) comprising a matrix of photosensitive elements and a plurality of lens elements (7) and / or filter elements (9) arranged in a pixel matrix in front of sub-pixel matrices of finer photosensitive elements, groups of photosensitive elements being placed behind a common lens element (7) and / or common filter element (9) to provide sub-pixel values ​​for each position in the pixel matrix, the raw image data comprising a matrix of pixel values ​​for the pixel matrix captured by the matrix of photosensitive elements, the matrix of pixel values ​​being divided into a set of views, each view comprising a matrix of sub-pixel values ​​comprising a view-specific sub-pixel of the sub-pixel matrix for each position in the matrix associated with a respective lens element (7) and / or filter element (9), The image processor unit (3) determining, for each captured image view, a first variation in subpixel values ​​in each view separately from other views of the same image; determining a second variation of sub-pixel values ​​associated with the same position in the pixel matrix behind each lens element (7) and / or filter element (9) in a set of views of the same image; processing the raw image data for the image using the determined first and second variations; and An image processor unit (3) characterized in that it is adapted to perform

14. Image processor unit (3) according to claim 13, characterized in that the image processor unit (3) is configured to process image data by performing the steps according to any one of claims 1 to 10.

15. A computer program comprising instructions which, when executed by a processing unit, cause the processing unit to perform the steps of the method of any one of claims 1 to 12.