Contrast-based autofocus
By dynamically selecting a subset of pixel values of the image sensor and performing bandpass filtering, the focus setting of the image capture device is optimized, which solves the problem of inaccurate focus setting in image sensor data processing in the prior art and achieves more efficient focus setting determination and improved image clarity.
Patent Information
- Application Number
- CN202010939650.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-10
- Filing Date
- 2020-09-09
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2040-09-09
AI Technical Summary
Existing contrast-based autofocus methods in image capture devices have room for improvement, especially when processing image sensor data, where it is difficult to efficiently and accurately determine the focus setting.
The focus setting determination process is optimized by dynamically selecting a subset of the image sensor's pixel values to generate contrast data and utilizing bandpass filtering and normalization techniques, including image processing using color filter arrays and filter combinations.
The accuracy and efficiency of determining focus settings for image capture devices are improved, enabling more flexibility in obtaining improved contrast-based features under different lighting conditions, and enhancing image clarity.
Smart Images

Figure CN112565587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to contrast-based autofocus of image capture devices. Background Art
[0002] It is known to include an image sensor in an image capture device, such as a smartphone camera or a digital camera, for capturing an image. In order to enhance the image quality of an image captured using the image sensor, the focus setting of the image capture device including the image sensor may be appropriately adjusted so that the image is in focus in the plane of the image sensor.
[0003] It is known to use a process known as contrast-based autofocus (AF) to control the focus setting of an image capture device. Contrast-based AF can be performed by measuring the contrast in images of a scene captured by an image sensor using multiple different focus settings. Contrast generally increases as the focus of the image capture device improves. The focus setting used to capture the image with the highest contrast can be used as the focus setting for subsequent images of the scene captured by the image capture device.
[0004] It is desirable to improve contrast-based autofocus methods and systems. Summary of the Invention
[0005] According to a first aspect, a contrast-based autofocus method for an image capture device is provided, the method comprising: acquiring sensor data representing an image captured by the image capture device, wherein the sensor data comprises pixel values for individual sensor pixels from an image sensor of the image capture device; dynamically selecting a subset of the pixel values to generate selected sensor data representing the subset of pixel values; processing the selected sensor data to generate contrast data representing a contrast-based characteristic of at least a portion of the image; and processing the contrast data to determine a focus setting for the image capture device.
[0006] According to a second aspect, a method for determining a focus setting of an image capture device is provided, the method comprising: for each image area of a plurality of image areas: obtaining a first value of a focus metric for the corresponding image area using a first image captured using a first focus setting of the image capture device; obtaining a second value of the focus metric for the corresponding image area using a second image captured using a second focus setting of the image capture device; and processing the first value and the second value to obtain an estimated focus setting for the corresponding image area; and determining the focus setting by performing a weighted summation of the estimated focus settings for at least two image areas of the plurality of image areas.
[0007] According to a third aspect, a method for determining a focus setting of an image capture device is provided, the method comprising: obtaining first focus data representing a first value of a focus metric from at least a portion of a first image captured using a first focus setting of the image capture device; normalizing the first value of the focus metric using a normalization coefficient to generate a normalized first value of the focus metric; obtaining second focus data representing a second value of the focus metric from at least a portion of a second image captured using a second focus setting of the image capture device; normalizing the second value of the focus metric using a normalization coefficient to generate a normalized second value of the focus metric; and processing the normalized first value of the focus metric and the normalized second value of the focus metric to determine the focus setting of the image capture device.
[0008] According to a fourth aspect, a system on chip is provided, configured to: obtain input data in a fixed-point data format; convert the format of the input data from the fixed-point data format to a floating-point data format to generate compressed data; store the compressed data in a local storage unit of the system on chip; extract the compressed data from the local storage unit; and convert the format of the compressed data from the floating-point format to the fixed-point format before processing the compressed data.
[0009] Further features will become apparent from the following description given by way of example only with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a flow chart illustrating a method of determining a focus setting of an image capture device according to an example;
[0011] Figure 2 is a graph schematically illustrating a relationship between a focus metric and a focus setting according to an example;
[0012] Figure 3 is a flow chart illustrating a contrast-based autofocus method of an image capture device according to an example;
[0013] Figure 4 is a schematic diagram illustrating a color filter array according to an example;
[0014] Figure 5 is to be applied to the Figure 4 Schematic diagram of a mask of pixel values obtained by a color filter array;
[0015] Figure 6 is a schematic diagram illustrating generation of filtered data according to an example;
[0016] Figure 7 is a schematic diagram illustrating generation of contrast data according to an example;
[0017] Figure 8 is a schematic diagram illustrating features of a bandpass filtering process according to an example;
[0018] Figure 9 is a flow chart illustrating a method of determining a focus setting of an image capture device according to an example;
[0019] Figure 10 is a schematic diagram illustrating an image area according to an example;
[0020] Figure 11 is shown according to some examples Figure 9 a flow chart of the features of the method;
[0021] Figure 12 is a schematic diagram of a finite state machine according to an example;
[0022] Figure 13 is shown according to some further examples Figure 9 a flow chart of the features of the method;
[0023] Figure 14 is shown according to some further examples Figure 9 a flow chart of the features of the method;
[0024] Figure 15 is a graph schematically illustrating a relationship between a focus metric and a lens position according to an example;
[0025] Figure 16 is a flow chart illustrating a method of determining a focus setting of an image capture device according to a further example;
[0026] Figure 17 is a schematic diagram of components of an image processing system according to an example; and
[0027] Figure 18 is a flowchart illustrating a data processing method according to an example. DETAILED DESCRIPTION
[0028] With reference to the accompanying drawings, details of the systems and methods according to the examples will become apparent from the following description. In this description, for the purpose of explanation, a number of specific details of certain examples are given. Reference in the specification to "example" or similar language means that the specific features, structures, or characteristics described in conjunction with the example are included in at least one example, but not necessarily in other examples. It should be further noted that in order to facilitate the description and understanding of the concepts implicit in the examples, certain examples are schematically described using certain omitted and / or somewhat simplified features.
[0029] Introduction to Contrast-Based Autofocus
[0030] To put the examples in this article into context, we will first refer to Figure 1 and Figure 2 An example of contrast-based autofocus (AF) is generally described.
[0031] Figure 1 is a flow chart illustrating a method of determining a focus setting of an image capture device according to an example. Figure 1 A method for contrast-based AF is provided, which is an iterative process for determining a focus setting for an image capture device, for example, based on the value of a focus metric. The image capture device can be, for example, a smartphone camera, a standalone digital camera, or a digital camera coupled or incorporated into a further electronic device or computer. The focus setting is, for example, the position of a lens of the image capture device relative to an image sensor of the image capture device. The sharpness of an image captured by the image capture device generally depends on the focus setting. Thus, by determining an appropriate focus setting, an image of a scene that appears sharper (rather than blurry or blurred) can be captured.
[0032] exist Figure 1 Item 100 of, obtaining a first value of a focus metric for a first image captured by an image capture device using a first focus setting. A variety of different focus metrics can be used. For contrast-based AF, the focus metric is, for example, indicative of, dependent upon, or based on contrast in the captured image. Contrast is, for example, the difference in brightness and / or color in an image (or portion of an image). The maximum contrast in an image (e.g., representing the difference between maximum brightness and minimum brightness) can be referred to as a contrast ratio or dynamic range. Contrast is typically related to relative differences in brightness and / or color rather than absolute values, because the human visual system is more sensitive to these relative differences in brightness and / or color rather than absolute values. Reference will be made to Figures 3 to 8 An example of generation of contrast data representative of contrast-based characteristics of at least a portion of an image for use in contrast-based AF is further described.
[0033] exist Figure 1 Item 102 of the present invention includes obtaining a second value of a focus metric for a second image captured by the image capture device using a second focus setting. For example, the first image is captured using a lens at a first position relative to the image sensor (e.g., a first distance between the vertical axis of the lens and the vertical axis of the image sensor). The second image may be captured using the lens at a second position relative to the image sensor that is different from the first position (e.g., a second distance between the vertical axis of the lens and the vertical axis of the image sensor).
[0034] This process can be repeated iteratively until multiple images are acquired using multiple different focus settings (e.g., multiple different distances between the lens and the image sensor). Figure 1Item 104 of the present invention involves obtaining an nth value of a focus metric for an nth image captured by an image capture device using an nth focus setting. n is an integer and can be predetermined or pre-set. Alternatively, the value of n can be determined during contrast-based AF processing. For example, once the value of the focus metric reaches a certain value or changes by a certain absolute or relative amount relative to one or more previous values, a decision can be made to stop capturing further images. For example, a hill climbing algorithm can be used to progressively obtain multiple focus metric values for multiple different focus settings while monitoring changes in the value of the focus metric. After the value of the focus metric drops by a certain percentage or more, after progressively passing through one or more different focus settings, it can be determined that the focus setting for subsequent image capture has been passed. Consequently, capturing subsequent images using further focus settings can be stopped.
[0035] exist Figure 1 Item 106 processes the value of the focus metric to determine a focus setting for subsequent captures of images of the scene by the image capture device. The value of the focus metric is maximized when an image captured by the image capture device using a particular focus setting is sharp. The particular focus setting can then be selected as the focus setting for subsequent image captures and can be considered the optimal focus setting for capture of the scene.
[0036] Figure 2 1 is a graph 108 schematically illustrating the relationship between a focus metric and a focus setting, where the focus setting in this case is the lens position of a lens of an image capture device. An x-axis 110 of graph 108 shows the lens position relative to the image sensor, and a y-axis 112 of graph 108 shows the value of the focus metric. Curve 114 is obtained by fitting a polynomial to multiple values of the focus metric obtained for multiple lens positions during contrast-based AF processing. Figure 2 The curve 114 is an ideal curve having a peak point 116 corresponding to the maximum value of the focus metric. The lens position of the image capture device can be adopted as the lens position corresponding to the peak point 116. The peak point 116 can be found analytically based on the curve 114. However, other cases may not include curve fitting for the acquired focus metric values. In this case, the focus setting for subsequent image capture can be adopted as the focus setting corresponding to the maximum focus metric value acquired during the contrast-based AF process. Figures 9 to 16 Further examples of determining focus settings are described in detail.
[0037] Generation of contrast data for contrast-based autofocus
[0038] Figure 3 is a flow chart illustrating a contrast-based AF method of an image capture device according to an example. Figure 3Item 118 of the present invention obtains sensor data representing an image captured by an image capture device. The image represented by the sensor data can be an entire image captured by the image capture device or a portion of a larger image. The sensor data includes pixel values for individual sensor pixels of an image sensor of the image capture device. The image sensor generally includes an array of sensor pixels, which can be any suitable light sensor for capturing an image. For example, a typical sensor pixel includes a light-sensitive element such as a photodiode that can convert incident light into an electronic signal or data. The sensor pixel can be, for example, a charge-coupled device (CCD) or a complementary metal oxide semiconductor (CMOS).
[0039] The pixel value represents at least one characteristic of light captured by, for example, an image sensor. For example, the sensor data may represent the intensity of light captured by each sensor pixel, which may be proportional to the number of photons captured by the sensor pixel. The intensity may represent the brightness of the captured light, which is a measure of the light intensity per unit area rather than the absolute intensity. In other examples, the sensor data may represent the brightness of the captured light, which may be considered to correspond to the perception of illuminance and may or may not be proportional to the illuminance. In general, the sensor data may represent any photometric quantity or characteristic that can be used to represent the visual appearance of an image represented by the sensor data. The sensor data may be in any suitable format, for example, a raw image format. For example, the sensor data may be streamed from the image sensor and saved or not saved to a frame buffer without saving the raw sensor data to a file. However, in this case, the sensor data obtained after processing the raw sensor data may be saved to a file.
[0040] exist Figure 3 Item 120 dynamically selects a subset of pixel values to generate selected sensor data representing the subset of pixel values. Dynamic selection of a subset of pixel values corresponds to, for example, selection of a subset of pixel values that may change over time or that is not fixed or constant. For example, a subset of pixel values that is believed to be most reliable or provide the most contrast information may be dynamically selected, which may improve contrast-based AF processing performed using those pixel values.
[0041] exist Figure 3 Item 122 processes the selected sensor data to generate contrast data representing a contrast-based characteristic of at least a portion of the image. A contrast-based characteristic is any feature of an image (or portion of an image) that, for example, represents contrast of the image or allows contrast to be determined or derived. There are many different contrast-based characteristics that can be used. For example, a contrast-based characteristic can represent variations in the image (e.g., edges of the image) in the time domain or frequency domain.
[0042] exist Figure 3Item 124 processes the contrast data to determine a focus setting of the image capture device. The contrast data may represent a value of a focus metric that may be used to obtain, for example, a reference Figure 1 and Figure 2 or reference Figures 9 to 16 Alternatively, the contrast data may be processed to obtain a value for a focus metric which in turn may be used to obtain, for example, a reference Figure 1 and Figure 2 or reference Figures 9 to 16 Describes the focus setting.
[0043] In some cases, the sensor data may be acquired by an image capture device that includes a color filter array. An example of a color filter array 126 is shown in FIG. Figure 4 It is schematically shown in FIG. Figure 4 The color filter array 126 includes an array (pattern) of color filter elements. The color filter elements correspond to individual sensor pixels of the sensor pixel array of the image sensor. For example, the color filter array 126 can be considered to form a mosaic or repeating pattern. The color filter elements generally allow light of a particular color to pass through the corresponding sensor pixels. In this way, the color filter array allows different sensor pixels of the sensor pixel array to receive incident light of different colors, thereby allowing a full-color image to be captured. Since typical light sensors are not sensitive to the wavelength of the incident light, typical sensor pixels will not be able to provide color information based on the detected light in the absence of a color filter array. However, by using a color filter array to divide the incident light into different wavelength ranges corresponding to different colors, the light intensity in these different wavelength ranges can be ascertained, and thus color information can be determined. Figure 4 , the red filter element is marked with “R”, the color filter element is marked with “G”, and the blue filter element is marked with “B”.
[0044] It will be appreciated that color can refer to any wavelength range of light. For example, a transparent or white filter element can be considered a color filter element in the sense that it allows certain wavelengths (e.g., all or substantially all wavelengths) in the visible spectrum to be transmitted to underlying sensor pixels. In other examples, some or all of the color filter elements can be non-white filter elements.
[0045] The color filter element array can be formed from repeating groups of color filter elements. Color filter element group 128 can include, for example, a red filter element, a blue filter element, and two green filter elements. Thus, the color filter array can correspond to a Bayer array, although other groups are possible in other examples. Each group can correspond to the sensor pixels necessary to obtain a full-color image of appropriate quality.
[0046] Can be used Figure 3 Methods to determine including Figure 4 The focus setting of the image capture device of the color filter array 126. Figure 4 In an example of , a subset of pixel values for generating contrast data may be dynamically selected based on intensity data representing the light intensity received by at least one sensor pixel. For example, if the light intensity acquired by the sensor pixels for a given image area is relatively low, e.g., less than or equal to a given threshold, then the image area may be considered unreliable or provide relatively little contrast information. Similarly, if analysis of the light intensity from these sensor pixels indicates that the image area suffers from a relatively high level of noise, then the image area may be considered unreliable. Accordingly, the pixel values of these sensor pixels may be deleted from the selected subset of pixel values. In Figure 4 Pixel values from the subset of sensor pixels corresponding to color filter element subset 130 can be selected for further processing. Because color filter element subset 130 includes color filter elements of multiple different colors (in this case, red, green, and blue), these pixel values are demosaiced before being processed to generate contrast data. Demosaicing can be applied to input image data representing multiple different color channels (e.g., acquired by sensor pixels corresponding to color filter array 126) to reconstruct a full-color image. In this case, the input image data includes intensity values for only one color channel at each pixel location, and the demosaicing process allows intensity values for each of at least one color channel to be acquired for each pixel location. The demosaicing process includes, for example, interpolating between adjacent pixels of the same color channel to acquire values for locations between these adjacent pixels (e.g., locations corresponding to pixels of different color channels). This process can be performed for each of the multiple color channels to acquire intensity values for each color channel at each pixel location. In some cases, grayscale demosaicing may be performed, wherein a grayscale intensity is acquired at each pixel location, indicating the intensity value of a single color channel (eg, from white (lightest) to black (darkest)).
[0047] By dynamically selecting a subset of pixel values, more accurate or reliable contrast-based characteristics can be obtained for an image or image portion, for example by excluding unreliable areas. For example, the subset of pixel values can change over time or with image capture devices used in different environments and lighting conditions. This allows for improved contrast-based characteristics to be obtained in a more flexible manner than would be possible if the sensor pixels used for contrast-based AF were fixed or unchanging. Additionally, by processing a subset of pixel values rather than the full set of pixel values, contrast data can be obtained more efficiently.
[0048] In some cases, the dynamically selected subset of pixel values is a first subset of pixel values from a first subset of sensor pixels corresponding to a color filter element of a first color. For example, the first subset of sensor pixels may be selected if the light intensity captured by the first subset of sensor pixels is relatively high, e.g., equal to or exceeding a certain threshold, or higher than the light intensity captured by other sensor pixels corresponding to color filter elements of different colors. This can be determined based on the light intensity captured by at least one sensor pixel in the first subset of sensor pixels and / or at least one sensor pixel from the other sensor pixels. For example, if the light intensity is higher, the contrast-based characteristic is more sensitive to changes in contrast. Therefore, using sensor pixels that utilize higher intensity captured light to generate the contrast-based characteristic can improve the accuracy of the contrast-based characteristic. This can be performed dynamically, thereby enabling focus settings to be accurately determined under a variety of different lighting conditions. For example, if the image capture device is used in an environment illuminated primarily by green light, the selected subset of pixel values may be pixel values captured by sensor pixels associated with a green filter element. If the same image capture device is subsequently used in an environment illuminated primarily by red light, the selected subset of pixel values may be the pixel values acquired by sensor pixels associated with red filter elements.
[0049] In some cases, the sensor data includes a second subset of pixel values from a second subset of sensor pixels corresponding to color filter elements of a second color and / or a third subset of pixel values from a third subset of sensor pixels corresponding to color filter elements of a third color. This can be, for example, Figure 4For example, the image capture device may include an RGB (red, green, blue) image capture device that has an image filter array 126 with a color filter array 126. In this case, the selected sensor data may be generated using a first subset of pixel values, but not a second subset of pixel values, and / or not a third subset of pixel values. For example, the subset of pixel values may be obtained only from sensor pixels associated with color filter elements of the same color. This simplifies contrast-based AF processing, for example, because contrast-based AF processing can be performed on the raw sensor data without requiring demosaicing of the sensor data. This allows for more efficient performance of contrast-based AF processing.
[0050] In an example, dynamically selecting a subset of pixel values may include setting a further subset of pixel values to a predetermined value. For example, the predetermined value may be zero. This may be considered to correspond to masking the sensor data to generate masked sensor data. Figure 5 Schematically shows the Figure 4 126 or a portion of the color filter array 126. In other examples, such as in Figure 5 In the example of , mask 132 may be smaller than the sensor pixel array. In this case, pixel values may be dynamically selected by sliding or moving mask 132 across the array of pixel values. Figure 5 In the example, mask 132 is a 2×2 mask whose dimensions correspond to Figure 4 Therefore, 4 bits can be used to represent the sensor pixel group 128 (e.g., corresponding to the sensor pixel repeating unit). Figure 5 The mask 132 of (although this is only an example). Figure 5 In the example of , the elements of mask 132 corresponding to the green filter elements of group 128 can be set to 1, and the elements of mask 132 corresponding to the blue and red filter elements of group 128 can be set to 0. Thus, when each element of mask 132 is multiplied by the pixel value acquired by the corresponding sensor pixel, the pixel value acquired by the sensor pixel associated with the green filter element remains unchanged and is selected as the subset of pixel values. However, the pixel values acquired by the sensor pixels associated with the red and blue filter elements are set to zero (which in this case is a predetermined value) and are therefore deselected and do not form part of the subset of pixel values. Although this is just an example, in other cases the subset of pixel values may be selected in different ways.
[0051] After dynamically selecting a subset of pixel values to generate the selected sensor data, Figure 3 In the example shown in Figure 2, selected sensor data is processed to generate contrast data.
[0052] Figure 6 is a schematic diagram illustrating generation of filtered data according to an example, which may correspond to or be used to generate contrast data (as further referenced). Figure 7 and Figure 8 described above).
[0053] exist Figure 6 For example, as referenced Figure 4 and Figure 5 As described above, selected sensor data 134 is acquired. The selected sensor data 134 is processed using a bandpass filtering process 36 to generate filtered data 138. The filtered data 138 may then be used as contrast data or may be processed to generate contrast data. The bandpass filtering process 136 may be used to suppress low-frequency and high-frequency components of the image represented by the selected sensor data 134. For example, low-frequency components of the image may not provide sufficient information about the contrast of the image. Conversely, when high-frequency components of the image are sensitive to high-contrast image features (e.g., edges in the image), these components may be affected by noise. Therefore, the signal-to-noise ratio may be highest in the intermediate frequency bands, where the intermediate frequency bands may be selected for processing the image using the bandpass filtering process 136.
[0054] exist Figure 6 In the example of , the bandpass filtering process 136 includes processing the selected sensor data 134 using at least one autoregressive (AR) filter 140 and a finite impulse response (FIR) filter 142. At the same time, the AR filter 140 and the FIR filter 142 can be considered to correspond to infinite impulse response (IIR) filters. The value output by the AR filter 140 depends on, for example, at least one previous output of the AR filter 140 and a random term (the random term is, for example, not completely predictable). Since the value output by the AR filter 140 depends on the previous output of the AR filter 140, the AR filter 140 can effectively capture information of an uncertain number of samples. The value output by the FIR filter 142 depends on, for example, a given number of recent samples of the input signal (e.g., the pixel values represented by the selected sensor data 134), which allows noise to be efficiently removed.
[0055] The filtered data 138 generated by the bandpass filtering process 136 indicates, for example, edges in an image represented by the selected sensor data 134. For example, the filtered data 138 may represent a filtered value for each corresponding pixel value input to the bandpass filtering process 136. A higher filtered value may be considered to indicate a higher gradient change associated with the sensor pixel from which the pixel value was obtained, indicating that the corresponding image portion includes a sharper edge. Thus, a higher filtered value may generally be considered to indicate greater contrast.
[0056] Figure 7 A pipeline 144 for generating contrast data according to an example is schematically shown. The contrast data may be processed to determine a focus setting of an image capture device.
[0057] For example, reference Figures 3 to 5 As described above, sensor data 146 is acquired. Sensor data 146 represents an image captured by the image capture device and includes pixel values for individual sensor pixels from an image sensor of the image capture device. The sensor pixels may be arranged in an array comprising rows and columns to correspond to color filter elements of a color filter array, although this is only one example.
[0058] In this case, the pixel values represented by sensor data 146 include base values, which are constant values that are added to pixel values during image capture processing to avoid negative pixel values. For example, an image sensor may record non-zero pixel values even in the absence of light (e.g., due to noise). To avoid reducing pixel values to less than zero, a base value may be added. Therefore, before further processing is performed, sensor data 146 is subjected to a base value. Figure 7 Black level removal 148 in , to remove the base data. It will be appreciated that in some cases black level removal may be performed at a different stage of the image processing pipeline, or in some cases may be omitted (e.g., if no base value is added).
[0059] exist Figure 7 Item 150 dynamically selects a subset of pixel values represented by sensor data 146, as shown in reference Figures 3 to 5 This generates selected sensor data 152 representing a subset of pixel values. Figure 61D filter 154 is configured to filter pixel values from a plurality of columns of the sensor pixel array. The vertical AR filter 156 is configured to filter pixel values from a plurality of rows of the sensor pixel array. The horizontal and vertical AR filters 154 and 156 are each, for example, one-dimensional (1D) filters in different dimensions. For example, the horizontal AR filter 154 may have a width of k (k is an integer) entries in a first dimension (e.g., the x or horizontal direction) corresponding to pixel values obtained from sensor pixels in k different columns, and a length of 1 in a second dimension (e.g., the y or vertical direction). Conversely, the vertical AR filter 156 may have a width of 1 in the first dimension and a length of m (m is an integer) entries in the second dimension corresponding to pixel values obtained from sensor pixels in m different columns. For example, the vertical AR filter 154 may be a 3×1 dimensional tensor, and the vertical AR filter 156 may be a 1×3 dimensional tensor. Using both the horizontal and vertical AR filters 154, 156 as part of the bandpass filtering process allows for more contrast information to be obtained from the image represented by the selected sensor data 152 (e.g., compared to using only the horizontal AR filter 154 or only the vertical AR filter 156).
[0060] exist Figure 7 In the example of FIG, a plurality of rows of selected sensor data 152 are processed sequentially in raster order using the horizontal AR filter 154. For example, the selected sensor data 152 relating to a plurality of rows or lines of pixel values may be processed row by row. In other words, the horizontal AR filter 154 may process the selected sensor data 152 one row of pixel values at a time. However, it will be appreciated that some pixel values of a given row may not be processed, for example, if those pixel values are not part of the subset of pixel values selected at item 150. Nevertheless, in other cases, pixel values that are not selected may be processed by the horizontal AR filter 154, but by utilizing the reference Figure 5 The mask processes these pixel values, which may take default or predetermined values (e.g., zero). Processing the selected sensor data 152 using the horizontal AR filter 154 may also include extracting the selected sensor data 152 (or sensor data 146) from a storage accessible to the pipeline 144 in raster order.
[0061] The output of the horizontal AR filter 154 may be obtained row by row. However, the vertical AR filter 156 generally includes processing subsets of pixel values from multiple rows, including, for example, accumulating the output of the vertical AR filter 156 for each of the multiple rows.
[0062] The data generated by the horizontal AR filter 154 may be stored in a storage unit such as local to the pipeline 144 or accessible to the pipeline 144. To reduce storage requirements, Figure 7 The example pipeline 144 includes sequentially processing a plurality of rows of selected sensor data 152 in raster order using a horizontal AR filter 154 to generate a plurality of sets of first data in a fixed-point data format. Each set of first data corresponds to a respective row in the plurality of rows. A fixed-point data format (sometimes referred to as a fixed-point data type) represents a number with a fixed number of digits after (or sometimes before) a decimal point. A fixed-point data format can represent the first data using a relatively large number of digits (e.g., 41 bits). While this can increase the accuracy of the processing of the first data, the storage requirements for such data can be correspondingly high. Therefore, in Figure 7 Item 160 can convert the format of multiple sets of first data from a fixed-point data format to a floating-point data format to generate second data. The floating-point data format approximates a number by representing the significand x the radix exponent format using a certain number of bits (e.g., 5) to represent the exponent and another number of bits (e.g., 10) to represent the significand. The radix can be fixed (e.g., radix 10), and one bit can also be used to represent the sign of the number, so that negative numbers can be represented. This makes it possible to efficiently represent very small and very large numbers. By converting the format of multiple sets of first data from a fixed-point data format to a floating-point data format, the number of bits used to represent the multiple sets of first data can be reduced. In other words, the second data can be smaller than the first data in size. For example, the size of the second data can be 16 bits, while the size of each set of first data can be 41 bits.
[0063] exist Figure 7 Item 162 stores the second data in a storage unit, such as a local storage unit of pipeline 144. The second data has reduced storage requirements compared to the first data and can be stored and retrieved more efficiently than if the first data were stored without converting the first data to a floating point data format.
[0064] At least a portion of the second data may be sequentially acquired from the storage unit. The portion of the second data may be acquired after processing the selected sensor data of at least the first row of the plurality of rows using the horizontal AR filter 154. Figure 7Item 164 converts the format of at least the portion of the second data from a floating-point data format to a fixed-point data format to generate third data. The third data can then be processed using the vertical AR filter 156. By processing the third data in a fixed-point data format, the output of the vertical AR filter can be obtained without rounding and improving processing accuracy. In some cases, the vertical AR filter 156 can be used to process multiple sets of first data output by the horizontal AR filter 154 for a given row row by row. In these cases, the output of the vertical AR filter 156 applied to the given set of first data can be converted from a fixed-point data format, stored in a storage unit, and subsequently extracted and used as input for subsequent iterations of the vertical AR filter 156 (the vertical AR filter 156 also receives, for example, a set of first data from a subsequent row as input). The output of the subsequent iteration of the vertical AR filter 156 itself can be converted from a fixed-point data format to a floating-point data format, stored in a storage unit, and then extracted from the storage unit and converted to a fixed-point data format. Similar processing can then be repeated by the vertical AR filter 156 for further processing iterations.
[0065] At item 166, the FIR filter 166 processes the data (which may be referred to as third data) output by the AR filters (in this case, the horizontal and vertical AR filters 154, 156) to generate filtered data. Figure 7 In the example, the output of FIR filter 166 is filter data, which is then further processed to obtain filtered data. However, in other examples, the output of FIR filter 166 can be the filtered data itself. FIR filter 166 is, for example, a two-dimensional filter. FIR filter 166 can include at least three lines (e.g., in the vertical dimension), so the dimension of FIR filter 166 is at least n×3, where n is an integer. Such a dimension has been found to generate a filter output that is sensitive to changes in contrast. In some cases, FIR filter 166 includes filter coefficients for multiple lines, and the filter coefficients for the second line between the first line and the nth line are all zero. By including some zero coefficients, the processing of FIR filter 166 data can be performed more efficiently. In addition, by including zero coefficients in the second line (e.g., the middle line), FIR filter 166 can appropriately act as a bandpass filter. In some cases, more than one line of FIR filter 166 (eg, at least two lines of FIR filter 166 that are not the topmost line and the bottommost line of FIR filter 166 ) may include zero coefficients in all or some of its entries.
[0066] The FIR filters may include horizontal FIR filters that filter pixel values from multiple columns and vertical FIR filters that filter pixel values from multiple rows (similar to the horizontal and vertical AR filters 154, 156). The order of the horizontal and vertical FIR filters may be the same or different from each other.
[0067] Examples of horizontal and vertical AR filters 154, 156 and FIR filter 166 are shown in Figure 8 Schematically illustrated in . The term x(m,n) denotes a subset of pixel values in the form of an array having m columns and n rows. The horizontal AR filter 154 is, in this case, a k-order horizontal AR filter, the vertical AR filter 156 is an m-order vertical AR filter, and the FIR filter 166 is an m×n-order 2DFIR filter, where k, m, and n are all integers. The values a1_h…ak_h represent the filter coefficients of the horizontal AR filter 154, the values a1_v…am_h represent the filter coefficients of the vertical AR filter 156, and the values b00…b0N, b10…b1N, bm0…bmn represent the filter coefficients of the FIR filter 166. Indicates horizontal delay, Indicates vertical delay.
[0068] from Figure 8 It can be seen that the first line processed by the FIR filter 166 in this example is received directly from the addition operation of the vertical AR filter 156 without buffering (where the addition operation is performed at Figure 8 154 and / or the vertical AR filter 156) before these outputs are processed by other components (although Figure 8 not shown).
[0069] Subsequent rows are retrieved from storage (e.g., from corresponding delay line buffers) for processing by FIR filter 166. These subsequent rows may be converted from a fixed-point data format to a floating-point data format prior to storage, and may be converted from a floating-point data format to a fixed-point data format after being retrieved from storage and prior to processing by FIR filter 166.
[0070] In an example where the filter coefficients of the FIR filter 166 for at least one row are zero (e.g., the filter coefficients of all rows between the first and last rows of the FIR filter 166 are zero), the outputs of the vertical AR filter 156 for the rows associated with these rows of the FIR filter 166 do not need to be retrieved from the storage unit because the outputs of the FIR filter 166 for these rows are assumed to be zero. This can reduce processing requirements.
[0071] Return Reference Figure 7 Before generating the contrast data, at least one least significant bit (LSB) of the filter output 168 can be discarded. This can be used to make the bit size of the filter output 168 (and the filtered data generated using the filter output 168) match the bit size of the selected sensor data 152, which can be processed using the filtered data to generate the contrast data. By discarding at least one LSB, the bit size of the filter output can be reduced without affecting or reducing the contrast information that can be obtained from the filter output.
[0072] At item 170, filter output 168 is clipped to a region of interest (ROI). For example, this may include retaining the filtered values of sensor pixels corresponding to a particular region (e.g., a region of sensor pixels corresponding to the ROI in the image) represented by filter output 168. However, item 170 may be omitted in some cases, for example, if sensor data 146 represents an ROI of an image rather than the entire image or if the ROI corresponds to the entire image.
[0073] The cropped filter outputs are squared in item 172 and summed in item 174. This includes, for example, squaring the filtered pixel values of the ROI (a dynamically selected subset of pixel values). The squared pixel values are then added together. The summation can be performed for the entire image, the entire ROI, or a subset of the image or ROI (e.g., an image region). A sum value can be obtained for each of multiple image regions of the image. By processing the filter outputs in this manner, filtered data can be generated. However, this is merely an example, and in other cases different processing can be applied to the filter outputs to generate filtered data. Figure 7 Item 176 may convert the data format of the filter data from a fixed-point data format to a floating-point data format. The filter data may then be stored in a storage unit accessible to the image capture device (e.g., a local storage unit of the image capture device). In some examples, the filter data may be considered to correspond to the contrast data itself. In these cases, the contrast data may be converted from a fixed-point data format to a floating-point data format before being stored in the storage unit accessible to the image capture device. Figure 7 In other cases of item 176), the contrast data may also be converted from a fixed-point data format to a floating-point data format before being stored in a storage unit accessible to the image capture device.
[0074] Processing similar to items 170, 172, 176, 177 may be applied to the selected sensor data 152 instead of the filter output 168. These items may be labeled with the same reference numerals and corresponding descriptions will be employed. By processing the selected sensor data 152 in this manner, intensity characteristic data representative of intensity-based characteristics of at least a portion of the image may be obtained. Figure 7 Item 176 may convert the intensity characteristic data into a floating point data format. However, it will be appreciated that in other cases, other processing may be applied to the selected sensor data 152 to generate intensity characteristic data (which may represent, for example, the intensity of an image or other intensity-based features).
[0075] exist Figure 7 , contrast data 178 is generated using the filtered data and the intensity characteristic data. In this case, the intensity characteristic data represents a subset of pixel values that are squared and then summed (which can be represented as I2), and the filtered data represents the output of the bandpass filtering process for each squared and then summed pixel value (which can be represented as E2). Contrast data is generated by dividing the filtered data by the intensity characteristic data, i.e., calculating E2 / I2. However, this is merely an example. In other cases, the generation of the intensity characteristic data can be omitted, and the filtered data can be used as the contrast data without the intensity characteristic data. In addition, it will be appreciated that contrast data can be generated separately for each image region of an image divided into image areas (referred to as image regions) or for the entire ROI or image.
[0076] In this example, the contrast data 178 may be generated in a floating point data format, which allows the contrast data 178 to be efficiently stored in a memory accessible to the pipeline 144 .
[0077] After generating contrast data 178 for a given focus setting of the image capture device, further contrast data may be generated for at least one further focus setting. The contrast data 178 and the other contrast data may be used to determine a focus setting for the image capture device, as described with reference to FIG. Figure 1 and Figure 2 or Figures 9 to 16 For example, the value represented by the contrast data 178 may be used as the value of the focus metric for a given focus setting, or the value represented by the contrast data 178 may be further processed to obtain the value of the focus metric.
[0078] Use contrast data to determine focus settings
[0079] Figure 9 is a flow chart illustrating a method of determining a focus setting of an image capture device according to an example. In item 168, Figure 9The method includes, for each image region of a plurality of image regions, obtaining a first value of a focus metric for the corresponding image region using a first image captured using a first focus setting of an image capture device. The first image may be divided into the plurality of regions in any suitable manner. Figure 10 An example of an image 170 divided into image areas 172 is schematically shown. Figure 10 A first image region 172a and a second image region 172b are labeled in FIG. 1 , but it will be appreciated that the image 170 is divided into a plurality of image regions (collectively referred to by reference numeral 172). Figure 10 In the example, each image region 172 is of equal size because image 170 is divided using a grid having square elements. However, in other examples, some image regions are different sizes than other image regions. For example, image 170 may be divided into image regions based on processing of image 170, e.g., to identify at least one ROI.
[0080] Return Reference Figure 9 , in item 174, for each image region of the plurality of image regions, obtaining a second value of the focus metric for the corresponding image region using a second image captured using a second focus setting of the image capture device. Figures 3 to 8 The acquired contrast data may represent, for example, a focus metric. However, in other cases, a different focus metric may be used (e.g., using a different focus metric than Figures 3 to 8 The focus metric derived from the contrast-based characteristics of Figure 9 method.
[0081] exist Figure 9 Item 176, for each of the plurality of image regions, processes the first value and the second value to obtain an estimated focus setting for the corresponding image region. Figures 11 to 15 As further described, the first and second values may be processed in a variety of different ways to obtain an estimated focus setting.
[0082] exist Figure 9 Item 178 of the present invention determines a focus setting by performing a weighted summation of estimated focus settings for at least two of the plurality of image regions. The weighted summation may, for example, take into account different image characteristics of the at least two of the plurality of image regions, e.g., such that an image region corresponding to the ROI receives a higher weight than other image regions.
[0083] Figure 11 is a flow chart illustrating features of performing a weighted summation of items 178 according to some examples. Figure 11 Item 180, obtains discrete measures. Discrete measures are expressed in Figure 9The dispersion of the estimated focus settings for the plurality of image regions obtained by item 176 is a measure of the dispersion of the distribution of the estimated focus settings, such as the standard deviation, variance, or interquartile range.
[0084] Figure 11 Item 182 includes identifying at least one estimated focus setting to be excluded from the weighted sum based on the dispersion measure. For example, estimated focus settings having a value that is greater than a given threshold distance from the center value of the distribution of estimated focus settings can be removed from the weighted sum. In this way, a subset of estimated focus settings can be used for the weighted sum, rather than all estimated focus settings. This can result in a more appropriate focus setting that is less sensitive to outliers.
[0085] At least two estimated focus settings (in some cases all of the plurality of estimated focus settings) may be used to obtain an average estimated focus setting, for example, corresponding to a center value of a distribution of estimated focus settings. The average estimated focus setting may be, for example, a mean, a mode, or a median. Figure 11 Item 182, the average estimated focus setting can be obtained before excluding outliers and then recalculated after these outliers are removed from the distribution. In other cases, the Figure 11 Item 182 of the embodiment. In any case, performing the weighted summing may include weighting the estimated focus settings of at least two estimated focus settings of at least two image regions of the plurality of image regions based on a difference between the corresponding estimated focus settings and the average estimated focus setting. In this way, outliers that are farther from the average estimated focus setting may be weighted with a smaller value than outliers that are closer to the average estimated focus setting. This further reduces any undue influence on the acquired focus setting by outliers that do not reflect an appropriate focus setting to be used (e.g., noisy image regions or image regions corresponding to image background).
[0086] In some cases, the estimated focus setting of at least one image region may be monitored over time. This allows unreliable image regions to be identified and removed or marked as unreliable. For example, Figure 2 As shown, the value of the focus metric increases to a peak value when the focus setting is changed in a given direction, and then begins to decrease when the focus setting is further changed in the given direction (the given direction is, for example, a direction of decreasing the distance between the lens and the image sensor). However, if the value of the focus metric increases, decreases, and then begins to increase again, this indicates that the image area is unreliable. Similarly, if the value of the focus metric continues to decrease without increasing, this also indicates that the image area is unreliable. Unreliable image areas such as this can be removed from Figure 9The estimated focus value of the image region may be excluded from the weighted sum of items 178 or may be weighted with a smaller weight value. For example, for a given image region, the estimated focus value of the image region may be determined to be excluded from the weighted sum based on a comparison between the first value and the second value.
[0087] It will be understood that Figure 9 The method can be extended to obtain multiple values of the focus metric for each of the multiple image regions using images captured using different corresponding focus settings of the image capture device. The multiple values of the focus metric can be used to obtain an estimated focus setting for each of the multiple image regions. By testing a greater number of different focus settings of the image capture device, the appropriate focus setting for future image capture can be determined more accurately. For example, multiple values of the focus metric can be obtained for a predetermined set of different focus settings of the image capture device. In other cases, Figure 9 The method can be applied to obtain focus metric values for different focus settings until estimated focus settings have been obtained for a predetermined proportion of image regions. For example, once estimated focus settings have been found for at least 70% of the image regions for the entire image or a sub-region (e.g., ROI) of the image, the method can stop obtaining focus metric values for further focus settings.
[0088] In some cases, monitoring of the estimated focus setting for an image region can be performed using a finite state machine. For example, a finite state machine can be used to determine that the estimated focus setting for an image region is to be excluded from the weighted solution. A finite state machine is a model of, for example, a system state that can be in one of a finite number of states at a given time. Figure 12 An example of a finite state machine 184 is schematically shown. Finite state machine 184 includes a start state 186, an end state 188, and various different states between the start and end states. In this case, the system transitions between different states based on changes in the value of a focus metric for a given image region. An increase in the value of the focus metric is indicated using a solid line, a decrease in the value of the focus metric that is less than or equal to a threshold decrease is indicated using a dotted line, and a decrease in the value of the focus metric that is greater than a threshold decrease is indicated using a dashed line.
[0089] After acquiring a first value of the focus metric, the system begins in a start state 186. If the second value of the focus metric increases, the system moves to a good 1 state 190. If the value increases further, the system moves to a good 2 state 192 and remains in the good 2 state while the value continues to increase. If the system is in the good 2 state and the value of the focus metric decreases by an amount greater than a threshold decrease, the system moves to an end state 188. If the system is in the good 2 state and the value of the focus metric decreases by an amount less than or equal to the threshold decrease, the system moves to a nearly state 194 and remains in the nearly state 194 while the value of the focus metric continues to decrease. If the system is in the nearly state 194 and the value of the focus metric decreases by an amount greater than the threshold decrease, the system moves to an end state 188. If the system reaches the end state 188, it can be determined that a reliable focus setting can be estimated from the acquired value of the focus metric.
[0090] However, if the system is in the start state 186 and the value of the focus metric decreases by an amount less than or equal to the threshold decrease, the system moves to the warning state 198 and remains in the warning state 198 if the value of the focus metric continues to decrease. If the system is in the warning state 198 and the value of the focus metric decreases by an amount greater than the threshold decrease, the system moves to the reject state 200. The image region may then be rejected as unreliable. If the system is in the warning state 198 and the value of the focus metric increases, the system moves to the OK state 202. If the system is in the OK state 202 and the value of the focus metric decreases, the system moves to the reject state 200 and the image region may be rejected. However, if the system is in the OK state 202 and the value of the focus metric increases, the system moves to the good 2 state 194.
[0091] If the value of the focus metric increases and then begins to decrease, or increases, decreases, and then begins to increase again, then the image region that appears to be reliable may subsequently be rejected. For example, if the system is in the good1 state 190 and the value of the focus metric subsequently decreases, then the system transitions to the warning state 198. Similarly, if the system is in the nearly state 196 and the value of the focus metric subsequently increases, then the system transitions to the reject state 200.
[0092] It will be understood that Figure 12 The finite state machine of is merely an example, and other finite state machines may be used in other examples.
[0093] Similar to Figures 9 to 12The method can be adapted to obtain a focus setting for capturing an image of a scene including objects at multiple depths. For example, such a scene may include an object in the foreground that is closer to the image capture device than other objects in the background of the scene. Figure 13 The flowchart of FIG. 1 shows an example method for obtaining the focus setting in this case.
[0094] In application Figure 13 Before the method, for example, use Figures 9 to 12 Using a method, estimated focus settings for multiple image regions are obtained. Increases and decreases in focus metric values for the image regions can be counted to obtain a fluctuation count for each image region. A fluctuation graph representing the fluctuation counts for the multiple image regions can then be obtained. The fluctuation count indicates, for example, the number of changes in the direction of change in the focus metric values. For example, if the focus metric value for a given image region increases and then decreases, the fluctuation count may be 1. Similarly, if the focus metric value for another image region increases, decreases, and then increases again, the fluctuation count may be 2.
[0095] A focus metric value may be obtained for each of a plurality of different focus settings (e.g., different lens positions) of the image capture device. This may be performed until a focus metric value is obtained for each of a predetermined plurality of different focus settings, or until an estimated focus setting is obtained for a predetermined number or proportion of the image area (e.g., at least 55% of the image area).
[0096] At this point, different image part geometries (e.g., rows, columns, or 3×3 image blocks) can be overlaid on multiple image areas. Each geometric figure that meets a certain criterion indicating that the geometric figure may correspond to the foreground of the scene (or a part of the scene that is at a different position relative to the image capture device than other parts) can be selected for further processing. For example, each geometric figure for which an estimated focus setting has not been acquired for a predetermined proportion of the image area (e.g., 2 / 3) can be selected. If any of these geometric figures are relatively dark (e.g., having a predetermined proportion, e.g., 2 / 3 of the image area, that is too dark, e.g., having an intensity value less than a predetermined threshold), these geometric figures can be deselected. Similarly, if the fluctuation count of a given geometric figure meets or exceeds a given threshold, for example, if the fluctuation count is similar to the number of different focus settings for which focus metric values have been acquired, the geometric figure can also be deselected.
[0097] If no geometry remains selected after this process, you can refer to Figures 9 to 12The focus settings of the plurality of image regions are obtained as described above. However, if at least one geometric figure is selected, the image regions corresponding to the selected at least one geometric figure may be combined into image sub-regions. The process for obtaining the estimated focus settings may be performed for the image sub-regions, for example, by continuing to obtain focus metric values for a further plurality of different focus settings of the image capture device. This process may be combined with the reference Figures 9 to 12 The process described is the same and may continue until an estimated focus setting is obtained for a larger proportion than before overlaying the image portion geometry on the image area (eg, 95% instead of 55%).
[0098] The estimated focus settings may then be clustered into a plurality of clusters, e.g., so that similar estimated focus settings can be grouped into clusters. In this way, a first image subregion comprising a first set of a plurality of image regions and a second image subregion comprising a second set of a plurality of image regions may be identified. The first and second image subregions are, e.g., non-overlapping and may be adjacent to each other or separated from each other by at least one further image region. After the first and second image subregions are identified, a Figure 13 method.
[0099] exist Figure 13 Item 202 determines a first average estimated lens position for a first image sub-region using a first set of estimated focus settings in the plurality of image regions.
[0100] exist Figure 13 Item 204 of the present invention includes determining a second average estimated lens position for a second image subregion using the second set of estimated focus settings in the plurality of image regions. Item 204 may be performed for a plurality of different image subregions, each image subregion corresponding to an estimated focus setting value for a respective cluster (if multiple clusters exist).
[0101] exist Figure 13 Item 206 determines that the first average estimated lens position is greater than the second average estimated lens position. For example, the first average estimated lens position may be determined to be greater than any of a plurality of other average estimated lens positions corresponding to other clusters. It will be appreciated that any suitable averaging method may be used, such as calculation of the mean, median, or mode. Thus, it may be determined that the first image sub-region corresponds to the foreground of the scene, which is closer to the image capture device than other portions of the scene.
[0102] exist Figure 13 Item 208 of the invention determines a focus setting by performing a weighted summation of the estimated focus settings for a first set of the plurality of image regions, as described with reference to FIG. Figure 9 as described in item 178 of the .
[0103] In some cases, item 208 may be performed if a first set of the plurality of image regions meets some criterion indicating that it is reliable and likely to accurately correspond to the foreground. For example, item 208 may not be performed if the first set includes fewer than two image regions. Item 208 may also be omitted if the distance between the lens positions of the first set and the lens positions of the second set is less than a given threshold, as this indicates that the first and second sets may include portions of the scene at similar distances from the image capture device. Item 208 may also be omitted if the degree of dispersion of the lens positions of the first set meets or exceeds a dispersion threshold, as this indicates that the lens positions may be unreliable or that the first set may include portions of the scene that are at a wide range of distances from the image capture device rather than corresponding to the foreground.
[0104] In the event that item 208 is omitted, focus setting of the image capture device can be performed by performing a weighted summation of the estimated focus setting for each image area of the image area including the image area outside the first set (e.g., the first and second sets combined) or a plurality of image areas divided from the image (or for each image area for which an estimated focus setting has been found).
[0105] The light intensity captured by a sensor pixel of an image capture device generally depends on the device's focus setting. Therefore, changing the focus setting can change the captured light intensity. This is due, for example, to the zoom effect of the image capture device's lens, which means that each image area tends to collect more photons when the lens is focused on an object at infinity (compared to when the lens is focused on an object closer to the lens). For example, as the lens moves away from the image capture device's image sensor, the image field of view increases. This means that light incident on the lens is projected onto a smaller area of the image sensor, thereby increasing the amount of light focused on a given sensor pixel. This effect may not be apparent in some focus metrics, such as the E2 metric discussed above, or other focus metrics that are relatively independent of light intensity, and / or under certain lighting conditions. For example, in a well-lit scene, due to the dominance of the E2 metric over the I2 metric, the focus metric E2 / I2 may not be perceptible. However, when the captured image is relatively dark or noisy (e.g., when I2 dominates over E2), this effect may be perceptible. This is perceptible because the potential light intensity captured by the image capture device increases with increasing lens position. Thus, a contrast-based feature can be seen as a bump on the increasing curve of the focus metric (where the gradual increase corresponds to the increasing light intensity with increasing lens position).
[0106] Figure 14is a flow chart illustrating an example of a method that may be used to counteract this effect, thereby compensating for changes in the value of the focus metric (rather than changes indicative of contrast changes) due to overall or global changes in captured light intensity.
[0107] exist Figure 14 Item 210 of the present invention involves obtaining intensity data representing a difference in light intensity captured by an image capture device using a first focus setting and a second focus setting. The difference may be captured in the form of a graph. For example, a focus metric graph may be obtained for at least a portion of a plurality of images using different respective focus settings. The value of the focus metric may indirectly represent the difference in light intensity captured using the first and second focus settings. In other cases, the light intensity captured using the first and second focus settings may be determined directly, for example, from the image intensity of the images captured using the first and second focus settings, respectively. Image intensity refers to, for example, the sum or average of pixel values of an image or image portion.
[0108] Figure 15 An example of a focus metric diagram 218 is schematically shown, where the focus metric in this example is E2 / I2. Figure 15 Graph 218 shows the dependence of the focus metric on light intensity as a function of lens position. For the focus metric E2 / I2, the E2 metric is invariant with respect to the position of the lens. However, the I2 metric (which corresponds to the sum of squared pixel values of at least a subset of pixels of an image or image portion) may be sensitive to lens position, at least in some cases (e.g., in low light conditions). For example, I2 may decrease with focus position, which means that the inverse of I2 increases. Dashed line 224 shows the contribution of the increase in the inverse of I2 to the actual focus metric E2 / I2, which is Figure 15 2 is shown by a solid line 226. The x-axis 220 of graph 218 corresponds to the lens position (in this example, the lens position is the focus setting), and the y-axis 222 of graph 218 corresponds to the value of the focus metric E2 / I2. Graph 218 has a bump 227 that corresponds to the optimal lens position for the image capture device. However, using the described method, due to the potential increase in image intensity as the lens position increases, the lens position obtained from graph 218 for future operation of the image capture device (e.g., the lens position corresponding to maximum image intensity) may be calculated as infinity rather than the position corresponding to bump 227.
[0109] Return Reference Figure 14 ,exist Figure 14Item 212 processes the intensity data to determine a compensation measure to be applied to the focus metric to compensate for the difference in light intensity. The compensation measure is, for example, an inverse function that counteracts changes in total image intensity due to changes in focus settings. This can include, for example, obtaining a graph of image intensity or fitting a straight line function to a straight line portion of the graph based on the image intensity and using the straight line function as the compensation measure or obtaining the compensation measure.
[0110] For example, if it is assumed that the decrease of I2 is linear, the value of I2 as a function of the lens position p can be expressed as follows:
[0111] I2(p)=(mp+1)I2′
[0112] Where I2 is the measured light intensity, p is the lens position, I2' is the corrected intensity (which is, for example, a constant or substantially constant intensity as the lens position changes), and m is the slope of the curve due to the effect of changing the lens position on the light intensity. Based on this, the corrected I2' value can be obtained as follows:
[0113]
[0114] This allows obtaining a corrected I2' value from the measured I2 value by dividing it by the determination factor.m can be obtained during calibration of a given image sensor or can be calculated using the separation distance between the lens and the image sensor.
[0115] It will be appreciated that this may be a simplification of the underlying relationship between the corrected and uncorrected image intensity values which may be easier to determine than the actual relationship. For example, the actual relationship may be non-linear which may be more expensive to compute.
[0116] exist Figure 14 Item 214 obtains uncorrected focus data representing an uncorrected value of a focus metric for at least the portion of a second image captured using a second focus setting of the image capture device.
[0117] exist Figure 14 Item 216 processes the uncorrected focus data using a compensation measure to generate second focus data representing a second value of a focus metric. In this way, the effect of a potential increase in image intensity due to an increase in the distance between the lens and the image capture device can be reduced or eliminated, thereby enabling a more accurate focus metric to be determined.
[0118] This can be performed, for example, during a calibration process of the image capture device by capturing a relatively dark scene. Figure 14Items 210 and 212. This is more efficient than repeatedly performing this process during operation of the image capture device (although this may be done in some cases). In other cases, the intensity data may be predicted intensity data, for example, obtained from predicted or expected performance of the image capture device. Although Figures 9 to 13 The context of obtaining an estimated focus setting of an image capture device is described in Figure 14 , but it will be appreciated that the focus setting of the image capture device may be performed using other methods of determining the focus setting of the image capture device and / or using other focus metrics than those described herein. Figure 14 method.
[0119] Figure 16 is a flow chart illustrating a method of determining a focus setting for an image capture device according to a further example. Figure 16 Item 228, capturing at least a portion of a first image using a first focus setting of an image capture device, and acquiring first focus data, the first focus data representing a first value of a focus metric. Figures 3 to 8 The acquired contrast data may represent, for example, a focus metric. However, in other cases, Figure 16 The method can utilize different focus metrics (e.g., using different Figures 3 to 8 The portion of the first image may be, for example, an image region among a plurality of image regions of the first image, or another portion of the first image such as an ROI. In some cases, the first focus data represents a first value of the focus metric for the entire first image.
[0120] exist Figure 16 Item 230 normalizes the first value of the focus metric using a normalization coefficient to generate a normalized first value of the focus metric. The normalization coefficient can be a fixed or constant number. However, in some cases, the normalization coefficient corresponds to the first value of the focus metric.
[0121] exist Figure 16 Item 232 of the invention obtains, for at least a portion of a second image captured using a second focus setting of the image capture device, second focus data representing a second value of the focus metric.
[0122] exist Figure 16 Item 234 normalizes the second value of the focus metric using a normalization coefficient to generate a normalized second value of the focus metric.
[0123] exist Figure 16 Item 236 processes the normalized first value of the focus metric and the normalized second value of the focus metric to determine a focus setting of the image capture device.
[0124] In some cases, the normalized first and second values of the focus metric may be found for each of the plurality of image regions. In this case, the normalized first and second values of the focus metric may be found for each of the plurality of image regions. Figures 9 to 13 The normalized first and second values are processed as described to determine a focus setting (eg, using a weighted sum of the normalized first and second values for at least some image regions).
[0125] In other cases, processing the normalized first value of the focus metric and the normalized second value of the focus metric can include fitting a polynomial function to the normalized first value and the normalized second value, and using the polynomial function to determine a focus setting of the image capture device. For example, the method can include determining normalized values of the focus metric for a plurality of different focus settings, e.g., until a predetermined number of focus settings are reached or the normalized values of the focus metric appear to be converging. For example, the method can include stopping investigating further focus settings once a difference between a most recently acquired normalized value of the focus metric and a previous normalized value of the focus metric meets or exceeds a given threshold. For example, if the most recently acquired normalized value is at least 3% less than the previous normalized value, the acquired normalized value can be used to determine the focus setting without acquiring further normalized values.
[0126] The polynomial function in these cases is a function of the focus setting. The order of the polynomial function can depend on the number of different focus settings for which standardized values of the focus metric have been obtained. For example, if standardized values have been obtained for at least four different focus settings, a third-order polynomial can be fitted to the standardized values. If standardized values have been obtained for three different focus settings, a second-order polynomial can be fitted to the standardized values. If standardized values have been obtained for two different focus settings, the focus setting for subsequent operation of the image capture device can be set to a default or predetermined value. For example, the lens position of a lens of the image capture device can be set so that the lens focuses on an object at an infinite distance from the lens. In other cases, if standardized values have been obtained at this stage for two different focus settings, the method can include obtaining further standardized values for additional different focus settings in an attempt to obtain standardized values for at least three different focus settings.
[0127] After fitting the polynomial function, the focus setting corresponding to the maximum value of the polynomial may be taken as the focus setting (eg, lens position).
[0128] In some cases, according to Figure 16The method may include accumulating a first normalized value for each of a plurality of image regions in an image region (e.g., a ROI). Normalized values for other focus settings different from the first focus setting within a given ROI may be similarly accumulated. A polynomial function may be fitted to the accumulated values rather than the normalized values for a single image region, which may more efficiently determine an appropriate focus setting.
[0129] Image processing system
[0130] The examples described in this article can use Figure 17 The image processing system 238 schematically shown in FIG.
[0131] Figure 10 The image processing system 238 includes an image sensor 240 such as described above. The image sensor 240 includes sensor pixels 242 for capturing light. The light received at the image sensor 240 is converted into image data. The image data is transmitted to an image signal processor 244, which is generally configured to generate image data representing at least a portion of an output image. The output image data may be encoded via an encoder 246 before being transmitted to other components for example for storage or further processing. The image signal processor 244 generally includes a plurality of units configured to perform various processing on the image data to generate the output image data. Figure 17 The image signal processor 244 may include a microprocessor, a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. The processor may also be implemented as a combination of computing devices, for example, a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration.
[0132] Figure 17The image signal processor 244 in the example of FIG is arranged to calculate the value of the focus metric as described herein, and thus can be considered to include a focus metric calculation unit 248. Data used in or generated as part of the focus metric calculation unit 248 can be stored in a storage 250 of the image processing system 238. The storage 250 can include at least one of a volatile memory (e.g., random access memory (RAM)) and a non-volatile memory (e.g., read-only memory (ROM) or a solid-state drive (SSD) such as flash memory). The storage 250 is, for example, an on-chip memory or buffer that can be accessed relatively quickly by the image signal processor 244. However, in other examples, the storage 250 can include further storage devices, such as magnetic, optical, or tape media, a compact disk (CD), a digital versatile disk (DVD), or other data storage media. The storage 250 can be removable or non-removable from the image processing system 238. Figure 17 The storage portion 250 in the image signal processor 244 is communicatively coupled to the image signal processor 244 so that data can be transferred between the storage portion 250 and the image signal processor 244. For example, the storage portion 250 can store image data representing at least a portion of an image (e.g., image data before demosaicing) and data generated during calculation of the focus metric value (e.g., data as described above).
[0133] The image signal processor 244 may also include a demosaicing system 252 for demosaicing the image data used in the focus metric calculation 248. The demosaicing system 252 is arranged to perform grayscale blanking to obtain grayscale intensities at corresponding pixel locations from data obtained from the image sensor (the data being, for example, Bayer data). For example, the demosaicing system 252 need not obtain RGB data and may instead obtain grayscale data. In this case, the sensor data used to determine the focus setting may be grayscale data obtained by the demosaicing system 252, which may be considered to correspond to values in the grayscale intensity plane Y. However, in other cases, the sensor data used to determine the focus setting may be the Bayer data itself and may be obtained from the sensor pixels before demosaicing is performed.
[0134] Figure 17 The image processing system 238 also includes a controller 254 for controlling features or characteristics of the image sensor 240. The controller 254 may include hardware or software components or a combination of hardware and software. For example, the controller 254 may include firmware 256, which includes software for controlling the operation of the controller 254. The firmware 256 may be stored in a non-volatile memory of the controller 254 or in a storage 250 that is accessible to the controller 254. Figure 17The controller 254 also includes an automatic image enhancement system 258 that is configured to perform processing such as determining whether adjustments need to be made to the image processing system 238 to improve image quality. For example, the automatic image enhancement system 258 may include an automatic exposure module (e.g., arranged to perform contrast-based autofocus), an automatic white balance module, and / or an automatic focus module. For example, the automatic image enhancement system 258 may include a focus controller such as a contrast-based autofocus controller. In other cases, the focus controller may be a separate unit from the controller 254, and / or the automatic image enhancement system 258 may be omitted. The controller 254 also includes a driver 260 for controlling the operation of the image sensor 240. For example, the driver 260 may control the configuration (e.g., lens position) of the image sensor 240 so that the image sensor 240 is in a configuration corresponding to a particular focus setting.
[0135] Data collection processing, which may be referred to as statistical data collection processing, may be performed using hardware such as the controller 254 or the image signal processor 244 (e.g., by obtaining statistical data (e.g., contrast data) based on image data obtained by an image capture device including the image signal processor 244).
[0136] Firmware such as the firmware 256 of the controller 254 or firmware associated with the image signal processor 244 may be used to perform, for example, reference Figures 9 to 16 The statistical information is processed to determine a focus setting for the image capture device. However, this is not intended to be limiting, and the focus setting may be determined using software, hardware, or a combination of software and hardware.
[0137] The components of image signal processor 244 may be interconnected using a system bus, which allows data to be transferred between the various components.
[0138] Data processing using system-on-chip
[0139] As reference Figure 17 As described above, the data format can be converted to improve the storage efficiency of the data. Figure 18 is a flow chart illustrating a data processing method according to such an example.
[0140] exist Figure 18 Item 262 acquires input data in a fixed-point format. The input data is obtained, for example, from image data representing an image. However, this is merely an example.
[0141] exist Figure 18 Item 264 converts the format of input data from a fixed-point format to a floating-point format to generate compressed data.
[0142] exist Figure 18 Item 266 stores the compressed data in a storage unit. The storage unit is, for example, a local storage unit (eg, an on-chip storage unit of a system on chip).
[0143] exist Figure 18 Item 268 extracts the compressed data from the storage unit.
[0144] exist Figure 18 Item 270 converts the format of the compressed data from a floating point format to a fixed point format before processing the compressed data. Processing the compressed data includes, for example, processing the compressed data as part of an image processing pipeline.
[0145] Figure 18 The approach is counter-intuitive, but can enable data to be stored and / or processed more efficiently. Additionally, in a typical system-on-chip, which may include an integrated circuit such as an application specific integrated circuit (ASIC), data to be processed may be stored in a fixed bit size format such that all data processed by the ASIC has the same bit size. This approach involves converting the data to a floating point format, which may have different bit sizes depending on the data to be converted. This allows the data to be stored more efficiently, but includes a further processing step of converting the data back to a fixed point format before processing. Nonetheless, Figure 18 The method is also more efficient because compressed data generally has lower storage requirements than data in fixed-point format. Therefore, compressed data can be stored and retrieved from storage more efficiently. For example, storage bandwidth and / or storage area can be reduced.
[0146] The above examples are to be understood as illustrative examples. Further examples are foreseeable.
[0147] It will be understood that any feature described in connection with any one example may be used alone or in combination with the other features described, and may also be used in combination with one or more features of any other example or any combination of any other example. In addition, equivalents and modifications not described above may also be adopted without departing from the scope of the appended claims.
Claims
1. A contrast-based autofocus method for an image capture device, the method comprising: acquiring sensor data representing an image captured by the image capture device, wherein the sensor data comprises pixel values for respective sensor pixels from an image sensor of the image capture device; dynamically selecting a subset of the pixel values to generate selected sensor data representing the subset of the pixel values; processing the selected sensor data to generate contrast data representing a contrast-based characteristic of at least a portion of the image, wherein processing the selected sensor data comprises: processing the selected sensor data using a bandpass filtering process to generate filtered data, and processing the filtered data to generate the contrast data, wherein the bandpass filtering process comprises processing the selected sensor data using at least one autoregressive (AR) filter and a finite impulse response (FIR) filter, and wherein the sensor pixels are arranged in an array comprising rows and columns, and the at least one AR filter comprises: a horizontal AR filter that filters pixel values from a plurality of the columns, and a vertical AR filter that filters pixel values from a plurality of the rows; and processing the contrast data to determine a focus setting for the image capture device, The method comprises: sequentially processing the selected sensor data of a plurality of rows among the rows in a raster order using the horizontal AR filter to generate a plurality of sets of first data in a fixed-point data format, each set of first data corresponding to a respective row among the plurality of rows; Converting the formats of the plurality of sets of first data from the fixed-point data format to a floating-point data format to generate second data; storing the second data in a storage unit; and After processing the selected sensor data of at least a first row of the plurality of rows using the horizontal AR filter: acquiring at least a portion of the second data from the storage unit; converting a format of at least the portion of the second data from the floating-point data format to the fixed-point data format to generate third data; and The third data is processed using the vertical AR filter.
2. The method according to claim 1, wherein Dynamically selecting the subset of the pixel values comprises setting a further subset of the pixel values to predetermined values.
3. The method according to claim 2, wherein: The predetermined value is zero.
4. The method according to claim 1, comprising: The subset of the pixel values is dynamically selected based on intensity data representing an intensity of light received by at least one of the sensor pixels.
5. The method according to any one of claims 1 to 4, wherein The image capture device includes a color filter array comprising an array of color filter elements corresponding to individual sensor pixels of the image sensor, and the subset of pixel values is a first subset of pixel values from a first subset of sensor pixels corresponding to color filter elements of a first color.
6. The method according to any one of claims 1 to 4, wherein The FIR filter is a two-dimensional filter including at least three lines, wherein the FIR filter includes filter coefficients for a plurality of lines, and filter coefficients for a second line between a first line and an nth line are all zero, and / or Wherein, the FIR filter includes: a horizontal FIR filter that filters pixel values from a plurality of the columns; and A vertical FIR filter filters pixel values from a plurality of the rows.
7. The method according to any one of claims 1 to 4, wherein Processing the selected sensor data to generate the contrast data includes: processing the selected sensor data to generate intensity characteristic data representative of intensity-based characteristics of at least the portion of the image; and The contrast data is generated using the filter data and the intensity characteristic data.
8. The method according to any one of claims 1 to 4, comprising: At least one least significant bit of a filter output of the bandpass filtering process is discarded before generating the contrast data.
9. The method according to any one of claims 1 to 4, wherein Determining the focus setting of the image capture device includes: For each of the multiple image regions: obtaining, using a first image captured with a first focus setting of the image capture device, a first value of a focus metric for a corresponding image region; obtaining a second value of the focus metric for a corresponding image region using a second image captured using a second focus setting of the image capture device; and processing the first value and the second value to obtain an estimated focus setting for the corresponding image region; and The focus setting is determined by performing a weighted summation of estimated focus settings for at least two image regions of the plurality of image regions.
10. The method according to claim 9, comprising: obtaining a dispersion measure representing a degree of dispersion of estimated focus settings for the plurality of image regions; as well as The estimated focus settings of the at least two image regions of the plurality of image regions are weighted using the dispersion measure.
11. The method according to claim 9, comprising: obtaining an average estimated focus setting using at least two of the estimated focus settings, Wherein performing the weighted summing comprises weighting the estimated focus settings of the at least two image regions of the plurality of image regions based on a difference between the corresponding estimated focus settings and the average estimated focus setting.
12. The method according to claim 9, comprising: For an image region of the plurality of image regions, an estimated focus value of the image region is determined to be excluded from the weighted sum based on the comparison of the first value and the second value.
13. The method according to claim 9, wherein A first set of the plurality of image regions corresponds to a first image sub-region, a second set of the plurality of image regions corresponds to a second image sub-region, and the method comprises: determining a first average estimated lens position for the first image sub-region using the first set of estimated focus settings in the plurality of image regions; determining a second average estimated lens position for the second image sub-region using the second set of estimated focus settings in the plurality of image regions; determining that the first average estimated lens position is greater than the second average estimated lens position; and The focus setting is determined by performing a weighted summation of estimated focus settings for the first set of the plurality of image regions.
14. The method according to claim 9, comprising: acquiring intensity data representing a difference in light intensity captured by the image capture device using the first focus setting and the second focus setting; processing the intensity data to determine a compensation measure to be applied to the focus metric to compensate for the difference in the light intensities; acquiring uncorrected focus data representing an uncorrected value of the focus metric for at least the portion of the second image captured using the second focus setting of the image capture device; as well as The uncorrected focus data is processed using the compensation measure to generate second focus data representing the second value of the focus metric.
15. The method according to any one of claims 1 to 4, wherein Determining the focus setting of the image capture device includes: acquiring first focus data for at least a portion of a first image captured using a first focus setting of the image capture device, the first focus data representing a first value of a focus metric; normalizing the first value of the focus metric using a normalization coefficient to generate a normalized first value of the focus metric; acquiring second focus data for at least a portion of a second image captured using a second focus setting of the image capture device, the second focus data representing a second value of the focus metric; normalizing the second value of the focus metric using the normalization coefficient to generate a normalized second value of the focus metric; The normalized first value of the focus metric and the normalized second value of the focus metric are processed to determine the focus setting of the image capture device.
16. The method according to claim 15, wherein The normalization coefficient corresponds to the first value of the focus metric.
17. The method according to any one of claims 1 to 4, further comprising: obtaining input data in the fixed-point data format, wherein the input data is derived from the sensor data; Converting the input data from the fixed-point format to the floating-point format to generate compressed data; storing the compressed data in a local storage unit of the system on chip; extracting the compressed data from the local storage; and The format of the compressed data is converted from the floating point data format to the fixed point data format before processing the compressed data as part of an image processing pipeline.
Citation Information
Patent Citations
Focus detection apparatus
US20120327291A1
Phase detection autofocus arithmetic
US20170090149A1