Method and image processor unit for processing image data
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- SHENZHEN GOODIX TECH CO LTD
- Filing Date
- 2022-11-30
- Publication Date
- 2026-08-01
AI Technical Summary
In the prior art, digital zooming limits spatial resolution due to the color filter array, resulting in the limitation of the digital zoom quality and the achievable magnification, and it is impossible to achieve high-quality arbitrary amplification effect without changing the optical lens system.
By activating a mechanical actuator to vibrate the image sensor, capture a series of images and align and fuse it, high-resolution images are generated using a multi-frame super-resolution algorithm, avoiding changes to the optical lens system.
It achieves high-quality digital zooming without increasing product cost and size, and maintains image quality at any magnification.
Smart Images

Figure TWG2TB001903417_001 
Figure TWG2TB001903417_002 
Figure TWG2TB001903417_003
Abstract
Description
Image processor unit and method for processing image data This invention relates to a method for processing image data from an image sensor, wherein the image data comprises a raw pixel matrix for each image, i.e., raw image data. The present invention further relates to an image processor unit for processing raw image data provided by an image sensor, the image sensor comprising a sensor array providing a raw pixel matrix for each image. Furthermore, this invention relates to a computer program configured to perform the steps of the aforementioned method. Digital imagers are widely used in everyday consumer products such as smartphones, tablets, laptops, cameras, cars, and personal devices. Using small image sensors is becoming a trend to maintain a small form factor and reduce production costs in lightweight products. Even when using image sensors with a large number of megapixels, color filter arrays (CFAs) (such as common Bayer color filter arrays) are often used to reduce costs. The use of color filter arrays limits or reduces spatial resolution because full-color images are generated through interpolation (de-mosaicing) of color channels. Digital zoom is a common feature that achieves the following: zooming to a region of interest (ROI) in an image, enlarging it to a larger size (even up to the full sensor size), or cropping content outside the ROI. Digital zoom is generally cheaper than optical zoom, which uses optical lens units, because it operates through image processing, not through complex optomechanical lens systems. However, unlike optical zoom, digital zoom does not always produce high-quality images. The quality and achievable magnification of digital zoom are further limited by the aforementioned resolution limiting factors, particularly the small sensor size and the use of color filter arrays. Therefore, achieving high-quality zoom at any magnification, overcoming the resolution limiting factors of the imaging system, and increasing the true optical resolution are highly desirable features that will satisfy consumers of the imaging products in question. The quality and achievable magnification factor of digital zoom by interpolating regions of interest (ROIs) within a single image are limited. Using information from multiple frames when generating the zoomed ROI can result in significantly better image quality (IQ). Multi-frame super-resolution (MFSR) is a known successful method for improving spatial resolution and enhancing detail in images. This method relies on fusing shifted frames of multiple subpixels to generate higher-resolution frames, making it a suitable option for digital zoom. B. Wronski, I. Garcia-Dorado, M. Ernst, D. Kelly, M. Krainin, C. Liang, M. Levoy, and P. Milanfar's "Handheld multi-frame super-resolution" (ACM Trans. Graph., Vol. 38, No. 4, Art, July 28, 2019) discloses a multi-frame super-resolution (MFSR) in which clusters of the original frames displaced by normal hand shaking or hand tremors at subpixel dimensions are fused to generate a higher-resolution frame. This is limited by the Bayer color filter array and relies solely on hand tremors, which are insufficient to generate zoomed images with arbitrary magnification factors, especially when the sensor color filter array has sparse sampling of color channels, such as Hexadeca Bayer CFA and RGBW CFA (RGBW = Red-Green-Blue-White). The complete RGB image is created directly from clusters of the original image from the color filter array, where the clusters of the original frames have a slight offset attributed to natural hand tremors. These frames are then aligned and merged to form a single image with red, green, and blue values at each pixel location. Original image demosaicing replaces the multi-frame super-resolution algorithm. N. El.-Yamany and P. Papamichalis's "Robust color image superresolution: An Adaptive M-Estimation Framework" (EURASIP Journal on Image and Video Processing, 2008, Art, 763254, 2008) reveals an adaptive M-estimation framework for robust color image superresolution that uses a robust error norm in the data fidelity term of the objective function and adapts the estimation procedure to each element in the low-resolution frame and each color component. This results in wrinkled details in the color superresolution image without the use of regularization and without artifacts. The paper "Adaptive framework for robust high-resolution image reconstruction in multiplexed computational imaging architectures" by N. El-Yamany, P. Papamichalis, and M. Christensen (Applied Optics, Vol. 47, No. 10, April 1, 2008, pp. B117-B126) discloses an adaptive algorithm for robustly reconstructing high-resolution images in multiplexed computational imaging architectures. P. Vandewalle, K. Krichane, D. Alleyson, and S. Süsstrunk's "Joint demosaicing and super-resolution imaging from a set of unregistered aliased images" (Proceedings Volume 6502, Digital Photograph III, Electronic Imaging, 2007) presents an algorithm for jointly performing demosaicing and super-resolution imaging from a set of raw images sampled by a color filter array. The combined approach allows for the calculation of alignment parameters between images in the raw camera data before introducing interpolation artifacts. Between the input images, there are small, unknown motions that can be modeled as planar motion. Motion is estimated using a frequency domain method for low-frequency information on the Bayer color filter array images. Luminosity and chromaticity are independent for each input. Higher-resolution luminosity and chromaticity images are calculated independently. These higher-resolution luminosity and chromaticity images are combined to form a higher-resolution color image. S. Farsiu, M. Elad, and P. Milanfar's paper, "Multi-frame demosaicing and super-resolution of color images" (IEEE Transactions on Image Processing, Vol. 15, No. 1, pp. 141-159, January 2006), discloses a hybrid method for super-resolution and demosaicing based on maximum a posteriori (MAP) estimation techniques by minimizing a multinomial cost function. The method determines the projection and estimation of the high-resolution image and the difference between it and each low-resolution image. Outliers in the data and errors attributed to potentially inaccurate motion estimations are removed. Bilateral regularization is used to spatially regularize the luminance component to improve edge sharpness and force interpolation along edges without crossing them. S. Farsiu, M. Elad, and P. Milanfar, “Multi-Frame Demosaicing and Super-Resolution from Under-Sampled Color Images” (Proc. of SPIE, The International Society for Optical Engineering, May 2004), explain the fusion of a set of low-resolution images with relative motion, resulting in a high-resolution image of 16 instances of low-resolution images with Bayer patterns. By fusing the high-resolution images together using the maximum probability estimate, such as low-resolution images after Bayer color filtering, the resolution is increased by a factor of four in each direction. This is a displacement and superposition of all the low-resolution images. The pattern of color distribution in the high-resolution image does not necessarily follow the original Bayer pattern, but depends on the relative motion of the low-resolution images. The field of view of the real-world low-resolution image varies frame by frame, so that the center and boundary patterns of red, green, and blue pixels differ in the resulting high-resolution image. The objective of this invention is to provide an improved method and an image processor unit for processing image data from an image sensor. This objective is achieved by: a method comprising the features described in claim 1, an image processor unit comprising the features described in claim 11, and a computer program comprising the features described in claim 13. Preferred embodiments are described in the dependent claims. To achieve image clustering with sufficient displacement of pixel positions, the method includes the following steps: a) activating a mechanical actuator coupled to an image sensor to cause vibration movement of the image sensor, and b) capturing image clustering when the mechanical actuator is activated. Then, the clusters of images, such as the clusters of the frame of a complete image, are fused by the following steps: c) aligning the original pixel matrix of the clusters of the captured image to a specific alignment, and d) combining the clusters of the image to achieve the resulting image by using multiple pixels of the original matrix that can be used at each pixel position in the matrix of the resulting image. This allows for increased true optical resolution during digital zoom without altering the existing optical-mechanical lens system in the camera setup, thus maintaining a small form factor and not increasing product costs. This method achieves high-quality digital zoom at any magnification factor. Controlled settings are not required. Furthermore, the captured scene does not need to meet any constraints. Activating the mechanical actuator has the following effect: regardless of hand tremor, or other factors, the movement of the image sensor is forced to cause a significantly greater displacement of the pixel position compared to when the sensor image is captured solely by hand tremor. Therefore, this method can be applied to various color filter array arrangements, especially when pixels of the same color are spaced further apart than in a Bayer CFA. By activating actuators in a handheld device, such as vibrator units, which are mechanical actuators configured primarily for other purposes, no additional hardware is required for the device. Vibrator units can be, for example, those used in smartphones, tablets, and portable devices for transmitting signals. Therefore, commonly available existing hardware / mechanical components and various camera products can be used without additional cost. These vibrator-like mechanical actuators also do not need to be coupled to an image sensor to move it relative to the camera device's housing or mounting frame. Therefore, the image sensor itself maintains its relative position to the image sensor's retaining frame or housing. Actuating the mechanical actuator has the effect of causing the retaining frame or housing to vibrate and move together with the image sensor. Before proceeding to steps a) and d), in order to control the acquisition of image clusters (including frames) when the mechanical actuator is activated, it is preferable to select a region of interest (ROI) of the image to be acquired, and determine the size and magnification factor of the region associated with the original image acquired by the image sensor. Once the ROI is selected, the selection of the ROI can be used as a trigger signal for the automatic activation of the mechanical actuator in step a). The size and magnification factor of the region of interest associated with the original image or original frame can be automatically determined based on the known size of the image captured by the image sensor and the selected region of interest of the complete image. The magnification factor is the relationship between the size of the original image captured by the image sensor and the selected size of the image to be captured (i.e., the region of interest). Selecting the region of interest allows for the use of reduced memory capacity, as only the ROI buffer is required, especially in memory-constrained imaging systems. The mechanical actuator can be dynamically activated, causing it to vibrate according to a predetermined or controlled trajectory by adapting to the intensity and duration of the vibration. Therefore, a sufficient number of sub-pixel displacements can be achieved for the target magnification factor, while also taking into account the color filter array arrangement. Due to the dynamic vibration, the image sensor also vibrates dynamically. It is possible to program the actuator to vibrate according to a predetermined trajectory, which can be adapted to the corresponding color filter array arrangement. This provides a sub-pixel displacement effect sufficient to guarantee the combination of cluster information for the image at each pixel location, so that for each pixel location, all color information can be obtained from the original image data of the corresponding color filter array. The alignment of the original matrix in step c) can be performed so that the captured data of the original frame is aligned to the selected base frame. By using the original pixel data, the pixel data of the aligned frame and the pixel data of the base frame are fused in the original domain. Alignment is also known as the "alignment" of the captured clusters of an image (e.g., the original frame). Alignment can be performed by aligning the clusters of the original frame to the first base frame to which the zoom is to be applied. Through a robust multi-frame super-resolution algorithm, the information of the aligned frame and the base frame is fused in the original domain, resulting in a zoom region of interest with higher resolution and full color. The selection of a base frame as the reference frame for alignment can be performed adaptively, particularly ensuring that only high-quality frames with usable information are used. This can be accomplished, for example, through two steps (performance): 1) Initially selecting candidate frames for global motion estimation based on the estimation of frame sharpness. Blurred frames attributable to optical and / or motion blur can be discarded in the selection step. This can be done by quantizing the optical and / or motion blur within the frame, where: - the gradient of the frame or frame ROI is calculated; - the percentage of edge pixels is calculated; - a predefined threshold is applied to the percentage of edge pixels to determine the frame's usability for motion estimation and its applicability for multi-frame fusion. 2) Determining the frame selection for global motion estimation and subsequent multi-frame fusion based on the estimation of motion between frames. Frames with significant regional motion can be discarded in the decision step. This can be done by quantifying motion between frames, where: - for example, the percentage of reliable motion regions within a frame is calculated from the percentage of reliable motion vectors in the motion vector field; - a predefined threshold is applied to the percentage of reliable motion vectors to determine the availability of the frame for motion estimation and its availability for multi-frame fusion. Then, the robust multi-frame super-resolution result of the selected region of interest (ROI) can be reproduced or generated for the user. Therefore, the image obtained in step d) can be reproduced for use in the zoom area of the original image captured by the image sensor. Step d) may include color interpolation of pixel data and steps to improve spatial resolution. The improvement in spatial resolution is primarily a result of fusing the original pixel data of the image cluster. The selected region of interest can also be the entire image, for example, a magnification factor of 1 or 100% of the image size. In this way, this method can be used to expand two optical resolutions without performing any digital zoom; that is, it can be used as a demosaic solution to compensate for the resolution loss introduced by interpolation of conventional color filter array data. The method used to process image data does not require any separate demosaic step, as this step can also be performed by fusing clusters of the original pixel data. The original frame can be divided into uniformly shaped regions. Then, the step of aligning the original matrix in step c) can be performed independently in each region. For global motion model parameter estimation, complete frame information can be used, while for regional motion model parameter estimation, the frame can be divided into uniformly shaped regions and independently aligned in each region. In step c), the original matrix can be aligned by using the color channel with the highest sample count, such as G (green) in a Bayer color filter array or W (white) in an RGBW color filter array. Preferably, the color channel with the highest sample count is used for the alignment / alignment procedure, provided that the color channel is unsaturated or non-zero. Selected color channels of each element in a low-resolution frame can be interpolated within a low-resolution grid to fill missing values attributed to sampling by the color filter array. Based on information from the selected color channels, the frame can be aligned to a selected base frame on the low-resolution grid. This method improves the accuracy and speed of image alignment / proximity because the highest-sampled color channel is used for frame matching. If more than one color channel in the color filter array has the same highest sampled value, all can be used for frame matching in alignment / proximity. The objective is further addressed by the following: an image processor unit, including the features described in claim 11 or 12; and a computer program, including instructions that, when executed by the processing unit, cause the processing unit to perform the steps of the aforementioned method. Figure 1 is an exemplary block diagram of electronic device 1, which includes a camera 2 and an image processor unit 3. The image processor unit 3 is used to process raw image data (IMG) provided by the image sensor 4 of the camera 2. RAW . Image sensor 4 includes a pixel array to enable the original image IMG. RAW This is the data set in the original pixel matrix of each image. To capture colors from the image, a color filter array (CFA) is placed in the optical path in front of the image sensor 4. The camera includes an optomechanical lens system 5, for example, a fixed, uncontrolled lens. The electronic device 1 further includes a mechanical actuator 6. The mechanical actuator 6 is provided in the electronic device primarily for other purposes, such as for sending messages to a user. This is a well-known feature of smartphones used for sending new messages or making calls. In this respect, electronic device 1 can be a handheld device, such as a smartphone, tablet, personal item, or camera, etc. Image processor unit 3 is arranged to process image data (IMG) from image sensor 4. RAW As described below, and while imposing movement on the electronic device 1 (including the image sensor 4) during the process of capturing the cluster information of an image / frame, the mechanical actuator 6 is controlled to capture the cluster information (e.g., the frame) of the image of each image / frame. The image processor unit 3 is arranged to align the pixel matrix of the cluster of the captured image to a specific alignment, and combines the cluster of the image by using a plurality of pixels of the original matrix available for each pixel location to achieve the resulting image IMG. FIN And the resulting image matrix. Figure 2 shows the different pixel arrays assigned to different color filter arrays. In a), a well-known low-order Bayer pixel array assigned to the Bayer color filter array CFA is shown. Numerous 2x2 blocks, including red (R), green (G), and blue (B) pixels, are repeated within the 8x8 pixel array. In b), a 6x6 color filter array pattern is shown, which forms a higher-order color filter array pattern compared to the Bayer color filter array pattern. The 6x6 blocks contain four 3x3 blocks composed of the corresponding colors red (R), green (G), and blue (B). The 3x3 blocks for a specific color RGB are arranged according to the color arrangement of the 2x2 blocks in the color filter array pattern (RGB). In c), the Quad Bayer color filter array pattern is shown. The higher-order pixel array contains pixels in an 8x8 matrix, with repeating 4x4 blocks, each 4x4 being formed by four 2x2 blocks of the corresponding color's R, G, and B. The 2x2 blocks of the corresponding color's RGB are arranged in the lower-order Bayer pattern RGGB. In d), a HexaDeca color filter array is presented, which is formed by an 8x8 matrix. Here, four 4x4 array-sized blocks are assigned to the R, G, and B of the corresponding colors. The four blocks of the corresponding colors are arranged in the same order as the low-order Bayer color filter array pattern shown in a). Figure 3 shows the raw image data (IMG) used to process image sensor 4. RAW The flowchart of the method is as follows: a) Activating the mechanical actuator 6 coupled to the image sensor 4 to cause the image sensor 4 to vibrate; b) Capturing the cluster information of the image when the mechanical actuator 6 is activated; c) Aligning the raw pixel matrix of the cluster information of the captured image to a specific alignment; and d) By using an IMG that can be used to obtain the image FIN The IMG image is obtained by combining multiple pixels of the original matrix at each pixel position in the matrix and combining the images. FIN . Step a) to activate mechanical actuator 6 can be triggered by the previous step of selecting the region of interest (ROI) in the image to which the user might want to zoom. The selection of the ROI can be simplified to the selection of the magnification factor. When a Region of Interest (ROI) is selected, such as manually on the screen, the programmed processor can calculate the size of the selected ROI and the target magnification factor. The target magnification factor is the relationship between the selected ROI and the full image size provided by the image sensor 4. When defining the magnification factor and the associated ROI, the available memory resources in the electronic device 1 can be taken into account. Electronic devices, and in particular image processor units, can be arranged in such a way that the user can repeatedly select the region of interest (ROI) in order to zoom to a gradually smaller ROI. Preferably, the selection of the region of interest (ROI) and / or the amplification factor triggers the start of the mechanical actuator, that is, serves as the trigger signal for step a). Since the movement of the mechanical actuator 6 was initiated in step a), the user receives tactile feedback as confirmation of the request for digital zoom features. This is because the mechanical actuator 6 of the electronic device 1 is coupled to the housing of the electronic device 1, causing the electronic device to vibrate in response to the action of the mechanical actuator 6. The movement of the mechanical actuator 6 causes the electronic device to vibrate, which in turn causes the image sensor 4 to vibrate. In this respect, the mechanical actuator 6 is not primarily intended and directly mounted to the image sensor 4 to cause controlled movement of the image sensor 4. More precisely, the mechanical actuator 4 is mounted to the housing or frame of the electronic device 1, wherein the image sensor 4 is also coupled to this frame or housing. Preferably, the housing of the mechanical actuator 6 is directly or indirectly coupled to the image sensor 4 and its retaining frame (e.g., a printed circuit board) to imply no relative movement between the mechanical actuator 6 and the image sensor 4. Driving the mechanical actuator 6 has the effect of vibrating the retaining frame of the mechanical actuator 6, including vibrating the housing 1 and the image sensor 4. After activating the mechanical actuator 6 in step a), the image cluster is captured in step b) while the mechanical actuator is still activated. Vibration of the electronic device 1, including the image sensor 4, causes sub-pixel displacement of the captured image. In the context of this invention, the image can also be understood as a frame of the entire image, that is, a portion of the entire image. The intensity and duration of the vibrations at which the electronic device 1 is activated by the mechanical actuator 6 can optionally be programmed to generate sufficient sub-pixel displacement for a target magnification factor, taking into account the color filter array arrangement. It is also possible to program the mechanical actuator 6 to vibrate according to a predetermined trajectory. As the mechanical actuator 6 moves, the image sensor 4 also vibrates, thereby imaging the same scene as the displacement rear view. When the image sensor 4 moves due to vibration caused by the mechanical actuator 6, it captures the cluster information of the original frame region of interest (ROI). The number of frames can be determined based on statistical analysis of the movement of the mechanical actuator 6, the corresponding sub-pixel displacements generated by the target color filter array arrangement and the required number of sub-pixel displacements, and the target magnification factor. Then, the captured cluster of information from the original frame or the region of interest (ROI) of the original frame is aligned (aligned) to the selected base frame (usually, the frame to which the zoom is first), and the information from the aligned frame and the base frame is fused in the original domain, thereby generating a full-color zoom (high-resolution) image of the ROI through robust multi-frame super-resolution MFSR. Then, the robust multi-frame super-resolution (MFSR) image obtained from the selected region of interest (ROI) can be processed into an IMG. FIN Reproduce / generate to the user. In this regard, the method provides the following steps: c) aligning the original pixel matrix of the captured image cluster to a specific alignment; and d) combining the image cluster by using a plurality of pixels of the original matrix available for each pixel location to achieve the resulting image, and the resulting image IMG. FIN The matrix. These steps c) and d) will be explained later. Figure 4 is a schematic diagram illustrating the steps of selecting a region of interest (ROI) on the display of an electronic device 1 using the user's hand 7. This causes the mechanical actuator 6 to move, that is, triggers the activation of the mechanical actuator 6, which causes the electronic device to vibrate and provides tactile feedback to the user, as indicated by the arrow pointing to the electronic device 1 shown in the first illustration. The activation of the mechanical actuator 6 has the effect of physically moving (e.g., vibrating) not only the housing of the electronic device but also the image sensor 4. The mechanical actuator 6 vibrates for a sufficient amount of time to allow enough raw image / frame to be captured. The intensity and duration of the vibration can be tuned to produce sufficient sub-pixel shift for the target zoom / magnification factor. Figure 5 is a schematic diagram of steps c) and d) of the method. The raw image / frame IMG captured by image sensor 4 RAW1...n The cluster contains a pixel matrix containing color information such as red (R), green (G), blue (B), and possibly white (W). Low-resolution frame IMG RAW1...n The sub-pixel displacements between frames are crucial for the effectiveness of the multi-frame super-resolution MFSR algorithm for arbitrary magnification factors r. This requires finding the relationship between each low-resolution frame IMG_i and the underlying low-aligned frame IMG. RAW An image alignment method for accurate displacement (and other motion model parameters) between 1 and 2 pixels. In addition to requiring subpixel accuracy, the image alignment algorithm should be as fast as possible to minimize the latency introduced by the combined operation of image alignment and fusion. Therefore, the low-resolution frame IMG is proposed. RAW1...n Alignment is performed within the original color filter array domain because the goal is to jointly perform both color interpolation (de-mosaicing) and spatial resolution improvement. Since alignment in the undersampled color filter array space may not yield the required accuracy, the following strategies are employed, as indicated in Figure 5: 1. Use the color channel with the highest sampling (e.g., G in the Bayer color filter array or W in the RGBW color filter array) in the alignment procedure, provided that the color channel is unsaturated or non-zero. 2. Low-resolution frame IMG. RAW1...n The selected color channels of each element are interpolated against a low-resolution grid to fill missing values attributed to sampling by the color filter array. Based on information from the selected color channels, the frame is aligned to a selected base frame in the low-resolution grid. This method improves the accuracy and speed of image alignment because a higher-sampled color channel (R, G, B) is used in frame matching. If more than one color channel in the color filter array has the same highest sampled value, all can be used for frame matching in alignment. 3. For global motion model parameter evaluation, the complete frame and information can be used. For regional motion model parameter estimation, the frame can be divided into uniformly shaped regions, and alignment can be performed independently in each region. Following the alignment / proximity of the original low-resolution image / frame in step c), in step d), robust adaptive multi-frame fusion is performed independently for each color (R, G, B). Image alignment errors can be attributed to inaccuracies in assumed motion models, occlusion, and regional motion, among other things. Furthermore, the presence of noise should be considered. Multi-frame super-resolution (MFSR) fusion should remain robust to these inaccuracies (including noise) in order to reproduce artifact-free and sharp content. Robust fusion can be based on known retransmission M-estimates from prior art. To account for real-world scenarios, both global and regional motion can exist within the captured low-resolution frame and the region of interest (ROI), which can be partitioned into uniformly shaped regions. This is the same partitioning that might be pursued in the alignment step. The super-resolution can be pursued by minimizing the proposed cost function as follows: in X = Unknown high-resolution frame X c,r = The region error r in the color channel c of the unknown high-resolution frame, where c is the index of the color channel and the color filter array, for example, R, G, and B in the standard Bayer color filter array. F k,r = The motion (shift) operator for region r in frame k; in the case of a global motion model only, F k,r = F k, k = 1, 2, ...N, where N is the number of low-resolution bounding boxes / regions of interest (ROIs). H r = The camera's point spread function (PSF), which allows for variation as a spatial variable, depending on the region r. D = Undersampling operator, which is the reciprocal of the target magnification operator, and is assumed to be the same for all drawing frames. S c = Color channel c filter array sub-operator Y k,c, r = Region r in color channel c of low-resolution frame #k ρ k,r = Robust cost function / robust estimator for region r in low-resolution bounding box k M = Number of color channels in the color filter array pattern R = The number of regions in the drawing frame. robust estimator function ρ k,r Each of them has an outlier threshold T. k, r The outlier threshold is based on the error term E. k, c, rThe calculation is dynamic, so that when the error term is small in region r, it is set to a high value, and when the error term is high for region r, it is set to a low value. Therefore, outliers from the fusion result are rejected by setting them to high values. To better understand steps c) and d) of aligning the original pixel matrix and combining the image clusters (fusion), the following key points are explained in more detail: A. Fusion of the original data compared to fusion of the de-mosaiced data Most prior art multi-frame super-resolution (MFSR) solutions developed and available in the prior art rely on fusing full-color (e.g., full RGB) low-resolution (LR) frames to generate full-color high-resolution frames. Full-color LR frames are typically generated in consumer cameras via demosaicing, with varying demosaicing sensor color filter array patterns and the detailed reconstruction capabilities of the demosaicing solution itself. Interpolation of undersampled color filter array data limits the resolution and detail in each of the LR frames used in the MFSR, thus limiting the final quality of the reconstructed high-resolution (HR) frames. Unlike the fusion of frames after demosaicing, the original MFSR fusion (i.e., fusion in the color filter array data space) results in higher image quality because the undersampling of the color filter array is compensated for by the rich content available from the low-resolution frames shifted from multiple subpixels. In this case, the MFSR is responsible for generating full-color high-resolution frames; that is, performing color interpolation and spatial resolution enhancement. B. Number of low-resolution frames in the MFSR One of the main factors limiting the magnification achievable in MFSR without control settings is the lack of sufficient subpixel shift between the captured frames that are fused to produce higher-resolution frames. Normal hand shaking (tremor) in handheld devices (such as smartphones, tablets, and portable devices) is insufficient to generate enough subpixel shift to produce zoomed images suitable for any magnification factor (from small to large), especially when the color filter array is undersampled. To elaborate on this, consider a monochrome imaging sensor. The table in Figure 6 lists 15 ideal integer displacements in the high-resolution grid (subpixel displacements in the low-resolution grid) to achieve four-fold (4x) magnification using an MFSR. Figure 6 shows a Bayer color filter array pattern consisting of a set of 4x4 pixel matrices with RGB colors. Figure 6 further illustrates the pattern obtained by resampling the low-resolution frames from the original high-resolution grid based on the original image captured by the Bayer color filter array pattern for 15 ideal integer displacements (i.e., frames with index numbers 1 to 15). For each pixel location, the resulting color R, G, B are indicated by the frame index number of that pixel location, which means a specific shift in the high-resolution grid. Low-resolution (LR) frame #1 is not shifted and is selected as the reference frame. All other LR frames #2 through #16 are aligned (positioned) to this reference frame #1. Each of the other frames #2 through #16 has its own specific shift relative to the reference LR frame #1. These shifts are indicated in the X (horizontal) and Y (vertical) directions, as depicted in the table. Optionally, the reference frames can be selected adaptively. In particular, the selection can be processed to ensure that only high-quality reference frames with usable information are used. The high-resolution grid is resampled to a low-resolution frame, where the spacing between adjacent pixels in the low-resolution grid is magnified by a factor of 4. This is highlighted for the RGGB pixels of frame #1 by using a thick perimeter. Sixteen shifts cover all necessary displacements within the 4x4 pixel block to achieve a 4x magnification in monochrome mode. Therefore, each pixel position of the 16x16 high-resolution frame, magnified by a factor of 4, is filled using samples from the 4x4 low-resolution grid. Therefore, by shifting and superimposing all low-resolution images, the low-resolution images after Bayer color filtering are fused together, increasing the resolution by a factor of four in each direction, as described in S. Farsiu, M. Elad, and P. Milanfar's "Multi-Frame Demosaicing and Super-Resolution from Under-Sampled Color Images" (Proc. of SPIE - The International Society for Optical Engineering, May 2004). Because each low-resolution frame is in the original format of a planar color filter array, the starting pixel in the upper left corner can be any color of the color filter array, that is, R, G, or B in this example, but not R. The resulting pattern of color distribution in the high-resolution image does not necessarily follow the original Bayer pattern, but depends on the relative motion of the low-resolution image. The field of view of the real-world low-resolution image varies from frame to frame, so that the center and boundary patterns of red, green, and blue pixels differ in the resulting high-resolution image. For example, frame #3, after alignment, has a starting top-left pixel at sample position B3 in the blue color, and the frame is shifted by 2 in the (horizontal) X direction and 0 in the (vertical) Y direction. The top-left sample B3 appears to be adjacent to the top-left sample R2 of frame #2. The subpixel displacement in a low-resolution grid caused by vibrations attributed to mechanical actuators is related to the pixel displacement in a high-resolution grid as follows. For Mx magnification, 1 / M pixel accuracy is required. Therefore, the displacement of the low-resolution frame relative to the selected reference frame in the X and Y directions must be a multiple of the defined subpixel accuracy (i.e., 1 / M). For 4x magnification, the displacement between low-resolution frames needs to be estimated with 0.25 pixel accuracy. For 2x magnification, 0.5 pixel accuracy is required. Based on the corresponding shift, the low-resolution frames of the color filter array are resampled into a magnified (e.g., 4x4) high-resolution grid, with each frame having its own color sample at the top left corner. An example of this is indicated in the table in Figure 6. Clearly, after resampling the original data within the high-resolution grid, the data becomes sparsely distributed and deviates from the original symmetrical Bayer pattern. With redundant low-resolution frames, some locations will have more color samples than others. Therefore, the reconstruction of the high-resolution frames will not have spatially consistent image quality. Figure 7 illustrates an example where fewer samples are available to fill a high-resolution grid. In this example, frame indices 9 and 14 are missing because the relevant pixel data R9, G9, B9, R14, G14, and B14 are missing in the resulting resampled high-resolution pattern. Therefore, the resulting reconstruction quality will be lower than that shown in Figure 6. This is indicated by the blank spaces in the matrix pattern. For high-quality MFSR reconstruction that successfully overcomes color undersampling of the color filter array and amplifies the true optical resolution by an arbitrary magnification factor r, the low-resolution frame in the better low-resolution grid has a sufficient number of sub-pixel displacements so that the original high-resolution grid will be densely filled (filled) by more color samples available at each location. The original color filter array data is undersampled in the high-resolution grid. For example, for a Bayer pattern, only one color sample is available for each location in the high-resolution grid. For perfect super-resolution results, at least three color samples (R, G, and B) are expected for each location in the high-resolution grid. Sixteen shifts are insufficient compared to the monochromatic case in Figure 6. Figure 8 shows an example with double magnification in a Bayer color filter array pattern, where the red channel R has 15 ideal (x, y) shifts, corresponding to the other color channels. R_i and B_i are the red and blue channel samples of frame #i, respectively, and frame #1 is the base low-resolution frame in this example. Using the shifts listed in the table of Figure 8, red, green, and blue samples can be used at every position in the high-resolution grid, achieving high-quality color interpolation and improved spatial resolution. This is indicated for a 4x4 matrix in the x and y directions. In this example, for each X-shift and Y-shift, the three frames are considered to have different colors at the top-left starting point of the low-resolution grid. This is listed here for the red (R) channel (X, Y) shift and the associated R / B channel in Table b) of Figure 8. Figure 9 presents a table of 36 ideal x, y displacements for a 4x magnification. Similarly, the number of pixel displacements in the x and y directions is listed and assigned to frame numbers #1...36. Figure 10 illustrates a four-fold magnified example of the Bayer color filter array pattern with 36 ideal x, y shifts as shown in Figure 9. It is evident that for each sample in the high-resolution grid, four samples are available, with each of the colors R, G, and B represented by a set of samples. Therefore, it is clear that for high-quality MFSR-based digital zoom or MFSR-based de-mosaic, low-resolution frame clusters, or low-resolution regions of interest (ROIs) with sufficient subpixel displacement within a low-resolution grid, should be available. This is protected by the mechanical actuator 6 of the activation electronics 1, as normal hand shaking (tremor) is insufficient for capturing the original image / original frame IMG. RAW This type of dense sub-pixel distribution is generated in the cluster of information. Therefore, by programming the mechanical actuator 6 with the correct strength and duration, the required sub-pixel displacement can be guaranteed to exist in the original frame IMG. RAW The image is captured in the image cluster (including the original full image and the original region of interest (ROI)). The vibrations generated by the electronic device create uniformly distributed sub-pixel displacements within an area defined by the target magnification factor r and the corresponding color filter array pattern of the camera 2. The image processor unit may be a digital signal processor appropriately programmed by a computer program containing instructions that, when executed by the image processing unit, cause the processing unit to perform steps a) to d) of the aforementioned method. The method, including computer programs and image processing units, can utilize existing hardware, particularly mechanical actuators, which are commonly used for other purposes in various camera products such as smartphones, tablets, and portable devices. The method amplifies true optical resolution during digital zoom without altering the existing optical lens system in the electronic device, thus maintaining a small form factor and not increasing product cost. The method achieves high-quality digital zoom at an arbitrary magnification factor *r* without assuming controlled settings or requiring the captured scene to meet certain constraints. The method is applicable to various color filter array arrangements, where only a region of interest (ROI) buffer is required in imaging systems with limited memory. The method can be used to amplify true optical resolution without performing any digital zoom; that is, it can be used as a de-mosaic solution to compensate for resolution loss introduced by conventional color filter array data interpolation. 1: Electronic device; 2: Camera; 3: Image processor; 4: Image sensor; 5: Optical-mechanical lens system; 6: Mechanical actuator; 7: Hand CFA; 8: Color filter array (IMG) RAW Original image data (IMG) FIN :The obtained image IMG RAW1 … n Low-resolution frame R: Color G: Color B: Color r: Region W: Color X: Direction Y: Direction The invention will be explained below with reference to the accompanying drawings and exemplary embodiments. In the drawings: Figure 1 - A block diagram of the electronic device including the camera, image processor unit, and mechanical actuator; Figure 2 - Examples of different color filter array patterns; Figure 3 - Flowchart of a method for processing image data from an image sensor; Figure 4 - Schematic diagram of a handheld device arranged to perform a method for processing image data; Figure 5 - Schematic diagram of the method steps for aligning and combining images; Figure 6 - An example of resampling a low-resolution frame from the original high-resolution grid using a Bayer color filter array pattern; Figure 7 - Instances with fewer available samples in the absence of figure frame indices 9 and 14; Figure 8 - An example of a Bayer color filter array image magnified twice by using image processing methods; Figure 9 - Table for 36 ideal x, y displacements magnified four times; Figure 10 - A four-fold magnified example of a Bayer color filter array frame with 36 ideal x, y shifts. Domestic storage information (please note in order of storage institution, date, and number): None. International storage information (please note in order of storage country, institution, date, and number): None. 1: Electronic devices 2: Camera 3: Image Processor Single 4: Image sensor 5: Optomechanical lens system 6: Mechanical actuator CFA: Color Filter Array IMG RAW Original image data IMG FIN :The obtained image
Claims
1. A method for processing image data (IMGRAW) of an image sensor (4), wherein the image data comprises a raw pixel matrix for each image, the method comprising: a) activating a mechanical actuator (6) coupled to the image sensor (4) to cause vibrational movement of the image sensor (4); b) capturing a cluster of the image when the mechanical actuator (6) is activated; c) aligning the raw pixel matrices of the cluster of the captured image (IMGRAW) to a specific alignment; and d) combining the cluster of the image (IMGRAW) to achieve a obtained image (IMGFIN) by using a plurality of pixels of the raw matrices available for each pixel position in the matrix of the obtained image (IMGFIN), wherein, Step a) dynamically actuates the mechanical actuator (6) so that the image sensor (4) vibrates according to a predetermined or controlled trajectory, and the predetermined or controlled trajectory is set to adapt to the target magnification factor and the corresponding color filter array arrangement.
2. The method as described in claim 1, wherein, Activating the mechanical actuator (6) is activating an actuator of a handheld device, such as a vibrator unit, as a mechanical actuator (6) configured primarily for other purposes.
3. The method as described in claim 1 or 2, wherein, Before proceeding to steps a) to d), it also includes selecting a region of interest (ROI) of an image to be captured, and determining the size and magnification factor (r) of the region of interest (ROI) associated with the original image (IMGRAW) captured by the image sensor (4).
4. The method as described in claim 3, wherein, The selection of the region of interest (ROI) serves as a trigger signal for the activation of the mechanical actuator (6) in step a).
5. The method as described in claim 1 or 2, wherein, In step c), the original matrices are aligned such that the captured clusters of the original image (IMGRAW) are aligned to a selected base image, and by using the original pixel data, in step d), the pixel data of the aligned corresponding pixel matrix and the pixel data of the related base image are fused in the original domain.
6. The method as described in claim 5 further includes initially selecting candidate frames for motion estimation based on an estimate of the image sharpness, and determining a frame selection for motion estimation and subsequently used as a reference frame or alignment based on an estimate of the motion between frames, adaptively selecting the base image, in particular the reference frame.
7. The method as described in claim 1 or 2 further includes reproducing a zoom area of the resulting image (IMGFIN) achieved in step d) or the original image (IMGRAW) captured by the image sensor.
8. The method as described in claim 1 or 2, wherein, Step d) involves color interpolation of pixel data and improvement of spatial resolution.
9. The method as described in claim 1 or 2 further comprises dividing the original image (IMGRAW) into uniformly shaped regions, wherein, Step c) also includes the original matrices that are independently aligned with the original images (IMGRAW) in each region.
10. The method as described in claim 1 or 2, wherein, In step c), the original matrices are aligned by using the color channel with the highest sampling.
11. An image processor unit (3) for processing raw image data (IMGRAW) provided by an image sensor (4), the image sensor (4) comprising a sensor array providing a raw pixel matrix for each image, characterized in that the image processor unit (3) is arranged to: a) activate a mechanical actuator (6) coupled to the image sensor (4) to cause the image sensor to vibrate; b) when the mechanical actuator is activated, capture a cluster of images (IMGRAW1...n); c) align the raw pixel matrices of the captured cluster of images (IMGRAW1...n) to a specific alignment; and d) combine the cluster of images (IMGRAW1...n) to achieve a obtained image (IMGFIN) by using a plurality of pixels of the raw matrices available for each pixel position in the matrix of the obtained image (IMGFIN), wherein, Step a) dynamically actuates the mechanical actuator (6) so that the image sensor (4) vibrates according to a predetermined or controlled trajectory, and the predetermined or controlled trajectory is set to adapt to the target magnification factor and the corresponding color filter array arrangement.
12. The image processor unit (3) as described in claim 11, wherein, The image processor unit (3) is arranged to process image data by performing the steps of the method described in any one of claims 1 to 9.
13. A computer program comprising instructions which, when executed by a processing unit, cause the processing unit to perform the steps of the method as described in any one of claims 1 to 10.