Image data processing method

The method addresses the inefficiency of existing image denoising and deflickering techniques by employing hierarchical multi-scale decomposition and weighted filtering, achieving fast and efficient noise reduction and flicker removal in image data.

JP2025515954APending Publication Date: 2025-05-20バエレザビエル +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024568499
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-05-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Existing image denoising and deflickering methods require significant computational power and are slow, especially for large images like video data, while maintaining image quality is a challenge.

Method used

A computer-implemented method using hierarchical multi-scale decomposition and weighted filtering to process image data, where pixel arrays are decomposed into low-frequency and high-frequency components, and edge-preserving convolutions are applied to reduce noise and flicker efficiently.

Benefits of technology

The method achieves fast denoising and deflickering with reduced computational requirements, preserving image quality by using weighted filtering based on multi-scale decomposition, particularly effective for low-light video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515954000001_ABST
    Figure 2025515954000001_ABST
Patent Text Reader

Abstract

1. A computer-implemented method for processing image data representing at least one image, the image data including at least one input pixel array, each pixel of the at least one input pixel array having an associated pixel value, the method comprising: recursively performing a hierarchical multi-scale decomposition of the image data into a multi-level pixel array hierarchy, wherein for each scale level of the multi-level hierarchy, the at least one input pixel array is decomposed into a low frequency pixel array and at least one high frequency pixel array.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates generally to methods for processing image data, and more particularly to methods for denoising and / or de-flickering image data, especially low light image data. [Background technology]

[0002] Digital images are inevitably degraded by noise, i.e. artifacts that do not result from the content of the original scene, which can reduce the visual quality of the image, especially in low-light images. In a sequence of images, such as a video image, the image may flicker due to changes in the overall light between frames. The problem of reducing image noise and flicker has been known and studied for a long time, but noise reduction methods are rather complex as they also lead to degradation of image quality, such as loss of edge sharpness, known as blurring, and / or the introduction of artifacts.

[0003] There are many different filtering methods, for example mean filtering or Wiener filtering, but these are all linear filtering methods in the spatial domain. Another example is bilateral filtering, a relatively widely used method for image denoising. It is a nonlinear technique that has the advantage of preserving edges relatively well. In this method, the intensity value of each pixel is replaced by a weighted average of the intensity values ​​of nearby pixels.

[0004] A problem associated with these methods is that the known methods require significant processing and / or computational power, which is particularly problematic for relatively large images, and thus, due to the power required, such methods can be relatively slow, especially for video images.

[0005] It is therefore an object of the present invention to solve or at least mitigate one or more of the problems mentioned above. In particular, it is an object of the present invention to provide an improved method for denoising and / or deflickering image data in a relatively fast manner while still maintaining efficiency. Summary of the Invention

[0006] For this purpose, a computer-implemented method for processing image data, characterized by the configuration of claim 1, is provided. In particular, the image data represents either at least one image, i.e., a single image, or a time-series image such as a video image. The image data can include, for example, low-light image data. The image data includes at least one input pixel array. A single image can be represented by a single input pixel array, and a series of images can be represented by a time-series pixel array. A pixel value is associated with each pixel of at least one input pixel array. The pixel value may be, for example, a one-dimensional pixel value representing the light intensity or depth of the pixel in the image. Alternatively, the pixel value may be a multi-dimensional pixel value such as the RGB intensity of the pixel in the image. The method for processing the image data is a computer-implemented method for processing image data representing at least one image, the image data including at least one input pixel array I(x, y, t), a pixel value being associated with each pixel of the at least one input pixel array, the method including the step of recursively performing a hierarchical multi-scale decomposition to convert the image data into a multi-level pixel array hierarchy, such that at each scale level of the multi-level hierarchy, at least one input pixel array is decomposed into a low-frequency pixel array and at least one high-frequency pixel array. The hierarchical multi-scale decomposition may include a wavelet decomposition (e.g., Haar wavelet decomposition), or a pyramid decomposition, or other suitable multi-scale decomposition including the execution of a discrete spectral transform. The recursiveness of the execution of the multi-scale decomposition is preferably applied only to the low-frequency pixel array, as is known to those skilled in the art, and only the low-frequency pixel array of the first scale level is preferably further decomposed into the low-frequency pixel array of the next scale level and at least one high-frequency pixel array.

[0007] The method further comprises the step of forming a first pixel array cluster of a scale level by selecting, for each scale level of a multi-level pixel array hierarchy, a plurality of said low frequency pixel arrays and / or said high frequency pixel arrays, in particular a plurality of said low frequency pixel arrays of a scale level of a multi-level hierarchy of time-series input pixel arrays, or a low frequency pixel array and at least one high frequency pixel array of a scale level of at least one input pixel array of a multi-level hierarchy. The first pixel array cluster may enable grouped processing of the selected pixel arrays. The selection of pixel arrays for forming the first cluster may depend on the type of desired processing, such as denoising of image data and / or deflickering of image data.

[0008] The method further comprises, for each scale level of the multi-level hierarchy, performing a first edge-preserving convolution on the low-frequency pixel array of said scale level of the multi-level pixel array hierarchy. In the method of the present invention, the first edge-preserving convolution uses weighted filtering in which the filtering weights for the pixels of the low-frequency pixel array depend simultaneously or jointly on the difference or distance between the pixel values ​​associated with the pixels in each of the pixel arrays of the first pixel array cluster of said scale level, meaning that corresponding pixel values ​​in each of the pixel arrays of the first cluster are jointly considered for grouped processing. The distance between pixel values ​​is understood as a mathematical distance, for example as a Euclidean distance, but not necessarily only. The dependency on the distance may include, for example, a function such that the weight decreases as the distance increases. Alternatively, other functional dependencies may be used. The dependency may include, for example, a combination of increasing weights as the distance increases and decreasing weights as the distance from a threshold distance increases. Compared to conventional bilateral filtering, this innovative weighted filtering uses weights based on multiple associated pixel arrays simultaneously, i.e. jointly, through a hierarchical decomposition into spatial, temporal, and low- and high-frequency pixel arrays. The weighting in conventional bilateral filtering only considers the spatial proximity and intensity difference of neighboring pixels within the image or within the associated pixel arrays themselves.

[0009] Finally, the method includes reconstructing an output pixel array by recursively performing an inverse transform of the hierarchical multi-scale decomposition on the filtered low-frequency pixel array and the high-frequency pixel array. As a result, the output pixel array can provide denoised and / or deflickered image data. Because the method relies on a clever combination of efficient hierarchical image decomposition and innovative weighted filtering, the method is relatively fast and requires less computational and processing time and / or capacity than known methods, which is particularly advantageous for low-light image data, especially for low-light video image data.

[0010] The first pixel array cluster may advantageously be formed by selecting a low frequency pixel array of said scale level of a multi-level hierarchy of said time-series input pixel arrays. Time-series input pixel arrays may be particularly applicable in the case of image data including video image data, where a relatively large number of image frames are captured over a reference time t 0 The time window [t min , t max

[0046] are acquired in a time series at various times t within the time series. Each image frame in such a time series can be represented by an input pixel array I(x,y,t), thus forming a time series input pixel array. Each input pixel array in the time series input pixel arrays is then multiplied in a recursive, multi-scale, hierarchical manner to produce a low frequency pixel array C LF (x,y,t) and at least one high frequency pixel array C HF The selection step for forming the first cluster, which is performed for each scale level, is performed for example at a reference time t 0 The time window [t min , t max ] may include low frequency pixel arrays of said scale levels at various times t.

[0011] Reference time t 0 The time window [t min , t max], the filtering weight w for the pixels of the low-frequency pixel array of the time series input pixel array at time t is TIFF2025515954000002.tif17170, where σ t is a parameter related to the amplitude of the flicker and / or noise. Absolute value |C LF (x,y,t)-C LF (x,y,t 0 )| is the distance on which the weights depend. The negative exponential function is the distance |C LF (x,y,t)-C LF (x,y,t 0 )| increases, the weight decreases. Other functions can be used depending on the desired dependencies and effects. Since the selection step and filtering are applied at each scale level of the multi-level hierarchy, this weight can be different at each scale level. This weight can be taken into account when performing the first edge-preserving convolution, resulting in a filtered low-frequency pixel array C' LF is generated. TIFF2025515954000003.tif17170, where [t min ,t max

[0036] is the time window over which filtering is applied. This time window may vary depending on the scale level of the multi-level hierarchy at which the first edge-preserving deconvolution is performed. In particular, the time window may be larger at higher scale levels of the multi-scale decomposition due to the lower resolution. Filtering image data comprising time-series image frames using the above method with the above filtering weights results in an output pixel array representing filtered image data in which inter-frame flicker of the time-series image frames is effectively minimized.

[0012] Alternatively, the first pixel array cluster can be formed by selecting a low frequency pixel array and at least one high frequency pixel array of the scale level of a multi-level hierarchy of at least one input pixel array. An input pixel array I(x,y) is selected from a low frequency pixel array C(x,y) and a high frequency pixel array C(x,y) in a recursive, multi-scale, hierarchical manner. LF (x, y) and at least one (preferably multiple) high frequency pixel array C HF (x,y) is decomposed into a time window [t min , t max For a set of input pixel sequences at time t in [, the hierarchical multi-scale decomposition may be performed for each time. The selection step performed for each scale level and each time to form a first cluster may include low frequency pixel sequences of said scale level at time t and all high frequency pixel sequences of said level at time t. Such selection into a first pixel sequence cluster may be particularly effective for denoising an image.

[0013] The filtering weight preferably depends on the distance between pixel values ​​in a neighborhood region associated with a pixel and surrounding said pixel. The neighborhood region may have the same size in the low frequency pixel array and at least one high frequency pixel array of said scale level. The size of the neighborhood region may be determined as a compromise between computation time and improved image quality.

[0014] The low frequency pixel array C LF The filtering weight W(i,j) for a pixel at (x,y) is, for example, The image can be given by TIFF2025515954000004.tif17170, where D is where k is the index of the high frequency pixel array of the scale level at which filtering is performed and (x+i,y+j) denotes the neighboring pixels around pixel (x,y). Again, the negative exponential function has the effect that the weights decrease as the distance D increases, although other functions could be used depending on the desired dependency and effect. With given weights, performing an edge-preserving convolution on the low frequency pixel array of said scale level produces a filtered low frequency pixel array C' LF is obtained. TIFF2025515954000006.tif17170

[0015] If the first pixel array cluster includes a low frequency pixel array and at least one high frequency pixel array of the scale level, the method may further include, for each scale level of the multi-level hierarchy, performing an edge-preserving convolution on at least one high frequency pixel array of the scale level of the multi-level pixel array hierarchy using weighted filtering, where the filtering weights for the pixels of the at least one high frequency pixel array depend simultaneously, i.e. jointly, on the distance between pixel values ​​associated with the pixels in each of the pixel arrays of the first pixel array cluster of the scale level. In that case, performing the edge-preserving convolution on the at least one high frequency pixel array of the scale level generates at least one filtered high frequency pixel array C' HF is obtained. TIFF2025515954000007.tif17170, where σ 1 is a parameter that depends on the scale level l, and σ 1 = α l ·σ, where σ depends on the average noise amplitude, and α is a constant (eg, α≈0.48).

[0016] If the edge-preserving convolution is also performed on at least one high-frequency pixel array, the step of reconstructing the output pixel array I'(x,y) may be performed by reconstructing the filtered low-frequency pixel array C' LF (x,y) and the filtered high frequency pixel array This is done by recursively performing an inverse transform of the hierarchical multi-scale decomposition on TIFF2025515954000008.tif9170. Filtering image data using the above method with the above filtering weights results in an output pixel array representing filtered image data in which noise in said image data is efficiently minimized. The efficiency, particularly the reduction in computation time and required computing power, is at least partially due to the fact that the weights used in filtering are directly based on the multi-scale decomposition into a low frequency pixel array and at least one high frequency pixel array, reducing the number of operations performed.

[0017] It may further be preferred that the filtering includes at least one coefficient configured to adjust the weights of the low frequency pixel array and the at least one high frequency pixel array of the first cluster. Such coefficients may be constant weights and depend on the noise level and the type of multi-scale decomposition. The coefficients include a coefficient K specific to the filtering of the low frequency pixel array. LF and a coefficient K for at least one high frequency pixel array. HF It could be.

[0018] Preferably, the method further comprises the steps of forming a second cluster of pixel arrays of said scale level by selecting said low frequency pixel arrays of said scale level of a multi-level hierarchy of time series input pixel arrays, and performing a second edge-preserving convolution step on the filtered low frequency pixel arrays. In this way, different selections can be performed for different purposes, e.g. a first selection step can form a first cluster for a first type of image processing and a second selection step can form a second cluster, which may be different from the first cluster, for a second type of image processing.

[0019] The first pixel array cluster may, for example, include a low-frequency pixel array and at least one high-frequency pixel array of the scale level of a multi-level hierarchy of at least one input pixel array (preferably of a time-series input pixel array), while the second pixel array cluster may include a low-frequency pixel array of the scale level of a multi-level hierarchy of the time-series input pixel array. The first pixel array cluster and the second pixel array cluster may include at least partially the same pixel array. In particular, a low-frequency pixel array of a given scale level may be part of the first cluster and the second cluster. In this way, the low-frequency pixel array may be subjected to two edge-preserving convolutions with different weights.

[0020] In this preferred embodiment of the method, as described above, a first edge-preserving convolution can be performed on the low frequency pixel array and at least one high frequency pixel array of the scale level of the multi-level pixel array hierarchy. In particular, the edge-preserving convolution uses weighted filtering, where the filtering weight for a pixel of the low frequency pixel array or at least one high frequency pixel array depends simultaneously, i.e. jointly, on the distance between the pixel values ​​associated with the pixel in each of the pixel arrays of the first pixel array cluster of the scale level. In particular, the weight can take into account, for example, adjacent pixels of both the low frequency pixel array and the high frequency pixel array of the given scale level at a given time point. A second edge-preserving convolution can then be performed on the filtered low frequency pixel array of the scale level of the multi-level pixel array hierarchy using weighted filtering, where the filtering weight for a pixel depends simultaneously, i.e. jointly, on the distance between the pixel values ​​associated with the pixel in each of the pixel arrays of the second pixel array cluster of the scale level. In particular, the weight can take into account, for example, adjacent pixels of both the low frequency pixel array and the high frequency pixel array of the given scale level at a given time point. 0 The pixel values ​​of the time series pixel array within a predetermined time range near the first edge-preserving convolution may be taken into account. It is also possible to take into account neighboring pixels in the weights of the second edge-preserving convolution, but the image quality may not be improved as much and the calculation time may be longer. The step of reconstructing the output pixel array may be performed by recursively performing an inverse transform of the hierarchical multi-scale decomposition on the low-frequency pixel array filtered by the second edge-preserving convolution and the high-frequency pixel array filtered by the first edge-preserving convolution. In this way, performing the first and second cluster formation steps, as well as the first and second edge-preserving deconvolution, may make it possible to first denoise the image data, especially on individual images, before reducing the flicker between the image data in the time series images.

[0021] The method may further comprise a step of post-processing the image data, in particular the image data of the reconstructed output pixel array. This step may preferably comprise performing a weighted filtering of the reconstructed output pixel array, said filtering weight for a pixel of said reconstructed output pixel array depending on the difference between the pixel value associated with said pixel and a local mean. The selection of such a post-processing step may depend on a hierarchical multi-scale decomposition. In some decompositions multi-scale induced artifacts such as the Gibbs phenomenon may occur, which may require correction in a post-processing step. This step of post-processing of image data may preferably be performed according to a computer-implemented method for post-processing of image data, which may be considered as an invention in itself. The image data is represented as an image pixel array, with each pixel of the image pixel array being associated with a value. The method comprises a step of determining an average pixel array by convolving the image pixel array, for example the output pixel array I'(x,y) of the method as described above, with an averaging kernel. The average kernel may for example be The kernel may be a Gaussian blur 3x3 kernel, such as TIFF2025515954000009.tif17170, or other known averaging kernel. The method further includes determining a difference δ of neighboring regions in the image pixel array (e.g., I'(x+i,y+j) and the average pixel array μ(x,y)). The method further includes determining a weighted filter result for said difference δ. The method of post-processing the image data finally adds the weighted filter result to the average pixel array μ(x,y), thereby generating post-processed image data I res As an example, a weighting filter may be applied to the reconstructed output pixel array I′(x,y), resulting in a post-processed pixel array that may be expressed as: TIFF2025515954000010.tif22170 where σ is a parameter that depends on the multiscale decomposition, in particular the type of decomposition and the number of scale levels, and δ(x+I,y+i)=I'(x,y)-μ(x,y), where μ(x,y) is the mean kernel for I'(x,y), e.g. This is the local average obtained by convolving with TIFF2025515954000011.tif17170.

[0022] The method may further include a step of pre-filtering the image data. Pre-filtering the image data may include normalizing the levels of the image data and / or removing statistical outliers among the pixels of the image data, for example due to dead pixels, burnt pixels, or locked pixels. This step of pre-filtering the image data may be performed using any known pre-filtering method.

[0023] Alternatively and preferably, the pre-filtering of the image data can be performed according to a computer implemented method for pre-filtering image data, which can be considered as an invention in itself, said image data being represented as an image pixel array, with each pixel of the image pixel array being associated with a value. The method comprises determining an average pixel array by convolving the image pixel array with an averaging kernel. The average kernel can be, for example, The kernel may be a Gaussian blur 3x3 kernel such as TIFF2025515954000012.tif17170, or other known averaging kernel. The method further comprises determining a variation pixel array v by convolving the absolute value of the difference |I(x,y)-μ(x,y)| between the image data and the average pixel array μ(x,y) with an averaging kernel, for example the aforementioned matrix M. The method then comprises determining a modified difference δ' between the image data and the average pixel array, i.e. δ=|I(x,y)-μ(x,y)|. The modified difference δ' comprises an exponential function of the difference that depends on the variation pixel array, such that the modified difference comprises a reduced value for differences for values ​​outside the distribution determined by the average pixel array and the variation pixel array. The method for pre-filtering the image data finally involves adding the modified difference δ' to the average pixel array μ(x,y), which returns the noise to normalized statistics and results in the pre-filtered image data I'(x,y).

[0024] The modified difference δ′ may preferably include a linear response to values ​​within the distribution determined by the mean pixel array and the variation pixel array, whereby the central pixel value may remain unmodified.

[0025] The modified difference δ′ may include a response factor ρ configured to adjust the modified difference to include reduced and amplified values, respectively, for differences in values ​​within the distributions determined by the average pixel array and the variation pixel array. In particular, the response factor ρ may be selected as follows: When ρ=1, we obtain a linear response in δ under local perturbations v. When ρ<1, a reduced response with respect to δ is obtained under local perturbations v. When ρ>1, an amplified response with respect to δ is obtained under local variations v.

[0026] Said corrected difference δ′ is advantageously The image size can be given by TIFF2025515954000013.tif13170, where N is a constant: TIFF2025515954000014.tif9170, where ρ is the response factor as described above. The modified difference is a function of the difference δ that includes two exponential functions. The interaction between the two exponential functions allows a single function to describe the different behavior inside and outside the distribution determined by the mean pixel array and the variation pixel array. Without the modified function, a similar behavior would be described and programmed using several different functions depending on the region, or in programming terms, using loops and conditional functions over the region. The modified difference can avoid such loops and conditional functions, simplifying and speeding up the calculations, thereby accelerating the pre-filtering of image data without compromising image quality.

[0027] The above computer-implemented method for pre-filtering image data can be used on image data independently of the method for processing image data as described above. However, the pre-filtering method can also be advantageously integrated into the method for processing image data, either as a separate pre-filtering step prior to the recursive performance of a hierarchical multi-scale decomposition of the image data, or as part of one or more edge-preserving deconvolutions in the processing method.

[0028] According to further aspects of the present invention there is provided a controller, a computer program product and a computer readable storage medium for carrying out the above method having the features of claims 12, 13 and 14 respectively and thus providing one or more of the aforementioned advantages. [Brief description of the drawings]

[0029] [Figure 1] FIG. 1 shows a schematic flow diagram illustrating a preferred embodiment of a computer-implemented method for processing image data according to an aspect of the present invention. [Diagram 2]FIG. 2 shows a series of graphs illustrating the effect of a preferred embodiment of a computer-implemented method for pre-filtering image data according to a further aspect of the present invention. [Diagram 3] FIG. 3 shows a schematic graph illustrating step 110 of the method shown in FIG. [Figure 4a] FIG. 4a shows a schematic graph illustrating step 120 of the method shown in FIG. [Figure 4b] FIG. 4b shows a schematic graph illustrating step 140 of the method shown in FIG. [Diagram 5] FIG. 5 shows a schematic graph illustrating step 160 of the method shown in FIG. [Figure 6] FIG. 6 illustrates a computing system suitable for performing various steps of methods according to example embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] FIG. 1 shows a schematic flow diagram illustrating a preferred embodiment of a computer-implemented method for processing image data according to an aspect of the present invention. After acquiring image data, e.g. an image frame, or more preferably a number of image frames, e.g. time-sequential image frames, such as in video image data, an optional pre-filtering step 100 can be performed. The image frame can be represented by a pixel array, and the time-sequential image frames can be represented as a time-sequential pixel array. In said one or more pixel arrays, each pixel can be associated with a uni- or multi-dimensional value, e.g. representing the intensity, depth or RGB value of that pixel. The pre-filtering 100 of the image data can normalize the shape and level of the image data and make it possible to remove statistical outliers due to dead pixels, burnt pixels, locked pixels, etc. This pre-filtering can be performed by any method known to the person skilled in the art. More preferably, a novel computer-implemented method for pre-filtering image data, which may be considered as an invention in itself, can be applied, which will be described in more detail with reference to FIG. 2. In the next step 110, the input pixel array is decomposed recursively into a multi-level pixel array hierarchy, and for each scale level of the multi-level hierarchy, at least one input pixel array is decomposed into a low frequency pixel array and at least one high frequency pixel array, which will be described in more detail with reference to FIG. 3. Steps 120, 130, 140 and 150 will be described in more detail with reference to FIG. 4a and FIG. 4b. These steps deal with a selection or cluster formation step 120, 140 to form a first and a second pixel array cluster, followed by an edge-preserving deconvolution 130, 150 including weights taking into account the pixel arrays of said first or second cluster. These steps are repeated (125, 145) and are performed for each scale level of the multi-level decomposition. In the first processing unit 135, noise of at least one pixel array can be removed, while in the second processing unit 155, flicker between image frames in the time series of image frames can be reduced.In step 160, at least one output pixel array is reconstructed by recursively performing the inverse transform of the hierarchical multi-scale decomposition on the filtered pixel arrays of steps 130 (165) and 150. Finally, optionally, the at least one output pixel array of step 160 may be post-processed (170) if the reconstructed image has multi-scale induced artifacts, such as Gibbs artifacts. Some of these steps may be performed simultaneously with others and / or the order of the steps may be reversed. As an example, a pre-filtering of the image data may be performed after the hierarchical multi-level decomposition.

[0031] Fig. 2 shows a series of graphs illustrating the effect of a preferred embodiment of a computer-implemented method for pre-filtering image data according to a further aspect of the present invention. According to this method, a mean pixel array μ(x,y) is determined by convolving the pixel array I(x,y) of the input image with an averaging kernel M. A variation pixel array v is then determined by convolving the absolute value of the difference δ between the input image pixel array and the mean pixel array |I(x,y)-μ(x,y)| with an averaging kernel, such as the aforementioned matrix M. In the method of the present invention, a modified difference δ' of the difference δ between the image data and the mean pixel array, i.e. δ=|I(x,y)-μ(x,y)|, is determined. The modified difference δ' comprises an exponential function of the difference that depends on the variation pixel array, whereby the modified difference comprises a reduced value 101 for the difference to values ​​outside the distribution determined by the mean pixel array and the variation pixel array. Instead of adding the difference δ to the average pixel array, the modified difference δ' is added to the average pixel array μ(x,y) to bring the noise back to normalized statistics, resulting in a pre-filtered image frame. The modified difference δ' may preferably be selected to have a linear response 102 for values ​​within the distribution determined by the average pixel array and the variation pixel array. More preferably, the modified difference δ' may include a response factor ρ configured to adjust the modified difference to have reduced and amplified values, respectively, for the difference of values ​​within the distribution determined by the average pixel array and the variation pixel array. In particular, in FIG. 2, graphs 103, 104, 106 represent the modified difference δ' in function of the unmodified difference δ for variance v=1. In the top graph 103, the response factor is ρ=1, so there is a linear response 102 with respect to δ for values ​​within the distribution determined by the average pixel array and the variation pixel array, and a reduced value 101 with respect to the difference for values ​​outside the distribution determined by the average pixel array and the variation pixel array.In the middle graph 104, the response factor is ρ<1 so there is a reduced response 105 in terms of δ for values ​​within the distribution determined by the average pixel array and the variation pixel array, and a reduced value 101 in terms of the difference for values ​​outside the distribution determined by the average pixel array and the variation pixel array. In the bottom graph 106, the response factor is ρ>1 so there is an amplified response 107 in terms of δ for values ​​within the distribution determined by the average pixel array and the variation pixel array, and a reduced value 101 in terms of the difference for values ​​outside the distribution determined by the average pixel array and the variation pixel array.

[0032] 3 is a schematic graph illustrating step 110 of the method depicted in FIG. 1. In step 110, a multi-level pixel array hierarchical decomposition is performed for all image frames represented as input pixel arrays I(x,y). This means that for each scale level 111, 112, 113 of the multi-level hierarchy, the input pixel array is decomposed into a low-frequency pixel array 115 and at least one high-frequency pixel array 116. This is done recursively for the low-frequency pixel arrays 115, i.e., the low-frequency pixel array 115 of the first level 111 can be further decomposed into a low-frequency pixel array 115 of the next level 112 and one or more high-frequency pixel arrays 116. Furthermore, the low-frequency pixel array 115 of the second level 112 can be further decomposed into a low-frequency pixel array 115 of the third level 113 and one or more high-frequency pixel arrays 116. This hierarchical multi-scale decomposition may include wavelet decomposition (e.g., Haar wavelet decomposition), or pyramidal decomposition, or other suitable multi-scale decomposition that involves performing a discrete spectral transform on the input pixel array and the low frequency pixel array.

[0033] Figures 4a and 4b show schematic graphs illustrating steps 120 and 140, respectively, of the method shown in Figure 1. Step 120 shown in Figure 4a is performed for each scale level 111, 112, 113 separately and comprises selecting a number of said low frequency pixel arrays and / or said high frequency pixel arrays of the respective scale level to form a first pixel array cluster of said scale level. The first pixel array cluster 121, 122, 123 may preferably comprise a low frequency pixel array 115 and at least one high frequency pixel array 116 of each scale level 111, 112, 113 of the multi-level hierarchy. Within a time window [t min ,t max For time-series image frames included in [ ], e.g. frames acquired at times t-1, t and t+1 (t being a reference time), this selection step 120 is performed not only for each scale level, but also for each time frame, in the same way that the hierarchical multilevel decomposition is performed for each time frame separately. In a next step 130, a first edge-preserving convolution is performed on the low frequency pixel array 115 and at least one high frequency pixel array 116 of each scale level 111, 112 of the multilevel pixel array hierarchy using weighted filtering. The filtering weight of a pixel depends simultaneously, i.e. jointly, on the distance between the pixel values ​​associated with the pixel in each of the pixel arrays of the first pixel array cluster of the scale level. As a result, pixel values ​​of both the high frequency pixel array and the low frequency pixel array of the first cluster of the respective scale level are considered simultaneously, i.e. jointly. The filtering weight may further depend on the difference between the pixel values ​​associated with the pixel in a neighborhood around the pixel. The filtering may also include at least one coefficient configured to adjust the weights of the low-frequency pixel array and the at least one high-frequency pixel array of the first cluster. Such coefficients are constant weights and depend on the noise level and the type of multi-scale decomposition. The coefficients include a coefficient K specific to filtering the low-frequency pixel array.LF and a coefficient K for at least one high frequency pixel array. HF This first edge-preserving convolution 130 can thus produce a filtered low-frequency pixel array C′ for each scale level. LF and at least one filtered high frequency pixel array C' HF Thus, steps 120 and 130 may be repeated for each scale level 125. This first selection of pixel arrays into first clusters and the associated first edge-preserving convolution may be considered as an image denoising step 135.

[0034] In the next step 140 shown in Fig. 4b, a number of pixel sequences are selected to form a second cluster 141, 142. The second pixel sequence cluster may include low frequency pixel sequences of scale levels 111, 112, 113 of the multi-level hierarchy of the time series input pixel sequences. The time window of this second cluster 141, 142 may vary depending on the scale level 111, 112, 113, and a larger time window may be selected for this second cluster 141, 142 for higher decomposition scale levels. As an example, at scale level 111, the second cluster 141 may include low frequency pixel sequences of t-1, t, and t+1, while at scale level 112, the second cluster 142 may include additional low frequency pixel sequences of earlier and / or later times. At scale level 113, even more low frequency pixel sequences may be included in the second cluster (not shown) of said level. In the next step 150, a second edge-preserving convolution is performed on the filtered low-frequency pixel array 115 at time t for each scale level 111, 112, 113 of the multi-level pixel array hierarchy using weighted filtering. The filtering weight for a pixel depends simultaneously, i.e. jointly, on the distance between the pixel values ​​associated with said pixel in each of the pixel arrays of the second pixel array clusters 141, 142 of the respective scale levels 111, 112, 113. Pixel values ​​in the neighborhood area around said pixel are preferably not taken into account in the weighting, since they are not expected to substantially affect the image quality. As can be seen in Fig. 4a and Fig. 4b, the first cluster 121 and the second cluster 141 contain at least partly the same pixel arrays, in particular the low-frequency pixel arrays of the respective levels of the frame at time t. As a result, a filtered low-frequency pixel array C' resulting in the step image 130. LF A second edge-preserving convolution may be performed on the low-frequency pixel array C" that is doubly filtered for each scale level. LFThe selection step 140 and deconvolution step 150 may be repeated again 145 for each scale level, which may be performed simultaneously with step 125 or sequentially. This second selection of pixel arrays forming the second cluster and the associated second edge-preserving convolution may be considered an image deflickering step 155.

[0035] FIG. 5 shows a schematic graph illustrating step 160 of the method shown in FIG. 1. In this step 160, an output pixel array I′(x,y) is reconstructed, which is the low-frequency pixel array C″ filtered by the second edge-preserving convolution 150. LF and a high-frequency pixel array C′ filtered by the first edge-preserving convolution 130 HF The method proceeds from step 130 directly to step 160 via arrow 165, since the high frequency pixel array is not selected to form the second cluster and is not considered in the second edge-preserving convolution. As mentioned above, the low frequency pixel array is preferably filtered twice, first by a first edge-preserving convolution and then by a second edge-preserving convolution. This recursion of step 160 is performed again only for the low frequency pixel array. In particular, starting from the highest scale level, for example scale level 113, the filtered low frequency pixel array and at least one high frequency pixel array of said level 113 can be reconstructed into a low frequency pixel array of a lower scale level (in particular scale level 112), and the low frequency pixel array and at least one high frequency pixel array of scale level 112 can be reconstructed into a low frequency pixel array of scale level 111. Therefore, the above method allows for obtaining final processed image data with significantly reduced noise and flicker without appreciably losing image details or significantly increasing image blur.

[0036] 6 illustrates a suitable computing system 800 including circuitry enabling execution of steps according to the described embodiments. The computing system 800 is generally formed as a suitable general-purpose computer and may include a bus 810, a processor 802, a local memory 804, one or more optional input interfaces 814, one or more optional output interfaces 816, a communication interface 812, a storage element interface 806, and one or more storage elements 808. The bus 810 may include one or more conductors enabling communication between the components of the computing system 800. The processor 802 may include any type of conventional processor or microprocessor that interprets and executes programming instructions. The local memory 804 may include a random access memory (RAM) or another type of dynamic storage device that stores information and instructions executed by the processor 802, and / or a read-only memory (ROM) or another type of static storage device that stores static information and instructions used by the processor 802. The input interface 814 may comprise one or more conventional mechanisms that allow an operator or user to input information into the computing device 800, such as, for example, a keyboard 820, a mouse 830, a pen, a voice recognition and / or biometric mechanism, a camera, etc. The output interface 816 may comprise one or more conventional mechanisms that output information to an operator or user, such as, for example, a display 840. The communication interface 812 may comprise transceiver-like mechanisms, such as, for example, one or more Ethernet interfaces, that allow the computing system 800 to communicate with other devices and / or systems (e.g., other computing devices 881, 882, 883). The communication interface 812 of the computing system 800 may be connected to such other computing systems via a local area network (LAN) or a wide area network (WAN), such as, for example, the Internet.The storage element interface 806 may comprise a storage interface, such as, for example, a Serial Advanced Technology Attachment (SATA) interface or a Small Computer System Interface (SCSI), for connecting the bus 810 to one or more storage elements 808, such as one or more local disks (e.g., SATA disk drives), and may control the reading and / or writing of data from and / or to these storage elements 808. Although the storage element(s) 808 above are described as local disks, generally other suitable computer-readable media may be used, such as, for example, removable magnetic disks, optical storage media such as CDs or DVDs, -ROM disks, solid state drives, flash memory cards, etc.

[0037] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) Hardware-only circuit implementations, such as analog and / or digital-only implementations; and (b) Combinations of hardware circuitry and software, such as (where applicable): (i) A combination of analog and / or digital hardware circuitry(s) and software / firmware; and (ii) a portion of hardware processor(s) (including digital signal processor(s)), software, and memory(s) with software that work together to cause a device, such as a mobile phone or a server, to perform various functions; and (c) Hardware circuitry(s) and / or processor(s) (e.g., microprocessor(s) or portion of microprocessor(s)) that require software (e.g., firmware) to operate, but the software may not be present if it is not necessary for operation. This definition of circuitry applies to all uses of the term in this application, including any claims. As another example, the term circuitry, as used in this application, covers implementations of only a hardware circuitry or processor (or processors), or implementations of portions of a hardware circuitry or processor and associated software and / or firmware. The term circuitry also covers, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices, if applicable to certain claim elements.

[0038] Although the present invention has been described with reference to specific embodiments, it will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that various changes and modifications can be made without departing from the scope of the present invention. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and the scope of the present invention is indicated by the appended claims, rather than by the foregoing description, and all changes that are within the meaning and range of equivalence of the claims are therefore intended to be included therein. In other words, it is contemplated to cover any and all modifications, variations, or equivalents that are within the scope of the basic principles and whose essential attributes are claimed in this patent application. Furthermore, readers of this patent application will understand that the words "comprising" or "compise" do not exclude other elements or steps, and the words "a" or "an" do not exclude a plurality, and that a single element, such as a computer system, processor, or other integrated unit, may perform the functions of several means recited in the claims. Any signs in the claims should not be interpreted as limiting the respective claims concerned. Terms such as "first," "second," "third," "a," "b," "c," etc., used in the description or claims are introduced to distinguish between similar elements or steps and do not necessarily describe a sequential or chronological order. Similarly, terms such as "top," "bottom," "upper," "lower," etc., are introduced for explanatory purposes and do not necessarily indicate a relative location. It should be understood that the terms so used are interchangeable under appropriate circumstances and that embodiments of the invention can operate in accordance with the invention in other orders or directions other than those described or illustrated above.

Claims

1. 1. A computer-implemented method for processing image data representing at least one image, the image data comprising at least one input pixel array I(x,y,t), each pixel of the at least one input pixel array having associated therewith a pixel value, the method comprising: Recursively performing a hierarchical multi-scale decomposition of the image data into a multi-level pixel array hierarchy, wherein for each scale level of the multi-level hierarchy, the at least one input pixel array is decomposed into a low frequency pixel array C LF (x, y) and at least one high frequency pixel array The steps are decomposed into For each scale level of the multi-level pixel array hierarchy, A plurality of low-frequency pixel arrays C of the scale level of the multi-level hierarchy of the time-series input pixel array I(x,y,t) LF (x, y, t), or the low frequency pixel array C of the scale level of the multi-level hierarchy of the at least one input pixel array LF (x, y) and the at least one high frequency pixel array forming a first pixel alignment cluster for the scale level by selecting one of the for each scale level of the multi-level hierarchy, performing a first edge-preserving convolution on the low frequency pixel array of the scale level of the multi-level pixel array hierarchy using weighted filtering, the filtering weight for a pixel simultaneously depending on the distance between the pixel values ​​associated with the pixels in each of the pixel arrays of the first pixel array cluster of the scale level; reconstructing an output pixel array by recursively performing an inverse transform of the hierarchical multi-scale decomposition on the filtered low-frequency pixel array and the high-frequency pixel array; A method comprising:

2. The method of claim 1 , wherein the first pixel array cluster is formed by selecting the plurality of low frequency pixel arrays at the scale level of the multi-level hierarchy of the time series input pixel arrays.

3. Reference time t 0 The time window [t min , t max ], the low-frequency pixel array C of the time-series input pixel array at time t LF The filtering weight w for a pixel of (x,y,t) is is given by Here, σ t The method of claim 2 , wherein σ is a parameter related to the amplitude of the flicker and / or noise.

4. The method of claim 1 , wherein the first pixel array cluster is formed by selecting the low frequency pixel array and the at least one high frequency pixel array of the scale level of the multi-level hierarchy of the at least one input pixel array.

5. The method according to any one of claims 1 to 4, wherein the filtering weights are further dependent on a distance between the pixel values ​​in a neighbourhood region around a pixel, associated with the pixel.

6. The low frequency pixel array C LF The filtering weight W for a pixel of (x,y) is D(i,j). It depends on the distance given by 6. The method of claim 4, wherein k is the index of the high frequency pixel array at the scale level where filtering is performed, and (x+i, y+j) denotes the neighboring pixels around pixel (x, y).

7. 7. The method of claim 4, further comprising, for each scale level of the multi-level hierarchy, performing an edge-preserving convolution on the at least one high frequency pixel array of the scale level of the multi-level pixel array hierarchy using weighted filtering, the filtering weight for a pixel simultaneously depending on the distance between the pixel values ​​associated with the pixels in each of the pixel arrays of the first pixel array cluster of the scale level.

8. 8. The method of claim 7, wherein the step of reconstructing the output pixel array is performed by recursively performing an inverse transform of the hierarchical multi-scale decomposition on the filtered low-frequency pixel array and the filtered high-frequency pixel array.

9. 9. The method according to claim 4, wherein the filtering weights include at least one coefficient configured to adjust a weight of each of the low frequency pixel array and the at least one high frequency pixel array of the first cluster.

10. 10. The method according to claim 4, further comprising the steps of: forming a second cluster of pixel arrays at the scale level by selecting the low frequency pixel arrays at the scale level of the multi-level hierarchy of time series input pixel arrays; and performing a second edge-preserving convolution step on the filtered low frequency pixel arrays.

11. the step of performing the first edge-preserving convolution is performed on the low frequency pixel array and the at least one high frequency pixel array of the scale level of the multi-level pixel array hierarchy; performing the second edge-preserving convolution on the filtered low frequency pixel array of the scale level of the multi-level pixel array hierarchy is performed using weighted filtering, the filtering weight for a pixel being simultaneously dependent on a distance between the pixel values ​​associated with the pixel in each of the pixel arrays of the second pixel array cluster of the scale level; 11. The method of claim 10, wherein the step of reconstructing the output pixel array is performed by recursively performing an inverse transform of the hierarchical multi-scale decomposition on the low-frequency pixel array filtered by the second edge-preserving convolution and the high-frequency pixel array filtered by the first edge-preserving convolution.

12. A controller comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the controller to perform a method according to any one of claims 1 to 11.

13. A computer program product comprising computer executable instructions for carrying out the method according to any one of claims 1 to 11, when the program is run on a computer.

14. A computer-readable storage medium comprising computer-executable instructions for carrying out the method according to any one of claims 1 to 11, when the program is run on a computer.