METHOD AND DEVICE FOR IMAGE CORRECTION
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- CARL ZEISS AG
- Filing Date
- 2016-11-29
- Publication Date
- 2026-04-23
AI Technical Summary
Existing image correction methods, such as speckle imaging, require significant computational effort or specialized hardware, and struggle to correct image disturbances like air streaks and unintentional movements in real-time terrestrial imaging, especially when distinctive patterns are absent.
A method involving image processing that divides images into tiles, transforms them into spatial frequency space, identifies shifts by evaluating the spectral density function, compensates for these shifts, and reassembles the tiles, allowing for robust motion compensation with sub-pixel accuracy and reduced computational effort.
Enables efficient removal of unwanted motion from image sequences with sub-pixel accuracy, even in the absence of distinctive patterns, while preserving moving objects and reducing computational complexity.
Description
[0001] This application relates to methods and devices for image correction and corresponding computer programs. In particular, this application relates to such methods and devices which can be used to correct the effects of air streaks and / or for image stabilization, especially for videos, i.e., image sequences. Such image corrections, in which the effects of disturbances such as air streaks are compensated for, are also referred to as image restoration or video restoration.
[0002] When optical systems, such as telescopes or cameras, are used in the atmosphere, atmospheric turbulence can disrupt the optical image. This turbulence can be caused, for example, by solar radiation, which in turn causes atmospheric turbulence. Such optical systems are often coupled with cameras or image sensors to enable image acquisition.
[0003] Such air streaks can lead to flickering, i.e., movement, in the image and / or blurring in the image that can be viewed through the optical system or recorded via the optical system.
[0004] Unwanted movement in the image can also occur if the optical system is unintentionally moved during image acquisition, for example, due to vibrations of a tripod. With scanning (probing) image acquisition devices, such as a line sensor that captures an object line by line, or a single-pixel sensor that scans an object, additional disturbances can arise from scanning movements. An example of a line sensor is a slit-scan laser scanning microscope (LSM), while an example of a single-pixel sensor combined with object scanning is an electron microscope or a conventional laser scanning microscope. Such movements can have translational and / or rotational components; that is, both displacement and rotation can occur.
[0005] Image disturbances that cause movement, such as the flickering mentioned above or unintentional movements of the optical system, are particularly disruptive in video recordings, i.e., in the case where a large number of successive images are taken.
[0006] With digital image capture using the image sensor commonly used today, it's possible to improve captured images through subsequent computer-aided processing. For example, there are various approaches to correcting unwanted motion or flickering through image processing. Some approaches reduce flickering by registering distinctive patterns, such as prominent edges, in the image. In other words, prominent edges are identified in a sequence of images, such as a video, and the images are corrected so that the patterns align in successive frames. This approach works well when the images contain such distinctive patterns, like edges, for example, when there are many houses with windows. It's also suitable for nature shots with bushes, grass, etc.This becomes difficult because distinctive edges or other suitable patterns are less common or nonexistent. If registering distinctive patterns fails in such a case, it can also negatively impact subsequent image sharpening if a sharpening algorithm is applied to multiple consecutive images.
[0007] Another approach is so-called "speckle imaging," which is described, for example, in C.J. Carrano, "Speckle Imaging over Horizontal Paths," Lawrence Livermore National Laboratory, July 8, 2002. Further information can be found, for example, in G. Weigelt and B. Wirnitzer, OPTICS LETTERS Vol. 8, No. 7, 1983, pp. 389ff., or Taylor W. Lawrence et al., OPTICAL ENGINEERING Vol. 31, No. 3, pp. 627ff., 1992. A major disadvantage of speckle imaging is that it requires significant computing time or correspondingly powerful computers. Speckle imaging was originally used primarily in astronomy, where, at least in many applications, sufficient time for image post-processing and / or corresponding computing power is available.For terrestrial imaging, which often needs to be done in real time, implementing such an approach based on speckle imaging is at least very complex, as it requires specially designed hardware (e.g., ASICs, application-specific integrated circuits), or even impossible. Furthermore, the speckle imaging algorithm processes moving objects, such as walking people or moving vehicles, as well as air currents. This can lead to artifacts in which these moving objects are typically blurred or even partially disappear.
[0008] WO 2009 / 108050 A1 discloses a method for image processing, in particular for correcting imaging errors, which uses a sequence of images. The images are divided into sub-images and transformed into the spatial frequency domain. Correction is then performed based on an argument of an exponential function. The sub-images are transformed into the spatial domain and reassembled.
[0009] Another method for image processing is known from WO 2011 / 139150 A1.
[0010] David J. Tolhurst et al, "Amplitude spectra of natural images", Ophthalmic and Physiological Optics, April 1, 1992, pages 229-232, shows amplitude spectra of images in which the amplitude decreases towards higher frequencies.
[0011] Therefore, it is an object of the present invention to provide methods, devices and corresponding computer programs with which such image corrections can be implemented with less computational effort.
[0012] For this purpose, a method according to claim 1, a device according to claim 12, and a computer program according to claim 15 are provided. The dependent claims define advantageous and preferred embodiments.
[0013] According to a first aspect, a method for image processing is provided, comprehensive: Providing a sequence of images, dividing the images into tiles (i.e., one or more tiles), transforming the tiles into spatial frequency space, identifying shifts by evaluating the argument of the spectral density function of the tiles in spatial frequency space by determining a factor by which the argument of the spectral density function changes frequency-proportional from image to image, compensating for at least some of the identified shifts, e.g., by averaging or by modifying the argument of the spectral density, back-transforming the tiles into spatial space, and reassembling the back-transformed tiles.
[0014] In this way, unwanted motion can be removed from a sequence of images with comparatively little computational effort. Because the process operates in the spatial frequency domain, for example, compared to methods working in spatial space, no registration of prominent image parts such as edges is necessary. This allows the method to operate robustly even in situations where such prominent image parts are absent or barely present. Furthermore, if the Nyquist-Shannon sampling theorem is observed, motion compensation can be performed with sub-pixel accuracy, for example, in the range of 0.01 pixels. This is often not possible with methods operating in spatial space.
[0015] The process can further include averaging over corresponding tiles of the sequence of images (which is equivalent to averaging over time) before assembly.
[0016] Averaging allows any remaining shifts to be factored out.
[0017] The process can further include sharpening over corresponding tiles of the sequence of images before or after assembly.
[0018] The tile size when dividing images into tiles can be determined depending on the size of air streaks during image capture and / or depending on the degree of camera shake during image capture.
[0019] The subdivision, transformation, identification, compensation, and inverse transformation can be repeated N times, N ≥ 1, where the size of the tiles can be gradually reduced during subdivision.
[0020] By gradually reducing the size, smaller air turbulence cells can also be taken into account and smaller shifts can be gradually removed, which further improves the image quality.
[0021] With different repetitions, different images of the sequence of images can be taken into account depending on the size of the tiles during transformation, identification, balancing and / or inverse transformation.
[0022] The factor by which the tiles are reduced in size can differ between at least two of the repetitions.
[0023] This can further reduce the computational effort by dynamically adapting the tiles used to the requirements.
[0024] The images can be color images. Subdivision, transformation, and inverse transformation can be performed separately for each color channel.
[0025] The identification of shifts can then be carried out using grayscale values, which are obtained, for example, by averaging the color channels, and the adjustment for the color channels can be performed for each color channel based on the result of the identification. The same can also apply to other calculations for image correction: Parameters for image correction can be obtained based on grayscale values, and then image correction can be carried out for each color channel based on these parameters.
[0026] This method can also be used to edit colored images.
[0027] According to the invention, identification is based on frequencies (in the spatial frequency domain) below a frequency threshold, i.e., at low frequencies. This facilitates identification, which is generally easier at low frequencies because the magnitude of the spectral density is comparatively high at lower frequencies. Determining based on low frequencies particularly simplifies the determination of shift parameters.
[0028] Identification can involve distinguishing between displacements caused by disturbances and displacements caused by moving objects, with the latter essentially being excluded during compensation. In this way, moving objects can be excluded from image processing and thus preserved without artifacts.
[0029] One or more steps of the process described above can be dynamically adjusted depending on the tiles and / or the content of the images. This allows for a flexible approach depending on the requirements.
[0030] One or more steps of the process described above can be performed in parallel to increase speed. For example, the transformation and / or other described calculations, which can be relatively time-consuming depending on the image size, can be performed in parallel. For example, the calculation for each tile or for groups of tiles can be performed in separate computing units, e.g., separate kernels. In other embodiments, the transformation can also be performed in parallel for different rows and / or columns of the image.
[0031] According to a second aspect, an image processing device is provided, comprising: a computing device with at least one processor and a memory, wherein a sequence of images can be stored in the memory, wherein the processor is configured to perform the following steps: Dividing the images into one or more tiles, transforming the tiles into the spatial frequency space, identifying shifts by evaluating the argument of the spectral density function of the tiles in the spatial frequency space by determining a factor by which the argument of the spectral density function changes proportionally to the frequency from image to image, compensating for at least some of the identified shifts, e.g. by modifying the argument of the spectral density or by averaging, back-transforming the tiles into the spatial space, and reassembling the back-transformed tiles.
[0032] The identification is based on frequencies below a frequency threshold.
[0033] The device can be configured to perform the procedures described above. The device can further include an optical system and a camera system coupled to the optical system for capturing the sequence of images. The camera system can, for example, have a 2D image sensor in which n x m pixels, where n, m is greater than 1, are captured simultaneously. Alternatively, it can have a 1D image sensor (line sensor) or a 0D image sensor (a single pixel) and operate in a scanning mode to, for example, capture a desired object.
[0034] According to a third aspect, a computer program is provided with program code which, when executed on a computing device, performs the method as described above. The computer program can, for example, be stored on a data carrier or in memory, or be provided via a network. The invention is explained in more detail below with reference to embodiments and the accompanying drawings. These show: Figure 1 a schematic representation of a device according to an exemplary embodiment, Figure 2 a flowchart to illustrate a process according to an exemplary embodiment, and Figures 3 and 4 Diagrams illustrating some steps of the exemplary embodiment of the Figure 2 .
[0035] Several exemplary embodiments are explained in detail below. It should be noted that these embodiments serve only for illustration and are not to be interpreted as limiting. For example, embodiments with a large number of elements or components are presented. This should not be interpreted to mean that all of these elements or components are essential for the implementation of the embodiments. Rather, other embodiments may have fewer elements or components than those shown or described, and / or alternative elements or components. In still other embodiments, additional elements or components may be provided, such as elements used in conventional optical systems or computing devices. Components or elements from different embodiments may be combined with one another unless otherwise specified.Variations and modifications described for one of the embodiments can also be applied to other embodiments.
[0036] The present invention relates to image correction. For the purposes of this application, an image is understood to be a single image or an image of a video composed of a plurality of successive images. In the case of videos, such an image is also referred to as a frame. For the purposes of this application, a sequence of images refers to a series of images recorded one after the other. Such a series of images can, for example, be recorded or played back as a video, or as a series of single images recorded in quick succession.
[0037] Figure 1 shows a device according to an exemplary embodiment. The device of Figure 1The system comprises a camera device 10 coupled to an optical system 16. The optical system 16 can be, for example, a telescope, a binocular telescope, or a microscope to which the camera device 10 is coupled. In another embodiment, the optical system 16 can simply be a camera lens of the camera device 10. One or more objects to be recorded are imaged onto an image sensor of the camera device 10 via the optical system 16, for example, a CCD sensor (charge-coupled device) or a CMOS sensor (complementary metal-oxide semiconductor). As already explained, the camera device can include a 2D image sensor. However, images can also be acquired line by line with a 1D image sensor or point by point (0D image sensor) by scanning.The terms camera device and image sensor are to be understood broadly here and refer to any type of device with which images of one or more objects can be captured, and also include, for example, electron microscopes and laser scanning microscopes (LSM). The camera device 10 can thus capture a sequence of digital images via the optical system 16. The device of the . Figure 1The system further comprises a computing unit 11, which may, for example, have one or more processors 13 and associated memory 14, such as non-volatile memory (e.g., flash memory), random access memory (RAM), or read-only memory (ROM). In particular, program code may be stored in the memory 14, which enables the execution of the image corrections described in more detail below when the program code is executed on the processor 13 in the computing unit 11. The program code can be supplied to the computing unit 11, for example, via a data carrier (memory card, CD or DVD, etc.) or via a network.
[0038] In order to process the images received from the camera unit 10, the computer unit 11 can receive the images via an interface 12 over a connection 15. This connection can be established directly via a wired or wireless connection, for example, if the camera unit 10 and the computer unit 11 are located in close proximity, such as within a microscope system. However, it is also possible for the images to be written to a memory card or other storage media in the camera unit 10 and then transferred from this storage medium to the computer unit 11. Transmission via computer networks, such as the internet, possibly with one or more intermediate stations, is also possible.The processing of the images in the computing unit can take place immediately after the images are captured, particularly in real time, but also at a later time. In other embodiments, the computing unit 11 can be integrated into the camera unit 10, for example, within a video camera. Thus, the representation of the... Figure 1 no special spatial relationship between the camera device 10 with the optical system 16 on the one hand and the computing device 11 on the other.
[0039] The computing device can perform a series of processing steps on a recorded sequence of images, which are described below with reference to the Figure 2These steps can be implemented, as already explained, by means of appropriate programming of the processor 13. The processor 13 can be a general-purpose processor such as a CPU, but it can also be, for example, a special graphics processor, another type of processor, or a combination thereof. Special graphics processors are frequently used within cameras, for example, while general-purpose processors are often used for image processing on a computer. The implementation of the present invention is not limited to any particular arrangement in this respect. In principle, it is also possible to implement the image processing described below using special hardware components, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0040] Figure 2 shows a flowchart of a process according to an exemplary embodiment, which, for example, is explained above in the Figure 1 The discussed device, in particular the computing unit 11, can be implemented. First, an overview of the method is given. Then, the various process steps are explained in detail. Finally, various variations, modifications, and application possibilities of the method are discussed.
[0041] The process involves using a sequence of images. To process a specific image within this sequence, a certain number of neighboring images are used, for example, 15 images before and 15 images after the specific image. The specific image can also be referred to as the image to be restored. In this example, restoring a specific image would therefore require a block of 31 images, preferably taken consecutively. The number of 15 images before and after the image to be restored is merely an example, and any other number can be used, depending on the available images and computing power. The number of images used before and after the image to be restored does not have to be the same. The images are preferably taken in quick succession with a shorter aperture time.
[0042] In a sequence of images, various events can lead to changes in the image content over time, i.e., from image to image. Some of these events are undesirable and can be corrected by the method according to the invention.
[0043] The camera can perform translational and / or rotational movements, which are undesirable in the case of vibrations and the like. This movement leads to a corresponding movement of the image content. An entire recorded object can also move. Furthermore, as already explained, flickering can be caused by atmospheric turbulence. Finally, only parts of the recorded image may move, for example, when a moving object such as a car or truck is being filmed.
[0044] The procedure of Figure 2As will be explained first, this method is particularly suitable for black and white or grayscale images. Its application to color images will be explained in more detail later.
[0045] In step 20 of the Figure 2 The images in the sequence are divided into tiles, i.e., into sub-areas. In step 21, a transformation into the frequency domain takes place, for example, by a Fourier transform.
[0046] In step 22, shifts of image components in the spatial frequency domain, which may have resulted from air turbulence, for example, are identified and compensated for or removed. In step 23, an inverse transformation is then performed, for example, by an inverse Fourier transform.
[0047] Step 24 checks whether a desired number of calculations has been completed. If not, smaller tiles are used in step 25 (for example, the tile size is halved), and steps 21 to 23 are repeated with the smaller tiles. If so, step 26 averages the sequence of images for each tile, for example, to average out any remaining image-to-image shifts, and / or sharpens the images. In step 27, the tiles are then assembled into a finished, corrected image, which can then be further processed or viewed.
[0048] Now the individual steps of the procedure will be described. Figure 2 This will be explained in more detail. First, the subdivision into tiles from step 20 will be explained with reference to the Figures 3 and 4 explained. Preferably, the tiles have one or more of the following properties: The size of the tiles should be as small as possible compared to the largest areas of disturbance in the image; for example, in the case of compensating for air turbulence, they should be smaller than the largest cells of the air turbulence. Image shifts caused by camera shake and similar factors, such as unstable telescope tripod mounts, should be significantly smaller than the size of the tiles—for example, at most half, at most a third, or at most a quarter of the tile size. The tiles should be easily reassembled at the end of the process.
[0049] A tile size for step 20 could, for example, have an edge length of one-quarter or one-eighth of the image size. The tiles can be square. For a typical image with 2048 x 1024 pixels, a tile size of 256 x 256 pixels could be chosen, although other sizes are also possible. Preferably, the tiles overlap so that each pixel not located at the edge is present in four tiles. For illustration, [the following is shown]. Fig. 3 such a division. For example, the Fig. 3 Image 30 is divided into nine tiles (3 x 3) 31A-31J. Each tile comprises four of the small squares, with the reference symbol in the center of each. To illustrate this, tile 31A is shown in Fig. 3 Shown in bold.
[0050] As can be seen, in this division only the pixels in the corners belong to one tile, other edge pixels belong to two tiles each, and pixels in the small middle squares belong to four tiles each. Figure 3 This is for illustrative purposes only, and depending on the image size, the number of tiles may be different, especially larger, for example as in the pixel example above.
[0051] Preferably, the tiles are multiplied by a weighting function. This suppresses artifacts during the transformation in step 21. Several suitable window functions for this weighting are known individually. The weighting in two dimensions for the case of Fig. 3 is schematically in Fig. 4The three tiles are represented as follows: each is multiplied in every direction by a weighting function 40, which decreases towards the edge of the tile. In the overlapping areas, the weighting functions preferably behave such that their sum is constant, e.g., 1 (indicated by a line 41), so that the tiles can later be recombined into a complete image by simply adding the values. A suitable window function is the Hanning window function, although other window functions can also be used.
[0052] The tile decomposition is applied to each image in the sequence of images (for example, the 31 images in the numerical example above).
[0053] It should be noted that if the movement to be compensated is the same or occurs throughout the entire image, then only a single tile can be used, i.e., this single tile then encompasses the entire image.
[0054] In step 21, a transformation from the spatial domain to the spatial frequency domain is performed on each of the tiles, whereby this transformation is a two-dimensional transformation corresponding to the two-dimensional image. In a preferred embodiment, a Fourier transform, in particular a fast Fourier transform (FFT), is used. However, other types of transformation can also be used, for example, other spectral transformations, on which the exact calculation is then performed in the subsequent step 22. For example, a wavelet transform can be used. Each image in spatial space then corresponds to a corresponding spectral density in spatial frequency space.
[0055] This transformation and / or the calculations described below, which can be relatively time-consuming depending on the image size, can be performed in parallel in some embodiments. For example, the calculation for each tile or for groups of tiles can be performed in separate computing units, e.g., separate kernels. In other embodiments, the transformation can also be performed in parallel for different rows and / or columns of the image. For example, a fast 2D Fourier transform (FFT) is performed in two sections, e.g., first in a first direction (x-direction, row-wise) and then in a second direction (y-direction, column-wise). In each of these two sections, such an FFT can be performed for the different rows or columns in different kernels or otherwise with different computing units, e.g.,with 1024 lines in 1024 kernels.
[0056] In step 22, shifts in the images—that is, shifts in an object's position from image to image, which can be caused, for example, by camera shake or air turbulence—are identified and corrected. The occurrence of air turbulence or camera shake leads, in particular, to a shift in each of the affected tiles according to the image content shift theorem. As explained above regarding tiling, it is advantageous for the tile size to be smaller than the corresponding air turbulence cells. In the frequency domain in which the tiles exist after the transformation of step 21, this shift manifests itself, according to the shift theorem, as a frequency-proportional change in the spectral density argument.
[0057] The displacement theorem for a displacement by x 0 in one dimension means: where represents the spectral transformation (for example, Fourier transformation), u ( x ) describes the image in the local area and U ( f ) describes the corresponding transform of the image, i.e., the spectral density, in spatial frequency space. exp(.) is the exponential function. As can be seen, the shift in spatial space by x 0 a change in the spectral density argument in the spatial frequency domain. The change in the spectral density argument corresponds to a shift. x 0 proportional to the frequency f. The same applies in two dimensions to two-dimensional images or two-dimensional tiles.
[0058] The possible changes within the sequence of images over time, explained above with reference to Figure 2, can be detected by corresponding changes in spectral density, in particular the spectral density argument.
[0059] Thus, a translational movement of the camera setup (for example, vibrations or the like) or a movement of the entire recorded object over time leads to a change in the spectral density argument from frame to frame over time, with this change being proportional to the frequency, as explained above. The factor by which the spectral density argument changes proportionally to the frequency, that is, for example, x The 0 in the equation above can be determined particularly easily and robustly at low spatial frequencies f.
[0060] Flicker caused by air striations can manifest in various ways depending on the tile size. With relatively large tiles compared to the size of the interfering air cells, flicker appears as stationary fluctuations in spectral density over time. In this case, depending on the design of the corresponding device and the dynamic conditions in the captured images, a time-averaged value can be used for compensation at each spatial frequency. This could be, for example, an arithmetically averaged value of the spectral density over the image sequence or a portion thereof, or an averaged argument and an averaged magnitude of the spectral density. This average value is then used in the corrected image. It has been found that using an averaged magnitude yields particularly good results, even though an arithmetic mean of the entire spectral density is probably mathematically more accurate.
[0061] For tiles that are small relative to the air cell size, flickering manifests as a steady-state displacement over time according to the displacement theorem explained above. However, as explained above, tiles that are large compared to the air cell size are preferred for correction.
[0062] Even when a part of an object moves, for example, the movement of a car or truck in the image, a distinction must be made between small and large tiles. For small tiles that lie entirely on the moving part of the image, the movement manifests itself as a factor with a frequency-proportional argument according to the displacement theorem, where the frequency-proportional argument, in particular, x 0, which changes linearly over time.
[0063] The situation is more complex with tiles that only partially overlap the moving object. Such tiles are therefore less suitable for detecting moving image parts.
[0064] The various image changes discussed above, particularly disturbances caused by camera movement or flicker caused by air currents, manifest themselves, at least in part, as a frequency-proportional change in the spectral density argument. The shift x 0 acts as a proportionality factor, as can be seen from the formula for the displacement theorem above.
[0065] To xTo determine 0 or other parameters that describe the frequency-proportional change due to shift, the spectral density associated with a specific number of images can be determined by averaging. The aforementioned parameters, which describe the frequency-proportional changes due to shifts, can then be determined relatively easily for each tile at low frequencies, relative to the average value. The parameters determined in this way at low frequencies are at least approximately valid for higher frequencies as well and can therefore be used for all frequencies.
[0066] The determination of the parameters, in particular of x Determining the magnitude of the spectral density at low frequencies is relatively simple because the magnitudes of the spectral density are comparatively high at low spatial frequencies. Therefore, the shift, for example, can be determined relatively robustly here. xDetermine 0. It should also be taken into account that the Shannon sampling theorem is violated in typical video recordings. This leads to errors, at least at high spatial frequencies and possibly even at medium spatial frequencies. Is the value x Once 0 has been determined for a tile or an image, this value is also valid for the spectral density at higher spatial frequencies and can also be used as a shift in the spatial domain.
[0067] Based on the parameters determined in this way, the shift-related, frequency-proportional component is calculated for each tile from the spectral density in the spatial frequency domain, relative to the mean value. This essentially corresponds to a compensation or equalization of the shift. In particular, by comparing the parameters of spatially and / or temporally adjacent tiles, it is possible to identify any image disturbances (flickering, shaking, etc.) that may be present, and such disturbances can be corrected. This shift-related, frequency-proportional component can also be saved for later use. After this calculation, the image content of each tile no longer shifts from image to image. The magnitude of the spectral density remains unchanged, as the change is reflected in the argument, as explained above.In another embodiment, the spectral density in step 22 can be averaged over all images in the sequence to reduce noise and suppress disturbances such as flicker caused by air streaks. In any case, after step 22, there is at least a reduced, if not eliminated, shift from image to image in the spectral domain.
[0068] Now, various examples of image correction, especially compensation for shifts, will be explained.
[0069] In some embodiments, averaging over time, for example by averaging the argument and the amount separately, is sufficient for relatively large tiles to compensate for translational movements of the camera setup and / or the entire recorded image and / or flickering caused by air streaks.
[0070] If, for example, only a single tile is used, as mentioned above, such a separate averaging of argument and magnitude can be performed at any desired frequency in the spatial frequency domain. In other embodiments, a plane fit with respect to the change in the argument according to the shift theorem from frame to frame in the spatial frequency domain can be performed to compensate for the interfering motion and, if necessary, to determine a motion to be separated (i.e., a motion that actually exists, e.g., due to the movement of a recorded object). This can be used particularly with 2D image sensors. A motion to be separated might be present, for example, when video recordings are made from an aircraft. The motion to be separated is then caused by the movement of the aircraft. The same applies to other vehicles. Such motion can also occur in other applications, e.g., with a surgical microscope.to a movement of a surgeon with his instruments from a movement of a microscope head over the resting patient (which is to be compensated for, for example).
[0071] If movement occurs in only some parts of the image, multiple tiles are preferably used. This allows image areas with moving objects to be treated differently than image areas without such objects. This preserves the movement of objects as much as possible while compensating for unwanted movements. The number and size of the tiles can be dynamically adjusted, for example, by using pattern recognition to identify where and in which tiles moving objects occur.
[0072] To account for movement of a portion of the captured image (for example, a car or truck), some implementations allow tiles where such an event is detected—for instance, based on a relatively large movement in the respective image area—to be left unchanged (unrestored). To account for such effects even more precisely, tiles can be used that are small in relation to the size of air turbulence. In this case, spatially and temporally adjacent tiles are used to detect movements of only parts of the image and subsequently perform compensation if necessary.
[0073] If it is detected that flickering occurs in a tile only due to air streaks, this can be compensated as described above, i.e. by averaging and / or by determining and subtracting the respective shift.
[0074] If a flicker is detected in a tile and a moving object is also fully contained within the tile, the two can be separated by examining the dependency over time, since flicker causes fluctuations, while, for example, the displacement in the case of movement increases linearly. This separation then allows only the fluctuation over time caused by flicker to be corrected.
[0075] If it is detected that a moving object in a tile is only partially present in the area of that tile, in addition to flickering, the time-shift caused by flickering in neighboring tiles can be determined by analyzing them. This shift can then be interpolated for the current tile and subsequently used for compensation. For example, the shift from neighboring tiles can be averaged. Compensation is again based on the shift theorem with a frequency-proportional factor.
[0076] In step 23, a reverse transformation corresponding to the transformation of step 21 is performed. If the transformation in step 21 was, for example, a Fourier transform, in particular a fast Fourier transform, then in step 23 an inverse Fourier transform is performed, in particular a fast inverse Fourier transform (IFFT).
[0077] These steps 20-23 can be repeated with progressively smaller tiles until the desired number of passes is reached. A single pass as described above may be sufficient, for example, to compensate for camera shake caused by vibrations of a tripod. If finer corrections are required, especially for minor air bubbles, more than one pass can be performed. For this, the tiles are reduced in size in step 25. For example, the side length of the tiles can be halved. The rules for tiling described above (see the description of step 20) should preferably be followed. In particular, these rules must be adhered to in both the preceding passes and the current pass, since a complete image will ultimately be reassembled.
[0078] For the example just mentioned of halving the side length, this means, for example, in some embodiments, that the multiplication with the Hanning window as described above... Fig. 4 The resizing of the larger tiles is not performed at their outer edges, but only towards the center, since the outer edges of the larger tiles were already weighted accordingly in a previous step by multiplication using the Hanning window. Steps 21-23 are then repeated with the reduced-size tiles. This resizing of the tiles can be performed several times, depending on the computing power, so that steps 21-23 are carried out with increasingly smaller tiles. The following points are given priority when resizing the tiles: If the image content within a tile shifts too far, for example, by more than approximately a quarter of the tile's side length, artifacts can occur because the image content then changes significantly, rather than just undergoing a comparatively small shift. Therefore, step 20 begins with sufficiently large tiles and then gradually reduces their size, compensating for correspondingly large shifts in each step, which are then no longer present in subsequent steps. In the calculation of step 22, as described above, only the translational shift of the image content within a tile is considered for restoration, using the shift theorem as explained above. Flicker caused by air turbulence cells that are smaller than the tile size is not taken into account. To account for this, a further pass with correspondingly smaller tiles must be performed.This is because a phase-modulated sine / cosine function in the spatial domain has a spectral density in the form of a Bessel function, i.e., a spectral density that, in addition to a fundamental frequency, exhibits a broad spectrum due to the phase modulation. This means that, for example, flickering of small cells (i.e., small compared to the tile size) caused by air turbulence manifests as a relatively complex spectral density, which is difficult to handle. By incrementally reducing the size of the tiles, such flickering, as described above for step 22, can be handled relatively easily when the tile size becomes smaller than a region of flickering or other shifts. Thus, the number of iterations and the degree of reduction performed can be adjusted depending on the available computation time and how small air turbulence cells or other effects are to be accounted for.
[0079] The embodiments described above can be used for various applications. For example, videos recorded with telescopes or other optical systems can be restored, particularly in terrestrial applications such as nature photography. However, applications in astronomy are also possible. By applying the described methods, the effective resolution can be increased, for example, doubled, in some embodiments. This can be used, for instance, in electronic zoom systems that enlarge small sections of an image.
[0080] It is important to note that when reducing the tile size from level to level, adjacent pixel values from different levels are not added together, i.e., from one iteration to the next. This distinguishes the approach presented here from pyramid algorithms, where pixel values from images are summed or averaged to ascend from level to level of the pyramid.
[0081] In some embodiments, the reduction in tile size in step 25 can be fixed from one iteration to the next. For example, the tile size can be halved in each iteration of step 25. In other embodiments, different reductions can be used. For example, a factor by which the side lengths of the tiles are divided can vary from iteration to iteration. The number of iterations and the tile reduction can be dynamic, for example, based on the current surface area and the amplitude of displacements caused, for example, by air currents. For example, the size adjustment can be based on displacements detected in step 22 during previous image sequences, for example, within the last second. The tile reduction can also vary across the image, i.e.,Not all tiles need to be resized to the same size.
[0082] For example, image areas containing objects close to the optical system may vary differently than other image areas containing more distant objects or objects above relatively warm objects, which generate more flicker. More distant objects exhibit more pronounced flicker than closer ones. Thus, for instance, in an image area with nearby objects, the tile size can be reduced more rapidly from pass to pass than in image areas where air turbulence plays a more significant role. In some image areas, certain reduction steps can also be omitted. The precise adjustment may depend on the specific application.
[0083] Once a desired number of iterations has been completed, averaging and / or sharpening is performed in step 26. For averaging, the spectral density argument of each tile is averaged over a certain number of images, for example, the aforementioned sequence of images. A corresponding averaging can also be performed for the magnitude of the spectral density. In this way, any remaining shifts can be averaged out. Instead of separately averaging the argument and magnitude, other approaches to averaging are also possible, for example, averaging the complex spectral density values in the Cartesian complex coordinate system. Similar to step 22, this averaging thus aims to compensate for shifts—in this case, remaining shifts that are not compensated for by step 22 in the iterations.
[0084] By averaging over time in step 26 and / or even earlier in step 22, noise, such as photon noise from the camera system, can be reduced. Averaging multiple images in the spatial frequency domain for this purpose is particularly useful when there is translational movement of the camera system or the entire image, or movement within a portion of the image. As explained above, such movements in the spatial frequency domain are relatively easy to detect. These movements can then be excluded from the averaging process, as otherwise, averaging moving objects could produce undesirable artifacts. In particular, movement in the spatial frequency domain can be easily described over time, for example, linearly or through a simple fit, and compensated for during the averaging process.Therefore, averaging over time is relatively easy in the spatial frequency domain, even for moving images. Such a method would be difficult in the time domain.
[0085] Furthermore, filtering, such as Wiener filtering, can be performed, which, during averaging, gives greater weight to values of the complex spectral density that exceed a predefined threshold, for example, a noise level. In this way, the influence of noise can be reduced more efficiently.
[0086] To compensate for loss of sharpness due to air schlieren interference, higher-frequency spectral density components can be amplified during sharpening, for example, components above a certain threshold frequency. This approach is known from the speckle imaging algorithm, for example, from D. Korff, Journal of the Optical Society of America, Vol. 63, No. 8, August 1973, pp. 971ff. According to this conventional approach, this amplification can be performed taking into account the shorter temporal aperture of an image, for example, 33 ms. It is helpful that it is not necessary to determine the gain factors for amplifying the higher-frequency spectral density components precisely. Depending on the properties of the air turbulence cells or air schlieren present, the amplification can vary from one tile to another, and can also be dynamic, i.e.,The settings can be adjusted to change over time from image sequence to image sequence.
[0087] Basically, any conventional sharpening algorithm can be used. The sharpening can be adjusted for each tile depending on the specific application, for example, each of the tiles in... Figure 3 The tile examples illustrate how sharpening can be performed differently from tile to tile and dynamically over time. Sharpening can be achieved, for example, by conventionally convolution over the image or tile in the spatial domain, or by amplifying the magnitudes in the frequency domain. In other words, a prior transformation can be used for sharpening, or the sharpening can be performed in the spatial domain.
[0088] It should be noted that blurriness can also be caused by atmospheric turbulence, and the blurriness increases with the severity of the turbulence. Therefore, in some implementation examples, sharpening can be performed depending on the presence of atmospheric turbulence.
[0089] In a simplified approach, the same sharpening can be applied to the entire field. However, in a preferred embodiment, the flicker is determined for each tile as explained above (for example, by determining the parameter). x 0 and its analysis over time), and the sharpening can be performed depending on the specific flicker. For example, stronger sharpening can be applied to stronger flicker, since higher flicker is generally accompanied by greater blurriness.
[0090] In step 27, the tiles are then reassembled into a single image. If the criteria mentioned above are taken into account during tiling, this can be done by simply summing the tiles. Otherwise—for example, if a suitable window function is not used—values in the overlapping area of tiles may need to be averaged. Other approaches are also possible.
[0091] Compared to the conventional speckle algorithm, the calculation in step 22 is simplified by using the spectral density argument and its frequency-proportional change, thus enabling faster computation. Furthermore, the conventional speckle imaging algorithm uses only a single tile size, and not multiple passes as is the case in some embodiments of the present application. With reference to Figure 2The described procedure can be varied and extended in various ways. Some of these variations and extensions are explained in more detail below.
[0092] The procedure of Figure 2 As mentioned, this can initially be applied to black and white or grayscale videos. However, it can be extended to color videos or other colored sequences of images. In one embodiment, this involves dividing the video into tiles and performing the transformation or inverse transformation (steps 20, 21, 23, and optionally 25 of the process). Figure 2The steps are performed separately for different color channels. For example, in a sequence of RGB (red, green, blue) images, these steps are performed separately for the red, green, and blue channels. Step 22 and the averaging in step 26 can then be performed on grayscale values calculated based on the red, green, and blue components. For example, in step 22, a grayscale spectral density is calculated by averaging the spectral densities for the red, green, and blue channels. Based on this grayscale spectral density, the necessary changes to the spectral density argument are then determined as described above. These changes can then be applied to the spectral densities for the red, green, and blue channels.
[0093] The averaging can also be performed as a weighted averaging. For example, it can be taken into account that the human eye is more sensitive to green light than to red or blue light, so the green channel can be weighted more heavily than the red or blue channel. A similar approach can be used for averaging and / or sharpening step 26.
[0094] Furthermore, in some embodiments, the method of Figure 2 This can be expanded to allow for the differentiation of "real" movements from unwanted movements (for example, those caused by the aforementioned air streaks). Such "real" movements could be objects moving through the recorded images, such as a moving vehicle or a person.
[0095] In some implementations, a moving object can be detected by observing that the spectral density argument varies more in one tile than in neighboring tiles during the same time interval, for example, by more than a predefined threshold. For instance, the spectral density argument in one tile might vary by + / - 90 degrees in a given time interval, while in neighboring tiles, the argument might only vary by + / - 30 degrees. This suggests that the tile in question is moving, whereas the change in the argument in neighboring tiles could be caused by atmospheric turbulence or similar phenomena.
[0096] If it is determined that a moving object is present in a particular tile, several approaches are possible. In some embodiments, such a tile can simply be displayed without restoration (i.e., without performing, for example, step 22 and / or step 26 for that tile). In other embodiments, the components of the changes, for example, in the spectral density argument for the moving object and for the flicker, can be separated to compensate for only the flicker component without distorting the moving object component. In some embodiments, this can be done based on the size of the spatial change. For example, the area of large-scale flicker caused by large-scale air streaks will often be larger than that caused by small, moving objects.If necessary, this can be done based on the tile sizes in the different passes.
[0097] In other examples, a distinction is made based on the direction of movement. A moving object often moves uniformly or at least always in the same direction over a period of time, whereas, for example, shimmering caused by air currents has a fluctuating character, so that the displacements can change direction. Other properties that distinguish disturbances caused by air currents or other unwanted movements from normally moving objects can also be used as a basis for this distinction.
[0098] As already explained, the above-described examples can be applied to a sequence of images, for example, frames from a video. In this way, disturbances such as air bubbles or vibrations can be compensated for. However, it is also possible to output only a single image as a result, as is common with conventional single-frame cameras. In this case, a camera device such as the one described in Figure 1The camera setup shown captures a sequence of images, including the final output image as well as a series of images before and after it, in order to perform the procedure described above. Such images can be captured at a high frame rate, for example, 300 frames per second. In this way, distortion or shifts caused by flicker can be compensated for in the final output image as described above, optionally taking moving objects into account as described. Sharpening can also be performed (step 26 in Figure 2 ), which is adapted to detected disturbances, such as air striations.
[0099] As already mentioned, the procedure of Figure 2It can also be used for image stabilization, i.e., to compensate for vibrations or camera shake. In this case, it may be sufficient to use only one tile size (i.e., one pass of steps 21 to 23 of the Figure 2 (to use). In this application, typically only the translational displacement, i.e., a displacement in a specific direction, is compensated. For this purpose, an analysis of the displacement (i.e., the change in the spectral density argument after the transformation at 21) over time can be performed as described, distinguishing displacements due to vibrations from object movements (as described above). Thus, in some implementations, the displacement can be compensated very accurately, sometimes with sub-pixel precision.
[0100] A multiple of the displacement can also be output as part of image stabilization. This is demonstrated using the example of an operating microscope as an optical system (for example, optical system 16 of the Figure 1This is explained below. Let's assume the operating microscope has ten times magnification and the image is displayed on a monitor via an eyepiece directly on the microscope head. In such a case, considering the movement of the microscope head on an arm that can be several meters long, it may be advantageous not to fully stabilize the image on the monitor, but rather to move it in the opposite direction to the movement of the microscope head. This ensures a stable image for the observer, who remains stationary relative to the patient and the floor. For example, depending on the optical design, the movement of the image on the monitor could be -1 / 10 or +1 / 10 times the movement captured by the camera in the microscope head. In this case, the entire displacement is not compensated for, but rather only a portion or a multiple thereof.In this example, this allows the movement of the object viewed by the operating microscope, caused by the surgeon's actions, to be reproduced flawlessly, and also enables the creation of a corresponding movement on the monitor image. This allows the surgeon to be presented with an intuitively understandable image. A slight rotation or change in magnification caused by the movement of the operating microscope head can also be stabilized, at least approximately, using appropriately small tiles, with each small tile receiving a slightly different translational shift compared to the other tiles.
[0101] The above example of an operating microscope serves only to further illustrate, in particular, the incomplete compensation of the displacement and is not to be interpreted as restrictive, since the illustrated embodiments generally refer to camera devices with optical systems as described with reference to Figure 1 can be described and applied.
Claims
1. Computer-implemented method for image processing, comprising: providing a sequence of images, subdividing the images into one or more tiles, transforming the tiles into the spatial frequency domain, identifying shifts by evaluating the spectral density function argument of the tiles in the spatial frequency domain by determining a factor by which the spectral density function argument changes, proportionally to the frequency, from image to image; compensating at least some of the identified shifts, transforming the tiles back into the spatial domain, and assembling the inverse-transformed tiles, wherein identification is based on frequencies below a frequency threshold value.
2. Method according to Claim 1, further comprising a focusing over tiles corresponding to one another in the sequence of images prior to the assembly.
3. Method according to Claim 1, further comprising an averaging over tiles corresponding to one another in the sequence of images prior to the assembly.
4. Method according to Claim 3, further comprising a focusing over tiles corresponding to one another in the sequence of images prior to the assembly and / or after the averaging.
5. Method according to any of Claims 1-4, wherein a tile size when subdividing the images into tiles is set depending on the size of air schlieren during the image recording and / or depending on a strength of blurring during the image recording.
6. Method according to any of Claims 1-5, wherein subdividing, transforming, identifying, compensating and inverse transforming is repeated N times, N > 1, wherein a size of the tiles is incrementally reduced during the subdivision.
7. Method according to Claim 6, wherein a factor by which the tiles are reduced in size differs between at least two of the repetitions.
8. Method according to any of Claims 1-7, wherein the images are colour images, wherein subdividing, transforming and inverse transforming are carried out separately for each colour channel, and / or wherein shifts are identified on the basis of greyscale values obtained by averaging the colour channels, wherein the compensation for the colour channels is carried out separately on the basis of the results of the identification.
9. Method according to any of Claims 1-8, wherein the identifying includes distinguishing between shifts from interferences and shifts by moving objects, wherein the shifts by moving objects are excluded during the compensation.
10. Method according to any of Claims 1-9, wherein one or more method steps of the method are adapted dynamically depending on the tiles and / or a content of the images.
11. Method according to any of Claims 1-10, wherein one or more method steps are carried out in parallelized fashion.
12. Apparatus for image processing, comprising: a computing device having at least one processor (13) and a memory (14), wherein a sequence of images is storable in the memory (14), wherein the processor is configured to carry out the following steps: subdividing the images into one or more tiles, transforming the tiles into the spatial frequency domain, identifying shifts by evaluating the spectral density function argument of the tiles in the spatial frequency domain by determining a factor by which the spectral density function argument changes, proportionally to the frequency, from image to image; compensating at least some of the identified shifts, transforming the tiles back into the spatial domain, and assembling the inverse-transformed tiles, wherein identification is based on frequencies below a frequency threshold value.
13. Apparatus according to Claim 12, wherein the apparatus is configured to carry out the method according to any of Claims 2-11.
14. Apparatus according to Claim 12 or 13, further comprising an optical system (16) and a camera device (10), coupled to the optical system (16), for recording the sequence of images.
15. Computer program having program code which, when executed on a computing device (11), carries out the method according to any of Claims 1-11.