System and method for fusing two or more source images into a target image
By adjusting the contribution of low-frequency and high-frequency components according to the local noise level of the main image in image fusion, the target image is generated, artifact and noise problems are solved, and image quality and adaptability are improved.
Patent Information
- Application Number
- CN202410178535.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-02-09
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-02-09
AI Technical Summary
Existing image fusion technology is prone to artifacts, affecting the accuracy of human viewing experience and machine perception systems, and is not effective in uneven lighting scenarios.
By determining the local noise level of the main image, the low-frequency and high-frequency components are derived, and the contribution of the high-frequency components is gradually increased according to the local noise level. Combined with the low-frequency components, the target image is generated, the threshold-type artifacts are avoided, and the noise content is controlled.
Reduce the appearance of artifacts, maintain the main image characteristics, reduce noise levels, improve image quality, and adapt to uneven lighting scenes.
Smart Images

Figure CN118505525B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing. Specifically, it proposes a method for fusing a primary image and a secondary image. The method can be applied to a visible light image to be combined with an infrared light image. Background Art
[0002] Image fusion is known to include the process of combining information from multiple source images into fewer target images (usually a single target image). The target image is expected to be more informative, more accurate, or of higher quality than each of the individual source images, and it is composed of all the necessary information. Image fusion is practiced for various purposes, one of which is to construct images that are more appropriate and more understandable to human viewers, or images that are more useful for machine perception. The source images can differ in depth of focus (multifocal image fusion), wavelength range used for photographic imaging (multispectral image fusion), and / or imaging modality (multimodal image fusion). Imaging modalities in this sense include imaging by reflection (e.g., conventional visual photography), penetration (e.g., X-ray imaging), and emission (e.g., magnetic resonance, ultrasound).
[0003] The combining operation at the heart of image fusion can be performed in either the spatial or transform domain. Broadly speaking, spatial domain image fusion can be described as stitching together pixels or groups of pixels at their respective locations. If image fusion is performed in the transform domain, the combining action is applied to different transformed components computed from the source images. For example, image fusion in the frequency domain (spatial frequency domain) involves combining the spectral elements of the images.
[0004] According to EP2570988A2 and similar disclosures, it is known to apply image fusion methods to combine visual images with infrared images. More specifically, EP2570988A2 proposes a method that receives a first image of a scene, wherein the first image is a visual image captured using a visual image sensor; receives a second image of the scene, wherein the second image is an infrared image captured using an infrared image sensor; extracts high spatial frequency content from the visual image, wherein the high spatial frequency content includes edges or contours of the visual image; and combines the high spatial frequency content extracted from the visual image with the infrared image to increase the level of detail in the infrared image. The method according to EP2570988A2 produces an infrared target image.
[0005] WO2018120936A1 discloses an image fusion method, comprising: obtaining a visible light image and an infrared image of the same scene; performing a first decomposition on the visible light image to obtain a first high-frequency component of the visible light image and a first low-frequency component of the visible light image; performing a first decomposition on the infrared image to obtain a first high-frequency component of the infrared image and a second low-frequency component of the infrared image; fusing the first high-frequency component of the visible light image and the first high-frequency component of the infrared image based on a first algorithm to generate a first fused high-frequency component; and reconstructing based on the first fused high-frequency component, the first low-frequency component of the visible light image, and the first low-frequency component of the infrared image to generate a fused image. Optionally, a pair of low-frequency components corresponding to the visible light image and the infrared image can be fused to generate a fused low-frequency component. The low-frequency components of the visible light image and the infrared image can be assigned weighting factors, which are set based on the brightness of the surrounding environment, the color contrast of the target scene, user preferences, etc.
[0006] EP3518179A1 discloses an image acquisition device for spectroscopic fusion, comprising: a spectrometer, a visible spectrum imaging module, a non-visible spectrum imaging module, a registration unit, a pre-processing and synthesis unit, and a fusion unit; wherein the spectrometer is configured to separate incident light into visible light and non-visible light; the visible spectrum imaging module is configured to perform photosensitive imaging based on the visible light separated by the spectrometer to form a first visible light image; the non-visible spectrum imaging module is configured to perform photosensitive imaging based on the non-visible light separated by the spectrometer to form a first non-visible light image; the registration unit is configured to perform positional registration on the first visible light image and the first non-visible light image to obtain a target visible light image and a second non-visible light image; the pre-processing and synthesis unit is configured to perform brightness adjustment on the second non-visible light image to obtain a target non-visible light image; and the fusion unit is configured to perform image fusion on the target visible light image and the target non-visible light image to obtain a fused image. Optionally, the fusion unit may be configured to: perform color space conversion on the target visible light image to obtain a luminance component and a color component of the target visible light image; perform low-pass filtering on the luminance component to obtain low-frequency information of the target visible light image; perform high-pass filtering on the target non-visible light image to obtain high-frequency information of the target non-visible light image; and perform weighting processing on the low-frequency information and the high-frequency information according to corresponding second-type weight values to obtain a luminance component corresponding to the fused image to be formed. The second-type weight value may be determined based on a luminance difference between the luminance components of the target non-visible light image and the target visible light image.
[0007] WO2013131929A1 discloses a method comprising the following steps: spatially interpolating frames of an input video frame sequence, thereby generating high-resolution, low-frequency (HRLF) spatial and temporal bands; performing cross-frame spatial high-frequency extrapolation of video frames of an input data sequence, thereby generating high-resolution, high-frequency (HRHF) spatial bands; and fusing the HRLF spatial and temporal bands and the HRHF spatial bands to obtain a spatiotemporal super-resolution video sequence.
[0008] WO2021184027A1 discusses a technique for fusing two images. Assume that a first image and a second image are captured simultaneously in a scene. In an example, the first image and the second image include an RGB image and a near-infrared (NIR) image. After the images are captured, they are processed and fused into a new, compact form of the image—the fused image—which contains the details of both images. The fused image is then tuned toward the color of the first image (e.g., the original input RGB image) while retaining the details of the fused image. An optional color correction operation can be applied, wherein, for example, the fused image is decomposed into a fused base component and a fused detail component using a first guided image filter. For example, the first image itself is decomposed into a first base component and a first detail component using a second guided image filter. In some embodiments, the fused base component and the fused detail component include low-frequency information and high-frequency information of the fused image, and the first base component and the first detail component include low-frequency information and high-frequency information of the first image. The first base component of the first image and the fused detail component of the fused image are combined into a final image, thereby maintaining the base color of the first image while having the details of the fused image.
[0009] It is also known that practical implementations of image fusion produce a certain amount of artifacts that may be noticeable to human viewers. As used herein, in the context of image fusion, an artifact (or visual artifact) is a visual feature in the target image that is not present in either of the source images. An artifact can be structural—it appears to add new lines, shapes, surfaces, or other graphical information to the image—or it may change the chromaticity, brightness, contrast, or other technical characteristics of the image or a portion thereof. Artifacts can be categorized as intra-image artifacts that can be noticed within the target image and inter-image artifacts that include gradual and sudden changes in the overall brightness or overall chromaticity of the image. While such changes may be minor in a single target image, they can be easily noticed during playback of a video sequence or when multiple images are rendered consecutively. Beyond these primarily aesthetic issues, it is conceivable that a fused image or video sequence with noticeable artifacts could cause machine perception systems to make erroneous decisions. For example, a video surveillance system might trigger an alarm based on the erroneous conclusion that a moving object is an intruder.
[0010] For these reasons, it is desirable to provide image fusion techniques that are less prone to artifacts. Summary of the Invention
[0011] It is an object of the present disclosure to provide a method for fusing two or more source images into a target image such that the target image can be expected to contain fewer artifacts than comparable prior art methods. A further object is to provide such a method having the property that the target image can be expected to be free of artifacts that are severe in the sense of being perceptible to a non-expert viewer and / or have the potential to mislead a machine perception system. A further object of the present disclosure is to provide a method for fusing a noisy source image with a further source image into a target image having controlled noise content. A further object is to provide a method by which a source image of an unevenly illuminated scene can be improved by fusing with a secondary source image. A further object is to provide such a method in a minimally destructive manner to the source image, in particular in a method that adds as little data as possible from the secondary source image. A still further object is to provide an apparatus and software program having these capabilities.
[0012] In a first aspect of the present disclosure, a method for fusing a primary image and at least one secondary image is provided. Generally speaking, a primary image can be identified by being noisier than the secondary image in a well-defined sense. For example, the local signal-to-noise ratio (SNR) of the primary image may be lower than the SNR of the secondary image across most of the image area, or the global SNR metric (e.g., mean, median, q-quantile) of the primary image may be lower than the global SNR metric of the same secondary image. Alternatively, the primary image can be identified by being associated with a richer visual representation than the secondary image. The representation may be richer in terms of the number of channels; for example, the primary image may be a color image expressed in a color space with multiple color channels (e.g., RGB, YCbCr, HSV), while the secondary image is a monochrome image (e.g., an infrared or near-infrared image, or a night mode image). Because the signal energy must be divided among the multiple channels in this type of primary image, a lower SNR can be expected. Further alternatively, the primary image can be identified as a user's preference over the secondary image. That is, by providing the special image to the image fusion method as a primary image, the user expresses a desire to preserve the characteristics of the special image to the greatest extent possible and to avoid introducing image data from the secondary image as much as possible.
[0013] The method according to the first aspect comprises: determining a local noise level of a region P(m,n) of a host image; deriving a low frequency (LF) component P from said region of the host image; LF (m,n), and derive the high frequency (HF) component S from the corresponding region S(m,n) of the secondary image HF(m, n); combine the LF component and the HF component into a target image region T(m, n), wherein the relative contribution of the HF component to the target image region gradually increases with the local noise level; repeat the above operation, and merge all the output target image regions thus obtained into a target image. It should be understood that the LF component and the HF component are generally referenced to a common cutoff frequency f c , and this cutoff frequency is applied with any suitable steepness; generally, a certain amount of roll-off is included to avoid certain artifacts, and the roll-off may depend on the local noise level or other characteristics of the image region, and also take into account the details of the use case at hand. For the avoidance of doubt, the cutoff frequency f c Refers to the spatial frequency of the image data, rather than the wavelength of the photographic light used to acquire the primary or secondary image. A region of the primary image is a pixel or group of pixels, and the corresponding region of the secondary image is identically located in terms of image coordinates.
[0014] An advantage associated with the first aspect of the present disclosure is that the relative contribution of the HF component to the target image region gradually increases with the local noise level, thereby avoiding thresholding artifacts (or quantization artifacts) to a certain extent and preserving the characteristics of the host image to the greatest extent possible. Because the contribution from the HF component is added according to the local noise level of the host image, the noise content of the target image can be controlled. In the language of the present disclosure, an increase relative to the local noise level is not gradual if the increase consists of a single step, i.e., a discontinuous or binary increase. Conversely, if the HF component continuously depends on the local noise level, the contribution from the HF component is said to increase gradually. This specifies the fact that strict mathematical continuity is impossible in the context of digital signal processing, where all variables must be represented in quantized form: no matter how smooth a function is, it cannot change in steps finer than the numerical resolution (floating-point precision) of the digital processing system—it can be described as quasi-continuous—yet for the purposes of the present disclosure it should be considered continuous.
[0015] The method according to the first aspect can be embodied in various ways. Three main groups of embodiments can be identified, which will be described in separate sections of this disclosure. Technical features can be combined across these embodiment groups, and in particular, different technical features belonging to different embodiment groups can be applied to different areas of a single master image.
[0016] In a first set of embodiments, the cutoff frequency is variable across the primary image. Embodiments in this set may provide that the image fusion method includes a further step of determining the cutoff frequency based on the local noise level of the region of the primary image. This determination may be repeated for one or more further regions of the primary image, so that different cutoff frequency values are applied when deriving the respective LF and HF components.
[0017] Specifically, a non-increasing function of the local noise level is used to determine the cutoff frequency. The noisier the region of the primary image, the larger the portion of the spectrum of the target image region originating from the secondary image will be. Conversely, for primary image regions with relatively low noise content, the characteristics of the primary image are preserved to a relatively large extent. The non-increasing function of the local noise level is preferably a continuous function consisting of multiple small increments (as previously described, this includes the case of quasi-continuous functions in digital signal processing). In this way, the relative contribution of the HF component to the target image region gradually increases with the local noise level, thereby avoiding threshold-type artifacts.
[0018] In a second set of embodiments of the first aspect, according to one or more weight coefficients that are variable across the main image to combine the LF and HF components (and possibly further components). Embodiments in this group may include the further step of determining weight coefficients based on the local noise level of said region of the main image. This determination may be repeated for one or more further regions of the main image so as to apply different weight coefficient values when combining the respective LF and HF components (and any further components).
[0019] Specifically, in some embodiments within the second group, the secondary coefficients to be applied to the components of the corresponding region of the secondary image are is a non-decreasing function of the local noise level. The noisier the area of the primary image, the more image data in the target image area will originate from the secondary image; conversely, for primary image areas with relatively low noise content, the characteristics of the primary image are preserved to a relatively large extent. Alternatively or additionally, the primary coefficients to be applied to the components of the area of the primary image are is a non-increasing function of the local noise level. Using primary coefficients with these properties has the corresponding effect of balancing the contributions from the primary and secondary images according to the local noise level.
[0020] In various embodiments within the second group, the cutoff frequency may be variable or constant throughout the main image. The use of a constant cutoff frequency may provide simplicity and robustness. c and weight coefficient When varying in response to the same or related local characteristics of the primary image, undesirable feedback behavior (e.g., oscillation) and / or undesirable mutual compensation effects can be further avoided. The constant cutoff frequency can be based on global characteristics of the primary image and / or the secondary image. Alternatively, the constant cutoff frequency can have a predetermined value.
[0021] In a third set of embodiments of the first aspect, the local noise level of the corresponding region of the secondary image is determined, i.e., in addition to the local noise level of the region of the primary image. When combining the LF and HF components, the relative contribution of the HF component can be determined so that the local noise level of the target image region remains below a threshold noise level. As already explained, the noise of a region of the primary image is generally greater than that of the corresponding region of the secondary image, and the noise is mitigated by adding the amount of data from the secondary image (or, more accurately, from the HF component derived therefrom). However, as the inventors have appreciated, the appropriate amount of data from the secondary image depends not only on the local noise level of the primary image, but also on the local noise level of the secondary image. In fact, the noisier the corresponding region of the secondary image, the more regions need to be added in order to reduce the local noise level of the primary image to an acceptable noise level. Based on the respective noise levels of the region of the primary image and the corresponding region of the secondary image, it is possible to make a reliable prediction of the resulting local noise level of the target image region. The prediction can rely on a linear combination of the noise levels, or on some type of regression model or trained model.
[0022] Based on these considerations, in some embodiments, the relative contribution of the HF component can be calculated to be a minimum or substantially minimum value such that the local noise level of the target image area is below a threshold noise level. This can be achieved by aiming to combine the HF and LF components so that the threshold noise level is precisely reached, and then increasing or boosting the HF component's contribution by a predetermined margin. Alternatively, minimization can be achieved by starting with a low value for the HF component's contribution and iteratively increasing it in small steps until the local noise level of the target image area falls below the threshold noise level. Equally, the calculation can be aimed at determining the relative contribution of the LF component such that the noise level of the target area is below the threshold noise level.
[0023] In some embodiments within the third group, the contribution of the HF component relative to the LF component is determined by the weighting factor to control, while in other embodiments it is controlled by the cut-off frequency f c It should be understood that the weight coefficient Or cutoff frequency f c The determination is then made independently across the primary image and for each region of the primary image. This determination is made taking into account a threshold noise level. The threshold noise level may be a predetermined value, particularly a globally applicable predetermined value, which the system owner may configure as they see fit. Alternatively, the threshold level may be calculated for each new primary image, each new secondary image, or each new combination of a primary and secondary image based on the global characteristics of the image.
[0024] In a specific embodiment of the third group, the local noise level of the LF component and the local noise level of the HF component are determined. With this information, the resulting local noise level of the target image area can be predicted particularly reliably. The prediction can rely on a linear combination of the noise levels or on some type of regression model or training model.
[0025] In each of the above embodiments of the image fusion method according to the first aspect, normalizing the secondary image to the primary image before combining is optional. Of course, normalization can be an integral part of the combining step, or can be performed before combining. In a broad sense, normalization can be rescaling the secondary image by a scaling factor so that the LF component of the primary image becomes comparable to the LF component of the secondary image (relative to the same cutoff frequency f). c ) are comparable or equal. It is possible to implement normalization as an operation that operates on the primary image instead of the secondary image, or as an operation that operates on both the primary and secondary images.
[0026] Furthermore, in each of the above embodiments of the first aspect, the local noise level of the region of the host image may be determined using a sensor noise model that depends on the local sensor readings. Alternatively, the local noise level of the region of the host image may be determined by direct calculation and / or by measurement.
[0027] In one envisaged use case of the image fusion method according to the first aspect, the primary image is a photographic image such as a color image acquired primarily using visible light and / or the secondary image is a photographic image such as an infrared or near-infrared image acquired primarily using non-visible light.
[0028] In a second aspect of the present disclosure, there is provided an apparatus having processing circuitry arranged to perform the method of the first aspect. The second aspect generally shares the effects and advantages of the first aspect, and it can be implemented with corresponding degrees of technical variation.
[0029] The present disclosure further relates to a computer program comprising instructions for causing a computer to perform the above-described method. The computer program may be stored or distributed on a data carrier. As used herein, a "data carrier" may be a transient data carrier such as a modulated electromagnetic wave or light wave or a non-transient data carrier. Non-transient data carriers include volatile and non-volatile memories such as permanent and non-permanent storage media of the magnetic, optical, or solid-state type. Such memories may be fixedly mounted or portable, still falling within the scope of a "data carrier."
[0030] In general, unless otherwise expressly defined herein, all terms used in the claims should be interpreted according to their ordinary meaning in the technical field. Unless expressly stated otherwise, all references to "a / an / the element, device, component, means, step, etc." should be interpreted openly as referring to at least one instance of that element, device, component, means, step, etc. Unless expressly stated otherwise, the steps of any method disclosed herein do not have to be performed in the exact order described. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Various aspects and embodiments will now be described by way of example with reference to the accompanying drawings, in which:
[0032] Figure 1 depicting transform components of a region in the primary image and transform components of a corresponding region in the secondary image in a two-dimensional frequency plane, deriving corresponding LF and HF components of these transform coefficients, and combining the LF and HF components into transform coefficients for the target image region;
[0033] Figure 2 and Figure 3 Alternative methods of deriving LF and HF components are mentioned;
[0034] Figure 4 shows segmentation of the primary and secondary images into image regions;
[0035] Figure 5 In its upper part are shown a panoramic main image PP of a scene and a secondary image S of a portion of the same scene, wherein the overlap of the panoramic main image PP and the secondary image S is identified as the main image P;
[0036] Figure 6 is a flow chart of an image fusion method according to an embodiment of the present invention;
[0037] Figure 7 shows a local arrangement with cameras for acquiring a primary image and a secondary image of a scene, and means configured to fuse the primary image and the secondary image into a target image; and
[0038] Figure 8 Pictured Figure 7 The arrangements seen in include possible distributed implementations of multiple networked entities. DETAILED DESCRIPTION
[0039] Various aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which specific embodiments of the invention are shown. However, these aspects may be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete and will fully convey the scope of all aspects of the invention to those skilled in the art. Like numbers refer to like elements throughout.
[0040] System Overview
[0041] Now we will continue to refer to Figure 6 An embodiment of an image fusion method 600 is described with reference to a flowchart of FIG. The method 600 operates on a primary image P and one or more secondary images S.
[0042] The method 600 may be executed by a general purpose computer. Figure 7 The method is performed by an apparatus of the type illustrated in FIG. Figure 7 The device 730 in comprises a processing circuit 731 such as a central processing unit or a graphics processing unit, data connections towards a source of the primary image P and the source of the secondary image S, and data connections towards a recipient of the target image T to be generated by the method 600. The device 730 may constitute or be part of a video management system (VMS) or a video authoring tool.
[0043] The processing circuit 731 is configured to execute a computer program 733 stored in the memory 732 of the apparatus 730. Figure 7In the example of , the source of the primary image is an RGB camera 710 and the source of the secondary image is an infrared camera 720. The infrared camera 720 can be a night camera associated with an active infrared light source directed at an area of a scene 790. Both cameras 710, 720, which can be digital still cameras or digital video cameras, are directed at the scene 790, which is illuminated in a (spatially) non-uniform manner at the relevant time, making it a challenging subject for photography. Specifically, the scene 790 is illuminated by a single downward light source that does not cover the entire horizontal extent of the scene 790, and furniture is present in a light cone that obscures the furniture and creates sharp shadows on the floor. It can be expected that the visible light image of the scene 790, especially the color image, has a high local noise content (i.e., a low SNR) in the darker areas. Recall that improving photographic images with heterogeneous lighting is the use case for which the image fusion method 600 is intended, i.e., by fusing a photographic image (the primary image) with a monochrome image (the secondary image). It will be appreciated that the cameras 710, 720 capturing the primary and secondary images are configured to do so approximately simultaneously so that the scene 790 does not change noticeably therebetween. It will be appreciated that despite this approximate simultaneity, the primary and secondary images may be captured using different exposure times.
[0044] like Figure 7 As illustrated in , device 730, as it generates target images T, outputs them to memory 740. A local or remote recipient may retrieve the target images from memory 740 and use them for playback (eg, using interface 750) or for further editing. Figure 7 At least some of the arrows in the figure represent a remote connection such as a data connection on the Internet or another wide area communication network. Each of the arrows may represent a wired or wireless data connection.
[0045] As an alternative to this localized implementation, Figure 8 8. A possible distributed implementation is illustrated. Here, a plurality of networked entities collaborate while exchanging information over a communication network 870. The networked entities (some of which may be characterized as clouds or edge devices) include an image fusion device 830 that is functionally equivalent to the device 730 described above. Further provided are a primary imaging device 810 and a plurality of secondary imaging devices 820, 821 for acquiring primary and secondary images, respectively. Among the networked entities, further seen is a server (or host computer) 860 for storing source images to be fed to the image fusion device 830 or target images generated by the image fusion device 830. Further provided is a user interface 850 configured to allow editing, playback, or other forms of use of the images.
[0046] Note that in some embodiments, the primary and secondary images can be frames from corresponding video sequences, and the target image can be combined into a target video sequence. Preferably, the cameras 710, 720 that acquire the video sequences from which the primary and secondary images originate are approximately synchronized. This approximate temporal synchronization can be considered satisfied if the appearance of the scene 790 does not change significantly in time between the acquisition of two corresponding frames in the corresponding video sequences. This justifies the use of image data from the secondary image to improve the primary image. Temporal synchronization can be achieved by feeding trigger signals to the cameras 710, 720 and / or by making available a common clock signal and a pre-agreed timing table.
[0047] As discussed in the previous section of this disclosure, a primary image can be identified as a noisier image compared to a secondary image. Noisiness can refer to a local signal-to-noise ratio (SNR), which can be low in the primary image in a local sense across most image regions, or can be low in a global sense, e.g., in terms of mean, median, or q-quantile. Primary / secondary identification can also refer to a metric similar to SNR, such as signal-to-interference plus noise ratio (SINR), a signal-to-noise statistic, a peak SNR, or an absolute noise measurement. It should be understood that noisiness corresponds to the inverse of SNR, where teachings such as "variable X gradually increases with local noise level" can be equivalently expressed as "variable X gradually decreases with decreasing local SNR."
[0048] As also mentioned previously, the primary image may alternatively be identified because it is represented with more channels than the secondary image (e.g., color versus monochrome), which generally increases its vulnerability to sensor noise. Still further, an input image may be considered a primary image simply because the user on whose behalf method 600 is being performed prefers to retain as much of the image data for that image as possible and to add as little image data from the secondary image as possible.
[0049] To simplify the following description, it is assumed that the primary and secondary images are of equal extent in the sense that they depict a common part of the scene. Figure 5 Note that this assumption does not exclude the case where the primary image P originates from a panoramic (or full-angle or overview) primary image PP of the scene, nor does it exclude the case where the secondary image S originates from a panoramic secondary image (not shown). Therefore, for the purpose of the image fusion method 600 to be described, the assumption that the primary and secondary images have the same extent remains valid without loss of generality. Figure 5 , illustrates a use case where the primary image P is taken from a panoramic primary image PP, where only the right part of the scene is covered by the secondary camera providing the secondary image S. In practical use cases, including video surveillance applications, it may not always be reasonable to arrange the secondary cameras to cover the entire scene; instead, areas that are expected to be well-lit may be excluded without causing harm. Figure 5As indicated, the auxiliary camera is directed towards an area of the garden that is partially obscured by trees and lacks nearby light sources.
[0050] At least some embodiments of the image fusion method 600 may be applied to primary and secondary images having unequal spatial resolutions (ie, images that, while having equal extents, include different numbers of pixels).
[0051] For the purpose of the following description of the method 600, it is assumed that the primary image and the secondary image are segmented into regions and that there is a correspondence between the regions of the primary image and the secondary image. Figure 5 In
[15] , the secondary image S is partitioned into a 3×2 matrix of rectangular regions that are indexed by an example numbering scheme on a table (rows, columns). Figure 5 In the upper portion of FIG, the primary image P is divided into a corresponding arrangement of three regions high and two regions wide. The upper left region of the primary image P and the upper left region of the secondary image S are corresponding regions. As briefly mentioned above, an image region in this sense can be a pixel or a group of adjacent pixels. In implementation, it may be advantageous to use image regions that coincide with macroblocks in video coding, particularly predictive video coding. For an example technical definition of the term "macroblock," reference is made to any video coding standard of ITU-T H.26x. Alternatively, image regions may be allocated in such a way that they coincide with coding blocks to be used for block-by-block transform coding of the primary or secondary image. This is convenient in implementation, particularly since re-encoding of a video frame can begin before the video frame has been fully processed according to image fusion method 600.
[0052] Figure 4 The segmentation of the primary image P and the secondary image S into 2×2 regions is shown, wherein the region index is explicitly indicated. This is a simplification for explanatory purposes. In a practical implementation of the method 600, the regions may be 8×8 pixels, 16×16 pixels, or 64×64 pixels so that their total number in the video frame is significantly higher. Without departing from the scope of the present disclosure, Figure 4 and Figure 5 The region segmentation seen in can vary significantly to include non-square arrangements and / or arrangements of regions that are not similar to the image and / or mixed arrangements of different regions having different sizes or shapes. It is recognized that some video coding formats support dynamic macroblock segmentation, i.e., the segmentation can be different for different video frames in a sequence, and the region segmentation can follow this change in macroblock segmentation. This is relatively easy to implement because the method 600 operates on a pair of source images at a time; therefore, if the method 600 adopts a different region segmentation starting from a particular new frame, the target image (target frame) that has already been generated is not affected.
[0053] Image fusion method
[0054] The method 600 starts with determining the local noise level of a region P(m,n) of a host image P. Step 602 Start. The notation P(m,n) can be considered to refer to image data representing the region identified by index (m,n) in the main image; the image data can be in pixel / plain format or in a transformed format. The transformed format can include coefficients obtained by a frequency domain transform.
[0055] The following is a calculation based on the read noise σ R and shot noise cS to estimate the effective noise σ of the sensor eff An example model:
[0056]
[0057] Among them, σ R Typically a constant, is the sensor signal (local sensor reading) and c is an empirical constant. Implicit σ eff is variable across images, i.e., it depends on at least the image region index (m,n). In an image region comprising multiple pixels, the sensor reading u is related to a preselected one of the pixels (e.g., the top left pixel, the middle pixel) or to a common value (e.g., the mean, the median). The constant σ R It is usually a reasonable assumption that the value of c is constant under different usage conditions and over the life of the sensor. Therefore, in practical implementation, it is sufficient to estimate the constant once for each sensor design or obtain the constant from the sensor manufacturer. If the actual sensor signal is a signed or complex quantity, or a vector, then u can be understood as the amplitude of the sensor signal. Next, the signal-to-noise ratio can be expressed as
[0058]
[0059] Note that in the simple sensor noise model discussed here, the sensor signal u is the only variable in expressions (1) and (2). That is, the sensor signal u is used as a proxy for the sensor noise, and therefore as a proxy for the local noise level of the host image. Figure 7 In the example use case illustrated in , the local noise level of the primary image can be estimated based on the sensor noise model for the RGB camera 710, and the local noise level of the secondary image can be estimated based on the independent sensor noise model for the infrared camera 720 (if any).
[0060] The local noise level can be expressed using direct or inverse noise metrics. Inverse metrics include SNR, SINR, signal-to-noise ratio statistics, and peak SNR, where the noise component is in the denominator. Example direct metrics can be noise measurements performed with a noisy sensor or noise estimates from a sensor noise model (1).
[0061] Execution of method 600 continues to Step 606 , wherein a low frequency (LF) component P is derived from said region of the main image LF (m,n). Further, a high frequency (HF) component S is derived from the corresponding region S(m,n) of the secondary image. HF (m,n). The LF component may include the region of the main image up to the cutoff frequency f c (spatial frequency), and the HF component may include all image data in the corresponding area of the secondary image up to an equal cutoff frequency f c As will be described in detail below, in some embodiments, the cutoff frequency f c It is constant throughout the main image, but variable in other embodiments.
[0062] In order to provide the cut-off frequency f c For a more precise understanding, recall that transform coding in a broad sense consists in Project image data on an orthogonal basis. With:
[0063]
[0064] in, are the transform coefficients for frequency (k1, k2), and is the image data for pixel (n1,n2), e.g. the intensity of a color channel. Note that The restriction to [0, N1] × [0, N2] can be used for the lowest frequency pair (k1, k2) corresponding to a single period or a constant value. Specifically, the basis can be composed of real-valued bi-periodic harmonic functions such as discrete cosine transform (DCT) functions, discrete sine transform (DST) functions, discrete Fourier transform (DFT) functions, and wavelet transform functions. In some embodiments, the orthogonal basis in (3) is composed of discrete cosine functions:
[0065]
[0066] The transform coefficients calculated in the projection operation constitute a discrete representation of the spectrum of the image data, and each transform coefficient corresponds to a frequency pair (k1, k2).
[0067] In step 606, by only keeping the c The LF components of the regions of the main image are derived from the following frequency-dependent transform coefficients. In some embodiments, the LF components are calculated so that the regions with l 2 Norm is less than f c The frequency-dependent transform coefficients is maintained, which can be conceptually written as:
[0068]
[0069] f c The remaining transform coefficients above are set to zero. Figure 1 In which the axes represent k1 and k2 respectively. This filtering operation corresponds to the process from P(m,n) to P LF (m,n). Conversely, by keeping only the cutoff frequency f c The frequency-dependent transform coefficients above To derive the HF component of the corresponding area of the secondary image,
[0070]
[0071] And f c The remaining ones below are set to zero.
[0072] In other embodiments, the LF component is calculated such that 1 Norm is less than f c The frequency pairs of related transform coefficients are kept, and the rest are set to zero, that is:
[0073]
[0074] This is Figure 2 In still other embodiments, the LF component is calculated such that ∞ Norm is less than f c The frequency pairs of related transform coefficients are kept, and the rest are set to zero, that is:
[0075]
[0076] This is Figure 3 As will be readily understood by those skilled in the art, these embodiments can be generalized to allow the use of l with arbitrary p≥1 p Norm of norm||(k1,k2)|| p to select the transform coefficients.
[0077] It is foreseeable that in most implementations of method 600, the cutoff frequency f c is applied with limited steepness. That is, a certain amount of roll-off is included to avoid certain artifacts, and the roll-off may depend on the local noise level or other characteristics of the image region. With respect to the three LF filtering options outlined above, if the norm of the frequency pairs of the transform coefficients is less than f c -∈, where ∈>0, the transform coefficients can be kept, and if the norm of the frequency pairs of the transform coefficients is greater than f c+∈, the transform coefficients can be set to zero. The smooth transition is applied to those values that lie in the range [f c -∈,f+∈]. The smoothness can correspond to C 0 ,C 1 ,C 2 Continuity or greater continuity.
[0078] Note that there are methods for calculating the LF and HF components without transforming the image data to a frequency plane representation and then back again. For example, convolution with an appropriate Gaussian kernel can be used to limit the frequency content as desired, as needed. In another example, real-time or continuous-time low-pass or high-pass filters can be utilized; the filters can be implemented as physical circuits or through software instructions. In the case of non-zero roll-off (finite steepness), the respective transfer functions of the low-pass (LP) filter and the high-pass (HP) filter can be of the following form:
[0079]
[0080] The "smooth transition" shape mentioned above may correspond to a shape such as |H LP The frequency-dependent magnitude of the transfer function of (2πif)|
[0081] After completing step 606, the execution flow of method 600 proceeds to Step 610 , where the LF component and the HF component are combined into a target image region T(m,n), where the relative contribution of the HF component to the target image region gradually increases with the local noise level. The target image region corresponds to the region of the main image, i.e., it has an index (m,n) equal to that of the region. Figure 1 The combination is conceptually illustrated in FIG, where the LF component P LF (m,n) is added to the HF component to form the target image region T(m,n):
[0082] T(m,n)=P LF (m,n)+S HF (m,n) (5)
[0083] exist Figure 1In , the transform coefficients originating from the secondary image are symbolized by solid dots and the transform coefficients originating from the primary image are symbolized by hollow circles. The combining operation (5) can be performed on a pixel representation or a frequency space representation of the image data. It will be appreciated that if the target image is to be represented in the same color space as the primary image, the combining operation (5) can be performed for multiple color channels. Accordingly, the image data of a single (monochrome) channel from the secondary image region is split into multiple sub-contributions of the corresponding primary image region. In RGB space, the contributions may be equal or may be adjusted to take into account applicable white balance settings in order to have the desired neutral appearance. In YCbCr space, it is sufficient to perform the combining operation (5) on the luminance channel Y.
[0084] In some embodiments, by using a variable cutoff frequency f c =f c (m,n) to control the relative contribution of the HF component. In other embodiments to be described in detail below, the relative contribution of the HF component is controlled by intentionally varying the primary and secondary weighting coefficients from one image region to another. Specifically, for the LF and HF components of the target image region T(m,n), the primary and secondary weighting coefficients can be independently varied over the interval [0,1].
[0085] The combining step 610 may optionally be performed after normalization Step 608 Before or with normalization Step 608 Normalization can be achieved by rescaling the secondary image by a scaling factor R(m,n) so that the LF component of the primary image becomes comparable to the LF component of the secondary image (relative to the same cutoff frequency f c ) are compared or equal. Accordingly, equation (5) can be modified as follows:
[0086] T(m,n)=P LF (m,n)+R(m,n)S HF (m,n) (6a)
[0087] in,
[0088]
[0089] where Y(·) represents a luminance function that represents the local brightness of the image, such as luminance or relative brightness. For a monochrome image, luminance is equal to the pixel intensity or average pixel intensity in the region. For a multi-channel image, luminance is a linear combination of the channels in linear units (e.g., according to a predetermined perceptual model). In the literature, there are various proposed or standardized luminance functions associated with different color spaces. In the particular case of the RGB color space (without gamma compression), luminance can be calculated as:
[0090] Y=0.2126R+0.7152G+0.0722B.
[0091] In a simple implementation, it is possible to use one color channel as a proxy for brightness (e.g., Y=G in the RGB color space). For the avoidance of doubt, it is emphasized that the scaling factor R(m,n) in equation (6b) depends on two LF components, namely the LF component of the primary image and the LF component of the secondary image. This can be done in a manner similar to how the LF component of the primary image is derived, and using equal cutoff frequencies f c , derive the LF component of the secondary image from the secondary image:
[0092]
[0093] The normalization operation (step 608) can be integrated into the combining operation (step 610) according to equation (6a). Alternatively, the normalization can be performed before combining, for example, by first replacing
[0094]
[0095] and then evaluate (5). Still further, it is possible to implement the normalization as an operation acting on the primary image instead of the secondary image, or as an operation acting on both the primary and secondary images.
[0096] Execution of method 600 continues by repeating step 612 above, along with any optional steps present in a particular embodiment, for any remaining image regions, i.e., for any remaining values of the region index (m, n). It is important to emphasize that in each of these repetitions, the relative contribution of the HF component can be different, so that the fusion of the primary and secondary images is truly variable across the entire primary image. As a result, some regions of the target image will have a more or less primary image appearance, while other regions may experience significant blending with image data from the secondary image. This can be likened to applying a local night mode to the latter regions.
[0097] In the final Step 614 Then the resulting target image regions are merged into the target image T.
[0098] The first set of embodiments
[0099] As described above, the image fusion method 600 encompasses various embodiments, which are organized into three main groups for clarity of this disclosure. Technical features across these groups can be combined to form new embodiments. As will be appreciated by those skilled in the art, different technical features belonging to different groups can also be applied to different regions of a master image.
[0100] In a first set of embodiments, the cutoff frequency is variable across the main image, f c =f c (m,n), where (m,n) is the image region index. Embodiments in this group may include a further sub-step 606.1 of determining a cutoff frequency for a region of the main image based on the local noise level of the region. This determination may be repeated for one or more further regions of the main image so that different cutoff frequency values are applied when deriving the respective LF and HF components. In this regard, the amount of roll-off to be used in step 606 may also be adjusted. For example, if the cutoff frequency f in the image region c The lower the value, the better. The lower the value, the more roll-off is applied in the image area.
[0101] Specifically, a non-increasing function of the local noise level f is used c =f c (σ eff ) to determine the cutoff frequency. (Equivalently, a non-decreasing function of SNR can be used The cutoff frequency is determined by the cutoff frequency. Consequently, the noisier the primary image region, the larger the portion of the spectrum of the target image region originating from the secondary image will be. Conversely, if the primary image region has a relatively low noise content, the characteristics of the primary image are preserved to a greater extent. The non-increasing function of the local noise level is preferably a continuous function (including, in digital signal processing, a quasi-continuous function consisting of multiple small increments). This allows the relative contribution of the HF component to the target image region to gradually increase with the local noise level, resulting in fewer or less noticeable threshold-type artifacts.
[0102] In one embodiment, the cutoff frequency f c It is calculated as an affine function based on the local noise level σ:
[0103]
[0104] Among them, σ1, σ2 are constants, and f max is the maximum frequency of the transform coefficients available in P(m,n) or S(m,n). It can be said that for the same p-norm as used in step 606, f max is the maximum value over all frequency pairs (k1, k2) for which equation (3) is evaluated. Instead of the affine function (7), smoother functions in the form of higher-order polynomials, exponential functions, logistic functions, etc. can be used.
[0105] Second set of embodiments
[0106] In a second set of embodiments, according to one or more weight coefficients that are variable across the main image to combine the LF component and the HF component (and possibly further components). Each weighting coefficient Can be a real number in [0,1]. In these embodiments, one of the following example expressions can be used to replace equation (5):
[0107]
[0108]
[0109]
[0110]
[0111] How to calculate the LF component of the corresponding area of the secondary image has been explained above. The HF component of the area of the primary image can be calculated similarly to the HF component of the corresponding area of the secondary image.
[0112] According to equation (8a), different weight coefficients are used for the LF component and the HF component of the target image region T(m,n). The LF component can be imagined as the sum of the first two terms in the equation, and the HF component is the sum of the last two terms.
[0113] According to equation (8b), the LF component of the target image region is derived only from the primary image, while the HF component of the target image region is a weighted combination of the HF component of the region P(m,n) of the primary image and the HF component of the corresponding region S(m,n) of the secondary image. This corresponds to assigning In words, the combining operation in equation (8b) can be described as follows: the HF component from the secondary image is pre-combined with the HF component from the primary image according to the weighting coefficients before being combined with the LF component from the primary image into the target image region. If it is expected that noise will mainly affect the higher part of the spectrum of the primary image, then combining according to equation (8b) may be appropriate.
[0114] According to equation (8c), the HF component of the target image region is derived only from the primary image, while the LF component of the target image region is a weighted combination of the LF component of the region P(m,n) of the primary image and the LF component of the corresponding region S(m,n) of the secondary image. This corresponds to assigning If the noise is expected to affect mainly the lower part of the spectrum of the main image, a combination according to equation (8c) may be appropriate.
[0115] According to equation (8d), the HF component of the target image region is derived only from the secondary image, while the LF component of the target image region is a weighted combination of the LF component of the region P(m,n) of the primary image and the LF component of the corresponding region S(m,n) of the secondary image. This corresponds to assigning If it is expected that the upper part of the spectrum of the main image will very often be heavily affected by noise, a combination according to equation (8d) may be appropriate.
[0116] Further, as shown in equation (6), each of the above expressions can be supplemented by a normalization factor.
[0117] Embodiments in this group may comprise determining a weight coefficient based on the local noise level of said region P(m,n) of the main image A further sub-step 610.1 of the main image may be performed on one or more further regions P(m ′ ,n ′ ) This determination is repeated so that different values of the weighting coefficients are applied when combining the respective LF and HF components and further components (if any) in step 610.
[0118] This determination may be subject to convexity constraints. Separate LF and HF convexity constraints may be applied. If the image data is represented in linear units, the convexity constraint may be formulated as: or The convexity constraint can be applied on a per-band basis; in practice, there may be one convexity constraint for the coefficients acting on the LF component and another convexity constraint for the coefficients acting on the HF component. Alternatively, the convexity constraint is the square of the weight coefficients with the following units:
[0119] In some embodiments of the second group, a sub-coefficient is used that is a non-decreasing function of the local noise level. The noisier the area of the primary image, the more image data in the target image area will originate from the secondary image, and vice versa. Alternatively or additionally, a primary coefficient that is a non-increasing function of the local noise level may be used. Using primary coefficients having these characteristics has the corresponding effect of balancing the contributions from the primary and secondary images according to the noise level.
[0120] In one embodiment, involving the LF spectrum, the sovereign weight coefficient is calculated as an affine function according to the local noise level σ:
[0121]
[0122] where the dependence of σ on (m,n) is implicit and the values of the constants σ1,σ2 are independent of the values in equation (7). The secondary weight coefficients can then be obtained as follows:
[0123]
[0124] This has the effect of giving progressively more weight to the secondary image as the noise level σ increases. Similarly, for the HF spectrum, the primary weight coefficient can be defined as follows:
[0125]
[0126] Among them, the constants σ3 and σ4 can be set independently of the above constants σ1 and σ2, and the secondary weight coefficient can be calculated as
[0127] In various embodiments within the second group, the cutoff frequency f c The cutoff frequency may be variable between image regions, or it may be constant across the entire primary image. The constant cutoff frequency may be based on a global characteristic of the primary image P and / or the secondary image S. A global characteristic of the primary or secondary image may be, for example, an average noise level or a peak noise level. Specifically, the constant cutoff frequency may be based on such a global characteristic of the panoramic primary image PP from which the primary image P is derived, or the panoramic secondary image from which the secondary image S is derived. Alternatively, the constant cutoff frequency may have a predetermined value, in particular a value configured by a user or system administrator.
[0128] The third set of embodiments
[0129] In a third set of embodiments of method 600, Step 604 In addition to determining the local noise level σ of the region P(m,n) of the main image P In addition to the local noise level σ of the corresponding area S(m,n) of the secondary image (simply denoted by σ in the description of the first and second groups of embodiments), the local noise level σ of the corresponding area S(m,n) of the secondary image is also determined. S (m,n). When combining LF components P LF (m,n) and HF component S HF (m,n) and any further image data components, then the local noise level σ of the target image region can be T (m,n) at the threshold noise level σ * The relative contribution of the HF component is determined in the following manner. Specifically, the relative contribution of the HF component can be determined so that the local noise level of the target image region just reaches (just below) the threshold noise level. In this way, the noise target expressed by the threshold noise level is met while maintaining the appearance of the main image to the greatest extent possible. Formulated in slightly different terms, the embodiments in the third group provide that the local noise level σ of the target image region should be T (m,n) at the threshold noise level σ * The relative contribution of the secondary image including the contribution derived from the HF component of the secondary image is determined in the following way.
[0130] As explained above, a region P(m,n) of the primary image generally has a higher local noise level than a corresponding region S(m,n) of the secondary image, and the noise of the primary region is mitigated by adding the amount of image data from the secondary image. The appropriate amount of data from the secondary image depends not only on the local noise level of the primary image, but also on the local noise level of the secondary image. Therefore, if the local noise level of the primary image region is to be reduced to a tolerable noise level, i.e., the threshold noise level σ * , then the larger the noise of the corresponding area S(m,n) of the secondary image, the more secondary images need to be added. Based on the noise levels of the respective areas of the main image and the corresponding areas of the secondary image, it is often possible to calculate the resulting local noise level σ of the target image area T(m,n) T (m,n) to make a reliable prediction. The prediction can rely on a linear combination of the various noise levels:
[0131] σ T (m,n)=γ P σ P (m,n)+γ S σ S (m,n),
[0132] Among them, γ P ,γ S is a positive constant. Alternatively, the squares of the individual noise levels are combined:
[0133] σ T (m,n) 2 =γ P σ P (m,n) 2 +γ S σ S (m,n) 2 .
[0134] Alternatively, a regression model or a training model can be used to predict the local noise level σ of the target image region T(m,n) T (m,n). Each of these options can be expressed as a "black box" function W:
[0135] σ T (m,n)=W(σ P (m,n),σ S (m,n)).
[0136] In some embodiments within the third group, the determination of the local noise level of a region P(m,n) of the primary image and a corresponding region S(m,n) of the secondary image is limited to the low-frequency portion and the high-frequency portion thereof, respectively. This may allow the determination of the local noise level σ of the target image region T(m,n) to be T(m,n) will make a more faithful prediction, especially since the combination in step 610 will not involve the entire region P(m,n), S(m,n). The separation into low-frequency and high-frequency parts for the purpose of local noise estimation can be based on the cutoff frequency f used in step 606. c the same cutoff frequency, especially if the cutoff frequency f c It can be constant, or it can correspond to a different preset frequency.
[0137] In some embodiments, the relative contribution of the HF component may be calculated as a minimum (ie, exact minimum or approximately minimum) value such that the local noise level σ of the target image region is T (m,n) at the threshold noise level σ * This can be achieved by aiming to combine the HF and LF components so that the threshold noise level is precisely reached, and then increasing or boosting the contribution of the HF component by a predetermined margin. Alternatively, this can be achieved by starting with a low value for the HF component's contribution and iteratively increasing it in small steps until the local noise level of the target image area falls below the threshold noise level. Equivalently, the calculation can aim to determine the relative contribution of the LF component so that the noise level of the target area is below the threshold noise level.
[0138] In some embodiments within the third group, the contribution of the HF component relative to the LF component is determined by a weighting factor to control, and in other embodiments by the cutoff frequency f c It should be understood that the weight coefficient Or cutoff frequency f c (m,n) are then variable across the main image, and the determination is made independently for each region of the main image. This determination is made taking into account the threshold noise level. In this regard, the amount of roll-off to be used in step 606 can also be adjusted. For example, at the cutoff frequency f c It may be preferable to apply relatively more roll-off in relatively lower image regions.
[0139] Threshold noise level σ * The threshold level may be a predetermined value that is configurable as the user or system owner deems appropriate. Alternatively, the threshold level may be calculated for each new primary image, each new secondary image, or each new combination of primary and secondary images based on a global characteristic of the image. For example, the global characteristic may correspond to an average noise level or a peak noise level.
[0140] In one example, the contribution of the HF component relative to the LF component is determined by a variable weight coefficient to be applied to each HF component. The relative contribution of the HF component from the secondary image should be controlled in such a way that it is the minimum that brings the local noise level of the target image area below the threshold noise level:
[0141] w HF (m,n)=min{0≤a≤1:σ T (m,n)≤σ *} (9)
[0142] Here, under the assumption that the variance is approximately linear, the local noise level of the HF component from the target image region is given by:
[0143] σ T (m,n)=αγ P σ P (m,n)+(1-a)γ s σ S (m,n) (10a)
[0144] or
[0145] σ T (m,n) 2 =aγ P σ P (m,n) 2 +(1-a)γ S σ S (m,n) 2 (10b)
[0146] or
[0147] σ T (m,n)=W(aσ P (m,n),(1-a)σ S (m,n)). (10c)
[0148] If (10a) or (10b) is used, then the minimization (9) has a closed-form solution. If expression (10c) is used instead, and for the general noise expression, then numerical methods can be relied upon, including iterative solvers.
[0149] Conclusion
[0150] Aspects of the present disclosure have been described above mainly with reference to a few embodiments. However, as will be readily appreciated by a person skilled in the art, other embodiments than the above disclosed embodiments are equally possible within the scope of the invention as defined by the appended patent claims. In particular, it is envisaged that image fusion is performed on a primary image and N≥2 secondary images. For this purpose, a common cutoff frequency may be applied and the weight coefficients w which are variable with respect to the image region index (m,n) may be used. LF(m, n), to control the relative contributions of the LF components derived from the primary image and the multiple secondary images.
Claims
1. A method (600) for fusing a primary image (P) and a secondary image (S), wherein: The primary image is noisier than the secondary image, and wherein the primary image and the secondary image are segmented into regions, the method comprising the steps of: a) determining (602) a local noise level of a region P(m,n) of the main image; b) deriving (606) a low frequency LF component P from said region of said main image LF (m,n), and derive (606) a high frequency HF component S from the corresponding region S(m,n) of the secondary image HF (m,n); wherein the LF component and the HF component refer to a common cutoff frequency (f c ); c) combining (610) the LF component and the HF component into a target image region T(m,n); d) repeating (612) the previous steps a) to c) for any remaining image regions; and e) merging (614) all output target image regions thus obtained into a target image (T), characterized in that f) the relative contribution of the HF component to the target image region gradually increases with the determined local noise level of the region of the host image.
2. The method according to claim 1, wherein The cut-off frequency is variable across the main image, wherein step b) of the method further comprises the sub-steps of: b1) Determining (606.1) the cutoff frequency for a region of the main image based on the local noise level of the region.
3. The method according to claim 2, wherein: The cutoff frequency is determined using a non-increasing function of the local noise level, which is a continuous non-increasing function of the local noise level.
4. The method according to claim 1, wherein The LF component and the HF component are combined according to one or more weight coefficients that are variable across the host image (610), wherein step c) of the method further comprises the sub-steps of: c1) Determining (610.1) the one or more weight coefficients based on the local noise level of the region of the main image.
5. The method according to claim 4, wherein The weight coefficients include: a main coefficient applied to components of the region of the host image and being a non-increasing function of the local noise level, and / or A secondary coefficient is applied to the component of the corresponding region of the secondary image and is a non-decreasing function of the local noise level.
6. The method according to claim 4, wherein: The cutoff frequency is constant throughout the main image.
7. The method according to claim 4, further comprising: The HF component P is derived from the region of the main image HF (m,n), The HF component from the secondary image is pre-combined with the LF component from the primary image according to the weight coefficient before being combined with the LF component from the primary image to form the target image region.
8. The method according to claim 1, wherein Step a) of the method is followed by: a1) determining (604) a local noise level of said corresponding area of said secondary image, Therein, the relative contribution of the HF component is determined such that the local noise level of the target image region is below a threshold noise level.
9. The method according to claim 8, further comprising: The relative contribution of the HF component is calculated as a minimum value such that the local noise level of the target image region is below the threshold noise level.
10. The method according to claim 8, wherein The LF component and the HF component are combined according to a weight coefficient that is variable across the main image (610), wherein step c) of the method further comprises the sub-steps of: c1) Determining (610.1) the weight coefficient based on the respective local noise levels of the regions of the primary image and the secondary image and based on the threshold noise level.
11. The method according to claim 8, wherein The cut-off frequency is variable across the main image, wherein step b) of the method further comprises the sub-steps of: b1) Determining (606.1) the cutoff frequency based on the respective local noise levels of the regions of the primary image and the secondary image and based on the threshold noise level.
12. The method according to claim 8, in, The threshold noise level is predetermined, and / or The threshold noise level is based on global characteristics of the primary image and / or the secondary image.
13. The method according to any one of the preceding claims, wherein Step c) of the method further comprises: The secondary image is normalized (608) to the primary image before completing the combining (610).
14. The method according to claim 1, wherein The local noise level of the region of the host image is determined using a sensor noise model that depends on local sensor readings.
15. An apparatus (730; 830) comprising a processing circuit (731) arranged to perform the method of claim 1.
Citation Information
Patent Citations
Resolution and contrast enhancement with fusion in IR images
EP2570988A2
Light-splitting combined image collection device
EP3518179A1
Method and apparatus for performing super-resolution
WO2013131929A1
Systems and methods for fusing infrared image and visible light image
WO2018120936A1
Tuning color image fusion towards original input color with adjustable details
WO2021184027A1