Image fusion method, computer program product, storage medium and electronic device
By employing frequency band fusion and fusion mask calculation methods, the problem of poor performance of black-and-white color image fusion technology in high-brightness environments was solved, achieving high-quality image fusion under different lighting conditions, eliminating black-and-white edge phenomena, and improving image clarity and overall quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing black-and-white color image fusion technology can improve image quality in low-light environments, but its effect is limited in high-brightness environments, and black-and-white edges often appear in the fused image, affecting image quality.
A frequency-band fusion method is adopted to decompose grayscale and color images into sub-images of multiple frequency bands. The frequency band characteristics are used to perform fusion mask calculation, and a non-edge-preserving smoothing filter is combined to eliminate black and white edges, adjust the brightness consistency of the image, and supplement the information of occluded areas.
It can significantly improve the quality of color-blended images in both low-light and high-light environments, eliminate black and white borders, and improve image clarity and overall quality.
Smart Images

Figure CN116263947B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image fusion method, a computer program product, a storage medium and an electronic device. BACKGROUND
[0002] In a large number of electronic image application fields, people's expectations for high-quality images are getting higher and higher, and the shooting performance of traditional single-camera image acquisition equipment cannot meet people's requirements for image quality, so in recent years, dual-camera image acquisition equipment has appeared.
[0003] The dual-camera image acquisition equipment includes a black-and-white camera and a color camera, and is mainly applied to a dark environment (such as night, indoor), first acquires a gray-scale image and a color image through the black-and-white camera and the color camera respectively, and then uses a black-and-white color image fusion technology to fuse the gray-scale image and the color image to obtain a fused image. The black-and-white color image fusion technology uses the feature that the detail information of the gray-scale image is more prominent than the detail information of the color image in a dark environment, and more uses the brightness information of the gray-scale image and fuses the color information of the color image in the fused image, so that the image quality of the dark color image obtained after fusion is higher than that of the dark color image obtained by the single-camera image acquisition equipment.
[0004] However, the existing black-and-white color image fusion technology is still relatively rough, and the quality of the fused image obtained is limitedly improved compared with the color image before fusion. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide an image fusion method, a computer program product, a storage medium and an electronic device to improve the above technical problems.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides an image fusion method, comprising: obtaining a grayscale image and a color image to be fused; wherein the color image comprises a channel image of a luminance channel and a channel image of a chroma channel; decomposing M frames of grayscale sub-images from the grayscale image, and decomposing M frames of luminance sub-images from the channel image of the luminance channel; wherein M is an integer greater than 1, the M frames of grayscale sub-images represent image information of the grayscale image in corresponding M frequency bands, and the M frames of luminance sub-images represent image information of the channel image of the luminance channel in the M frequency bands; fusing the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fused sub-images; wherein each frame of grayscale sub-image is used for fusion with one frame of luminance sub-image of a corresponding frequency band; superimposing the M frames of fused sub-images to obtain a luminance fused image; and combining the luminance fused image and the channel image of the chroma channel to obtain a color fused image.
[0008] The above method fuses the color image and the corresponding grayscale image in frequency bands. Since the sub-images in different frequency bands have different characteristics (for example, the sub-images in a high frequency band represent image details, and the sub-images in a low frequency band represent the overall brightness and contrast of the image), the fusion method in frequency bands makes the image fusion process more precise, and improves the quality of the color fused image.
[0009] In an implementation form of the first aspect, the fusing the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fused sub-images comprises: calculating M frames of fusion masks corresponding to the M frequency bands according to the grayscale sub-images and the luminance sub-images; wherein a pixel value in the fusion mask represents a fusion weight of a pixel value at a same position in a corresponding luminance sub-image and a corresponding grayscale sub-image; and performing weighted fusion on the M frames of grayscale sub-images and the M frames of luminance sub-images by using the M frames of fusion masks to obtain the M frames of fused sub-images; wherein each frame of fusion mask is used for fusion of a corresponding frame of grayscale sub-image and a corresponding frame of luminance sub-image.
[0010] Existing black-and-white color image fusion technologies are basically only applicable to dark light environments. By taking advantage of the feature that the detail information of a grayscale image is more prominent than the detail information of a color image in a dark light environment, the luminance information of the grayscale image is more used in a fused image, and the color information of the color image is fused, so that the image quality of a dark light color image obtained after fusion is higher than that of a dark light color image obtained by a single camera image acquisition device. However, for a high light environment, the detail information of the grayscale image is not absolutely more prominent than the detail information of the color image, and therefore the existing black-and-white color image fusion technology is not applicable to the high light environment.
[0011] In the implementation manner, the fusion mask is calculated according to the gray-scale sub-image and the brightness sub-image, that is, the pixel value (representing the fusion weight) in the fusion mask is adaptively determined by the content of the gray-scale sub-image and the brightness sub-image, so that the fusion according to the fusion mask does not necessarily tend to use more information in the gray-scale sub-image, thereby the implementation manner can obtain a high-quality color fusion image in a high-light environment, and the technical blank of the prior art is made up. Of course, for a dark-light environment, the implementation manner can effectively improve the image fusion quality relative to the prior art due to the frequency-division fusion.
[0012] In an implementation manner of the first aspect, M=3, and the M frequency bands are a high frequency band, a middle frequency band and a low frequency band respectively.
[0013] The implementation manner gives a simple and effective way of dividing frequency bands, and the number of frequency bands is neither too many nor too few, and the decomposed sub-images can well express the characteristics of the image in the frequency domain: the parts in the high frequency band and the middle frequency band represent the details of the image (the high frequency band is small details, and the middle frequency band is larger details), and determine the definition of the image; the part in the low frequency band represents the overall brightness change of the image, and cannot determine the definition of the image. Therefore, different fusion strategies can be adopted for the sub-images according to the characteristics of the three frequency bands.
[0014] In an implementation manner of the first aspect, the calculating the M fusion masks corresponding to the M frequency bands according to the gray-scale sub-image and the brightness sub-image comprises: calculating a fusion mask of the middle frequency band according to the gray-scale sub-image of the middle frequency band and the brightness sub-image of the middle frequency band; calculating a fusion mask of the high frequency band according to the gray-scale sub-image of the high frequency band and the brightness sub-image of the high frequency band; and calculating a fusion mask of the low frequency band according to the fusion mask of the middle frequency band and the fusion mask of the high frequency band.
[0015] The implementation manner gives a calculation manner of the fusion masks of different frequency bands. The sub-images of the high frequency band and the middle frequency band represent the details of the image, and determine the definition of the image, so the fusion masks of the high frequency band and the middle frequency band can be directly calculated according to the pixel value in the sub-image. The sub-image of the low frequency band does not contain image details, and cannot determine the definition of the image, so the fusion mask of the low frequency band calculated according to the sub-image of the low frequency band is not accurate enough.
[0016] In an implementation form of the first aspect, the fusion mask is a binary image, if a pixel value in the fusion mask takes a first numerical value, a pixel value of a corresponding fusion sub-image at a same position adopts a pixel value in a corresponding gray sub-image, and if the pixel value in the fusion mask takes a second numerical value, the pixel value of the corresponding fusion sub-image at the same position adopts a pixel value in a corresponding brightness sub-image; and the calculating the fusion mask of the low frequency band according to the fusion mask of the medium frequency band and the fusion mask of the high frequency band comprises: if pixel values of the fusion mask of the medium frequency band and the fusion mask of the high frequency band at a same position both take the first numerical value, setting a pixel value of the fusion mask of the low frequency band at the position as the first numerical value, otherwise setting the pixel value of the fusion mask of the low frequency band at the position as the second numerical value.
[0017] The implementation form above gives a calculation manner of the fusion mask of the low frequency band when the fusion mask is a binary image.
[0018] In an implementation form of the first aspect, the calculating the fusion mask of the medium frequency band according to the gray sub-image of the medium frequency band and the brightness sub-image of the medium frequency band comprises: taking absolute values of pixel values in the gray sub-image of the medium frequency band to obtain a first mask calculation image; and taking absolute values of pixel values in the brightness sub-image of the medium frequency band to obtain a second mask calculation image; and calculating the fusion mask of the medium frequency band according to a size relationship between pixel values at a same position of the first mask calculation image and the second mask calculation image.
[0019] In the implementation form above, for the sub-images of the medium frequency band, the pixel value size represents the detail information (i.e. the definition) in the original image, so the fusion mask of the medium frequency band can be calculated directly by the size relationship between the pixel values at the same position of the mask calculation images. Meanwhile, when the gray image and the color image are decomposed, the pixel values in the sub-images generated by the decomposition may appear as negative numbers, so the pixel values in the sub-images need to be taken as absolute values before the pixel value comparison, to avoid errors in the pixel value size relationship judgment due to the negative numbers.
[0020] In an implementation form of the first aspect, the calculating the fusion mask of the medium frequency band according to the gray sub-image of the medium frequency band and the brightness sub-image of the medium frequency band comprises: taking absolute values of pixel values in the gray sub-image of the medium frequency band to obtain a first mask calculation image; and taking absolute values of pixel values in the brightness sub-image of the medium frequency band to obtain a second mask calculation image; filtering the first mask calculation image and the second mask calculation image respectively by using a non-edge-preserving smoothing filter to obtain a third mask calculation image and a fourth mask calculation image; and calculating the fusion mask of the medium frequency band according to a size relationship between pixel values at a same position of the third mask calculation image and the fourth mask calculation image.
[0021] The inventors have found that black-white edges are prone to occur near some edges of color images, while gray images do not have this problem. In the fusion results of the prior art, these black-white edges are retained, which seriously affects the image quality. In the implementation mode described above, the non-edge-preserving smoothing filter can widen the edges in the image, and the widened edges cover the areas where the black-white edges are located, so that when image fusion is performed in these areas, there is a high probability that the image information of the gray image containing no black-white edges will be used instead of the image information of the color image, so that the black-white edges in the color fusion image are eliminated or at least weakened.
[0022] In an implementation mode of the first aspect, the fusion mask is a binary image, if a pixel value in the fusion mask takes a first numerical value, a pixel value of a corresponding fusion sub-image at the same position adopts a pixel value in a corresponding gray sub-image, and if the pixel value in the fusion mask takes a second numerical value, the pixel value of the corresponding fusion sub-image at the same position adopts a pixel value in a corresponding luminance sub-image; the calculation of the size relationship between the pixel values of the third mask calculation image and the fourth mask calculation image at the same position, and the calculation of the fusion mask of the middle frequency band, comprises: for any one pixel position in the fusion mask of the middle frequency band, if the pixel value of the third mask calculation image at the position is greater than the product of the pixel value of the fourth mask calculation image at the position and an adjustment threshold, the pixel value of the fusion mask of the middle frequency band at the position is set to the first numerical value, otherwise the pixel value of the fusion mask of the middle frequency band at the position is set to the second numerical value.
[0023] In the implementation mode described above, considering that the quality of imaging of gray images and color images in different environments is different, the direct comparison of pixel values may not be accurate, so the adjustment threshold can be set to reflect the difference in the image acquisition environment, and then the fusion mask is calculated by comparing the adjusted pixel value of the luminance sub-image with the pixel value of the gray sub-image, so that the calculation accuracy of the fusion mask is higher. In an alternative, the adjustment threshold can also be set for the gray sub-image.
[0024] In an implementation form of the first aspect, after the gray image and the color image to be fused are obtained, and before the M frames of gray sub-images are decomposed from the gray image, the method further comprises: adjusting the brightness of the gray image to be consistent with the color image; and the calculating the M frames of fusion masks corresponding to the M frequency bands according to the gray sub-images and the brightness sub-images comprises: calculating the fusion mask of the middle frequency band according to the gray sub-image of the low frequency band and the brightness sub-image of the low frequency band; calculating the fusion mask of the middle frequency band according to the gray sub-image of the middle frequency band and the brightness sub-image of the middle frequency band; and calculating the fusion mask of the high frequency band according to the gray sub-image of the high frequency band and the brightness sub-image of the high frequency band.
[0025] In the implementation form described above, the color image is taken as the fusion reference, that is, the gray image is fused to the color image, so the brightness of the gray image is adjusted to the same level as the color image, and then the fusion masks of the high, middle and low frequency bands are calculated respectively by using the gray sub-images and the brightness sub-images of the high, middle and low frequency bands, so that the calculation precision decline (especially for the calculation of the fusion mask of the low frequency band) caused by the inconsistency of the brightness of the gray image and the color image can be improved.
[0026] In an implementation form of the first aspect, the decomposing the M frames of gray sub-images from the gray image comprises: filtering the gray image by using a first low-pass filter to obtain the gray sub-image of the low frequency band; wherein the cutoff frequency of the first low-pass filter is the demarcation line between the low frequency band and the middle frequency band; filtering the gray image by using a second low-pass filter to obtain a temporary gray image, and calculating the gray sub-image of the middle frequency band according to the gray image and the temporary gray image; wherein the cutoff frequency of the second low-pass filter is the demarcation line between the middle frequency band and the high frequency band; and calculating the gray sub-image of the middle frequency band according to the gray image, the gray sub-image of the low frequency band and the gray sub-image of the high frequency band.
[0027] In the implementation form described above, the gray image is first filtered by using the first low-pass filter to obtain the gray sub-image of the low frequency band, and then the gray image is filtered by using the second low-pass filter to obtain the temporary gray image, and the temporary gray image is equivalent to the gray sub-image containing the middle and low frequency bands, so that the gray sub-image of the high frequency band can be calculated according to the gray image and the temporary gray image, and finally the gray sub-image of the middle frequency band can be calculated by using the gray image, the gray sub-image of the low frequency band and the gray sub-image of the high frequency band, so that this image decomposition method is simple and efficient.
[0028] In an implementation form of the first aspect, the decomposing the M frames of grayscale sub-images from the grayscale image comprises: calculating an occlusion region mask according to the grayscale image and the channel image of the luminance channel; wherein a pixel value in the occlusion region mask represents a probability that a position corresponding to the pixel value belongs to an occlusion region, the occlusion region being a region in the captured scene that only exists in the color image but not in the grayscale image; performing weighted fusion on the grayscale image and the channel image of the luminance channel by using the occlusion region mask to obtain a grayscale occlusion fusion image; and decomposing the M frames of grayscale sub-images from the grayscale occlusion fusion image.
[0029] For the case that the grayscale image and the color image are captured by a black-and-white camera and a color camera of the same electronic device respectively, since the positions of the black-and-white camera and the color camera are different, there is a disparity between the grayscale image and the color image, that is, there is a part of the region in the captured scene that only exists in the color image but not in the grayscale image (or in other words, this part of the region is occluded in the grayscale image), and there is another part of the region that only exists in the grayscale image but not in the color image (or in other words, this part of the region is occluded in the color image).
[0030] If the color image is taken as the fusion reference, the region in the captured scene that only exists in the color image but not in the grayscale image can be defined as the occlusion region, the image information of the grayscale image in the occlusion region is supplemented by using the image information of the color image in the occlusion region, and then the color image and the grayscale image after the supplement (i.e., the grayscale occlusion fusion image) are fused, so that the scene regions included by the color image and the grayscale image are consistent, and thus the subsequent image fusion effect can be improved.
[0031] In the implementation form described above, the pixel value in the occlusion region mask not only describes the position of the occlusion region, but also can be used as a fusion weight to supplement the image information of the grayscale image in the occlusion region.
[0032] In a second aspect, an embodiment of the present application provides an image fusion device, comprising: an image acquisition component, configured to acquire a grayscale image and a color image to be fused; wherein the color image comprises a channel image of a luminance channel and a channel image of a chroma channel; an image decomposition component, configured to decompose M frames of grayscale sub-images from the grayscale image, and decompose M frames of luminance sub-images from the channel image of the luminance channel; wherein M is an integer greater than 1, the M frames of grayscale sub-images represent image information of the grayscale image in corresponding M frequency bands, and the M frames of luminance sub-images represent image information of the channel image of the luminance channel in the M frequency bands; an image fusion component, configured to fuse the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fused sub-images; wherein each frame of grayscale sub-image is used to fuse one frame of luminance sub-image of a corresponding frequency band; and an image superposition component, configured to superimpose the M frames of fused sub-images to obtain a luminance fused image; and a channel splicing component, further configured to combine the luminance fused image and the channel image of the chroma channel to obtain a color fused image.
[0033] In a third aspect, an embodiment of the present application provides a computer program product, comprising computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation manner of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer program instructions, and the computer program instructions, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation manner of the first aspect.
[0035] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores computer program instructions, and the computer program instructions, when read and executed by the processor, perform the method provided in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0037] Figure 1 The steps of the image fusion method provided by the embodiments of the present application are shown;
[0038] Figure 2 The steps of the image fusion method provided by the embodiments of the present application are shown;Figure 1 the step S120 in the method 100 can include the following sub-steps:
[0039] Figure 3 A schematic diagram of the decomposition process of a grayscale image is shown.
[0040] Figure 4 A comparison diagram of a color image, a grayscale image and a color fusion image is shown.
[0041] Figure 5 A functional component included in the image fusion device provided by the embodiment of the application is shown.
[0042] Figure 6 A structure of the electronic device provided by the embodiment of the application is shown. DETAILED DESCRIPTION
[0043] In recent years, important progress has been made in the research of computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence. Artificial intelligence (AI for short) is a new science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks and many other technology categories. Computer vision, as an important branch of artificial intelligence, is specifically to let machines recognize the world. Computer vision technology usually includes face recognition, live detection, fingerprint recognition and anti-forgery verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, character recognition, video processing, video content recognition, behavior recognition, three-dimensional reconstruction, virtual reality, augmented reality, simultaneous localization and mapping, computational photography, robot navigation and positioning, and other technologies. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security, city management, traffic management, building management, park management, face passage, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile imaging, cloud services, smart home, wearable devices, driverless vehicles, autonomous driving, intelligent medical care, face payment, face unlocking, fingerprint unlocking, face and certificate verification, smart screens, smart televisions, cameras, mobile Internet, network live broadcast, beauty, makeup, medical cosmetology, intelligent temperature measurement and other fields. The image fusion method in the embodiment of the application also utilizes image processing and other technologies.
[0044] Black and white color image fusion technology is widely used in dual-camera image acquisition devices, but the inventors have found through long-term research that although the existing black and white color image fusion technology can improve the quality of images in dark light environments, the technology still has many defects, for example:
[0045] First, the prior art generally directly fuses color images and grayscale images, and the quality of the obtained fused image is improved very limitedly relative to the color image before fusion;
[0046] Second, the prior art mainly utilizes the characteristic that the quality of a grayscale image is higher than that of a color image in a dark environment to design a fusion strategy, and the quality of each region in the grayscale image is not necessarily higher than that of the color image in a high-light environment, so the original fusion strategy cannot be directly applied to the high-light environment;
[0047] Third, the fusion result of the prior art has a black-and-white edge at the edge of the image, and a specific example is referred to Figure 4 .
[0048] The image fusion method provided by the embodiments of the present application improves the above technical defects by performing frequency band fusion, calculating a fusion mask and other technical means. It should be understood that, in addition to the new scheme proposed by the embodiments of the present application, the above defects existing in the prior art are also the conclusion of the inventors after practice and careful study, and therefore should also be regarded as the contribution of the inventors in the process of invention.
[0049] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0050] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms “first”, “second” and the like are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0051] The terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. The terms "an embodiment", "one embodiment", or the like, do not necessarily refer to the same embodiment or embodiments.
[0052] Figure 1 Steps of the image fusion method provided by the embodiments of the present application are shown, which can be executed by the image fusion device provided by the embodiments of the present application, but are not limited thereto. Figure 6 The electronic device shown is executed, and the structure of the electronic device can be referred to the description of the electronic device in the following. Figure 6 Referring to Figure 1 , the image fusion method comprises:
[0053] Step S110: Obtain a gray-scale image and a color image to be fused.
[0054] The image to be fused refers to the image that needs to be fused with each other. The image to be fused can include a frame of gray-scale image and a frame of color image. The target of fusion is to calculate a new frame of color image according to the image to be fused, and the image quality of the new frame of color image is higher than that of the color image before fusion. The color image includes a channel image of a luminance channel and a channel image of a chroma channel. For example, the directly obtained image is a YUV image, which is the "color image" required by step S110. The Y channel is the luminance channel, and the UV channel is the chroma channel. For another example, the directly obtained image is an RGB image. Since the RGB image does not have a luminance channel, it is not the "color image" required by step S110. The RGB image needs to be converted into a YUV image to obtain the "color image" required by step S110. The gray-scale image has only one channel, which can be considered as having only a luminance channel. Therefore, the gray-scale image should be fused with the channel image of the luminance channel of the color image.
[0055] As to how to obtain the to-be-fused images satisfying the above requirements, the present application is not limited: for example, the to-be-fused images can be real-time images received from a camera of an electronic device (such as a mobile phone, a wearable device, a robot, etc., which does not necessarily have to be the same device as the electronic device performing step S110), the camera of the electronic device including a black-and-white camera (which can be a natural light camera or an infrared camera) and a color camera, and the black-and-white camera and the color camera are used to obtain a grayscale image and a color image, respectively; for another example, the to-be-fused images can be color images and grayscale images previously saved in a memory of the electronic device, etc. In addition, it should be understood that the color image and the grayscale image are directed to the same scene in terms of content, otherwise it is meaningless to fuse them. Note that although it can be known from the following content that part of the implementation of the image fusion method in the embodiments of the present application has good performance in a high-light environment, the present application is not limited to the lighting conditions for collecting the grayscale image and the color image: for example, the color image and the grayscale image can be taken in a dark environment, or the color image and the grayscale image can be taken in a high-light environment, etc.
[0056] Step S120: decompose M frames of grayscale sub-images from the grayscale image, and decompose M frames of luminance sub-images from the channel image of the luminance channel.
[0057] Wherein, M is an integer greater than 1, the M frames of grayscale sub-images represent image information of the grayscale image in the corresponding M frequency bands, and the M frames of luminance sub-images represent image information of the channel image of the luminance channel in the M frequency bands. The present application is not limited to the value of M: for example, if the value of M is 3, the M frequency bands can be a high frequency band, a medium frequency band and a low frequency band; for another example, if the value of M is 2, the M frequency bands can be a high frequency band and a low frequency band. Those skilled in the art can determine the value of M according to actual conditions.
[0058] It should be understood that the image in the spatial domain can be converted into the frequency domain, and the "decomposition" in step S120 is to split the complete image signal (for example, the grayscale image or the channel image of the luminance channel) in the frequency domain into M sub-signals (for example, the grayscale sub-image or the luminance sub-image). Wherein, the M sub-signals occupy different frequency bands in the frequency domain, and the frequency ranges of the M frequency bands can have overlapping parts or no overlapping parts, but after combining these frequency bands together, they can at least cover the distribution range of the complete image signal in the frequency domain. It should be noted that the decomposition of the complete image signal does not mean that the original image signal does not exist after the decomposition, and the original image signal can exist or can be retained.
[0059] For simplicity, the case where the frequency range of M frequency bands has no overlapping part and the M frequency bands together cover the whole frequency band (covering the whole frequency band necessarily covers the whole frequency range of the image signal) is taken as an example to illustrate how to decompose the whole image signal in the frequency domain into sub-signals in the M frequency bands.
[0060] In an implementation, M-1 frequency points separating adjacent frequency bands can be determined first, for example, M=3, then two different frequency points a and b can be determined in the frequency domain, and a>b. Then, the whole frequency band can be further divided into M frequency bands according to the M-1 frequency points, for example, three frequency bands can be divided according to a and b, a high frequency band with a frequency greater than a, a middle frequency band with a frequency between a and b, and a low frequency band with a frequency less than b (the relationship between the two points a and b and the frequency bands is not important). Finally, a filter can be designed according to the frequency band division, and the whole image signal is filtered to decompose the sub-signals in the corresponding frequency bands, for example, a low-pass filter with a cutoff frequency of b can be designed to obtain the sub-signals in the low frequency band, and so on. Of course, the steps of determining the frequency points and the frequency bands can only be considered when designing the filter, and the actual operation on the image is only filtering using the filter.
[0061] It should be understood that the specific values of the frequency points in the above example can be determined by those skilled in the art according to actual conditions, and the present application does not limit this. The high frequency band, the middle frequency band and the low frequency band do not necessarily correspond to fixed frequency ranges, but can be a relative concept (meaning that the frequency range of the high frequency band is greater than that of the middle frequency band, and the frequency range of the middle frequency band is greater than that of the low frequency band).
[0062] The following will further illustrate how to decompose the gray image using the filter according to steps S121-S123 in Figure 2
[0063] Step S121: filtering the gray image using a first low-pass filter to obtain a gray sub-image in the low frequency band.
[0064] The cutoff frequency of the first low-pass filter is the boundary between the low frequency band and the middle frequency band, that is, the frequency point b in the above example. The first low-pass filter allows image signals with a frequency lower than the boundary between the low frequency band and the middle frequency band to pass through, while image signals with a frequency higher than the boundary between the low frequency band and the middle frequency band cannot pass through, thereby obtaining a gray sub-image (denoted as mono_low) in the low frequency band.
[0065] Step S122: filtering the gray image using a second low-pass filter to obtain a temporary gray image, and calculating a gray sub-image in the high frequency band according to the gray image and the temporary gray image.
[0066] The cutoff frequency of the second low-pass filter is the boundary between the medium frequency band and the high frequency band, i.e. the frequency point a in the above example, a > b. The second low-pass filter allows the image signal with a frequency lower than the boundary between the medium frequency band and the high frequency band to pass through, while the image signal with a frequency higher than the boundary between the medium frequency band and the high frequency band cannot pass through, obtaining a temporary image (denoted as mono_temp).
[0067] Since the second low-pass filter retains the image signal with a frequency lower than the boundary between the medium frequency band and the high frequency band, it can be known that the temporary gray scale image retained after the second low-pass filter filtering contains the image information of the gray scale sub-image of the medium frequency band and the image information of the gray scale sub-image of the low frequency band, while the gray scale image contains the image information of the gray scale sub-image of the high frequency band, the image information of the gray scale sub-image of the medium frequency band and the image information of the gray scale sub-image of the low frequency band. Therefore, the gray scale image and the temporary gray scale image can be subtracted to obtain the gray scale sub-image of the high frequency band (denoted as mono_high), for example, mono_high = mono - mono_temp. Here, the subtraction can refer to the subtraction of the pixel values at the corresponding positions in the images.
[0068] Step S123: calculating the gray scale sub-image of the medium frequency band according to the gray scale image, the gray scale sub-image of the low frequency band and the gray scale sub-image of the high frequency band.
[0069] The gray scale sub-image of the medium frequency band (denoted as mono_mid) can be obtained by subtracting the gray scale sub-image of the high frequency band and the gray scale sub-image of the low frequency band from the gray scale image, for example, mono_mid = mono - mono_high - mono_low.
[0070] Please refer to Figure 3 , Figure 3 The above decomposition process of the gray scale image is shown in the schematic diagram, in which the horizontal axis is the frequency (the frequency gradually decreases to the right side), and the horizontal axis also marks two frequency points a and b (the boundary between the high frequency band and the medium frequency band, and the boundary between the medium frequency band and the low frequency band). In the diagram, the frequency response curves of the first low-pass filter (the cutoff frequency is b) and the second low-pass filter (the cutoff frequency is a) are also marked, as well as the signal curve of the gray scale image mono. From the diagram, it can be known that the first low-pass filter retains the image signal with a frequency lower than the boundary between the medium frequency band and the low frequency band, while the second low-pass filter retains the image signal with a frequency lower than the boundary between the medium frequency band and the high frequency band. Figure 3The decomposition process of the low-pass filter on the gray image can be seen as follows: the low-frequency gray sub-image mono_low is obtained by filtering through the first low-pass filter; the temporary gray image (mono_low+mono_mid) is obtained by filtering through the second low-pass filter; then the high-frequency gray sub-image (mono_high=mono-mono_temp) and the medium-frequency gray sub-image (mono_mid=mono-mono_high-mono_low) can be calculated according to the gray image, the temporary gray image and the low-frequency gray sub-image, and this image decomposition method is simple and efficient.
[0071] In an alternative, on the basis of steps S121-S122, the medium-frequency gray sub-image can also be calculated based on the temporary gray image and the low-frequency gray sub-image, without performing step S123, for example, mono_mid=mono_temp-mono_low.
[0072] In addition, the selection of the first low-pass filter and the second low-pass filter in steps S121-S123 is not limited in the present application, for example, a Gaussian filter can be used, or a Butterworth filter (butterworth), a box filter (box filter) and the like can be used.
[0073] Taking the Gaussian filter as an example, the cutoff frequency of the Gaussian filter can be adjusted by adjusting the sigma value of the Gaussian filter, and the larger the sigma value is, the lower the cutoff frequency is. For example, for the first low-pass filter in step S121, sigma=3 can be set, and for the second low-pass filter in step S12, sigma=1 can be set. It should be noted that when designing the filter, the person skilled in the art can also not necessarily know exactly what the cutoff frequency corresponding to the Gaussian filter with sigma=3 is, but only needs to roughly know that this filter can filter out low-frequency signals.
[0074] Steps S121-S123 use low-pass filters to decompose the gray image, and in other alternative schemes, high-pass filters can also be used for image decomposition, for example:
[0075] Step A1: filtering the gray image by using a first high-pass filter to obtain a high-frequency gray sub-image.
[0076] The cutoff frequency of the first high-pass filter is the demarcation line between the medium-frequency band and the high-frequency band. The first high-pass filter allows image signals with a frequency higher than the demarcation line between the medium-frequency band and the high-frequency band to pass through, while image signals with a frequency lower than the demarcation line between the medium-frequency band and the high-frequency band cannot pass through, thereby obtaining the high-frequency gray sub-image.
[0077] Step A2: filter the gray-scale image with a second high-pass filter to obtain a temporary gray-scale image, and calculate the low-frequency gray-scale sub-image according to the gray-scale image and the temporary gray-scale image.
[0078] The cut-off frequency of the second high-pass filter is the demarcation line between the low frequency band and the medium frequency band. The second high-pass filter allows the image signals with a frequency higher than the demarcation line between the low frequency band and the medium frequency band to pass through, while the image signals with a frequency lower than the demarcation line between the low frequency band and the medium frequency band cannot pass through, thereby obtaining the temporary gray-scale image (still denoted as mono_temp for simplicity, but note that the meaning is different from the temporary gray-scale image in step S122).
[0079] Then, subtract the gray-scale image from the first temporary gray-scale image to obtain the low-frequency sub-image, for example, mono_low = mono - mono_temp.
[0080] Step A3: calculate the medium-frequency gray-scale sub-image according to the gray-scale image, the low-frequency gray-scale sub-image, and the high-frequency gray-scale sub-image.
[0081] Similar to step S123 described above, the medium-frequency gray-scale sub-image can be obtained by the formula: mono_mid = mono - mono_high - mono_low.
[0082] Obviously, the decomposition of the gray-scale image can also be performed by using a combination of low-pass filters and high-pass filters at the same time, or other filters (such as band-pass filters) can also be used for the decomposition of the gray-scale image, which will not be described one by one.
[0083] It should be understood that the way of decomposing M frames of luminance sub-images from the channel image of the luminance channel is similar to the way of decomposing M frames of gray-scale sub-images from the gray-scale image, which will not be described here.
[0084] The inventors have found that the sub-images (which can be gray-scale sub-images or luminance sub-images) in different frequency bands have different characteristics, so that different fusion strategies can be adopted for the sub-images in different frequency bands generated by the decomposition, that is, the fusion is performed in different frequency bands (see step S130 for details). This may be better than directly fusing the original image (which can be a gray-scale image or a channel image of a luminance channel). For example, for the case of M = 3, the high-frequency sub-image mainly represents the small details (such as edges and textures) in the original image, the medium-frequency sub-image mainly represents the slightly larger details in the original image, and the low-frequency sub-image mainly represents the places where the changes are gentle in the original image, which determines the overall brightness and contrast of the image. In addition, since dividing into 3 frequency bands will not cause the number of frequency bands to be too large to greatly increase the complexity of calculation, the scheme of the present application is mainly described by taking the case of M = 3 as an example.
[0085] Step S130: fusing the M frames of gray sub-images and the M frames of brightness sub-images to obtain M frames of fused sub-images.
[0086] Each frame of gray sub-image is used to fuse with a frame of brightness sub-image corresponding to a frequency band, i.e. the image fusion is performed in a frequency band manner on the M frequency bands.
[0087] In an alternative solution, step S130 can be performed by calculating M frames of fusion masks and using the M frames of fusion masks to fuse the M frames of gray sub-images and the M frames of brightness sub-images, which specifically includes the following steps:
[0088] Step B1: calculating M frames of fusion masks corresponding to the M frequency bands according to the gray sub-images and the brightness sub-images.
[0089] Step B2: using the M frames of fusion masks to perform weighted fusion on the M frames of gray sub-images and the M frames of brightness sub-images to obtain M frames of fused sub-images.
[0090] The fusion mask is a kind of weight image, i.e. each pixel value in the fusion mask represents a weight. Each frame of fusion mask corresponds to a frequency band and is used to fuse a frame of gray sub-image and a frame of brightness sub-image corresponding to the frequency band.
[0091] Specifically, the size of the fusion mask, the corresponding gray sub-image and the corresponding brightness sub-image is the same, and each pixel value in the fusion mask represents the fusion weight of the pixel values at the same position in the corresponding brightness sub-image and gray sub-image. Thus, the corresponding brightness sub-image and gray sub-image can be weighted fused (weighted summation) pixel by pixel using the fusion mask.
[0092] Optionally, if each pixel value in the fusion mask is a single value, it can only directly represent the fusion weight of one of the brightness sub-image and the gray sub-image, but the fusion weight of the other can be calculated according to the pixel value, so it can also be regarded as that the fusion weights of both are represented by the pixel value in the fusion mask.
[0093] For example, if the pixel value m (value range [0, 1]) at a pixel position in the fusion mask directly represents the fusion weight of the gray sub-image at the pixel position, the fusion weight of the brightness sub-image at the pixel position can be 1-m. At this time, the fusion process can also be expressed by formula (M=3):
[0094] fusion_high = colour_y_high * (1-mask_high) + mono_high * mask_high
[0095] fusion_low = colour_y_low * (1 - mask_low) + mono_low * mask_low
[0096] fusion_low = colour_y_low * (1 - mask_low) + mono_low * mask_low
[0097] wherein fusion_high is the fusion sub-image of the high frequency band, colour_y_high is the luminance sub-image of the high frequency band, mask_high is the fusion mask of the high frequency band (the pixel value in the mask_high is the fusion weight corresponding to mono_high); fusion_mid is the fusion sub-image of the middle frequency band, colour_y_mid is the luminance sub-image of the middle frequency band, mask_mid is the fusion mask of the middle frequency band (the pixel value in the mask_mid is the fusion weight corresponding to mono_mid); fusion_low is the fusion sub-image of the low frequency band, colour_y_low is the luminance sub-image of the low frequency band, mask_low is the fusion mask of the low frequency band (the pixel value in the mask_low is the fusion weight corresponding to mono_low). Note that the above formula is a pixel-by-pixel operation, meaning that for the pixels in the same position in the images participating in the calculation, the weighted sum is performed according to the above formula.
[0098] The pixel value in the fusion mask can take a continuous value, for example, a value in [0, 1]. For example, if the pixel value at a certain pixel position in the fusion mask directly represents the fusion weight of the grayscale sub-image at the pixel position, the probability that the clarity of the grayscale sub-image is higher than that of the luminance sub-image at the pixel position can be taken as the pixel value, meaning that if the grayscale sub-image at the pixel position is clearer, the pixel value of the fusion sub-image at the pixel position should more adopt the information in the grayscale sub-image, otherwise the pixel value of the fusion sub-image at the pixel position should more adopt the information in the luminance sub-image.
[0099] Alternatively, the pixel values in the fusion mask can also take discrete values. For example, the fusion mask can be a binary image, in which the pixel values can only take two values, a first value (e.g., 1) or a second value (e.g., 0), representing different meanings in the weighted fusion. For example, taking 1 means that the pixel value of the corresponding fusion sub-image at the same position as the fusion mask adopts the pixel value in the corresponding gray-scale sub-image (corresponding to the above formula, the fusion weight corresponding to the luminance sub-image is 0, so the luminance sub-image does not affect the weighted result, and the weighted result is completely from the gray-scale sub-image), and taking 0 means that the pixel value of the corresponding fusion sub-image at the same position adopts the pixel value in the corresponding luminance sub-image (corresponding to the above formula, the fusion weight corresponding to the gray-scale sub-image is 0, so the gray-scale sub-image does not affect the weighted result, and the weighted result is completely from the luminance sub-image).
[0100] In particular, since the discrete values 0 and 1 also belong to [0, 1], 0 and 1 can also be regarded as the probability that the clarity of the gray-scale sub-image is higher than that of the luminance sub-image.
[0101] As for how to calculate the M frames of fusion masks corresponding to the M frequency bands according to the gray-scale sub-images and the luminance sub-images, there are at least two cases: the first case is to calculate the M frames of fusion masks using all the gray-scale sub-images and the luminance sub-images; and the second case is to calculate the M frames of fusion masks using part of the gray-scale sub-images and the luminance sub-images. For the specific calculation methods of the above two cases, please refer to the description below.
[0102] It should be understood that there is also a way of not using fusion masks to fuse the M frames of gray-scale sub-images and the M frames of luminance sub-images, which will be described in detail below.
[0103] Step S140: superimposing the M frames of fusion sub-images to obtain a luminance fusion image.
[0104] Since the M frames of gray-scale sub-images and the M frames of luminance sub-images are all decomposed, each frame of gray-scale sub-image and luminance sub-image in them only contains part of the image information of the gray-scale image and the channel image of the luminance channel, and correspondingly, each frame of fusion sub-image after fusion also only contains part of the image information of the luminance fusion image. For example, the gray-scale sub-image of the low frequency band and the luminance sub-image of the low frequency band only contain the image information of the low frequency band of the gray-scale image and the channel image of the luminance channel, and the fusion sub-image of the low frequency band also only contains the image information of the low frequency band of the luminance fusion image.
[0105] Therefore, it is necessary to superimpose the M frames of fusion sub-images to obtain a luminance fusion image containing complete image information. Here, "superimposition" can be understood as an operation of merging different frequency band image information, which can be adding the pixel values of the M frames of fusion sub-images at the corresponding positions to obtain the pixel value of the luminance fusion image at the same position, which can be expressed by the formula:
[0106] fusion = fusion_high + fusion_mid + fusion_low
[0107] wherein fusion is the luminance fusion image, fusion_high is the fusion sub-image of high frequency band, fusion_mid is the fusion sub-image of middle frequency band, and fusion_low is the fusion sub-image of low frequency band. Note that the above formula is a pixel-by-pixel operation.
[0108] The inventors have found that some ways of decomposing the gray-scale image and the luminance channel image can result in the gray-scale sub-image and / or the luminance sub-image having pixel values that exceed the preset pixel value range (e.g., [0, 255]), thus resulting in the luminance fusion image also having pixel values that exceed the range, which does not meet the requirements of some image formats.
[0109] For example, M = 3, in the step S123, the gray-scale sub-image of the middle frequency band can be calculated by the formula mono mid = mono - mono high - mono low, and in the process of subtraction, the pixel value in mono mid can be negative. Further, the fusion sub-image of the middle frequency band is calculated by the formula fusion mid = colour y mid * mask mid + mono mid * (1 - mask mid), and since the pixel value in mono mid can be negative, the pixel value in fusion mid can also be negative. According to the superposition formula above, the luminance fusion image fusion = fusion high + fusion mid + fusion low, since the pixel value in fusion mid can be negative, it can also result in the pixel value in fusion exceeding the pixel value range [0, 255].
[0110] In an implementation, the pixel value range of the luminance fusion image can be controlled within the preset pixel value range by the following method.
[0111] For example, M = 3, and the preset pixel value range is [0, 255], the superposition formula above can be appropriately improved as follows:
[0112] fusion = Min(Max(fusion high + fusion mid + fusion low, 0), 255)
[0113] Wherein, the Max function is used to control the pixel value of the luminance fusion image to be above 0 (including 0), and the Min function is used to control the pixel value of the luminance fusion image to be below 255 (including 255). Note that the above formula is a pixel-by-pixel operation.
[0114] Step S150: combine the luminance fusion image and the channel image of the chroma channel to obtain a color fusion image.
[0115] Since the gray-scale image and the channel image of the luminance channel are both single-channel images, the luminance fusion image generated after the fusion of the two is also a single-channel image, containing only the luminance information of the image, and therefore the luminance fusion image needs to be combined with the channel image of the chroma channel to obtain a color fusion image containing both luminance information and color information. The "combination" here can be understood as channel splicing.
[0116] The image fusion method in Figure 1 provides a way of frequency-band fusion of gray-scale images and color images. Since the sub-images in different frequency bands have different characteristics, the frequency-band fusion makes the image fusion process more precise and improves the quality of the color fusion image.
[0117] In an implementation of the method, the gray-scale image and the color image can also be fused in a frequency-band manner using a fusion mask of different frequency bands. The fusion mask is calculated according to the gray-scale sub-image and the luminance sub-image, that is, the pixel value (representing the fusion weight) in the fusion mask is determined adaptively by the content of the gray-scale sub-image and the luminance sub-image. Therefore, fusion according to the fusion mask does not necessarily tend to use more information in the gray-scale sub-image, so that this implementation can also obtain a color fusion image with high quality in a high-light environment, filling the gap in the prior art. Of course, for a dark-light environment, the above method can also effectively improve the quality of image fusion due to the use of frequency-band fusion.
[0118] Next, based on the above embodiments, how to calculate the M-frame fusion mask when fusing M-frame gray-scale sub-images and M-frame luminance sub-images using a fusion mask is described in detail.
[0119] Taking the case of M = 3 as an example, in an implementation, the fusion masks of the high, medium, and low frequency bands can be calculated by performing steps D1 and D2.
[0120] Step D1: calculate the fusion mask of the medium frequency band according to the gray-scale sub-image of the medium frequency band and the luminance sub-image of the medium frequency band; and calculate the fusion mask of the high frequency band according to the gray-scale sub-image of the high frequency band and the luminance sub-image of the high frequency band.
[0121] First, how to calculate the fusion mask of the medium frequency band is described:
[0122] Step E1: taking absolute values of pixel values in the gray sub-image of the mid-frequency band to obtain a first mask calculation image; and taking absolute values of pixel values in the brightness sub-image of the mid-frequency band to obtain a second mask calculation image.
[0123] According to the foregoing description, it can be known that, when the to-be-fused image is decomposed, the pixel values in the gray sub-image and the brightness sub-image of the mid-frequency band may appear negative, therefore, in order to facilitate the comparison of the pixel values subsequently, the pixel values in the gray sub-image and the brightness sub-image of the mid-frequency band need to be taken absolute values before the fusion mask of the mid-frequency band is calculated, so as to ensure that the pixel values in the image are all converted into non-negative numbers. In addition, the pixel values can also be converted into non-negative numbers through other ways, for example, the pixel values in the gray sub-image of the mid-frequency band and the brightness sub-image of the mid-frequency band can be respectively squared to obtain a first mask calculation image and a second mask calculation image.
[0124] Step E2: calculating the fusion mask of the mid-frequency band according to the size relationship between the pixel values of the first mask calculation image and the second mask calculation image at the same position.
[0125] For example, if the fusion mask of the mid-frequency band is a binary image, that is, the pixel value in the fusion mask of the mid-frequency band takes 1 (one possible value of the first number), then the pixel value of the fusion sub-image of the mid-frequency band at the same position adopts the pixel value in the gray sub-image of the mid-frequency band, if the pixel value in the fusion mask of the mid-frequency band takes 0 (one possible value of the second number), then the pixel value of the fusion sub-image of the mid-frequency band at the same position adopts the pixel value in the brightness sub-image of the mid-frequency band. At this time, the fusion mask of the mid-frequency band can be calculated according to the following rules:
[0126] For the pixel value at a certain pixel position in the first mask calculation image, if the pixel value is greater than the pixel value at the same pixel position in the second mask calculation image, then the pixel value of the fusion mask of the mid-frequency band at the pixel position takes 1; otherwise, if the pixel value is not greater than the pixel value at the same pixel position in the second mask calculation image, then the pixel value of the fusion mask of the mid-frequency band at the pixel position takes 0. This rule can be expressed by the following pseudo code:
[0127]
[0128] Wherein, |mono_mid| is the first mask calculation image, |colour_y_mid| is the second mask calculation image, and the symbol || represents taking absolute value. It is noted that the above pseudo code is executed pixel by pixel, which means that the above pseudo code is executed for the pixels at the same position in the images participating in comparison one by one.
[0129] The principle of such calculation is that the mid-frequency gray sub-image and the mid-frequency brightness sub-image respectively represent the detail information in the channel image of the gray image and the brightness channel, and the first mask calculation image and the second mask calculation image are only the absolute values of the mid-frequency gray sub-image and the mid-frequency brightness sub-image, thus also respectively representing the detail information in the channel image of the gray image and the brightness channel, and these detail information reflects the definition of the channel image of the gray image and the brightness channel. The greater the pixel value is, the more significant the image detail is, and the higher the image definition is.
[0130] For example, at a certain pixel position, if the pixel value in the first mask calculation image is large, it indicates that the definition of the gray image at the pixel position is higher than that of the channel image of the brightness channel, so the pixel value of the fusion mask of the mid-frequency band at this position should be set to 1, indicating that the pixel value in the mid-frequency gray sub-image is selected during fusion. If the pixel value in the second mask calculation image is large, it indicates that the definition of the channel image of the brightness channel at the pixel position is higher than that of the gray image, so the pixel value of the fusion mask of the mid-frequency band at this position should be set to 0, indicating that the pixel value in the mid-frequency brightness sub-image is selected during fusion.
[0131] In an alternative of steps E1-E2, considering that the quality of imaging of the gray image and the color image is different in different environments, the direct comparison of the pixel values may not be accurate, so the difference in the image acquisition environment can be reflected by setting an adjustment threshold, and the pixel value of the brightness sub-image is adjusted according to the adjustment threshold, and then the adjusted result is compared with the pixel value of the gray sub-image to calculate the fusion mask (the pixel value of the gray sub-image can also be adjusted, the method is similar), so that the calculation accuracy of the fusion mask is higher.
[0132] For example, for the pixel value of the first mask calculation image at a certain pixel position, if the pixel value is greater than the product of the pixel value of the second mask calculation image at the position and the adjustment threshold, the pixel value of the fusion mask of the mid-frequency band at the pixel position is 1; otherwise, if the pixel value is not greater than the product of the pixel value of the second mask calculation image at the position and the adjustment threshold, the pixel value of the fusion mask of the mid-frequency band at the pixel position is 0. The adjustment threshold can be a constant, for example, 0.5, 0.6, etc. The adjustment threshold can be a preset value or a temporarily calculated value. The above rule is represented by pseudo code as follows:
[0133]
[0134] wherein thr mid is the adjustment threshold value corresponding to the mid-frequency band, and in particular, since 1 can also be taken as the value of the adjustment threshold value, the above scheme of calculating the fusion mask of the mid-frequency band without setting the adjustment threshold value can be regarded as the case of thr mid = 1 in the above scheme of calculating the fusion mask of the mid-frequency band with setting the adjustment threshold value. Note that the above pseudo code is executed pixel by pixel.
[0135] If the fusion mask of the mid-frequency band is not a binary image, the fusion mask of the mid-frequency band can also be calculated in a manner similar to the above manner, for example, in step E2, the difference between the pixel values of the first mask calculation image and the second mask calculation image at the same position can be calculated, and according to a preset mapping rule, the difference value is mapped to a value in [0, 1] (but not only 0 and 1) as the pixel value of the fusion mask of the mid-frequency band at this position. For example, if the difference between the pixel values of the first mask calculation image and the second mask calculation image at a certain position is 200, it is specified that the pixel value of the fusion mask of the mid-frequency band at this position is 0.9, if the difference between the pixel values of the first mask calculation image and the second mask calculation image at a certain position is 0, it is specified that the pixel value of the fusion mask of the mid-frequency band at this position is 0.5, and so on.
[0136] In addition, if the manner of some decomposition image ensures that the pixel values in the gray sub-image and the brightness sub-image of the mid-frequency band will not appear negative, the pixel values in the gray sub-image and the brightness sub-image of the mid-frequency band can also be directly compared to determine the pixel value in the fusion mask of the mid-frequency band, without the need to calculate the first mask calculation image and the second mask calculation image.
[0137] The calculation manner of the fusion mask of the high-frequency band can refer to the calculation manner of the fusion mask of the mid-frequency band, for example, refer to the above steps E1-E2 or the alternative scheme (setting the adjustment threshold value) of steps E1-E2, and if the calculation manner in the alternative scheme is adopted, it should be noted that the value of the adjustment threshold value corresponding to the high-frequency band can be the same as or different from the value of the adjustment threshold value corresponding to the mid-frequency band.
[0138] Step D2: calculating the fusion mask of the low-frequency band according to the fusion mask of the mid-frequency band and the fusion mask of the high-frequency band.
[0139] For example, if the fusion mask of the mid-frequency band is a binary image, which is defined with reference to the example in step D1, the fusion mask of the low-frequency band can be calculated according to the following rules: if the pixel values of the fusion mask of the mid-frequency band and the fusion mask of the high-frequency band at the same position are both 1, the pixel value of the fusion mask of the low-frequency band at this position is set to 1 (the fusion sub-image of the low-frequency band adopts the pixel value of the gray sub-image of the low-frequency band at this pixel position), otherwise the pixel value of the fusion mask of the low-frequency band at this position is set to 0 (the fusion sub-image of the low-frequency band adopts the pixel value of the brightness sub-image of the low-frequency band at this pixel position).
[0140] As mentioned above, the pixel value size of the sub-images (or the images further calculated based on the sub-images, such as the first mask calculation image) located in the high frequency band and the medium frequency band reflects the image definition, so the pixel value in the fusion mask of the high frequency band and the medium frequency band can be directly determined according to the size relationship between the pixel values (who is clear to take who); and the pixel value size of the sub-images (or the images further calculated based on the sub-images) located in the low frequency band only represents the overall brightness or contrast of the image, and basically does not determine the definition of the image, so it is not accurate to directly calculate the fusion mask of the low frequency band by comparing the pixel values. Therefore, the fusion mask of the low frequency band in step D2 is calculated according to the fusion mask of the medium frequency band and the fusion mask of the high frequency band, rather than directly calculated according to the sub-image of the low frequency band.
[0141] In addition, in the above example, the color image is taken as the fusion reference, that is, the grayscale image is fused into the color image, so only when the image definition of the grayscale image in the high frequency and medium frequency parts is higher than that of the color image (the pixel values of the fusion masks of the medium frequency band and the high frequency band at the same position are both 1), the low frequency part adopts the information of the grayscale image (the pixel value of the fusion mask of the low frequency band at the same position is 1), otherwise, the information of the color image is adopted (the pixel value of the fusion mask of the low frequency band at the same position is 0), that is, the information in the color image is more inclined to be used.
[0142] Of course, this rule can also be adjusted, for example, the following rule can also be set: as long as one of the pixel values of the fusion masks of the medium frequency band and the high frequency band at the same position is 1, the pixel value of the fusion mask of the low frequency band at the position is set to 1, otherwise, the pixel value of the fusion mask of the low frequency band at the position is set to 0.
[0143] In addition, if the grayscale image is taken as the fusion reference, that is, the color image is fused into the grayscale image, a corresponding set of rules can also be set, which will not be described again.
[0144] In an alternative, in order to facilitate the statistics of the pixel values of the fusion masks of the medium frequency band and the high frequency band at the same position, a confidence flag can be set for each frequency band, which takes a Boolean value. If the pixel value of the grayscale sub-image at a certain pixel position is greater than the pixel value of the luminance sub-image at the pixel position, the confidence flag is set to true, otherwise, it is set to false. If the values of the confidence flags of the high frequency band and the medium frequency band are both true, the pixel value of the fusion mask of the low frequency band at the position is set to 1, otherwise, the pixel value of the fusion mask of the low frequency band at the position is set to 0. Taking the calculation method containing the adjustment threshold as an example, the above rule can be represented by the pseudo code as follows:
[0145]
[0146] wherein |mono_high| is an image obtained by taking absolute value of pixel values in the high-frequency sub-image of the monochrome image, |colour_y_high| is an image obtained by taking absolute value of pixel values in the high-frequency sub-image of the luminance sub-image of the color image, flag_mid is the confidence flag of the mid-frequency band, flag_high is the confidence flag of the high-frequency band, thr_mid is the adjustment threshold value corresponding to the mid-frequency band, thr_high is the adjustment threshold value corresponding to the high-frequency band, && represents logical and operation, and no confidence flag is needed for the low-frequency band. Note that the above pseudo code is executed pixel by pixel.
[0147] Further, the inventors find that some color images have white edges before fusion due to the algorithm of the image signal processor (ISP) and the like, and the prior art cannot well process these white edges when performing fusion of black-and-white and color images, so that the white edges still exist in the fusion result, seriously affecting the image quality. The possible solutions are described below mainly taking white edges as an example, and the case of black edges is similar.
[0148] Reference Figure 4 The uppermost is a color image, and there is a white edge at the black vertical bar indicated by the arrow at the upper left corner of the color image; the middle is a monochrome image, and no white edge appears at the same position compared with the color image.
[0149] If the monochrome image and the color image are fused directly according to the manner introduced in steps D1-D2, for the convenience of explaining the principle, it is assumed that the monochrome image is clearer than the color image (this condition is basically met in a dark light environment), that is, the pixel values (taking absolute value) in the mid-high frequency sub-image of the monochrome image are generally larger than those in the luminance sub-image, but for the color fusion image, only in a narrow range at the edge of the black vertical bar, the image information of the monochrome image is probably used (because the edge belongs to mid-high frequency, and this part of the monochrome image is clearer than the color image), and the rest still probably uses the image information of the color image (because the rest belongs to low frequency, and the difference in clarity between the monochrome image and the color image is not obvious), which will cause the color fusion image to still have a white edge at the periphery of the black vertical bar, because the white edge has a certain width in the color image and does not only exist in the narrow range at the edge of the black vertical bar.
[0150] Therefore, still taking the case of the mid-frequency band as an example, in an optional solution, after obtaining the first mask calculation image and the second mask calculation image in step D1, a non-edge-preserving smoothing filter can be used to smooth the first mask calculation image and the second mask calculation image to improve the problem of white edges appearing in the color fusion image, and the result can be referred to as followsFigure 4 The bottom image in the middle row shows that the white border at the black vertical bar is eliminated in the color fusion image. The detailed implementation of this scheme is as follows:
[0151] Step D3 (continue Step D1, do not perform Step D2): filter the first mask calculation image and the second mask calculation image respectively using a non-edge-preserving smoothing filter to obtain a third mask calculation image and a fourth mask calculation image.
[0152] Step D4: calculate the fusion mask of the middle frequency band according to the size relationship between the pixel values at the same position of the third mask calculation image and the fourth mask calculation image.
[0153] The non-edge-preserving smoothing filter is a kind of smoothing filter algorithm, which is different from the edge-preserving smoothing filter algorithm (for example, guided filter, bilateral filter). The non-edge-preserving smoothing filter algorithm does not preserve the edges of the image during filtering, that is, it performs indiscriminate smoothing processing on the image content. For example, box filter and Gaussian filter are both non-edge-preserving smoothing filters.
[0154] The method for calculating the fusion mask of the middle frequency band in Step D4 according to the third mask calculation image and the fourth mask calculation image is similar to the method for calculating the fusion mask of the middle frequency band in Step D2 according to the first mask calculation image and the second mask calculation image. Only the first mask calculation image is replaced by the third mask calculation image and the second mask calculation image is replaced by the fourth mask calculation image. Therefore, it is not repeated here.
[0155] The principle of eliminating the white border by Steps D3-D4 is analyzed as follows:
[0156] The non-edge-preserving smoothing filter can blur the edges in the image, or in other words, can widen the distribution range of the edges in the image. After the edges are widened, although the frequency components are reduced, they still belong to the middle-high frequency band in the image. In Step D3, the edge part of the black vertical bar in the third mask calculation image and the fourth mask calculation image is widened. The widened edge covers the area where the white border is located, and in this part of the widened edge area, the pixel values in the third mask calculation image are generally larger (the gray image is clearer than the color image, which is manifested in that the pixel values in the middle frequency are larger). Therefore, there is a high probability that the image information of the gray image will be used instead of the image information of the color image in the area where the white border is located, so that the white border in the color fusion image is eliminated or at least weakened.
[0157] It should be understood that the fusion mask of the high frequency band can also be calculated in a similar manner to Steps D1, D3, and D4 to eliminate or weaken the black-white border in the color image. Therefore, it is not repeated here.
[0158] The method of calculating the fusion mask introduced above does not use all of the M-frame gray sub-images and M-frame brightness sub-images: for example, the low-frequency gray sub-images and the low-frequency brightness sub-images are not used, and the reason has been analyzed above. However, there is also a scheme of calculating the M-frame fusion mask using all of the M-frame gray sub-images and M-frame brightness sub-images, for example, the low-frequency fusion mask can also not be calculated by the method of step D2, but can be calculated by the same method as that of calculating the high-frequency fusion mask.
[0159] However, it should be noted that since the low-frequency gray sub-images and the low-frequency brightness sub-images do not basically contain the detail information of the image, the pixel value size mainly represents the overall brightness and contrast of the image. Therefore, after step S110 is executed and before step S120 is executed, the brightness of the gray image and the brightness of the color image also need to be adjusted to be consistent, so as to eliminate or weaken the influence of brightness on the pixel value size, so that the result of the pixel value comparison in the low frequency band is more meaningful, thereby making the calculation result of the low-frequency fusion mask more accurate.
[0160] The brightness consistency adjustment here includes at least two cases: the first case is that if the color image is taken as the fusion reference, the brightness of the color image is taken as the reference for brightness adjustment, the brightness of the gray image is adjusted so that the brightness of the gray image is consistent with the brightness of the color image. The second case is that if the gray image is taken as the fusion reference, the brightness of the gray image is taken as the reference for brightness adjustment, the brightness of the color image is adjusted so that the brightness of the color image is consistent with the brightness of the gray image. As to how to adjust the brightness, the present application does not limit it, for example, a histogram matching algorithm can be used to adjust the brightness to be consistent, etc.
[0161] In particular, since only the pixel value in the brightness channel of the color image represents the brightness, the brightness adjustment of the color image can also be performed only on the brightness channel.
[0162] In addition, it should be pointed out that the so-called brightness "consistency" does not mean that the brightness of the two frames of images is exactly the same, but means that the brightness of the two frames of images is basically at the same level, which can be the same or can have a certain difference.
[0163] Some fusion mask calculation methods are introduced above, and the fusion mask is calculated in order to fuse the gray sub-images and the brightness sub-images, but in some optional schemes, the fusion mask can also not be calculated, but the fusion can be directly performed according to the size relationship between the pixel values of the gray sub-images and the brightness sub-images to obtain the M-frame fusion sub-image.
[0164] For example, taking the case of the middle frequency band, the fusion can be performed according to the following rules: if the pixel value in the gray sub-image of the middle frequency band is greater than the pixel value in the luminance sub-image of the middle frequency band at a certain pixel position, the pixel value of the fusion sub-image of the middle frequency band at the pixel position adopts the pixel value of the gray sub-image of the middle frequency band, otherwise, the pixel value of the fusion sub-image of the middle frequency band at the pixel position adopts the pixel value of the luminance sub-image of the middle frequency band, so as to achieve the purpose of selecting the gray sub-image of the middle frequency band and the luminance sub-image of the middle frequency band to obtain a clearer sub-image as the fusion result.
[0165] Further, if the pixel value in the gray sub-image of the middle frequency band is greater than the pixel value in the luminance sub-image of the middle frequency band, the pixel value of a frame of the binary image at the position is set to 1, otherwise, the pixel value of the binary image at the position is set to 0. After the pixel values in the binary image are set, the pixel values are regarded as fusion weights for fusing the gray sub-image of the middle frequency band and the luminance sub-image of the middle frequency band. As can be seen from the foregoing, the binary image is a binary middle frequency band fusion mask. Therefore, the above example of not calculating the fusion mask can achieve the same effect as the case of calculating the fusion mask and the fusion mask being a binary image.
[0166] Next, based on the above embodiments, how to decompose M gray sub-images from the gray image in the case of considering occlusion is introduced as follows:
[0167] Since the positions of the black-and-white camera and the color camera on the electronic device are different, there is a disparity between the gray image and the color image collected by the two cameras, that is, there is a part of the area in the shooting scene that exists only in the color image but does not exist in the gray image (or in other words, this part of the area is occluded in the gray image), and there is another part of the area that exists only in the gray image but does not exist in the color image (or in other words, this part of the area is occluded in the color image). How to calculate the occluded area is not limited in the present application, for example, the LR-Check algorithm can be used to calculate the occluded area.
[0168] If the color image is taken as the fusion reference, the area in the shooting scene that exists only in the color image but does not exist in the gray image can be defined as the occluded area, and the image information of the color image in the occluded area is used to supplement the image information of the gray image in the occluded area, so as to ensure that the color image and the gray image include consistent scene areas, thereby improving the subsequent image fusion effect.
[0169] To achieve the above purpose, in one implementation manner, the decomposition process of the gray image in step S120 can further include the following sub-steps:
[0170] Step C1: calculating an occlusion area mask according to the gray image and the channel image of the luminance channel.
[0171] Step C2: weighted fusion of the gray-scale image and the channel image of the luminance channel using the occlusion region mask to obtain a gray-scale occlusion fusion image.
[0172] Step C3: decompose M gray-scale sub-images from the gray-scale occlusion fusion image.
[0173] The occlusion region mask is an image describing the position of the occlusion region, or in other words, calculating the occlusion region mask can also be considered as calculating the position of the occlusion region. The size of the occlusion region mask is the same as that of the gray-scale image and the color image. Optionally, each pixel value in the occlusion region mask represents the probability (taking a value in [0, 1]) that the position of the pixel value belongs to the occlusion region. In particular, if the occlusion region mask is a binary image, for example, the pixel value only takes 0 and 1, taking 0 means that the current pixel does not belong to the occlusion region, and taking 1 means that the current pixel belongs to the occlusion region, and the set of pixels with value 1 is an accurate description of the position of the occlusion region.
[0174] Further, in order to supplement the image information of the gray-scale image in the occlusion region with the image information of the color image in the occlusion region, the occlusion region mask can also be regarded as a weight image, that is, each pixel value in the occlusion region mask also represents the fusion weight of the channel image of the luminance channel at the pixel position. Thus, the channel image of the luminance channel and the gray-scale image can be weighted fused (weighted summation) using the occlusion region mask to achieve the effect of supplementing the image information of the gray-scale image in the occlusion region with the channel image of the luminance channel.
[0175] The principle can be understood as follows: if the pixel value in the occlusion region mask is calculated reasonably, then in the real occlusion region, the pixel value (fusion weight) in the occlusion region mask is larger (close to or equal to 1), and outside the occlusion region, the pixel value (fusion weight) in the occlusion region mask is smaller (close to or equal to 0), so that in the gray-scale occlusion fusion image obtained after fusion, the image information in the occlusion region will come from or mainly come from the channel image of the luminance channel, and the image information outside the occlusion region will still retain or basically retain the original image information in the gray-scale image, that is, the image information in the occlusion region that the gray-scale image fails to collect is effectively supplemented, thereby ensuring that the scene region included by the channel image of the luminance channel and the gray-scale image is consistent, and further improving the subsequent image fusion effect.
[0176] In the process of weighted fusion, for the fusion weight corresponding to the gray image, although it is not directly represented by the pixel value in the unoccluded region mask, it can be calculated according to the pixel value in the unoccluded region mask. For example, if the pixel value n (value range [0, 1]) at a certain pixel position in the unoccluded region mask represents the fusion weight of the channel image of the luminance channel at the pixel position, then the fusion weight corresponding to the gray image at the pixel position can be 1-n. At this time, the fusion process can also be represented by the formula: mono' = colour_y * n + (1-n) * mono, wherein colour_y is the channel image of the luminance channel, mono is the gray image, and mono' is the gray occlusion fusion image. Note that the above formula is a pixel-by-pixel operation.
[0177] As an alternative, each pixel value in the unoccluded region mask can also represent the probability that the position where the pixel value is located does not belong to the occluded region. For how to use the unoccluded region mask in this scheme, please refer to the foregoing description, which will not be repeated here.
[0178] The above introduces some calculation methods of the unoccluded region mask. The unoccluded region mask is calculated in order to realize the complement of the image information of the gray image in the occluded region by the image information of the color image in the occluded region through the fusion of the channel image of the luminance channel and the gray image. However, in some optional schemes, the unoccluded region mask can also not be calculated, but the fusion of the two can be directly performed according to whether the pixel position belongs to the occluded region to obtain the gray occlusion fusion image.
[0179] For example, the fusion can be performed according to the following rules: if it is judged that a certain pixel position belongs to the occluded region, then the pixel value of the gray occlusion fusion image at the pixel position adopts the pixel value of the channel image of the luminance channel, otherwise the pixel value of the gray image.
[0180] Further, if a certain pixel position belongs to the occluded region, the pixel value of a binary image at the position is set to 1, otherwise the pixel value of the binary image at the position is set to 0. After the pixel values in the binary image are set, the pixel values are regarded as fusion weights for fusing the gray sub-image and the channel image of the luminance channel. As can be seen from the foregoing content, the binary image is a binary unoccluded region fusion mask. Therefore, the above example of not calculating the unoccluded region fusion mask can achieve the same effect as the case of calculating the unoccluded region fusion mask and the unoccluded region fusion mask being a binary image.
[0181] The specific implementation of step C3 can refer to the foregoing description, such as steps S121-S123 or steps A1-A3, which will not be repeated here.
[0182] In another implementation, if the grayscale image is taken as the fusion reference, the area in the shooting scene that only exists in the grayscale image but does not exist in the color image can be defined as an occlusion area, and the image information of the grayscale image in the occlusion area is supplemented to the channel image of the luminance channel, so that the channel image of the luminance channel and the grayscale image are consistent in the scene area.
[0183] At this time, each pixel value in the occlusion area mask represents the probability that the position where the pixel value is located belongs to the occlusion area (definition of the occlusion area is different from the previous one), or the fusion weight corresponding to the grayscale image at the pixel position. The weighted fusion of the grayscale image and the channel image of the luminance channel by using the occlusion area mask can obtain a luminance occlusion fusion image, and then M frames of luminance sub-images are decomposed from the luminance occlusion fusion image to serve as the fusion basis in the subsequent step S130 (fusion of M frames of grayscale sub-images and M frames of luminance sub-images).
[0184] Figure 5 The functional components included in the image fusion device provided by the embodiments of the present application are shown. Referring to Figure 5 , the image fusion device 200 includes:
[0185] The image acquisition component 210 is configured to acquire a grayscale image and a color image to be fused; the color image includes a channel image of a luminance channel and a channel image of a chroma channel.
[0186] The image decomposition component 220 is configured to decompose M frames of grayscale sub-images from the grayscale image, and decompose M frames of luminance sub-images from the channel image of the luminance channel; M is an integer greater than 1, the M frames of grayscale sub-images represent image information of the grayscale image in a corresponding M frequency bands, and the M frames of luminance sub-images represent image information of the channel image of the luminance channel in the M frequency bands.
[0187] The image fusion component 230 is configured to fuse the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fusion sub-images; each frame of grayscale sub-image is used to fuse one frame of luminance sub-image of a corresponding frequency band.
[0188] The image superposition component 240 is configured to superimpose the M frames of fusion sub-images to obtain a luminance fusion image.
[0189] The channel splicing component 250 is further configured to combine the luminance fusion image and the channel image of the chroma channel to obtain a color fusion image.
[0190] In an implementation form of the image fusion device 200, the image fusion component 230 is configured to calculate M fusion masks corresponding to the M frequency bands according to the gray sub-images and the brightness sub-images, wherein a pixel value in the fusion mask represents a fusion weight of pixel values at a same position in a corresponding brightness sub-image and a corresponding gray sub-image; and perform weighted fusion on the M gray sub-images and the M brightness sub-images by using the M fusion masks to obtain the M fused sub-images, wherein each fusion mask is used for fusion of a corresponding gray sub-image and a corresponding brightness sub-image.
[0191] In an implementation form of the image fusion device 200, M = 3, and the M frequency bands are respectively a high frequency band, a middle frequency band and a low frequency band.
[0192] In an implementation form of the image fusion device 200, the image fusion component 230 is configured to calculate a fusion mask of the middle frequency band according to a gray sub-image of the middle frequency band and a brightness sub-image of the middle frequency band, and calculate a fusion mask of the high frequency band according to a gray sub-image of the high frequency band and a brightness sub-image of the high frequency band, and calculate a fusion mask of the low frequency band according to the fusion mask of the middle frequency band and the fusion mask of the high frequency band.
[0193] In an implementation form of the image fusion device 200, the fusion mask is a binary image, if a pixel value in the fusion mask takes a first value, a pixel value at a same position in a corresponding fused sub-image adopts a pixel value in a corresponding gray sub-image, and if the pixel value in the fusion mask takes a second value, the pixel value at the same position in the corresponding fused sub-image adopts a pixel value in a corresponding brightness sub-image; and the image fusion component 230 is further configured to, if pixel values at the same position in the fusion mask of the middle frequency band and the fusion mask of the high frequency band both take the first value, set a pixel value of the fusion mask of the low frequency band at the position to the first value, or set the pixel value of the fusion mask of the low frequency band at the position to the second value.
[0194] In an implementation form of the image fusion device 200, the image fusion component 230 is configured to take absolute values of pixel values in the gray sub-image of the middle frequency band to obtain a first mask calculation image, and take absolute values of pixel values in the brightness sub-image of the middle frequency band to obtain a second mask calculation image, and calculate the fusion mask of the middle frequency band according to a size relationship between pixel values at a same position in the first mask calculation image and the second mask calculation image.
[0195] In an implementation of the image fusion apparatus 200, the image fusion component 230 takes absolute values of pixel values in the gray-scale sub-images in the middle frequency band to obtain a first mask calculation image; and takes absolute values of pixel values in the brightness sub-images in the middle frequency band to obtain a second mask calculation image; filters the first mask calculation image and the second mask calculation image respectively by using a non-edge-preserving smoothing filter to obtain a third mask calculation image and a fourth mask calculation image; and calculates the fusion mask of the middle frequency band according to a size relationship between pixel values at the same positions in the third mask calculation image and the fourth mask calculation image.
[0196] In an implementation of the image fusion apparatus 200, the fusion mask is a binary image, if a pixel value in the fusion mask takes a first numerical value, a pixel value at the same position in a corresponding fusion sub-image adopts a pixel value in a corresponding gray-scale sub-image, and if the pixel value in the fusion mask takes a second numerical value, the pixel value at the same position in the corresponding fusion sub-image adopts a pixel value in a corresponding brightness sub-image; the image fusion component 230 is further configured to, for any pixel position in the fusion mask of the middle frequency band, if a pixel value of the third mask calculation image at the position is greater than a product of a pixel value of the fourth mask calculation image at the position and an adjustment threshold, set a pixel value of the fusion mask of the middle frequency band at the position to the first numerical value, otherwise set the pixel value of the fusion mask of the middle frequency band at the position to the second numerical value.
[0197] In an implementation of the image fusion apparatus 200, the image fusion apparatus 200 further comprises: a brightness adjustment component configured to, after obtaining the gray-scale image and the color image to be fused, and before decomposing M frames of gray-scale sub-images from the gray-scale image, adjust brightness of the gray-scale image to be consistent with the color image; the image fusion component 230 is configured to calculate the fusion mask of the middle frequency band according to the gray-scale sub-images in the low frequency band and the brightness sub-images in the low frequency band, calculate the fusion mask of the middle frequency band according to the gray-scale sub-images in the middle frequency band and the brightness sub-images in the middle frequency band, and calculate the fusion mask of the high frequency band according to the gray-scale sub-images in the high frequency band and the brightness sub-images in the high frequency band.
[0198] In an implementation manner of the image fusion apparatus 200, the image decomposition component 220 filters the gray-scale image by using a first low-pass filter to obtain a low-frequency gray-scale sub-image; a cutoff frequency of the first low-pass filter is a boundary between a low-frequency band and a medium-frequency band; filters the gray-scale image by using a second low-pass filter to obtain a temporary gray-scale image, and calculates a high-frequency gray-scale sub-image according to the gray-scale image and the temporary gray-scale image; a cutoff frequency of the second low-pass filter is a boundary between the medium-frequency band and the high-frequency band; and a medium-frequency gray-scale sub-image is calculated according to the gray-scale image, the low-frequency gray-scale sub-image and the high-frequency gray-scale sub-image.
[0199] In an implementation manner of the image fusion apparatus 200, the image decomposition component 220 calculates an occlusion region mask according to the gray-scale image and the channel image of the luminance channel; a pixel value in the occlusion region mask represents a probability that a position of the pixel value belongs to an occlusion region; the occlusion region refers to a region in a shooting scene that exists only in the color image but does not exist in the gray-scale image; the gray-scale image and the channel image of the luminance channel are fused by using the occlusion region mask to obtain a gray-scale occlusion fusion image; and the M frames of gray-scale sub-images are decomposed from the gray-scale occlusion fusion image.
[0200] The image fusion apparatus 200 provided by the embodiments of the present application has the implementation principle and the technical effects introduced in the foregoing method embodiments, and for brief description, the part not mentioned in the device embodiments can refer to the corresponding content in the method embodiments.
[0201] Figure 6 The structure of the electronic device 300 provided by the embodiments of the present application is shown. Referring to Figure 6 , the electronic device 300 includes a processor 310, a memory 320 and a communication interface 330, and these components are interconnected and communicate with each other through a communication bus 340 and / or other forms of connection mechanism (not shown).
[0202] The processor 310 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with the processing capability of signals. The processor 310 described above can be a general-purpose processor, including a central processing unit (CPU), a micro controller unit (MCU), a network processor (NP) or other conventional processors; it can also be a special-purpose processor, including a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. And when the processor 310 is multiple, part of them can be general-purpose processors, and the other part can be special-purpose processors.
[0203] The memory 320 includes one or more (only one is shown in the figure), which can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM) and the like.
[0204] The processor 310 and other possible components can access the memory 320 to read and / or write data therein. In particular, one or more computer program instructions can be stored in the memory 320, and the processor 310 can read and run these computer program instructions to implement the image fusion method provided by the embodiments of the present application.
[0205] The communication interface 330 includes one or more (only one is shown in the figure) and can be used for direct or indirect communication with other devices to interact data. The communication interface 330 can include an interface for wired and / or wireless communication.
[0206] It can be understood that, Figure 6 The structure shown is only schematic, and the electronic device 300 can further include more or less components than those shown in the figure, or have a different configuration from that shown in the figure. For example, the electronic device 300 can further include a camera (a color camera and a black-and-white camera) for taking images or videos, and the taken images or frames in the videos can be used as the to-be-fused images in step S110; for another example, if the electronic device 300 does not need to communicate with other devices, the communication interface 330 can also not be provided. Figure 6 Figure 6 The components shown in the figure can be implemented in hardware, software or a combination thereof. The electronic device 300 can be a physical device such as a mobile phone, a wearable device, a video camera, a camera, a PC, a notebook computer, a tablet computer, a server, a robot, etc., or a virtual device such as a virtual machine, a container, etc. In addition, the electronic device 300 is not limited to a single device, but can also be a combination of multiple devices or a cluster of a large number of devices.
[0207] Figure 6 The electronic device 300 can be a physical device such as a mobile phone, a wearable device, a video camera, a camera, a PC, a notebook computer, a tablet computer, a server, a robot, etc., or a virtual device such as a virtual machine, a container, etc. In addition, the electronic device 300 is not limited to a single device, but can also be a combination of multiple devices or a cluster of a large number of devices.
[0208] The embodiments of the present application also provide a computer readable storage medium, which stores computer program instructions. When the computer program instructions are read and run by a processor, the image fusion method provided by the embodiments of the present application is executed. For example, the computer readable storage medium can be implemented as Figure 6 the memory 320 in the electronic device 300 in the figure.
[0209] The embodiments of the present application also provide a computer program product, which includes computer program instructions. When the computer program instructions are read and run by a processor, the image fusion method provided by the embodiments of the present application is executed.
[0210] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An image fusion method, characterized by, The method comprises the following steps: obtaining a grayscale image and a color image to be fused; wherein the color image comprises a channel image of a luminance channel and a channel image of a chroma channel; decomposing M frames of grayscale sub-images from the grayscale image, and decomposing M frames of luminance sub-images from the channel image of the luminance channel; wherein M is an integer greater than 1, the M frames of grayscale sub-images represent image information of the grayscale image in corresponding M frequency bands, and the M frames of luminance sub-images represent image information of the channel image of the luminance channel in the M frequency bands; fusing the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fused sub-images; wherein each frame of grayscale sub-image is used for fusing with one frame of luminance sub-image of a corresponding frequency band; superimposing the M frames of fused sub-images to obtain a luminance fused image; combining the luminance fused image and the channel image of the chroma channel to obtain a color fused image.
2. The image fusion method of claim 1, wherein, The step of fusing the M frames of grayscale sub-images and the M frames of luminance sub-images to obtain M frames of fused sub-images comprises the following steps: calculating M frames of fusion masks corresponding to the M frequency bands according to the grayscale sub-images and the luminance sub-images; wherein a pixel value in the fusion mask represents a fusion weight of a pixel value at a same position in a corresponding luminance sub-image and a corresponding grayscale sub-image; performing weighted fusion on the M frames of grayscale sub-images and the M frames of luminance sub-images by using the M frames of fusion masks to obtain the M frames of fused sub-images; wherein each frame of fusion mask is used for fusion of a corresponding frame of grayscale sub-image and a corresponding frame of luminance sub-image.
3. The image fusion method of claim 2, wherein, M=3, and the M frequency bands are high, medium and low frequency bands respectively.
4. The image fusion method of claim 3, wherein, The step of calculating M frames of fusion masks corresponding to the M frequency bands according to the grayscale sub-images and the luminance sub-images comprises the following steps: calculating a fusion mask of the medium frequency band according to the grayscale sub-image of the medium frequency band and the luminance sub-image of the medium frequency band, and calculating a fusion mask of the high frequency band according to the grayscale sub-image of the high frequency band and the luminance sub-image of the high frequency band; calculating a fusion mask of the low frequency band according to the fusion mask of the medium frequency band and the fusion mask of the high frequency band.
5. The image fusion method of claim 4, wherein, The fusion mask is a binary image, if a pixel value in the fusion mask takes a first numerical value, a pixel value at a same position in a corresponding fused sub-image adopts a pixel value in a corresponding grayscale sub-image, and if the pixel value in the fusion mask takes a second numerical value, the pixel value at the same position in the corresponding fused sub-image adopts a pixel value in a corresponding luminance sub-image. The step of calculating a fusion mask of the low frequency band according to the fusion mask of the medium frequency band and the fusion mask of the high frequency band comprises the following steps: if pixel values at a same position in the fusion mask of the medium frequency band and the fusion mask of the high frequency band both take the first numerical value, a pixel value of the fusion mask of the low frequency band at the position is set to the first numerical value, otherwise, the pixel value of the fusion mask of the low frequency band at the position is set to the second numerical value.
6. The image fusion method of claim 4 or 5, characterized in that, The step of calculating a fusion mask of the medium frequency band according to the grayscale sub-image of the medium frequency band and the luminance sub-image of the medium frequency band comprises the following steps: taking absolute values of pixel values in the gray scale sub-image of the middle frequency band to obtain a first mask calculation image, and taking absolute values of pixel values in the brightness sub-image of the middle frequency band to obtain a second mask calculation image; calculating the fusion mask of the middle frequency band according to a size relationship between pixel values at the same position in the first mask calculation image and the second mask calculation image.
7. The image fusion method of claim 4 or 5, characterized in that, The method for calculating the fusion mask of the middle frequency band according to the gray scale sub-image of the middle frequency band and the brightness sub-image of the middle frequency band comprises: taking absolute values of pixel values in the gray scale sub-image of the middle frequency band to obtain a first mask calculation image, and taking absolute values of pixel values in the brightness sub-image of the middle frequency band to obtain a second mask calculation image; filtering the first mask calculation image and the second mask calculation image respectively by using a non-edge-preserving smoothing filter to obtain a third mask calculation image and a fourth mask calculation image; calculating the fusion mask of the middle frequency band according to a size relationship between pixel values at the same position in the third mask calculation image and the fourth mask calculation image.
8. The image fusion method of claim 7, wherein, The fusion mask is a binary image, if a pixel value in the fusion mask takes a first numerical value, a pixel value of a corresponding fusion sub-image at the same position adopts a pixel value in a corresponding gray scale sub-image, if the pixel value in the fusion mask takes a second numerical value, the pixel value of the corresponding fusion sub-image at the same position adopts a pixel value in a corresponding brightness sub-image. The method for calculating the fusion mask of the middle frequency band according to a size relationship between pixel values at the same position in the third mask calculation image and the fourth mask calculation image comprises: for any pixel position in the fusion mask of the middle frequency band, if a pixel value of the third mask calculation image at the position is greater than a product of a pixel value of the fourth mask calculation image at the position and an adjustment threshold, a pixel value of the fusion mask of the middle frequency band at the position is set to the first numerical value, otherwise the pixel value of the fusion mask of the middle frequency band at the position is set to the second numerical value.
9. The image fusion method of claim 3, wherein, After the gray scale image and the color image to be fused are obtained, and before the M frames of gray scale sub-images are decomposed from the gray scale image, the method further comprises: adjusting brightness of the gray scale image to be consistent with the color image; The method for calculating the M frames of fusion masks corresponding to the M frequency bands according to the gray scale sub-images and the brightness sub-images comprises: calculating the fusion mask of the middle frequency band according to the gray scale sub-image of the low frequency band and the brightness sub-image of the low frequency band, calculating the fusion mask of the middle frequency band according to the gray scale sub-image of the middle frequency band and the brightness sub-image of the middle frequency band, and calculating the fusion mask of the high frequency band according to the gray scale sub-image of the high frequency band and the brightness sub-image of the high frequency band.
10. The image fusion method of any one of claims 3-9, wherein, The method for decomposing the M frames of gray scale sub-images from the gray scale image comprises: filtering the gray scale image by using a first low-pass filter to obtain the gray scale sub-image of the low frequency band; wherein a cutoff frequency of the first low-pass filter is a demarcation line between the low frequency band and the middle frequency band. filtering the gray-scale image by using a second low-pass filter to obtain a temporary gray-scale image, and calculating a high-frequency gray-scale sub-image according to the gray-scale image and the temporary gray-scale image; wherein a cutoff frequency of the second low-pass filter is a demarcation line between a middle frequency band and a high frequency band; calculating a middle-frequency gray-scale sub-image according to the gray-scale image, the low-frequency gray-scale sub-image and the high-frequency gray-scale sub-image.
11. The image fusion method of any one of claims 1-10, wherein, the decomposing the M frames of gray-scale sub-images from the gray-scale image comprises: calculating an occlusion region mask according to the gray-scale image and the channel image of the luminance channel; wherein a pixel value in the occlusion region mask represents a probability that a position where the pixel value is located belongs to an occlusion region, and the occlusion region refers to a region in a shooting scene that exists only in the color image but does not exist in the gray-scale image; performing weighted fusion on the gray-scale image and the channel image of the luminance channel by using the occlusion region mask to obtain a gray-scale occlusion fusion image; decomposing the M frames of gray-scale sub-images from the gray-scale occlusion fusion image.
12. A computer program product, characterised in that, computer program instructions, which are read and run by a processor, perform the method in any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, computer program instructions, which are read and run by a processor, perform the method in any one of claims 1-11.
14. An electronic device, comprising: comprise: a memory and a processor, the memory storing computer program instructions, which are read and run by the processor, perform the method in any one of claims 1-11. a memory and a processor, the memory storing computer program instructions, which are read and run by the processor, perform the method in any one of claims 1-11.
Citation Information
Patent Citations
Image processing method and apparatus for terminal, and terminal
CN107534735A
Picture partition method and device
WO2020043136A1