Video fusion method and image processing device
Patent Information
- Application Number
- CN202610654002.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]本申请实施例提供了一种视频融合方法和图像处理设备,以解决现有技术存在的考虑不够全面,难以满足实际需求的问题
本申请实施例提供的一种视频融合方法,通过对可见光视频图像的第一特征信息和红外视频图像的第二特征信息进行解算,得到三维环境向量;三维环境向量包括能见度指数、照度指数以及粉尘指数;基于三维环境向量进行参数映射处理,得到图像融合参数组;图像融合参数组包括多个图像增强参数;基于图像融合参数组对可见光视频图像和红外视频图像进行融合处理,得到视频融合图像。本申请通过解算可见光与红外特征信息,构建包含能见度指数、照度指数、粉尘指数的三维环境向量,不仅可全方位量化雾天、弱光、粉尘、沙尘等各类复杂环境的实际场景特征;还利用了红外图像穿透粉尘的现有特性,即粉尘区域在可见光下模糊但在红外下清晰,从而可以抑制可见光中的伪粉尘干扰,进而实现了对环境状态的精细化感知。此外,最后的融合处理充分结合可见光图像色彩真实、纹理细节丰富与红外图像抗强光、弱光可视性强、穿透粉尘雾气的双重优势,从而提高了视频图像的融合质量。
Smart Images

Figure CN122656869A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to a video fusion method and an image processing device. Background Technology
[0002] In practical applications, video surveillance scenarios in mining areas are extremely complex and variable. For example, regarding dust / smoke interference, the dust in the mining area causes severe attenuation of visible light image contrast, resulting in a completely white image; regarding lighting, the conditions vary drastically from bright open light to complete darkness underground, and searchlights at night can cause localized overexposure; regarding atmospheric visibility, rain and fog blur distant targets. Therefore, it is necessary to fuse different video images to output a clearer video image.
[0003] However, existing technologies are usually based on image statistical features (such as pixel grayscale, gradient, neighborhood distribution, etc.) for fusion, which is not comprehensive enough and difficult to meet actual needs. Summary of the Invention
[0004] This application provides a video fusion method and an image processing device to address the problem that existing technologies are not comprehensive enough and cannot meet practical needs.
[0005] In a first aspect, embodiments of this application provide a video fusion method, including: The first feature information of the visible light video image and the second feature information of the infrared video image are calculated to obtain a three-dimensional environment vector; the three-dimensional environment vector includes the visibility index, the illuminance index and the dust index. The image fusion parameter set is obtained by performing parameter mapping processing based on the three-dimensional environment vector; the image fusion parameter set includes multiple image enhancement parameters. Based on the image fusion parameter set, visible light video images and infrared video images are fused to obtain a fused video image.
[0006] Optionally, the first feature information includes a brightness value, and the second feature information includes a grayscale value; the first feature information of the visible light video image and the second feature information of the infrared video image are calculated to obtain a three-dimensional environment vector, including: The dust index is determined based on brightness and grayscale values; The visibility index is determined based on the luminance value and the percentile difference. The illuminance index is calculated based on the luminance value and the image size of the visible light video image.
[0007] Optionally, the dust index can be determined based on brightness and grayscale values, including: For any pixel in a visible light video image, the luminance variance is calculated for the set of pixels in the local window corresponding to that pixel to obtain the local variance corresponding to that pixel; the local window is constructed with the arbitrary pixel as the center and a distance set as the window size. For any pixel in an infrared video image, the gradient of the grayscale value of any pixel is calculated using an edge detection algorithm to obtain the gradient magnitude of any pixel; any pixel in an infrared video image and any pixel in a visible light video image are the same pixel. The local variance and gradient magnitude of any pixel are linearly normalized and multiplied to obtain the pixel-level dust intensity of any pixel. The dust index is determined based on the pixel-level dust intensity of each pixel.
[0008] Optionally, parameter mapping is performed based on the 3D environment vector to obtain an image fusion parameter set, including: The environmental severity level is obtained by weighted fusion of three-dimensional environment vectors. Hysteresis control is applied to the three-dimensional environment vector based on the environmental severity level, the environmental level lock value, and the dead zone threshold to obtain the target environment vector. The image fusion parameter set is obtained by performing parameter mapping processing based on the target environment vector.
[0009] Optionally, parameter mapping is performed based on the 3D environment vector to obtain an image fusion parameter set, including: The initial fusion parameter set is obtained by performing parameter mapping processing based on the three-dimensional environment vector. The initial fusion parameter set is temporally smoothed based on the smoothing coefficient and the previous frame fusion parameter set to obtain the image fusion parameter set; the previous frame fusion parameter set refers to the image fusion parameter set associated with the visible light video image acquired at the moment before the acquisition time corresponding to the visible light video image.
[0010] Optionally, the smoothing coefficient is determined as follows: Based on the brightness image corresponding to the visible light video image and the previous brightness image corresponding to the visible light video image at the previous moment, the motion intensity at the acquisition moment is determined. Numerical truncation is performed based on motion intensity, mapping coefficient, and smoothing threshold to obtain the smoothing coefficient; the mapping coefficient is used to characterize the mapping relationship between motion intensity and smoothing coefficient.
[0011] Optionally, multiple image enhancement parameters include infrared weights, infrared detail injection intensity, dehazing intensity, gamma correction coefficients, and color saturation coefficients; based on the image fusion parameter set, visible light video images and infrared video images are fused to obtain a fused video image, including: Dehazing enhancement processing is performed on visible light video images based on dehazing intensity and gamma correction coefficient to obtain dehazing enhanced images; The initial fused image is obtained by weighting and fusing the dehazing enhanced image and the infrared video image based on infrared weights. Based on the infrared detail injection intensity and infrared video image, high-frequency infrared injection is performed on the initial fused image to obtain a high-frequency injected image; The high-frequency injected image is processed based on the color saturation coefficient to obtain the video fusion image.
[0012] Optionally, the visible light video image is dehazed and enhanced based on the dehazing intensity and gamma correction coefficient to obtain a dehazed enhanced image, including: Brightness conversion is performed on visible light video images to obtain brightness images; The brightness image is processed based on the atmospheric scattering model to obtain an initial fog-free image; The initial fog-free image and the visible light video image are weighted and fused based on the defogging intensity to obtain the target fog-free image; Gamma correction is performed on the target haze-free image based on the gamma correction coefficient to obtain a dehaze-enhanced image.
[0013] Optionally, based on the infrared detail injection intensity and the infrared video image, infrared high-frequency injection is performed on the initial fused image to obtain a high-frequency injected image, including: Gaussian filtering is applied to the infrared video image to obtain the low-frequency component image; Infrared video images are processed based on low-frequency component images to obtain high-frequency component images; Based on the infrared detail injection intensity and high-frequency component image, infrared high-frequency injection is performed on the initial fused image to obtain a high-frequency injected image.
[0014] Optionally, the high-frequency injected image is processed based on the color saturation coefficient to obtain a video fused image, including: The high-frequency injected image is processed to obtain a grayscale image; Color space conversion is performed on the visible light video image to obtain its chromaticity information; The high-frequency injected image is inversely converted to color space based on chromaticity information to obtain a color reconstructed image; The grayscale image and the color reconstructed image are weighted and fused based on the color saturation coefficient to obtain the video fused image.
[0015] Optionally, parameter mapping is performed based on the 3D environment vector to obtain an image fusion parameter set, including: Trilinear interpolation is performed based on 3D environment vectors and a pre-stored environment parameter mapping table to obtain image fusion parameter sets; the environment parameter mapping table is used to store the mapping relationship between different 3D environment vectors and different image fusion parameter sets. or, The 3D environment vector is input into the trained parameter mapping model for processing to obtain the image fusion parameter set.
[0016] Secondly, embodiments of this application provide a video fusion apparatus, including: The calculation unit is used to calculate the first feature information of the visible light video image and the second feature information of the infrared video image to obtain a three-dimensional environment vector; the three-dimensional environment vector includes the visibility index, the illuminance index and the dust index; The parameter mapping unit is used to perform parameter mapping processing based on the three-dimensional environment vector to obtain the image fusion parameter set; the image fusion parameter set includes multiple image enhancement parameters. The fusion unit is used to fuse visible light video images and infrared video images based on the image fusion parameter set to obtain a fused video image.
[0017] Thirdly, embodiments of this application provide an image processing apparatus, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video fusion method as described in any one of the first aspects above.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video fusion method as described in any one of the first aspects above.
[0019] Fifthly, embodiments of this application provide a computer program product that, when run on an image processing device, enables the image processing device to execute the video fusion method described in any one of the first aspects.
[0020] The beneficial effects of the embodiments in this application compared with the prior art are: This application provides a video fusion method that calculates a three-dimensional environment vector by solving the first feature information of a visible light video image and the second feature information of an infrared video image. The three-dimensional environment vector includes a visibility index, an illuminance index, and a dust index. Based on the three-dimensional environment vector, parameter mapping processing is performed to obtain an image fusion parameter set, which includes multiple image enhancement parameters. Based on the image fusion parameter set, the visible light video image and the infrared video image are fused to obtain a fused video image. This application constructs a three-dimensional environment vector containing a visibility index, an illuminance index, and a dust index by solving the visible light and infrared feature information. This not only allows for comprehensive quantification of the actual scene characteristics of various complex environments such as fog, low light, dust, and sandstorms, but also utilizes the existing characteristic of infrared images penetrating dust, i.e., dust areas are blurred under visible light but clear under infrared light, thereby suppressing false dust interference in visible light and achieving refined perception of environmental conditions. Furthermore, the final fusion process fully combines the advantages of visible light images, such as true color and rich texture details, with infrared images, such as strong resistance to strong light and low light visibility, and the ability to penetrate dust and fog, thereby improving the fusion quality of video images. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating the implementation of a video fusion method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the specific implementation of step S101 in a video fusion method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the specific implementation of step S103 in a video fusion method provided in an embodiment of this application; Figure 4 This is a flowchart illustrating the implementation of a video fusion method provided in another embodiment of this application; Figure 5 This is a diagram illustrating the implementation of hysteresis control according to an embodiment of this application; Figure 6 This is a flowchart illustrating the implementation of a video fusion method provided in another embodiment of this application; Figure 7 This is a flowchart illustrating the overall implementation of a video fusion system provided in one embodiment of this application. Figure 8 This is a schematic diagram of the structure of a video fusion device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application. Detailed Implementation
[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0024] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0025] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0026] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0027] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0029] In practical applications, video surveillance scenarios in mining areas are extremely complex and variable. For example, regarding dust / smoke interference, the dust in the mining area causes severe attenuation of visible light image contrast, resulting in a completely white image; regarding lighting, the conditions vary drastically from bright open light to complete darkness underground, and searchlights at night can cause localized overexposure; regarding atmospheric visibility, rain and fog blur distant targets. Therefore, it is necessary to fuse different video images to output a clearer video image.
[0030] Existing technologies typically rely on image statistical features (such as pixel grayscale, gradient, and neighborhood distribution) for fusion. However, these low-level features often fail to accurately distinguish between environmental degradation and scene content. For example, the bright areas of a white mining truck may be misjudged as dust obscuring the view; noise in low-light conditions at night may be misjudged as high-frequency texture details. Furthermore, most existing technologies adjust for single factors (such as adjusting gain only based on brightness), lacking the comprehensive perception and decoupling ability for complex environmental conditions such as dust + low light or fog + strong light. This results in repeated oscillations of fusion parameters under complex conditions, leading to flickering and unstable video output.
[0031] Therefore, this application provides a video fusion method that can achieve refined perception of environmental conditions to improve the fusion quality of video images.
[0032] Please see Figure 1 , Figure 1 This is a flowchart illustrating the implementation of a video fusion method according to an embodiment of this application. In this embodiment, the video fusion method is executed by an image processing device. The image processing device can be a terminal device or an intermediate processing device.
[0033] It should be noted that the aforementioned terminal devices can be devices capable of independently acquiring, processing, and outputting business results, including but not limited to: smart cameras and artificial intelligence (AI) surveillance cameras, etc.
[0034] Intermediate processing devices can be processing modules / boards used only for image signal relay and computation acceleration, including but not limited to: Field-Programmable Gate Array (FPGA) image processing boards, independent image signal processors (ISPs) and other image processing chips, as well as computing cards with graphics processing units (GPUs) and digital signal processors (DSPs).
[0035] like Figure 1 As shown, a video fusion method provided in one embodiment of this application may include steps S101 to S103, which are described in detail below: In S101, the first feature information of the visible light video image and the second feature information of the infrared video image are calculated to obtain a three-dimensional environment vector; the three-dimensional environment vector includes visibility index, illuminance index and dust index.
[0036] In one implementation of this application, an image processing device can acquire visible light video images in real time through a visible light camera device that is communicatively connected to it.
[0037] In another implementation of this application, the image processing device can acquire infrared video images in real time through an infrared camera device that is communicatively connected to it.
[0038] It should be noted that the visible light video images and infrared video images obtained above can be video images captured at the same time on the same scene.
[0039] In some possible embodiments, after obtaining the visible light video image and the infrared video image, the image processing device can perform image preprocessing on the visible light video image and the infrared video image to make the temporal sequence and spatial dimension of the two video images correspond one-to-one.
[0040] In this embodiment, image preprocessing includes, but is not limited to: frame synchronization alignment, size normalization, grayscale correction, noise and bad pixel removal, and Gaussian filtering smoothing.
[0041] It should be noted that the first feature information may include the brightness value, and the second feature information may include the grayscale value.
[0042] In this embodiment, the image processing device can input the aforementioned first feature information and second feature information into a trained feature calculation model for processing to obtain a three-dimensional environment vector. The feature calculation model can be constructed from a pre-built neural network model.
[0043] Specifically, the feature calculation model can be obtained by training a pre-built neural network model based on a first preset sample set. Each sample data in the first preset sample set includes sample feature information (including feature information corresponding to the sample's visible light image and feature information corresponding to the sample's infrared image) and a corresponding three-dimensional environment vector. When training the pre-built neural network model, the sample feature information of each sample is used as the input to the neural network model, and the corresponding three-dimensional environment vector is used as the output. Through training, the neural network model can learn the correspondence between all possible feature information and the three-dimensional environment vector, and the trained neural network model becomes the feature calculation model.
[0044] In this embodiment of the application, the three-dimensional environment vector includes visibility index, illuminance index, and dust index.
[0045] The dust index is used to characterize the concentration of suspended particulate matter that is opaque to visible light but transparent to infrared light.
[0046] The visibility index is used to reflect the degree of atmospheric transparency.
[0047] The illuminance index describes the illuminance level of the environment in which an image is captured, i.e., the light intensity of the scene being photographed.
[0048] In one embodiment of this application, when the first feature information includes a brightness value and the second feature information includes a grayscale value, in order to improve the accuracy of determining each index, the image processing device can specifically use methods such as... Figure 2 Steps S201 to S203 shown implement step S101, as detailed below: In S201, the dust index is determined based on the brightness and grayscale values.
[0049] In this embodiment, the image processing device can calculate the average brightness value of each pixel in the visible light video image to obtain the average brightness value of visible light. Then, the image processing device can normalize the average brightness value of visible light to normalize it to the range [0, 1].
[0050] Meanwhile, the image processing device can calculate the average gray value of each pixel in the infrared video image to obtain the infrared gray average value. Then, the image processing device can normalize the infrared gray average value to normalize it to the range [0, 1].
[0051] Subsequently, by utilizing the difference between visible light attenuation due to dust scattering and infrared light attenuation due to dust penetration, the image processing equipment can calculate the joint deviation between brightness and grayscale, and perform linear mapping and amplitude limiting normalization on the joint deviation to obtain the final dust index.
[0052] Where, joint deviation = .
[0053] In one embodiment of this application, since infrared light has the existing characteristic of penetrating dust (i.e., dust areas are blurred under visible light but clear under infrared light), in order to suppress pseudo-dust interference in visible light, the image processing device can specifically implement step S201 according to the following steps, detailed below: For any pixel in a visible light video image, the luminance variance is calculated for the set of pixels in the local window corresponding to that pixel to obtain the local variance corresponding to that pixel; the local window is constructed with the arbitrary pixel as the center and a distance set as the window size. For any pixel in an infrared video image, the gradient of the grayscale value of any pixel is calculated using an edge detection algorithm to obtain the gradient magnitude of any pixel; any pixel in an infrared video image and any pixel in a visible light video image are the same pixel. The local variance and gradient magnitude of any pixel are linearly normalized and multiplied to obtain the pixel-level dust intensity of any pixel. The dust index is determined based on the pixel-level dust intensity of each pixel.
[0054] In this embodiment, for any pixel in a visible light video image, the image processing device can construct a local window corresponding to that pixel, centered on that pixel and setting the distance as the window size, and determine the set of pixels within that local window. Then, the image processing device can calculate the local mean value of that arbitrary pixel based on the brightness values of each pixel in the pixel set and the number of pixels within the local window.
[0055] Specifically, the image processing device can calculate the local mean of any pixel according to the following formula: ; in, Represents pixels x The local mean, Represented in pixels The set of pixels centered on a local window. Represented in pixels The number of pixels in the local window centered on the target. This represents the brightness value of pixel y in the pixel set.
[0056] Subsequently, the image processing device can calculate the original value of the local variance of any pixel based on the local mean of any pixel, the brightness value of each pixel in the aforementioned pixel set, and the number of pixels in the local window of any pixel.
[0057] Specifically, the image processing device can calculate the original value of the local variance of any pixel according to the following formula: ; in, Represents pixels x The original values of local variance, Represents pixels x The local mean, Represented in pixels The set of pixels centered on a local window. Represented in pixels The number of pixels in the local window centered on the target. This represents the brightness value of pixel y in the pixel set.
[0058] Finally, the image processing device can normalize the original value of the local variance of any pixel based on a linear normalization function to obtain the final local variance of any pixel.
[0059] Specifically, the image processing device can calculate the final local variance of any pixel according to the following formula: ; in, Represents pixels x The local variance, Represents pixels x The original values of local variance, This represents the linear normalization function.
[0060] It should be noted that the denser the dust, the fewer pixels... x The smaller the local variance.
[0061] In some possible embodiments, the linear normalization function can be applied to any scalar field. (as mentioned above) Scaling is performed using the quantile range. Therefore, the linear normalization function can be specifically expressed as follows: ; ; in, This represents the 5% quantile operator. Represents the 95th percentile operator. This represents the truncation function, which is... Limited to the range , Represents extremely small positive numbers, used to avoid denominators of 0 (dimensionless, values can be any...). ).
[0062] In this embodiment, for any pixel in an infrared video image, the image processing device can perform gradient calculation on the grayscale value of any pixel using an edge detection algorithm to obtain the gradient magnitude of any pixel.
[0063] When the edge detection algorithm is the Sobel operator, the image processing device can calculate the gradient magnitude of any pixel according to the following formula: ; ; ; in, This represents the Sobel level operator kernel. This represents the Sobel vertical operator kernel. This represents the convolution operation. Represents pixels The horizontal gradient response, Represents pixels The vertical gradient response, Represents pixels The gradient magnitude.
[0064] It should be noted that the clearer the infrared structure, i.e. the larger the gradient amplitude, the more likely it is not a completely flat background.
[0065] It should be understood that any pixel in the aforementioned infrared video image is the same pixel as any pixel in the visible light video image.
[0066] In this embodiment, since areas with blurred visible light (small variance) and clear infrared light (large gradient) are more likely to be dust-covered areas, the image processing device can linearly normalize and multiply the local variance and gradient magnitude of any pixel to obtain the pixel-level dust intensity of any pixel.
[0067] Specifically, the image processing device can calculate the pixel-level dust intensity of any pixel using the following formula: ; in, Represents pixels x The pixel-level dust intensity (dimensionless, with a value range of [0, 1]). Represents a linear normalization function. Represents pixels xThe local variance, Represents pixels x The gradient magnitude.
[0068] In this embodiment, after obtaining the pixel-level dust intensity of each pixel, the image processing device can sort the pixel-level dust intensity in ascending order to obtain an intensity sequence, and determine the pixel-level dust intensity corresponding to the median of the intensity sequence as the final dust index.
[0069] This embodiment utilizes the characteristic that dusty areas are blurred under visible light but clear under infrared light to construct a dust index, reducing the probability of white, bright objects being misjudged as dust, and making environmental recognition in dusty scenes more stable.
[0070] In S202, the visibility index is determined based on the luminance value and the percentile difference.
[0071] In this embodiment, the image processing device can convert the red, green and blue (RGB) three-channel image corresponding to the visible light video image according to the luminance weighting to obtain the luminance channel, that is, the luminance image of a single channel.
[0072] Specifically, the brightness image is represented as follows: ; in, Represents a brightness image. This represents the R channel image corresponding to a visible light video image. This represents the G-channel image corresponding to a visible light video image. This represents the B-channel image corresponding to a visible light video image.
[0073] In this embodiment, after obtaining the brightness image corresponding to the visible light video image, the image processing device can calculate the visibility index using the quantile difference of brightness contrast to reduce extreme value interference.
[0074] In practical applications, quantile difference refers to the difference between two quantiles after removing the extreme values at both ends, which measures the degree of dispersion of data. It is a measure of the range (maximum value). Improvements to the minimum value.
[0075] Specifically, the image processing device can calculate the visibility index according to the following formula: ; in, This represents the visibility index (dimensionless, with a value range of [0, 1]). Represents the 95th percentile operator. This represents the 5% quantile operator. This represents a brightness image.
[0076] It should be noted that a higher visibility index indicates poorer visibility (i.e., heavier fog).
[0077] In S203, the illuminance index is calculated based on the luminance value and the image size of the visible light video image.
[0078] In this embodiment, the image processing device can calculate the average brightness of the visible light video image based on the brightness value of each pixel and the image size of the visible light video image.
[0079] Specifically, the image processing device can calculate the above average brightness according to the following formula: ; in, Indicates average brightness. Indicates the image height. Indicates the image width. This represents the brightness value of the pixel at coordinates (x, y) in the brightness image corresponding to the visible light video image.
[0080] In this embodiment, the image processing device can calculate the proportion of dark area pixels in a visible light video image based on the brightness value of each pixel, the image size of the visible light video image, and the dark area threshold. The dark area pixel proportion describes the proportion of pixels with brightness less than the dark area threshold.
[0081] Specifically, the image processing device can calculate the aforementioned percentage of dark area pixels using the following formula: ; in, Indicates the percentage of pixels in the dark area. Indicates the image height. Indicates the image width. This represents the brightness value at pixel coordinates (x, y) in the brightness image corresponding to the visible light video image. Indicates an indicator function, This represents the dark area threshold (dimensionless, with a value range of [0, 1], and a value of 0.2).
[0082] In this embodiment, the image processing device can calculate the illuminance index based on the above-mentioned average brightness and the percentage of pixels in the dark area.
[0083] Specifically, the image processing device can calculate the illuminance index according to the following formula: ; in, represents the illuminance index (dimensionless, with a value range of [0, 1]). This represents the weighting coefficient (dimensionless, with a value of 0.5). Indicates the percentage of pixels in the dark area. This indicates the average brightness.
[0084] It should be noted that a higher illuminance index indicates lower illuminance.
[0085] In S102, parameter mapping is performed based on the three-dimensional environment vector to obtain an image fusion parameter set; the image fusion parameter set includes multiple image enhancement parameters.
[0086] It should be noted that the image processing device can pre-divide the three-dimensional environment space (represented by a three-dimensional environment vector) into multiple typical states (i.e., anchor points), meaning that different three-dimensional environment vectors can map to different typical states. Furthermore, the values of multiple image enhancement parameters are not entirely the same in different typical states.
[0087] The aforementioned typical states include, but are not limited to: normal state (used to describe a good environment), heavy dust state (used to describe severe dust), low illumination state (used to describe low illumination at night or underground), and low visibility state (used to describe low visibility due to rain, fog, etc.).
[0088] It should be noted that the various image enhancement parameters include, but are not limited to: infrared weighting, infrared detail injection intensity, dehazing intensity, gamma correction coefficient, and color saturation coefficient.
[0089] For example, the following are the values of several image enhancement parameters for different typical states of different 3D environment vector mappings: Normal state ( Examples of several image enhancement parameters are as follows: , , , , ; Heavy dust condition ( Examples of several image enhancement parameters are as follows: , , , , ; Low light conditions ( Examples of several image enhancement parameters are as follows: , , , , ; Low visibility conditions ( Examples of several image enhancement parameters are as follows: , , , , .
[0090] in, Indicates the infrared weight (dimensionless, with a value range of [0, 1]). Indicates the infrared detail injection intensity (dimensionless, value range can be [0, 1]). This represents the defogging intensity (dimensionless, with a value range of [0, 1]). This represents the gamma correction factor (dimensionless, with a range of values that can be any). ), This represents the color saturation coefficient (dimensionless, with a value range of [0, 1]).
[0091] Please refer to Table 1, which is a typical anchor point state example table in the environment-parameter mapping space provided in an embodiment of this application.
[0092] Table 1:
[0093] In this embodiment of the application, after obtaining the three-dimensional environment vector, the image processing device can perform parameter mapping processing on the three-dimensional environment vector at this time based on the mapping relationship between different three-dimensional environment vectors and different typical states, so as to obtain the image fusion parameter set.
[0094] In some possible implementations, parameter mapping includes two methods: look-up table (LUT) and model mapping, which facilitates deployment on different computing platforms and balances real-time performance with fusion quality.
[0095] Therefore, in this embodiment, the image processing device can determine the corresponding implementation method based on its own deployed processing chip, and implement parameter mapping processing based on the implementation method.
[0096] In one embodiment of this application, when step S102 is implemented by the FPGA / DSP through a LUT, the image processing device can specifically implement step S102 according to the following steps, detailed below: Based on the three-dimensional environment vector and the pre-stored environment parameter mapping table, trilinear interpolation is performed to obtain the image fusion parameter set; the environment parameter mapping table is used to store the mapping relationship between different three-dimensional environment vectors and different image fusion parameter sets.
[0097] It should be noted that image processing devices can discretize the three-dimensional environment space (represented by three-dimensional environment vectors) into... A grid is constructed, with each grid pre-stored as a 5-dimensional parameter vector (i.e., multiple image enhancement parameters), thus obtaining an environment parameter mapping table. This environment parameter mapping table stores the mapping relationships between different 3D environment vectors and different image fusion parameter sets.
[0098] In this embodiment, the image processing device can perform trilinear interpolation based on the three-dimensional environment vector and the above-mentioned environment parameter mapping table to obtain the image fusion parameter set.
[0099] In practical applications, trilinear interpolation is a linear interpolation method on a three-dimensional regular mesh. It uses the values of the eight vertices of a cube to estimate the value of any point inside it.
[0100] Specifically, LUT trilinear interpolation can be expressed as: ; in, This represents the image fusion parameter set output by the mapping. This represents the parameter vector stored at a discrete grid point in the environmental parameter mapping table. This represents the interpolation weight (dimensionless, with a value range of [0, 1]). This represents the index of the dust index coordinate in the discrete grid within the environmental parameter mapping table. This represents the index of the visibility index coordinates within the discrete grid in the environmental parameter mapping table. This represents the index of the illuminance index coordinate in the discrete grid within the environmental parameter mapping table.
[0101] In another embodiment of this application, when step S102 is implemented by the NPU / GPU through model mapping, the image processing device can specifically implement step S102 according to the following steps, detailed below: The 3D environment vector is input into the trained parameter mapping model for processing to obtain the image fusion parameter set.
[0102] In this embodiment, the parameter mapping model can be trained by a multi-layer perceptron (MLP).
[0103] The image processing device can train a small neural network based on a multilayer perceptron to fit the nonlinear relationship between the 3D environment vector and the fusion parameter set. Specifically, the image processing device can first construct a sample set, with the 3D environment vector Z=[D, V, L] as the input and the fusion parameter set as the output. Fusion parameter group Multiple image enhancement parameters are taken from the typical states in step S102 above. To cover continuous environmental changes between typical states, the image processing device can perform multi-dimensional interpolation on the three-dimensional environment vector between different typical states to generate intermediate state samples, so that the training data covers normal, heavy dust, low illumination, low visibility and their transition ranges. After completing the construction of the sample set, backpropagation is used to train the network weights and biases. After training, the trained model parameters are exported and fixed.
[0104] In this embodiment, the implementation of the image processing device through the parameter mapping model can be represented as follows: ; ; in, This represents the hidden layer vector of the parameter mapping model. This represents the weight matrix from the input layer to the hidden layer of the parameter mapping model. This represents the weight matrix from the hidden layer to the output layer of the parameter mapping model. This represents the hidden layer bias vector of the parameter mapping model. This represents the output layer bias vector of the parameter mapping model. This represents the activation function of the parameter mapping model. This represents the image fusion parameter set of the mapping output.
[0105] In S103, the visible light video image and the infrared video image are fused based on the image fusion parameter group to obtain the fused video image.
[0106] In this embodiment of the application, after obtaining the image fusion parameter set, the image processing device can first perform a weighted summation of the visible light video image and the infrared video image, that is, perform basic fusion to obtain an initial fused image. Then, the image processing device can perform adaptive enhancement processing on the initial fused image based on the above image fusion parameter set to obtain the final video fused image.
[0107] In some possible embodiments, the image processing device can output the video fusion image in real time after obtaining it.
[0108] In one embodiment of this application, when multiple image enhancement parameters include infrared weights, infrared detail injection intensity, dehazing intensity, gamma correction coefficients, and color saturation coefficients, in order to improve infrared information contribution in dusty and low-light environments, preserve visible light color and texture in good environments, and compensate for visible light detail degradation through detail injection, the image processing device can specifically achieve the following: Figure 3 Steps S301 to S304 shown implement step S103, as detailed below: In S301, the visible light video image is dehazed and enhanced based on the dehazing intensity and gamma correction coefficient to obtain a dehazed and enhanced image.
[0109] In this embodiment, the image processing device can first perform dehazing processing on the visible light video image by adjusting the dehazing intensity to obtain an initial haze-free image. Then, the image processing device can perform gamma correction on the initial haze-free image based on the gamma correction coefficient to obtain a dehazing-enhanced image.
[0110] In one embodiment of this application, the image processing device may specifically implement step S301 according to the following steps, as detailed below: Brightness conversion is performed on visible light video images to obtain brightness images; The brightness image is processed based on the atmospheric scattering model to obtain an initial fog-free image; The initial fog-free image and the visible light video image are weighted and fused based on the defogging intensity to obtain the target fog-free image; Gamma correction is performed on the target haze-free image based on the gamma correction coefficient to obtain a dehaze-enhanced image.
[0111] In this embodiment, the image processing device can perform luminance conversion on the red, green and blue (RGB) three-channel image corresponding to the visible light video image according to luminance weighting to obtain a luminance channel, that is, a single-channel luminance image.
[0112] In some possible embodiments, the brightness image is represented as follows: ; in, Represents a brightness image. This represents the R channel image corresponding to a visible light video image. This represents the G-channel image corresponding to a visible light video image. This represents the B-channel image corresponding to a visible light video image.
[0113] In practical applications, the core of the atmospheric scattering model is to describe the process by which light reaches the camera after being scattered / absorbed by molecules and aerosols in the atmosphere using single scattering and radiative transfer.
[0114] In this embodiment, the image processing device can process the brightness image based on a simplified atmospheric scattering model to obtain an initial fog-free image.
[0115] Specifically, the atmospheric scattering model can be simplified as follows: ; in, Represents the pixels in the initial fog-free image x pixel values, This represents the estimated value of atmospheric light (dimensionless, with a range of [0, 1]). This represents transmittance (dimensionless, with a value range of [0, 1]). This represents the lower limit of transmittance (dimensionless, and can take a value of 0.1). Represents pixels x The brightness value.
[0116] In some possible embodiments, the image processing device can calculate the above-mentioned atmospheric light estimate and transmittance based on the brightness values of each pixel in the visible light video image and the quantile difference of the brightness image.
[0117] Specifically, the image processing device can calculate the atmospheric light estimate and transmittance according to the following formula: ; ; in, This represents the estimated atmospheric light value. Indicates transmittance. Represents the 95th percentile operator. Represents a brightness image. Represented in pixels The set of pixels centered on a local window. Represents pixels x brightness value, This represents the transmittance estimation coefficient (dimensionless, and its value can be 0.95).
[0118] In this embodiment, the image processing device can perform weighted fusion of the initial fog-free image and the visible light video image based on the defogging intensity to obtain the target fog-free image.
[0119] Specifically, the image processing device can obtain the target haze-free image according to the following formula: ; in, Represents pixels in the target haze-free image x pixel values, Indicates the defogging intensity. Represents pixels in a fog-free image x pixel values, Represents pixels x The brightness value.
[0120] In this embodiment, after obtaining the target haze-free image, the image processing device can perform gamma correction on the target haze-free image based on the gamma correction coefficient and gamma correction function to obtain a dehazing enhanced image.
[0121] It should be noted that the gamma correction function can be defined as: ; in, Represents the gamma correction function. Indicates the gamma correction factor. This indicates the parameters that require gamma correction.
[0122] In this embodiment, the image processing device can obtain the target haze-free image according to the following formula: ; in, Indicates a fog-free image of the target. Represents the gamma correction function. Represents pixels in the target haze-free image x pixel values, Indicates the defogging intensity. Indicates the gamma correction factor. This represents a brightness image.
[0123] In S302, the dehazing enhanced image and the infrared video image are weighted and fused based on infrared weights to obtain the initial fused image.
[0124] In this embodiment, after obtaining the dehazing enhanced image, the image processing device can perform basic fusion of the dehazing enhanced image and the infrared video image, that is, perform weighted fusion of the dehazing enhanced image and the infrared video image based on infrared weights to obtain an initial fused image.
[0125] Specifically, the image processing device can obtain the initial fused image according to the following formula: ; in, Represents the initial fused image. Represents infrared video images, Indicates infrared weighting, This indicates a fog-free image of the target.
[0126] In S303, based on the infrared detail injection intensity and the infrared video image, the initial fused image is subjected to infrared high-frequency injection to obtain a high-frequency injected image.
[0127] In this embodiment, after obtaining the initial fused image, in order to improve the contribution of infrared information and compensate for the degradation of visible light details in dusty and low-light environments, the image processing device can perform high-frequency infrared injection on the initial fused image based on the infrared detail injection intensity and the infrared video image to obtain a high-frequency injected image.
[0128] In one embodiment of this application, the image processing device may specifically implement step S303 according to the following steps, as detailed below: Gaussian filtering is applied to the infrared video image to obtain the low-frequency component image; Infrared video images are processed based on low-frequency component images to obtain high-frequency component images; Based on the infrared detail injection intensity and high-frequency component image, infrared high-frequency injection is performed on the initial fused image to obtain a high-frequency injected image.
[0129] In this embodiment, the image processing device can specifically obtain the low-frequency component image according to the following formula: ; ; in, Indicates Gaussian filtering. Represents the Gaussian function. Represents pixels in low-frequency component images x pixel values, Indicates the pixel in an infrared video image x Any pixel in a local window, Represented in pixels The set of pixels centered on a local window. This represents the Gaussian kernel scale parameter.
[0130] Subsequently, the image processing device processes the infrared video image based on the aforementioned low-frequency component image to obtain the high-frequency component image, namely the infrared high-frequency component.
[0131] Specifically, the image processing device can obtain the high-frequency component image according to the following formula: ; in, Represents high-frequency component images. Indicates Gaussian filtering. This represents an infrared video image.
[0132] In this embodiment, after obtaining the high-frequency component image, the image processing device can perform infrared high-frequency injection on the initial fused image based on the infrared detail injection intensity and the high-frequency component image to obtain a high-frequency injected image.
[0133] Specifically, the image processing device can obtain a high-frequency injected image according to the following formula: ; in, Indicates a high-frequency injection image. Represents the initial fused image. Represents high-frequency component images. Indicates the infrared detail injection intensity.
[0134] In S304, the high-frequency injected image is processed based on the color saturation coefficient to obtain the video fusion image.
[0135] In this embodiment, after obtaining the high-frequency injected image, in order to preserve the visible light color and texture under good conditions, the image processing device can process the high-frequency injected image based on the color saturation coefficient to reconstruct the fused brightness and visible light chromaticity components, and control the mixing ratio of color and grayscale through the color saturation coefficient, thereby obtaining the final video fusion image.
[0136] In one embodiment of this application, the image processing device may specifically implement step S304 according to the following steps, as detailed below: The high-frequency injected image is processed to obtain a grayscale image; Color space conversion is performed on the visible light video image to obtain its chromaticity information; The high-frequency injected image is inversely converted to color space based on chromaticity information to obtain a color reconstructed image; The grayscale image and the color reconstructed image are weighted and fused based on the color saturation coefficient to obtain the video fused image.
[0137] It should be noted that for specific grayscale processing methods, please refer to existing image grayscale techniques, which will not be elaborated here.
[0138] In this embodiment, color space conversion specifically refers to YUV (luminance + chrominance) color space conversion.
[0139] YUV color space conversion refers to the mathematical transformation between the RGB and YUV color models. U represents the blue chromaticity component, and V represents the red chromaticity component.
[0140] In this embodiment, after obtaining the chromaticity information of the visible light video image, the image processing device can perform inverse color space conversion on the high-frequency injected image based on the chromaticity information to obtain a color reconstructed image. Then, the image processing device can perform weighted fusion of the grayscale image and the color reconstructed image based on the color saturation coefficient to control the mixing ratio of color and grayscale, thereby obtaining the final video fused image.
[0141] Specifically, the image processing device can obtain the video fused image according to the following formula: ; in, Indicates video-fused images, Represents a color-reconstructed image. Represents a grayscale image. This represents the color saturation coefficient.
[0142] As can be seen from the above, the video fusion method provided in this application calculates a three-dimensional environment vector by solving the first feature information of a visible light video image and the second feature information of an infrared video image. The three-dimensional environment vector includes a visibility index, an illuminance index, and a dust index. Based on the three-dimensional environment vector, parameter mapping processing is performed to obtain an image fusion parameter set. The image fusion parameter set includes multiple image enhancement parameters. Based on the image fusion parameter set, the visible light video image and the infrared video image are fused to obtain a video fused image. This application constructs a three-dimensional environment vector containing a visibility index, an illuminance index, and a dust index by solving visible light and infrared feature information. This not only comprehensively quantifies the actual scene characteristics of various complex environments such as fog, low light, dust, and sandstorms, but also utilizes the existing characteristic of infrared images penetrating dust, i.e., dust areas are blurred under visible light but clear under infrared light, thereby suppressing false dust interference in visible light and achieving refined perception of environmental conditions. Furthermore, the final fusion process fully combines the advantages of visible light images, such as true color and rich texture details, with infrared images, such as strong resistance to strong light and low light visibility, and the ability to penetrate dust and fog, thereby improving the fusion quality of video images.
[0143] Please see Figure 4 , Figure 4 This is a flowchart illustrating the implementation of a video fusion method according to another embodiment of this application. Relative to... Figure 1 In the corresponding embodiment, step S102 may include S401~S403, as detailed below: In S401, the three-dimensional environment vectors are weighted and fused to obtain the environmental severity level.
[0144] In this embodiment, the image processing device can specifically calculate the environmental severity level based on the following formula: ; in, Indicates time Environmental severity level Indicates time Dust index, Indicates time Visibility index, Indicates time Illuminance index, Indicates dust weight. Indicates visibility weight. This represents the illuminance weight.
[0145] It should be noted that the weights satisfy... .
[0146] In S402, hysteresis control is performed on the three-dimensional environmental vector based on the environmental severity level, environmental level lock value, and dead zone threshold to obtain the target environmental vector.
[0147] In this embodiment, after obtaining the environmental severity level, the image processing device can subtract the environmental severity level from the environmental level lock value to calculate the change in the environment. Then, the image processing device can compare this change with a dead zone threshold to achieve hysteresis control of the three-dimensional environment vector. That is, the dead zone threshold determines whether to update the lock input state to obtain the final target environment vector.
[0148] Specifically, the image processing device can obtain the target environment vector according to the following formula: ; ; in, Indicates time 3D environment vector , Indicates time The locked environment vector (i.e., the target environment vector). Indicates time The locked environment vector, Indicates time The severity level is locked (i.e., the environmental level lock value). Indicates time The severity level is locked. Indicates the dead zone threshold. Indicates time The environmental severity level.
[0149] In S403, parameter mapping is performed based on the target environment vector to obtain the image fusion parameter set.
[0150] In this embodiment, after obtaining the target environment vector, the image processing device can perform parameter mapping processing based on the target environment vector to obtain an image fusion parameter set.
[0151] It should be noted that the implementation process of step S403 can be found in the implementation process of step S102, and will not be elaborated here.
[0152] Please see Figure 5 , Figure 5 This is a diagram illustrating the implementation of hysteresis control according to an embodiment of this application. Figure 5As shown, the image processing device can input the environmental severity level to the hysteresis control logic. The hysteresis control logic then compares the change in the environment (calculated from the environmental severity level and the environmental level lock value) with a dead-zone threshold to determine whether to update the lock input state. Specifically, when the image processing device detects a change less than the dead-zone threshold, it can lock the current state, defining the previously acquired 3D environmental vector as the target environmental vector. When the image processing device detects a change greater than or equal to the dead-zone threshold, it can update the mapping input, defining the locked environmental vector from the previous moment as the target environmental vector. The hysteresis control logic can then obtain the image fusion parameter set based on this target environmental vector to output stable parameters.
[0153] In some possible embodiments, after updating the mapping input, the hysteresis control logic can also perform a smooth transition on the initial fusion parameter set obtained based on the target environment vector corresponding to the update, so as to obtain the final image fusion parameter set.
[0154] Please also refer to Table 2, which is an example table of hysteresis control provided in one embodiment of this application.
[0155] Table 2:
[0156] As can be seen from the above, the video fusion method provided in this embodiment, by weighting and fusing the three-dimensional environment vectors to obtain a unified environmental severity level, can comprehensively reflect the severity of the scene, avoiding environmental judgment distortion caused by sudden changes in a single environmental indicator or local anomalies, and improving the robustness and accuracy of environmental perception. Subsequently, hysteresis control based on the environmental level lock value and dead zone threshold can filter out frequent switching caused by small fluctuations in the environmental index and noise jitter, preventing the fusion parameters from repeatedly oscillating near the critical value. This effectively solves the problems of inconsistent brightness, inconsistent detail, and flickering / jittering in the video fusion image, achieving continuous and stable video output. Finally, the target environment vector output after hysteresis control has completed fluctuation filtering and state stabilization. Based on this, parameter mapping can achieve smooth, continuous, and controllable changes in the image fusion parameter set, effectively solving problems such as discontinuities, abrupt changes, and sudden quality variations in the fused image caused by parameter mutations.
[0157] Please see Figure 6 , Figure 6 This is a flowchart illustrating the implementation of a video fusion method provided in another embodiment of this application. Compared to... Figure 1 In the corresponding embodiment, step S102 may include S501~S503, as detailed below: In S501, parameter mapping is performed based on the three-dimensional environment vector to obtain the initial fusion parameter set.
[0158] In this embodiment, the implementation process of step S501 can be referred to the implementation process of step S102, and will not be repeated here.
[0159] In S502, the initial fusion parameter set is subjected to temporal smoothing based on the smoothing coefficient and the previous frame fusion parameter set to obtain the image fusion parameter set.
[0160] It should be noted that the previous frame fusion parameter set refers to the image fusion parameter set associated with the visible light video image acquired at the moment preceding the acquisition time corresponding to the visible light video image.
[0161] In this embodiment, the image processing device can perform motion-adaptive infinite impulse response (IIR) filtering on the initial fusion parameter set output by the mapping to obtain the final image fusion parameter set.
[0162] In practical applications, the output of an IIR filter depends not only on the current input and past inputs, but also on past outputs.
[0163] In some possible embodiments, the image processing device can obtain the image fusion parameter set according to the following formula: ; in, This represents the image fusion parameter group after smoothing the current frame. This represents the previous frame's blending parameter group after smoothing. This represents the initial fusion parameter set output by the current frame mapping. This represents the smoothing coefficient (dimensionless, with a value range of [0, 1], for example, when the image is static). During strenuous exercise ).
[0164] In one embodiment of this application, the smoothing coefficient can be adaptively adjusted according to the motion intensity. Therefore, the smoothing coefficient can be determined according to the following steps, detailed below: Based on the brightness image corresponding to the visible light video image and the previous brightness image corresponding to the visible light video image at the previous moment, the motion intensity at the acquisition moment is determined. Numerical truncation is performed based on motion intensity, mapping coefficient, and smoothing threshold to obtain the smoothing coefficient; the mapping coefficient is used to characterize the mapping relationship between motion intensity and smoothing coefficient.
[0165] In this embodiment, the image processing device can specifically obtain the motion intensity at the acquisition time according to the following formula: ; in, Indicates time (i.e., the intensity of motion at the time of data collection) Indicates the image height. Indicates the image width. Let represent the brightness value of the pixel at coordinates (x, y) in the brightness image at time t. This represents the brightness value of the pixel at coordinate (x, y) in the brightness image at time t-1.
[0166] Subsequently, the image processing device can perform numerical truncation based on motion intensity, mapping coefficients, and a smoothing threshold to obtain a smoothing coefficient. The mapping coefficients characterize the mapping relationship between motion intensity and the smoothing coefficients.
[0167] Specifically, the image processing device can obtain the smoothing coefficient according to the following formula: ; in, Indicates time The intensity of exercise, This represents the lower limit of the smoothing coefficient (dimensionless, and its value can be 0.5). This represents the upper limit of the smoothing coefficient (dimensionless, and can take the value 0.9). Represents the mapping coefficient, indicating This represents the truncation function, which is... Limited to the range .
[0168] As can be seen from the above, the video fusion method provided in this embodiment first completes parameter mapping based on the three-dimensional environment vector, so that the initial fusion parameters are accurately matched with the current environment state. Then, the continuity of parameters is optimized through temporal smoothing processing, which can reduce environmental perception distortion and effectively solve the problem of abnormal picture caused by parameter mutation. In addition, by using the fusion parameter group of the previous frame and the smoothing coefficient to perform temporal filtering on the initial fusion parameters of the current frame, the frequent jumps of parameters caused by small fluctuations in the environmental index and noise disturbances can be effectively eliminated, so that the fusion parameters transition smoothly frame by frame. This makes the brightness, contrast and enhancement intensity of the video picture continuous and stable, effectively improving the comfort of human eye viewing.
[0169] It should be noted that the above Figure 4 and Figure 6 The respective implementations can be executed together or one of them can be executed separately.
[0170] when Figure 4 and Figure 6When the respective embodiments are executed together, the image processing device can perform hysteresis control on the three-dimensional environment vector based on the environmental severity level, environment level lock value and dead zone threshold corresponding to the three-dimensional environment vector to obtain the target environment vector; then, parameter mapping processing is performed based on the target environment vector to obtain the initial fusion parameter set; finally, the initial fusion parameter set is subjected to temporal smoothing processing based on the smoothing coefficient and the fusion parameter set of the previous frame to obtain the image fusion parameter set.
[0171] It should be noted that the above Figure 4 and Figure 6 The corresponding implementations are executed together, which is an input locking + output filtering mechanism that prevents parameter breathing effect in static scenes and ensures smooth transition in sudden change scenes.
[0172] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0173] Please see Figure 7 , Figure 7 This is a flowchart illustrating the overall implementation of a video fusion system provided in one embodiment of this application. Figure 7 As shown, the video fusion system includes: an environment perception module, a parameter mapping module, a timing stability control module, and a fusion module.
[0174] Specifically, the environmental perception module is used to receive dual light inputs (i.e., visible light video images and infrared video images) and calculate the illuminance index, visibility index, and dust index. The three indices are used to obtain a three-dimensional environment vector. Among them, the key point in calculating the above three indices is to use infrared information to correct the dust index (see the specific implementation process of step S201 for details).
[0175] The parameter mapping module is used to maintain the mapping engine (i.e., the pre-stored 3D LUT (i.e., the environment parameter mapping table) and the trained parameter mapping model, to take 3D environment vectors as input and output the initial fusion parameter set (see the specific implementation process of step S102 for details).
[0176] The temporal stabilization control module is used to perform motion adaptive filtering on the three-dimensional environment vector and apply a hysteresis strategy to the initial fusion parameters to achieve hysteresis and smooth control, thereby outputting stable parameters, namely the image fusion reference group (see the specific implementation process of steps S401~S403 and steps S501~S502 for details).
[0177] The fusion module performs dehazing enhancement, infrared high-frequency injection, and weighted fusion of various images according to the image fusion parameter group to obtain the final video fusion image (see the specific implementation process of steps S301~S304 for details).
[0178] Corresponding to the video fusion method described in the above embodiments, Figure 8 A schematic diagram of a video fusion apparatus according to an embodiment of this application is shown. For ease of explanation, only the parts relevant to the embodiment of this application are shown. (Refer to...) Figure 8 The video fusion device 800 includes: a solution unit 81, a parameter mapping unit 82, and a fusion unit 83. Wherein: The calculation unit 81 is used to calculate the first feature information of the visible light video image and the second feature information of the infrared video image to obtain a three-dimensional environment vector; the three-dimensional environment vector includes the visibility index, the illuminance index and the dust index.
[0179] The parameter mapping unit 82 is used to perform parameter mapping processing based on the three-dimensional environment vector to obtain an image fusion parameter set; the image fusion parameter set includes multiple image enhancement parameters.
[0180] The fusion unit 83 is used to perform fusion processing on visible light video images and infrared video images based on the image fusion parameter group to obtain a fused video image.
[0181] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0182] Figure 9 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application. Figure 9 As shown, the image processing device 9 of this embodiment includes: at least one processor 90 ( Figure 9(Only one is shown in the diagram), memory 91, and computer program 92 stored in said memory 91 and executable on said at least one processor 90, which, when executed, implements the steps in any of the above video fusion method embodiments.
[0183] The image processing device 9 may include, but is not limited to, a processor 90 and a memory 91. Those skilled in the art will understand that... Figure 9 The image processing device 9 is merely an example and does not constitute a limitation on it. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0184] The processor 90 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0185] In some embodiments, the memory 91 may be an internal storage unit of the image processing device 9, such as the RAM of the image processing device 9. In other embodiments, the memory 91 may be an external storage device of the image processing device 9, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the image processing device 9. Furthermore, the memory 91 may include both internal and external storage units of the image processing device 9. The memory 91 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 91 can also be used to temporarily store data that has been output or will be output.
[0186] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0187] This application provides a computer program product that, when run on an image processing device, enables the image processing device to implement the steps described in the above-described method embodiments.
[0188] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an image processing device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.
[0189] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0190] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A video fusion method, characterized in that, include: The first feature information of the visible light video image and the second feature information of the infrared video image are calculated to obtain a three-dimensional environment vector; the three-dimensional environment vector includes visibility index, illuminance index and dust index. Based on the three-dimensional environment vector, parameter mapping processing is performed to obtain an image fusion parameter set; the image fusion parameter set includes multiple image enhancement parameters. The visible light video image and the infrared video image are fused based on the image fusion parameter set to obtain a fused video image.
2. The video fusion method as described in claim 1, characterized in that, The first feature information includes a brightness value, and the second feature information includes a grayscale value; the step of solving the first feature information of the visible light video image and the second feature information of the infrared video image to obtain a three-dimensional environment vector includes: The dust index is determined based on the brightness value and the grayscale value; The visibility index is determined based on the brightness value and the percentile difference; The illuminance index is calculated based on the luminance value and the image size of the visible light video image.
3. The video fusion method as described in claim 2, characterized in that, Determining the dust index based on the brightness value and the grayscale value includes: For any pixel in the visible light video image, the brightness variance of the set of pixels in the local window corresponding to the arbitrary pixel is calculated to obtain the local variance corresponding to the arbitrary pixel; the local window is constructed with the arbitrary pixel as the center and a distance set as the window size; For any pixel in the infrared video image, the gradient of the grayscale value of the pixel is calculated using an edge detection algorithm to obtain the gradient magnitude of the pixel; any pixel in the infrared video image and any pixel in the visible light video image are the same pixel. The local variance and gradient magnitude of any given pixel are linearly normalized and multiplied to obtain the pixel-level dust intensity of that given pixel. The dust index is determined based on the pixel-level dust intensity of each pixel.
4. The video fusion method as described in claim 1, characterized in that, The parameter mapping process based on the three-dimensional environment vector yields an image fusion parameter set, including: The environmental severity level is obtained by weighted fusion of the three-dimensional environmental vectors. Based on the environmental severity level, environmental level lock value, and dead zone threshold, hysteresis control is performed on the three-dimensional environmental vector to obtain the target environmental vector. The image fusion parameter set is obtained by performing parameter mapping processing based on the target environment vector.
5. The video fusion method as described in claim 1, characterized in that, The parameter mapping process based on the three-dimensional environment vector yields an image fusion parameter set, including: Based on the three-dimensional environment vector, parameter mapping processing is performed to obtain the initial fusion parameter set; The initial fusion parameter set is subjected to temporal smoothing based on the smoothing coefficient and the previous frame fusion parameter set to obtain the image fusion parameter set; the previous frame fusion parameter set refers to the image fusion parameter set associated with the visible light video image acquired at the moment before the acquisition time corresponding to the visible light video image.
6. The video fusion method according to any one of claims 1-5, characterized in that, The multiple image enhancement parameters include infrared weighting, infrared detail injection intensity, dehazing intensity, gamma correction coefficient, and color saturation coefficient; the fusion processing of the visible light video image and the infrared video image based on the image fusion parameter set to obtain a fused video image includes: Based on the dehazing intensity and the gamma correction coefficient, the visible light video image is subjected to dehazing enhancement processing to obtain a dehazing enhanced image; Based on the infrared weights, the dehazing enhanced image and the infrared video image are weighted and fused to obtain an initial fused image; Based on the infrared detail injection intensity and the infrared video image, high-frequency infrared injection is performed on the initial fused image to obtain a high-frequency injected image; The high-frequency injected image is processed based on the color saturation coefficient to obtain the video fusion image.
7. The video fusion method as described in claim 6, characterized in that, The process of performing dehazing enhancement processing on the visible light video image based on the dehazing intensity and the gamma correction coefficient to obtain a dehazing enhanced image includes: The visible light video image is subjected to brightness conversion to obtain a brightness image; The brightness image is processed based on the atmospheric scattering model to obtain an initial fog-free image; Based on the defogging intensity, the initial fog-free image and the visible light video image are weighted and fused to obtain the target fog-free image; The target haze-free image is subjected to gamma correction based on the gamma correction coefficient to obtain the dehaze-enhanced image.
8. The video fusion method as described in claim 6, characterized in that, The step of performing high-frequency infrared injection on the initial fused image based on the infrared detail injection intensity and the infrared video image to obtain a high-frequency injected image includes: The infrared video image is subjected to Gaussian filtering to obtain a low-frequency component image; The infrared video image is processed based on the low-frequency component image to obtain the high-frequency component image; Based on the infrared detail injection intensity and the high-frequency component image, infrared high-frequency injection is performed on the initial fused image to obtain the high-frequency injected image.
9. The video fusion method as described in claim 6, characterized in that, The process of processing the high-frequency injected image based on the color saturation coefficient to obtain the video fused image includes: The high-frequency injected image is processed to obtain a grayscale image; The visible light video image is subjected to color space conversion to obtain the chromaticity information of the visible light video image; Based on the chromaticity information, the high-frequency injected image is subjected to inverse color space transformation to obtain a color reconstructed image; The grayscale image and the color reconstructed image are weighted and fused based on the color saturation coefficient to obtain the video fused image.
10. An image processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video fusion method as described in any one of claims 1 to 9.