Image Fusion Method and Apparatus, Computer Readable Storage Medium, and Terminal
Through multi-round alignment and fusion method and deep learning technology, inter-frame offset and noise reduction problems in image fusion are solved, and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202210143422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-02-16
AI Technical Summary
The prior art fails to effectively solve the inter-frame offset problem and noise reduction processing during image fusion, resulting in low quality of the fused image and sudden changes in manual processing traces and noise levels.
The multi-round alignment and fusion method is used to perform multiple rounds of alignment processing on the features of the reference frame and the matching frame, and the final fusion features are determined, and feature extraction and fusion processing are performed through the residual convolutional neural network and the deformable convolutional neural network.
Effectively solve the problem of inter-frame offset, enhance the dynamic range of the image, improve the quality of the fusion image, ensure consistent noise levels, and improve the fusion effect.
Smart Images

Figure CN114511487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image fusion method and apparatus, a computer-readable storage medium, and a terminal. Background Art
[0002] With the development of computer technology, there is a wide demand for high-quality images in various fields such as security monitoring, vehicle-mounted imaging, medical imaging, and art photography. High-quality images often have a very high dynamic range and can provide rich information and a real visual experience. However, during the image acquisition process, affected by factors such as image acquisition devices, acquisition environments, and noise, single-exposure images often have poor quality, low image dynamic range, and cannot record all the information in the scene. Therefore, it is necessary to use multi-exposure synthesis technology to generate high-dynamic-range images with enhanced details. The purpose of multi-exposure synthesis is to achieve a high dynamic range (High Dynamic Range, HDR) effect while making the image look free of artificial processing traces.
[0003] In the existing image fusion processing technology, when fusing multiple frames of images, the situation of inter-frame offset that often exists in each frame of images collected for the same scene in reality is often not considered. For example, the global inter-frame offset caused by the camera moving more or less during the shooting process, or the local movement caused by the expression change or posture change of the object being photographed. And the problem of inter-frame offset has not been effectively solved, which may reduce the effect of image fusion. In addition, in the existing technology, many HDR algorithms use images after various processes as algorithm inputs. For example, they are processed in the 8-bit (Bit) red-green-blue domain (Red-Green-Blue, RGB) or the luminance-chrominance YUV domain (where "Y" represents luminance Luma, and "U" and "V" represent chroma). Compared with the original RAW domain (original image), these color domains themselves lack some information, and such input images have undergone many non-linear processes, making it impossible to determine the noise intensity in the image. Thus, it increases the difficulty of the algorithm for noise reduction, and the process of various transformations of the processed image is more complex, with a greater computational overhead and lower image fusion efficiency. Moreover, in the existing technology, some also use a frame-by-frame noise reduction method during image fusion, which is not only inefficient and difficult, but also the noise reduction effects of each frame are inconsistent, easily leading to a sudden change in the noise level on the final fused image, seriously affecting the fusion effect.
[0004] Therefore, there is an urgent need for an image fusion method that can effectively solve the inter-frame offset problem and perform noise reduction processing when fusing multiple frames of noisy raw images with different exposure degrees, enhance the image dynamic range, and obtain high-quality fused images. Summary of the Invention
[0005] One of the objectives achieved by the present invention is to provide an image fusion method, which can effectively solve the inter-frame offset problem and perform noise reduction processing when fusing multiple frames of noisy original images with different exposure degrees, and enhance the dynamic range of the image while obtaining a high-quality fused image.
[0006] To achieve the above objective, an embodiment of the present invention provides an image fusion method, including the following steps: respectively performing feature extraction on a reference frame original image and a matching frame original image to obtain a reference frame feature and a matching frame feature; performing multiple rounds of alignment and fusion processing according to the reference frame feature and the matching frame feature to determine a final fusion feature; decoding the final fusion feature to obtain a fused image.
[0007] Optionally, before respectively performing feature extraction on the reference frame original image and the matching frame original image, the method further includes: determining multiple frames of original images with different exposure times and the same noise label collected for the same scene; grouping the original images according to the exposure time of each frame of the original image to obtain multiple groups of grouped images; selecting a group of images from the multiple groups of grouped images as the reference frame original image, and selecting at least one group of images from the remaining images as the matching frame original image; wherein, among the multiple groups of grouped images, the exposure times of the images in each group are different, and the exposure time of each frame of the image in each group is the same.
[0008] Optionally, the algorithm used for feature extraction is a residual convolutional neural network algorithm; the step of respectively performing feature extraction on the reference frame original image and the matching frame original image includes: respectively inputting the reference frame original image and the matching frame original image into a residual block composed of multiple residual convolutional neural networks for feature extraction to obtain the reference frame feature and the matching frame feature.
[0009] Optionally, based on the reference frame features and the matching frame features, multiple rounds of alignment and fusion processing are performed to determine the final fusion features, including: in the first round of alignment and fusion processing, the downsampled reference frame features of the first round and the downsampled matching frame features of the first round obtained by respectively downsampling the reference frame features and the matching frame features by a preset multiple are subjected to alignment and fusion processing to obtain the fusion features of the first round; in each subsequent round of alignment and fusion processing, the upsampled fusion features obtained by upsampling the fusion features obtained in the previous round and the downsampled reference frame features obtained by downsampling the reference frame features are subjected to alignment and fusion processing to obtain the fusion features of this round; until the number of rounds of alignment and fusion processing reaches the preset number of rounds, the fusion features obtained after alignment and fusion processing are used as the final fusion features; wherein, in the multiple rounds of alignment and fusion processing, the multiple of downsampling the reference frame features gradually decreases starting from the preset multiple, and the multiple of upsampling the fusion features obtained in the previous round is equal to the difference between the multiple of downsampling the reference frame features in the previous round and the multiple of downsampling the reference frame features in the current round.
[0010] Optionally, during the alignment of the downsampled reference frame features of the first round and the downsampled matching frame features of the first round, the aligned features of the first round are obtained, and the offset of the first round is determined; each subsequent round of alignment and fusion processing includes: based on the upsampled offset obtained by upsampling the offset determined in the previous round, the upsampled fusion features obtained by upsampling the fusion features obtained in the previous round and the downsampled reference frame features obtained by downsampling the reference frame features are subjected to alignment processing to obtain the aligned features of this round, and the offset of this round is determined; the aligned features of this round and the downsampled reference frame features in this round are subjected to fusion processing to obtain the fusion features of this round; wherein, the multiple of upsampling the offset determined in the previous round is the same as the multiple of upsampling the fusion features obtained in the previous round.
[0011] Optionally, in the multiple rounds of alignment and fusion processing, the algorithm used for alignment processing is the deformable convolutional neural network algorithm.
[0012] Optionally, in the multiple rounds of alignment and fusion processing, the fusion processing performed in each round includes: using the connection function Contact to connect the aligned features of this round and the downsampled reference frame features in this round to obtain the connected features of this round; inputting the connected features of this round into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fusion features of this round.
[0013] Optionally, in each round of the multi-round alignment and fusion process, after performing ghost removal processing on the aligned features obtained in that round to obtain the ghost-removed features of that round, the fusion process is then performed.
[0014] Optionally, performing ghost removal processing on the aligned features obtained in that round to obtain the ghost-removed features of that round includes: using a convolutional neural network algorithm to obtain a convolutional result based on the aligned features obtained in that round and the downsampled reference frame features in that round; using an activation function to determine a ghost removal weight value based on the convolutional result; multiplying the aligned features obtained in that round by the ghost removal weight value to obtain the ghost-removed features of that round.
[0015] Optionally, the activation function is selected from: the sigmoid function Sigmoid, the hyperbolic tangent function Tanh, and the rectified linear unit function ReLU.
[0016] Optionally, in the multi-round alignment and fusion process, the fusion process performed in each round includes: using a connection function Contact to perform a connection process on the ghost-removed features obtained in that round and the downsampled reference frame features in that round to obtain the connected features of that round; inputting the connected features obtained in that round into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fusion features of that round.
[0017] Optionally, after decoding the final fusion features to obtain a fused image, the method further includes: performing image signal processing on the fused image to obtain a color image.
[0018] An embodiment of the present invention further provides an image fusion device, including:
[0019] A feature extraction module, configured to respectively perform feature extraction on a reference frame original image and a matching frame original image to obtain a reference frame feature and a matching frame feature; an alignment and fusion module, configured to perform a multi-round alignment and fusion process according to the reference frame feature and the matching frame feature to determine a final fusion feature; a decoding module, configured to decode the final fusion feature to obtain a fused image.
[0020] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the steps of the above image fusion method are executed.
[0021] An embodiment of the present invention further provides a terminal, including a memory and a processor, where a computer program capable of running on the processor is stored on the memory, and when the processor runs the computer program, the steps of the above image fusion method are executed.
[0022] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:
[0023] In the embodiment of the present invention, feature extraction is respectively performed on the reference frame original image and the matching frame original image to obtain the reference frame feature and the matching frame feature; according to the reference frame feature and the matching frame feature, multiple rounds of alignment and fusion processing are performed to determine the final fusion feature; the final fusion feature is decoded to obtain the fused image. Compared with the prior art, when image fusion is performed, the problem of inter-frame offset is often not effectively solved, resulting in obvious artificial processing traces and low overall quality in the fused image. The embodiment of the present invention adopts a multi-round alignment and fusion method, and before each image fusion process, alignment processing is first performed, which can effectively solve the offset problem between the input images of each frame, and achieve the effect of improving the quality of the fused image while obtaining a high-dynamic-range image. In addition, compared with many existing HDR algorithms that use images after various non-linear processes as algorithm inputs, such images themselves lack a lot of information compared with the original images, and it is difficult to determine the noise intensity, increasing the difficulty of noise reduction, or using a frame-by-frame noise reduction method during image fusion, which is prone to sudden changes in the noise level, ultimately resulting in low-quality fused images. The embodiment of the present invention simultaneously uses multiple original images as algorithm inputs. The original images contain more image information and are easy to calibrate the noise intensity to provide a reference for noise reduction processing, so that the image detail information can be enriched while effectively solving the noise reduction problem and ensuring a consistent noise level, significantly improving the fusion effect.
[0024] Furthermore, the embodiment of the present invention adopts a multi-scale method. First, the input reference frame feature and the feature of the matching frame are respectively downsampled, and then alignment and fusion are performed on a small scale to obtain a fusion feature; then in each round of alignment and fusion processing, the fusion feature obtained in the previous round is upsampled and then aligned and fused with the reference frame feature on a large scale to obtain the fusion feature of the current round. In this way, alignment and fusion processing from small to large scales are performed round by round until a preset number of rounds is reached to determine the final fusion feature. Adopting the above technical solution can improve the accuracy of alignment and the quality of the finally output fused image.
[0025] Further, during the process of aligning the downsampled reference frame features of the first round and the downsampled matching frame features of the first round, an offset of the first round is also determined; in each subsequent round of alignment and fusion processing, by using the upsampled offset obtained by upsampling the offset determined in the previous round (the offset at a small scale) as the initial value, the upsampled fusion features and the downsampled reference frame features in the current round can be aligned to obtain the aligned features of this round, and an offset of the current round (the offset at a large scale) is determined during the alignment process to be used as the initial value for the next round of alignment processing, so that precise alignment processing can be achieved in each round, improving the final fusion effect.
[0026] Further, in the multi-round alignment and fusion processing, the algorithm used for alignment processing is the deformable convolutional neural network algorithm. Among them, the offset added in the deformable convolution unit is a part of the network structure. After adding the learning of this offset, the size and position of the deformable convolution kernel can be dynamically adjusted according to the image content to be recognized currently. Its intuitive effect is that the sampling point positions of the convolution kernels at different positions will adaptively change according to the image content, so as to adapt to geometric deformations such as the shape and size of different objects. Especially for irregular images, a good alignment effect can also be achieved, thereby improving the quality of the finally obtained fused image.
[0027] Further, in each round of the multi-round alignment and fusion processing, after performing ghost removal processing on the aligned features obtained in this round to obtain the ghost-removed features of this round, the fusion processing is then carried out. Among them, the ghost removal processing includes: using the convolutional neural network algorithm to obtain the convolution result according to the aligned features obtained in this round and the downsampled reference frame features in this round; using the activation function to determine the ghost removal weight value according to the convolution result; multiplying the aligned features obtained in this round by the ghost removal weight value to obtain the ghost-removed features of this round. Thus, the appearance of ghosts can be effectively suppressed during the image fusion process, further improving the quality of the obtained fused image. Description of the Drawings
[0028] Figure 1 is a flowchart of an image fusion method in an embodiment of the present invention;
[0029] Figure 2 is Figure 1 a flowchart of a specific implementation manner of step S12 in
[0030] Figure 3 is a schematic diagram of the overall framework of an image fusion model in an embodiment of the present invention;
[0031] Figure 4 is Figure 3Schematic diagram of the basic composition of the image fusion model in
[0032] Figure 5 It is a partial schematic diagram of an image fusion model using a multi-scale method in an embodiment of the present invention;
[0033] Figure 6 It is a flowchart of another image fusion method in an embodiment of the present invention;
[0034] Figure 7 It is a schematic structural diagram of an image fusion device in an embodiment of the present invention. Detailed implementation manners
[0035] As mentioned above, due to the widespread demand for high-quality images in various fields, it is necessary to adopt multi-exposure synthesis technology to generate high-dynamic-range images with enhanced details.
[0036] In the existing image fusion processing technology, when fusing multiple frames of images, the frame-to-frame offset of each frame of images collected for the same scene in reality is often not considered. Therefore, the obvious traces of manual processing of the images may be caused, reducing the image fusion effect; in addition, many HDR algorithms use the images after various processes as the algorithm input, for example, processing in the 8-bit RGB domain or YUV domain. Compared with the original RAW domain (original image), these color spaces have missing detail information, greater difficulty in noise reduction, and complex processing processes. Therefore, the noise reduction effect is poor, and a high-dynamic-range and high-quality fused image cannot be obtained.
[0037] In an embodiment of the present invention, feature extraction is respectively performed on the reference frame original image and the matching frame original image to obtain the reference frame features and the matching frame features; according to the reference frame features and the matching frame features, multiple rounds of alignment and fusion processing are performed to determine the final fusion features; the final fusion features are decoded to obtain the fused image. Compared with the prior art, when image fusion is performed, the problem of inter-frame offset is often not effectively solved, resulting in obvious artificial processing traces and low overall quality in the fused image. The embodiment of the present invention adopts a multi-round alignment and fusion method, and before each image fusion process, alignment processing is first performed, which can effectively solve the offset problem between the input images of each frame, and achieve the effect of obtaining a high-dynamic-range image while improving the quality of the fused image. In addition, compared with many existing HDR algorithms that use images after various non-linear processes as algorithm inputs, such images themselves lack a lot of information compared with the original images, and it is difficult to determine the noise intensity, increasing the difficulty of noise reduction, or using a frame-by-frame noise reduction method during image fusion, which is prone to the phenomenon of sudden change in noise level, ultimately resulting in low quality of the fused image. The embodiment of the present invention simultaneously uses multiple original images as algorithm inputs. The original images contain more image information and are easy to calibrate the noise intensity to provide a reference for noise reduction processing, so that while enriching the image detail information, the noise reduction problem can be effectively solved and the noise level can be ensured to be consistent, significantly improving the fusion effect.
[0038] To make the above objects, features, and beneficial effects of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.
[0039] Refer to Figure 1 , Figure 1 which is a flowchart of an image fusion method in an embodiment of the present invention. The image fusion method may include steps S11 to S14:
[0040] Step S11: Feature extraction is respectively performed on the reference frame original image and the matching frame original image to obtain the reference frame features and the matching frame features;
[0041] Step S12: According to the reference frame features and the matching frame features, multiple rounds of alignment and fusion processing are performed to determine the final fusion features;
[0042] Step S13: The final fusion features are decoded to obtain the fused image.
[0043] In the specific implementation of step S11, the original image can also be referred to as a "RAW image file". The reason for this naming is that such images have not been processed, printed, or used for editing. Usually, the original image stores the data captured from the image sensor in an uncompressed format and has a wide color gamut internal color, allowing for precise adjustment, modification, or processing to generate higher-quality images.
[0044] The reference frame original image and the matching frame original image are respectively used as the initially input images in the embodiments of the present invention. Among them, the reference frame image and the matching frame image are captured for the same scene or the same object and have different exposure times (or exposure degrees); there may be an inter-frame offset between the reference frame original image and the matching frame original image, that is, the global motion and / or local motion between frames. The global motion is the displacement between different frame images caused by a slight change in the spatial position of the image sensor such as a camera or a camera during the image shooting or acquisition process, and the local motion is the displacement between different frame images caused by a slight change in the acquisition object during the image shooting or acquisition process (such as a person's facial expression change, a vehicle's movement, a leaf being blown by the wind, etc.).
[0045] Furthermore, before respectively performing feature extraction on the reference frame original image and the matching frame original image, the method further includes: determining multiple frames of original images captured for the same scene with different exposure times and the same noise label; grouping the original images according to the exposure times of each frame of the original image to obtain multiple groups of grouped images; selecting one group of images from the multiple groups of grouped images as the reference frame original image, and selecting at least one group of images from the remaining images as the matching frame original image; among the multiple groups of grouped images, each group of images has a different exposure time, and the exposure time of each frame of image in each group of images is the same.
[0046] In the specific implementation, the device for capturing the original image can be a smart phone, a tablet computer, a desktop computer, a security monitoring camera, a vehicle-mounted camera, a medical imaging device, etc.
[0047] Among them, the reference frame original image can be one frame of the original image or a group of images containing multiple frames of the original image; the matching frame original image can be one frame of the original image, or a group of images containing multiple frames of the original image, or multiple groups of images respectively containing several frames of the original image. Among them, each group of images has a different exposure time, and the exposure time of each frame of image in each group of images is the same.
[0048] Among them, the exposure time mainly refers to the photosensitive time of the negative film. The longer the exposure time, the brighter the photo generated on the negative film; the shorter the exposure time, the darker the photo generated on the negative film. The multi-exposure fusion technology is to fuse the original images with different exposure times to enhance the information in the images and obtain a high dynamic range effect.
[0049] As a non-limiting embodiment, the reference frame original image may be: an original image with a medium exposure time; the matching frame original images may be: an original image with a long exposure time and an original image with a short exposure time.
[0050] In the embodiments of the present invention, compared with many existing HDR algorithms that use images after various non-linear processes as algorithm inputs, such images lack a lot of information compared with the original images, and it is difficult to determine the noise intensity, increasing the difficulty of noise reduction. Or compared with the frame-by-frame noise reduction method used in image fusion, it is easy to produce a sudden change in the noise level, ultimately resulting in a low quality of the fused image. The embodiments of the present invention simultaneously use multiple frames of original images as algorithm inputs. The original images contain more image information and are easy to calibrate the noise intensity to provide a reference for noise reduction processing. Thus, while enriching the detailed image information, it can effectively solve the noise reduction problem and ensure the consistency of the noise level, significantly improving the fusion effect.
[0051] Furthermore, the algorithm used for feature extraction is the residual convolutional neural network algorithm; the feature extraction of the reference frame original image and the matching frame original image respectively includes: inputting the reference frame original image and the matching frame original image into a residual block composed of multiple residual convolutional neural networks for feature extraction to obtain the reference frame features and the matching frame features.
[0052] Among them, the residual convolutional neural network algorithm is a classic improved algorithm of the convolutional neural network algorithm. The convolutional neural network (CNN) is a type of feedforward neural network (FNN) that contains convolutional calculations and has a deep structure. It is one of the representative algorithms of deep learning. The convolutional neural network has the ability of feature learning and can perform translation-invariant classification on the input information according to its hierarchical structure. Therefore, it is also called the "Shift-Invariant Artificial Neural Networks (SIANN)". The convolutional neural network can be composed of three basic structures: the input layer, the hidden layer, and the output layer. Among them, the hidden layer can include three common architectures: the convolutional layer, the pooling layer, and the fully connected layer. However, in the traditional convolutional neural network, as the depth of the network increases, problems such as gradient explosion and gradient disappearance will occur, making it increasingly difficult to train, and the training error will also increase as the network depth increases. The proposed residual convolutional neural network makes it easier to train deep networks. Its principle is to add the output of a certain previous layer after a linear module and before a non-linear module at a certain layer. This operation is also called skip connection.
[0053] In the specific implementation of step S12, based on the reference frame feature and the matching frame feature, multiple rounds of alignment and fusion processing are performed to determine the final fusion feature.
[0054] Refer to Figure 2 , Figure 2 is Figure 1 a flowchart of a specific implementation manner of step S12 in . The performing multiple rounds of alignment and fusion processing based on the reference frame feature and the matching frame feature to determine the final fusion feature may include steps S21 to S23, which are described below.
[0055] In step S21, in the first round of alignment and fusion processing, the first-round downsampled reference frame feature and the first-round downsampled matching frame feature obtained by respectively performing downsampling on the reference frame feature and the matching frame feature by a preset multiple are subjected to alignment and fusion processing to obtain the first-round fusion feature.
[0056] In step S22, in each subsequent round of alignment and fusion processing, the upsampled fusion feature obtained by upsampling the fusion feature obtained in the previous round and the downsampled reference frame feature obtained by downsampling the reference frame feature are subjected to alignment and fusion processing to obtain the fusion feature of this round.
[0057] Among them, the downsampling may refer to an operation of reducing an image, that is, generating a thumbnail of the corresponding image. The result of downsampling is that the number of pixel points in the image decreases and the scale of the image decreases; the upsampling may refer to an operation of enlarging an image. The result of upsampling is that the number of pixel points in the image increases and the scale of the image increases. Specifically, if a frame of image is regarded as a data set composed of many two-dimensional pixel points, then the purpose of downsampling is to extract a part of the data from the data set to obtain a subset of the data set, that is, the number of two-dimensional pixel points in the downsampled image will be less than the number of two-dimensional pixel points in the original image, while the number of two-dimensional pixel points in the upsampled image will be greater than the number of two-dimensional pixel points in the original image.
[0058] Among them, the scale of the image refers to the resolution of the image. Usually, different features of the image can be observed at different scales, so as to complete different tasks. Generally speaking, finer / denser sampling can see more details, and coarser / sparser sampling can see the overall trend.
[0059] It should be noted that in the multi-round alignment and fusion process, the downsampling multiple of the reference frame features decreases round by round starting from the preset multiple, and the upsampling multiple of the fusion features obtained in the previous round is equal to the difference between the downsampling multiple of the reference frame features in the previous round and the downsampling multiple of the reference frame features in the current round.
[0060] Among them, the gradual decrease starting from the preset multiple may be: the downsampling multiple of the reference frame features in each round decreases by a fixed value compared with the downsampling multiple of the reference frame features in the previous round, or may decrease by different values, as long as it is ensured that in the same round: the upsampled fusion features obtained after upsampling the fusion features obtained in the previous round and the downsampled reference frame features obtained after downsampling the reference frame features have the same scale.
[0061] It should be noted that in the specific implementation, it should be ensured that the upsampling multiple of the fusion features obtained in the previous round is equal to the difference between the downsampling multiple of the reference frame features in the previous round and the downsampling multiple of the reference frame features in the current round. The purpose is to make the scale of the upsampled fusion features in the same round consistent with the scale of the downsampled reference frame features.
[0062] In a specific implementation, the step of downsampling the reference frame may be performed before each round of alignment and fusion processing; alternatively, different multiples of downsampling may be performed on the reference frame features before the first round of alignment and fusion processing to obtain multiple downsampled reference frame features with different scales for backup. In the first round of alignment and fusion processing, the downsampled reference frame feature with the smallest scale is selected for the alignment and fusion processing in the first round. In each subsequent round of alignment and fusion processing, a downsampled reference frame feature with a larger scale than that in the previous round is selected as the downsampled reference frame feature in this round for the alignment and fusion processing in this round.
[0063] In the embodiments of the present invention, compared with the prior art when performing image fusion, the problem of inter-frame offset is often not effectively solved and noise reduction processing is carried out, resulting in obvious artificial processing traces and low overall quality in the fused image. The embodiments of the present invention adopt a multi-scale method. First, the input reference frame features and the features of the matching frame are respectively downsampled, and then aligned and fused at a small scale to obtain fused features; then, in each round of alignment and fusion processing, after upsampling the fused features obtained in the previous round, they are aligned and fused with the reference frame features at a large scale to obtain the fused features of the current round. In this way, the alignment and fusion processing from small to large scales is carried out round by round until the preset number of rounds is reached to determine the final fused features. By adopting the above scheme, the accuracy of alignment can be significantly improved, while obtaining a high-dynamic-range image, effectively solving the problem of inter-frame offset and improving the quality of the fused image.
[0064] Further, in the process of aligning the downsampled reference frame features in the first round and the downsampled matching frame features in the first round, the aligned features in the first round are obtained, and the offset in the first round is determined; each subsequent round of alignment and fusion processing includes: based on the upsampled offset obtained by upsampling the offset determined in the previous round, the upsampled fused features obtained by upsampling the fused features obtained in the previous round and the downsampled reference frame features obtained by downsampling the reference frame features are subjected to alignment processing to obtain the aligned features in this round, and the offset in this round is determined; the aligned features obtained in this round and the downsampled reference frame features in this round are subjected to fusion processing to obtain the fused features in this round.
[0065] In the embodiments of the present invention, by using the upsampled offset obtained by upsampling the offset determined in the previous round (the offset at a small scale) as the initial value, the upsampled fused features and the downsampled reference frame features in the current round can be aligned to obtain the aligned features in this round, and the offset in the current round (the offset at a large scale) is determined during the alignment process to be used as the initial value for the next round of alignment processing. Thus, precise alignment processing can be achieved in each round, improving the final fusion effect.
[0066] Among them, the upsampling multiple of the offset determined in the previous round is the same as the upsampling multiple of the fused feature obtained in the previous round, and the purpose is to make the scale of the upsampled fused feature in the same round consistent with the scale of the upsampled offset.
[0067] Specifically, the offset determined in the first round refers to the offset between the downsampled reference frame feature in the first round and the downsampled matching frame feature in the first round; the offset determined in each subsequent round of alignment and fusion processing refers to the offset between the upsampled fused feature in this round and the downsampled reference frame feature.
[0068] Among them, the offset can refer to the distance of the features of different frame images in the spatial coordinates;
[0069] In a specific implementation, by using the upsampled offset after upsampling the offset determined in the previous round (the offset at a small scale) as the initial value, the upsampled fused feature in the current round and the downsampled reference frame feature can be aligned to obtain the aligned feature in this round, and the offset in the current round (the offset at a large scale) is determined during the alignment process as the initial value for the next round of alignment processing.
[0070] In step S23, when the number of rounds of alignment and fusion processing reaches the preset number of rounds, the fused feature obtained after alignment and fusion processing is used as the final fused feature.
[0071] It can be understood that the relationship between the preset number of rounds and the preset multiple is specifically as follows: when the downsampling multiple of the reference frame feature in each round is reduced by 1 times compared with the downsampling multiple of the reference frame feature in the previous round (the downsampling multiple decreases by 1 times round by round), the preset number of rounds is equal to the preset multiple + 1; when the downsampling multiple of the reference frame feature in each round is reduced by a different value compared with the downsampling multiple of the reference frame feature in the previous round, the preset number of rounds should be less than the preset multiple + 1.
[0072] In a specific implementation manner, the preset multiple is set to 4 times, and the downsampling multiple of the reference frame feature decreases by 1 times round by round, then the preset number of rounds should be set to 5 rounds:
[0073] In the first round, the downsampling multiple of the reference frame feature is 4 times;
[0074] In the second round, the downsampling multiple of the reference frame is 3 times, and the upsampling multiple of the fused feature obtained in the first round is 4 - 3 = 1 time;
[0075] In the third round, the downsampling multiple of the reference frame is 2 times, and the upsampling multiple of the fused feature obtained in the second round is 3 - 2 = 1 time;
[0076] In the fourth round, the downsampling multiple of the reference frame is 1 time, and the upsampling multiple of the fused feature obtained in the third round is 2 - 1 = 1 time;
[0077] In the fifth round, the downsampling multiple of the reference frame is 0 time, and the upsampling multiple of the fused feature obtained in the fourth round is 1 - 0 = 1 time; At this point, the preset number of rounds is reached, and the last round of alignment and fusion processing is ended.
[0078] In another specific implementation, the preset multiple is set to 4 times, and the downsampling multiples of the reference frame features decrease by different multiples round by round (for example: decrease by 1 time, decrease by 2 times, decrease by 1 time respectively), then the preset number of rounds should be set to 4 rounds:
[0079] In the first round, the downsampling multiple of the reference frame features is 4 times;
[0080] In the second round, the downsampling multiple of the reference frame is 4 - 1 = 3 times, and the upsampling multiple of the fused feature obtained in the first round is 4 - 3 = 1 time;
[0081] In the third round, the downsampling multiple of the reference frame is 3 - 2 = 1 time, and the upsampling multiple of the fused feature obtained in the second round is 3 - 1 = 2 times;
[0082] In the fourth round, the downsampling multiple of the reference frame is 1 - 1 = 0 time, and the upsampling multiple of the fused feature obtained in the third round is 1 - 0 = 1 time; At this point, the preset number of rounds is reached, and the last round of alignment and fusion processing is ended.
[0083] Furthermore, in the multi-round alignment and fusion processing, the algorithm used for alignment processing is the deformable convolutional neural network algorithm.
[0084] Among them, the deformable convolutional neural network algorithm (Deformable Convolution Neural Networks, DCNN) is a convolutional neural network algorithm that can perform complex geometric transformation modeling. Since the geometric structure in the module used to construct the convolutional neural network is fixed, its ability to model geometric transformations is essentially limited, while the deformable convolutional neural network improves the ability of the convolutional neural network to model geometric transformations. It is based on the principle of further displacement adjustment of the position information of spatial sampling in the module, and this displacement can be learned in the target task without additional supervision signals.
[0085] In the embodiments of the present invention, through the use of the deformable convolutional neural network algorithm for alignment processing, the added offset in the deformable convolutional unit is a part of the network structure. After adding the learning of this offset, the size and position of the deformable convolutional kernel can be dynamically adjusted according to the image content to be recognized currently. Its intuitive effect is that the sampling point positions of the convolutional kernels at different positions will adaptively change according to the image content, so as to adapt to geometric deformations such as the shape and size of different objects. Especially for irregular images, good alignment effects can also be obtained, thereby improving the quality of the finally obtained fused image.
[0086] Further, in the multi-round alignment and fusion processing, the fusion processing performed in each round includes: using the connection function Contact to connect the aligned features obtained in this round and the downsampled reference frame features in this round to obtain the connected features in this round; inputting the connected features obtained in this round into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fused features in this round.
[0087] Further, in each round of the multi-round alignment and fusion processing, after performing ghost removal processing on the aligned features obtained in this round to obtain the ghost-removed features in this round, the fusion processing is then performed.
[0088] Furthermore, performing ghost removal processing on the aligned features obtained in this round to obtain the ghost-removed features in this round includes: using the convolutional neural network algorithm to obtain a convolutional result according to the aligned features obtained in this round and the downsampled reference frame features in this round; using an activation function to determine a ghost removal weight value according to the convolutional result; multiplying the aligned features obtained in this round by the ghost removal weight value to obtain the ghost-removed features in this round.
[0089] Among them, as some non-limiting embodiments, the activation function can be selected from: the Sigmoid function, the hyperbolic tangent function Tanh, and the rectified linear unit function ReLU.
[0090] Among them, Sigmoid and Tanh are two widely used activation functions, both of which are S-shaped saturation functions. Among them, Sigmoid is used for the output of hidden layer neurons, and its value range is (0, 1). It can map a real number to the interval (0, 1) and can be used for binary classification. The advantage of Sigmoid is that it is smooth and easy to differentiate. Since the output of Sigmoid is always positive and not centered around zero, this will cause the weight update to only occur in one direction, thus affecting the convergence speed. Tanh is an improved version of Sigmoid, derived from two basic hyperbolic functions, hyperbolic sine and hyperbolic cosine. It is a symmetric function centered around zero, with a fast convergence speed and is not prone to fluctuations in the loss value. Also known as the rectified linear unit, it is a commonly used activation function (activation function) in artificial neural networks. ReLU (Linear rectification function) usually refers to non-linear functions represented by the ramp function and its variants. In a neural network, ReLU, as the activation function of a neuron, defines the non-linear output result of the neuron after a linear transformation.
[0091] Further, in the multi-round alignment and fusion process, the fusion process performed in each round includes: using the connection function Contact to connect the deghosted feature obtained in this round and the downsampled reference frame feature in this round to obtain the connected feature in this round; inputting the connected feature obtained in this round into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fusion feature in this round.
[0092] In the embodiment of the present invention, in each round of the multi-round alignment and fusion process, the aligned feature obtained in this round is subjected to deghosting processing to obtain the deghosted feature in this round, and then the obtained deghosted feature and the downsampled reference frame feature in this round are subjected to fusion processing, so that the appearance of ghosts can be effectively suppressed during the image fusion process, and the quality of the obtained fused image can be further improved.
[0093] In specific implementation, for more detailed content regarding steps S21 to S23, please refer to the previous text and Figure 1 perform according to the step descriptions therein, which will not be elaborated here.
[0094] Continue to refer to Figure 1 In step S13, the final fusion feature is decoded to obtain the fused image.
[0095] Among them, decoding is briefly the reverse process of image encoding. Image encoding is the process of extracting features from an image, and image decoding is the process of restoring the extracted features to the image before feature extraction.
[0096] Reference Figure 3 , Figure 3 is a schematic diagram of the overall framework of an image fusion model in an embodiment of the present invention.
[0097] Among them, the inputs of the image fusion model are: the original image with short exposure time, the original image with medium exposure time, and the original image with long exposure time. Among them, the original image with medium exposure time is used as the reference frame original image, and the original image with short exposure time and the original image with long exposure time are used as the matching frame original images; then, the original images with the above three different exposure times are respectively encoded (i.e., feature extraction of the images); then, the reference frame features and the matching frame features extracted are successively subjected to alignment, ghost removal, fusion, and decoding processing; the output of the image fusion model is the fused image, that is: a frame of high-dynamic-range, clean, and ghost-free RAW image.
[0098] It should be noted that in specific implementation, before using the image fusion model for image fusion, it is necessary to train the image fusion model using a training sample set, specifically including:
[0099] Construct a training sample set, the training sample set contains multiple groups of original images. Among them, each frame image in each group of original images is an image collected for the same scene with different exposure times and the same noise label; set a loss function, and use the training sample set to train the image fusion model until the loss function converges and then stop training to obtain the trained image fusion model.
[0100] Further, the setting of the loss function and using the training sample set to train the image fusion model until the loss function converges and then stop training to obtain the trained image fusion model includes: dividing the training sample set into a preset number of training sample subsets; using the L1 norm loss function as the loss function, and inputting the training sample subsets into the image fusion model one by one for training until the loss function converges and then end the training to obtain the trained image fusion model.
[0101] After obtaining the trained image fusion model, input the reference frame original image and the matching frame original image into the trained neural network model to obtain the fused image.
[0102] In specific implementation, for the detailed processes of each processing stage and the input and output of relevant data, refer to the previous text and Figure 1 、 Figure 2 for the relevant descriptions, which will not be elaborated here.
[0103] Reference Figure 4 ,Figure 4 is Figure 3 a schematic diagram of the basic composition of the image fusion model in
[0104] The image fusion model includes an encoding / decoding module 41, an alignment module 42, a deghosting module 43, and a fusion module 44.
[0105] Among them, the encoding / decoding module 41 is composed of residual blocks formed by multiple residual convolutional neural networks (for example, convolutional neural network 1, convolutional neural network 2, convolutional neural network 3, convolutional neural network 4), and is used to extract features from the input original reference frame image and original matching frame image, and output reference frame features and matching frame features;
[0106] The core of the alignment module 42 is a deformable convolutional neural network, which is used to perform alignment processing on the reference frame features and the matching frame features, and output aligned features as the input of the deghosting module 43;
[0107] The deghosting module 43 includes a convolutional neural network and draws on the idea of the attention mechanism: first, the convolutional neural network is used to perform convolution on the reference frame features and the aligned features to obtain a convolutional structure; then, the sigmoid function is used to determine the deghosting weight value according to the convolution result; then, the aligned features are multiplied by the deghosting weight value, and the deghosted features are output as the input of the fusion module 44;
[0108] The fusion module 44 includes multiple residual blocks. First, the reference frame features and the deghosted features are connected by contact to obtain connected features, and then the residual blocks are used to fuse the connected features to output fused features.
[0109] In specific implementation, for the detailed running processes of the various modules of the image fusion model, refer to the previous text and Figures 1 to 3 the description in
[0110] Refer to Figure 5 , Figure 5 is a partial schematic diagram of an image fusion model using a multi-scale method in an embodiment of the present invention.
[0111] Among them, S is used to represent the total number of pixels in the image. Each time the downsampling is doubled, the total number of pixels in the image becomes 1 / 4 of the original; W is used to represent the number of pixels in the horizontal direction of the image, H is used to represent the number of pixels in the vertical direction of the image, C is used to represent the color depth of the image, and n is used to represent the sampling multiple (the maximum value of n is the preset multiple).
[0112] Specifically, for each doubling of downsampling of the image, the number of pixel points in the image becomes 1 / 4 of the original. For example, for an image represented by W×H×C, the image obtained after one-time downsampling of the image is W / 2 1 ×H / 2 1 ×C×2 2 ; the image obtained after two-time downsampling of the image is W / 2 2 ×H / 2 2 ×C×2 4 ; the image obtained after three-time downsampling of the image is W / 2 3 ×H / 2 3 ×C×2 6 ……; the image obtained after n-time downsampling of the image is W / 2 n ×H / 2 n ×C×2 2n .
[0113] Among them, in each round of its fusion processing, the multiple of downsampling the reference frame features is reduced by 1 times compared with the multiple of downsampling the reference frame features in the previous round (the multiple of downsampling is reduced by 1 times round by round). When the preset multiple is n, the preset number of rounds is n + 1.
[0114] Among them, in each round of its fusion processing, the upsampled fusion feature obtained by upsampling the fusion feature obtained in the previous round and the downsampled reference frame feature obtained by downsampling the reference frame feature always have the same scale.
[0115] It should be noted that Figure 5 the image fusion model adopting the multi-scale method shown is only a non-limiting embodiment of the present invention. In specific implementation, the multiple of downsampling the reference frame features in each round can be reduced by a fixed value (such as Figure 5 reduced by 1 time round by round as shown), or different values can be reduced, as long as it is ensured that in the same round, the multiple of upsampling the fusion feature obtained in the previous round is equal to the difference between the multiple of downsampling the reference frame features in the previous round and the multiple of downsampling the reference frame features in the current round, so that the scale of the upsampled fusion feature in the same round is consistent with the scale of the downsampled reference frame feature.
[0116] Referring to Figure 6 , Figure 6 is a flowchart of another image fusion method in an embodiment of the present invention. The another image fusion method may include steps S61 to S67, which will be described below.
[0117] In step S61, determine the original images with different exposure times and the same noise label collected for the same scene in multiple frames, and then group the original images according to the exposure times of each frame of the original images to obtain multiple groups of grouped images.
[0118] In step S62, select a group of images from the multiple groups of grouped images as the reference frame original image, and select at least one group of images from the remaining images as the matching frame original image.
[0119] Among the multiple groups of grouped images, the exposure times between the groups of images are different, and the exposure time of each frame of image in each group of images is the same.
[0120] In step S63, perform feature extraction on the reference frame original image and the matching frame original image to obtain a reference frame feature and a matching frame feature.
[0121] In step S64, in the first-round alignment and fusion process, perform alignment and fusion processing on the first-round downsampled reference frame feature and the first-round downsampled matching frame feature obtained by respectively performing downsampling on the reference frame feature and the matching frame feature by a preset multiple to obtain the first-round fusion feature.
[0122] In step S65, in each subsequent round of alignment and fusion process, perform alignment and fusion processing on the upsampled fusion feature obtained by upsampling the fusion feature obtained in the previous round and the downsampled reference frame feature obtained by downsampling the reference frame feature to obtain the fusion feature of this round.
[0123] In step S66, until the number of rounds of alignment and fusion processing reaches the preset number of rounds, use the fusion feature obtained after alignment and fusion processing as the final fusion feature.
[0124] Among them, in the above-mentioned multiple rounds of alignment and fusion processing, the multiple of downsampling the reference frame feature gradually decreases starting from the preset multiple, and the multiple of upsampling the fusion feature obtained in the previous round is equal to the difference between the multiple of downsampling the reference frame feature in the previous round and the multiple of downsampling the reference frame feature in the current round.
[0125] In step S67, decode the final fusion feature to obtain a fused image.
[0126] In step S68, perform image signal processing on the fused image to obtain a color image.
[0127] Among them, Image Signal Processing (ISP) is generally used to process the output data of an image sensor, such as performing functions like automatic exposure control, automatic gain control, automatic white balance, color correction, and removing bad pixels.
[0128] In specific implementation, for more detailed content regarding steps S61 to S68, please refer to the foregoing and Figures 1 to 5 perform according to the step descriptions therein, which will not be elaborated here.
[0129] Refer to Figure 7 , Figure 7 is a schematic structural diagram of an image fusion device in an embodiment of the present invention. The image fusion device may include:
[0130] A feature extraction module 71, configured to perform feature extraction on the reference frame original image and the matching frame original image respectively to obtain a reference frame feature and a matching frame feature;
[0131] An alignment and fusion module 72, configured to perform multiple rounds of alignment and fusion processing according to the reference frame feature and the matching frame feature to determine a final fusion feature;
[0132] A decoding module 73, configured to decode the final fusion feature to obtain a fused image.
[0133] Regarding the principle, specific implementation, and beneficial effects of this image fusion device, please refer to the foregoing and Figures 1 to 6 the related descriptions of the image fusion method shown therein, which will not be elaborated here.
[0134] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above image fusion method. The computer-readable storage medium may include non-volatile memory or non-transitory memory, and may also include optical discs, mechanical hard disks, solid-state drives, etc.
[0135] Specifically, in the embodiments of the present invention, the processor may be a central processing unit (CPU for short), and the processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0136] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory. The volatile memory may be a random access memory (RAM for short), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM for short) are available, such as static random access memory (SRAM for short), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM for short), double data rate synchronous dynamic random access memory (DDR SDRAM for short), enhanced synchronous dynamic random access memory (ESDRAM for short), synchlink dynamic random access memory (SLDRAM for short), and direct rambus random access memory (DR RAM for short).
[0137] An embodiment of the present invention further provides a terminal, including a memory and a processor. A computer program capable of running on the processor is stored on the memory. When the processor runs the computer program, the steps of the above image fusion method are executed. The terminal may include, but is not limited to, terminal devices such as mobile phones, computers, and tablets, and may also be a server, a cloud platform, etc.
[0138] It should be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article indicates that the associated objects before and after are in an "or" relationship.
[0139] In the embodiments of this application, "a plurality of" refers to two or more.
[0140] In the embodiments of this application, the descriptions such as first and second are only for schematic and distinguishing the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation to the embodiments of this application.
[0141] It should be noted that the sequence numbers of the steps in this embodiment do not represent the limitation of the execution sequence of each step.
[0142] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. An image fusion method, characterized in that, it includes: Performing feature extraction on the reference frame original image and the matching frame original image respectively to obtain the reference frame features and the matching frame features; According to the reference frame features and the matching frame features, performing multiple rounds of alignment and fusion processing to determine the final fusion features; Decoding the final fusion features to obtain the fused image; Among them, according to the reference frame features and the matching frame features, performing multiple rounds of alignment and fusion processing to determine the final fusion features includes: In the first round of alignment and fusion processing, performing alignment and fusion processing on the first-round downsampled reference frame features and the first-round downsampled matching frame features obtained by respectively downsampling the reference frame features and the matching frame features by a preset multiple to obtain the first-round fusion features; In each subsequent round of alignment and fusion processing, performing alignment and fusion processing on the upsampled fusion features obtained by upsampling the fusion features obtained in the previous round and the downsampled reference frame features obtained by downsampling the reference frame features to obtain the fusion features of this round; Until the number of rounds of alignment and fusion processing reaches the preset number of rounds, taking the fusion features obtained after alignment and fusion processing as the final fusion features; Among them, in the multiple rounds of alignment and fusion processing, the multiple of downsampling the reference frame features decreases round by round starting from the preset multiple, and the multiple of upsampling the fusion features obtained in the previous round is equal to the difference between the multiple of downsampling the reference frame features in the previous round and the multiple of downsampling the reference frame features in the current round.
2. The method according to claim 1, characterized in that, Before performing feature extraction on the reference frame original image and the matching frame original image respectively, the method further includes: Determining multiple original images collected for the same scene with different exposure times and the same noise label annotation; Grouping the original images according to the exposure times of each frame of the original images to obtain multiple groups of grouped images; Selecting a group of images from the multiple groups of grouped images as the reference frame original image, and selecting at least one group of images from the remaining images as the matching frame original image; Among them, among the multiple groups of grouped images, the exposure times of each group of images are different, and the exposure time of each frame of image in each group of images is the same.
3. The method according to claim 1 or 2, characterized in that, The algorithm used for feature extraction is the residual convolutional neural network algorithm; The performing feature extraction on the reference frame original image and the matching frame original image respectively includes: Respectively inputting the reference frame original image and the matching frame original image into a residual block composed of multiple residual convolutional neural networks for feature extraction to obtain the reference frame features and the matching frame features.
4. The method according to claim 1, characterized in that, During the alignment process of the first-round downsampled reference frame features and the first-round downsampled matching frame features, obtaining the first-round aligned features and determining the first-round offset; Each subsequent round of alignment and fusion processing includes: Based on the upsampled offset obtained by upsampling the offset determined in the previous round, the upsampled fused feature obtained by upsampling the fused feature obtained in the previous round and the downsampled reference frame feature obtained by downsampling the reference frame feature are aligned to obtain the aligned feature of this round, and the offset of this round is determined; The aligned feature obtained in this round and the downsampled reference frame feature in this round are fused to obtain the fused feature of this round; Among them, the multiple of upsampling the offset determined in the previous round is the same as the multiple of upsampling the fused feature obtained in the previous round.
5. The method according to claim 4, wherein, In the multi-round alignment and fusion processing, the algorithm used for alignment processing is the deformable convolutional neural network algorithm.
6. The method according to claim 4, wherein, In the multi-round alignment and fusion processing, the fusion processing performed in each round includes: Using the connection function Contact, the aligned feature obtained in this round and the downsampled reference frame feature in this round are connected to obtain the connected feature of this round; The connected feature obtained in this round is input into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fused feature of this round.
7. The method according to claim 4, wherein, In each round of the multi-round alignment and fusion processing, after performing ghost removal processing on the aligned feature obtained in this round to obtain the ghost-removed feature of this round, the fusion processing is performed.
8. The method according to claim 7, wherein, Performing ghost removal processing on the aligned feature obtained in this round to obtain the ghost-removed feature of this round includes: Using the convolutional neural network algorithm, based on the aligned feature obtained in this round and the downsampled reference frame feature in this round, a convolutional result is obtained; Using an activation function, based on the convolutional result, a ghost removal weight value is determined; The aligned feature obtained in this round is multiplied by the ghost removal weight value to obtain the ghost-removed feature of this round.
9. The method according to claim 8, wherein, The activation function is selected from: Sigmoid function Sigmoid, hyperbolic tangent function Tanh, rectified linear unit function ReLU.
10. The method according to claim 7, wherein, In the multi-round alignment and fusion processing, the fusion processing performed in each round includes: Using the connection function Contact, the ghost-removed feature obtained in this round and the downsampled reference frame feature in this round are connected to obtain the connected feature of this round; The connected feature obtained in this round is input into a residual block composed of multiple residual convolutional neural networks for fusion to obtain the fused feature of this round.
11. The method according to claim 1, wherein, After decoding the final fused feature to obtain a fused image, the method further includes: Performing image signal processing on the fused image to obtain a color image.
12. An image fusion device, wherein, Comprising: A feature extraction module, configured to perform feature extraction on the reference frame original image and the matching frame original image respectively to obtain a reference frame feature and a matching frame feature; An alignment and fusion module, configured to perform multiple rounds of alignment and fusion processing based on the reference frame feature and the matching frame feature to determine a final fusion feature; A decoding module, configured to decode the final fusion feature to obtain a fused image; Wherein, the alignment and fusion module performs the following steps: In the first round of alignment and fusion processing, perform alignment and fusion processing on the first-round downsampled reference frame feature and the first-round downsampled matching frame feature obtained by respectively downsampling the reference frame feature and the matching frame feature by a preset multiple to obtain a first-round fusion feature; In each subsequent round of alignment and fusion processing, perform alignment and fusion processing on the upsampled fusion feature obtained by upsampling the fusion feature obtained in the previous round and the downsampled reference frame feature obtained by downsampling the reference frame feature to obtain the fusion feature of this round; Until the number of rounds of alignment and fusion processing reaches a preset number of rounds, use the fusion feature obtained after alignment and fusion processing as the final fusion feature; Wherein, in the multiple rounds of alignment and fusion processing, the multiple of downsampling the reference frame feature gradually decreases starting from the preset multiple, and the multiple of upsampling the fusion feature obtained in the previous round is equal to the difference between the multiple of downsampling the reference frame feature in the previous round and the multiple of downsampling the reference frame feature in the current round.
13. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when run by a processor, executes the steps of the image fusion method according to any one of claims 1 to 11.
14. A terminal, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, wherein, when the processor runs the computer program, it executes the steps of the image fusion method according to any one of claims 1 to 11.
Citation Information
Patent Citations
High-dynamic image reconstruction method based on neural network
CN111986106A
System and method for compositing high dynamic range images
US20200267300A1