MSF-Net network model training method, image fusion method and system

Through the MSF-Net network model training method, the target MSF-Net network model is trained using the total loss value of the sum of the pixel intensity loss value and the fusion quality loss value, which solves the problem that the fusion image is not close to the actual high dynamic range scene in the prior art, and achieves a high dynamic range image fusion effect closer to the actual scene.

CN119251615BActive Publication Date: 2025-05-23WUHAN UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411305602.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-23
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

The fusion images obtained by existing exposure fusion algorithms are closer to the image set to be fused, rather than actual high dynamic range scenarios. Especially on mobile devices with limited computing resources, it is difficult to effectively capture and retain detailed information of high-brightness areas and low-brightness areas.

Method used

A MSF-Net network model training method is proposed. By taking a pair of first image sets and second image sets as samples, multiple samples are obtained as training sets, and inputting pre-trained MSF-Net network model for training, the sum of pixel intensity loss values ​​and fusion mass loss values ​​is used as the total loss value, and the network parameters are adjusted until the total loss value converges, and the target MSF-Net network model is obtained.

Benefits of technology

Through the new loss function, the fused high dynamic range image not only includes the information of the image set to be fused, but also learns information of other different exposed images in the same high dynamic range scene, making the fused image closer to the actual high dynamic range scene, and maintains the scene depth and local contrast to avoid dizziness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251615B_ABST
    Figure CN119251615B_ABST
Patent Text Reader

Abstract

The present invention relates to a MSF-Net network model training method, an image fusion method and a system. The method comprises: taking a pair of first image sets and second image sets as a sample, obtaining multiple samples as training sets; inputting the training sets into a pre-trained MSF-Net network model for training, and the training process is as follows: inputting each of the second image sets into the pre-trained MSF-Net network model for image fusion, and obtaining fused high dynamic range images corresponding to each second image set; obtaining each pair of pixel intensity loss values ​​and fusion quality loss values; taking the sum of each pair of pixel intensity loss values ​​and fusion quality loss values ​​as the total loss value of the pair of images; adjusting the network parameters in the pre-trained MSF-Net network model based on each total loss value until the target MSF-Net network model is obtained. The fused high dynamic range image obtained by the method is closer to the actual high dynamic range scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an MSF-Net network model training method, an image fusion method and a system. Background Art

[0002] A high-contrast natural scene has a high dynamic range, with brightness ranging from 0.1 to 0.1, and the dynamic range can reach up to 10 orders of magnitude. However, the dynamic range that can be captured by a single exposure is very limited, and image data is usually recorded using 8 bits, resulting in low dynamic range images and inevitable over / under exposure areas. Therefore, in very bright or dark situations, a lot of detail information will be lost, which seriously affects the performance of machine vision tasks such as intelligent driving and navigation. High dynamic range imaging has been introduced to effectively solve this problem. Due to the limited information captured by a single image, high dynamic range methods based on a single image usually perform poorly. Capturing multiple images with different exposures provides an effective solution for high dynamic range imaging, which can retain rich details and vivid colors. Although camera motion and moving objects are common problems with multiple images, images with different exposures can be aligned.

[0003] At present, multi-exposure fusion algorithms are mainly divided into model-driven algorithms and data-driven algorithms. However, these algorithms only make the fused image close to the set of images to be fused, rather than the actual high dynamic range scene. When these algorithms are used, they all assume that each high dynamic range scene can capture enough low dynamic range images with different exposures with a normalized exposure ratio. However, this assumption is usually incorrect, especially for mobile devices with limited computing resources. Generally speaking, a high dynamic range scene only captures a few low dynamic range images with different exposures. For example, when the input is two images with large exposure ratios, the exposure fusion algorithm will produce serious light and dark inversion. If the input is three images with large exposures, the detail information of the high brightness area and the low brightness area in the high dynamic range scene is difficult to retain.

[0004] Therefore, although the prior art proposes a variety of exposure fusion algorithms, the fused images obtained by these exposure fusion algorithms are closer to the image set to be fused rather than the actual high dynamic range scene. Summary of the invention

[0005] In order to overcome the technical problem that the fused image obtained by the traditional exposure fusion algorithm is closer to the image set to be fused rather than the actual high dynamic range scene, the present invention provides an MSF-Net network model training method, an image fusion method and a system.

[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a MSF-Net network model training method, comprising:

[0007] A pair of a first image set and a second image set is taken as a sample, and a plurality of samples are obtained as a training set; wherein the first image set includes a plurality of low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of a plurality of low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set;

[0008] The training set is input into the pre-trained MSF-Net network model (Multi-Scale Fusion Network) for training. The training process is as follows:

[0009] S121, inputting each of the second image sets into the pre-trained MSF-Net network model for image fusion, to obtain a fused high dynamic range image corresponding to each of the second image sets;

[0010] S122, obtaining a pixel intensity loss value between each pair of the high dynamic range images and the corresponding first image set, and obtaining a fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set;

[0011] S123, taking the sum of the pixel intensity loss value and the fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set as the total loss value;

[0012] S124, adjusting the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeating steps S121 to S124 until each of the total loss values ​​converges, and / or the number of iterations reaches a preset number, and then using the current pre-trained MSF-Net network model as the target MSF-Net network model.

[0013] In a second aspect, in order to solve the above technical problem, the present invention provides an image fusion method, which is applied to the MSF-Net network model training method shown above, comprising:

[0014] Obtain multiple low dynamic range images to be fused at different exposures in the same high dynamic range scene;

[0015] The multiple low dynamic range images to be fused are input into the trained target MSF-Net network model to obtain a fused target high dynamic range image.

[0016] In a third aspect, in order to solve the above technical problems, the present invention provides a MSF-Net network model training system, comprising:

[0017] An acquisition module, configured to take a pair of a first image set and a second image set as a sample, and acquire multiple samples as a training set; wherein the first image set includes multiple low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of multiple low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set;

[0018] A training module, used for inputting the training set into a pre-trained MSF-Net network model for training, wherein the training module includes a fusion unit, a first loss value unit, a second loss value unit and a target unit;

[0019] The fusion unit is used to input each of the second image sets into the pre-trained MSF-Net network model for image fusion, so as to obtain a fused high dynamic range image corresponding to each of the second image sets;

[0020] The first loss value unit is used to obtain a pixel intensity loss value between each pair of the high dynamic range images and the corresponding first image set, and to obtain a fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set;

[0021] The second loss value unit is used to take the sum of the pixel intensity loss value and the fusion quality loss value between each pair of the high dynamic range image and the corresponding first image set as the total loss value;

[0022] The target unit is used to adjust the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeatedly execute steps S21 to S24 until each of the total loss values ​​converges, and / or the number of iterations reaches a preset number, and then the current pre-trained MSF-Net network model is used as the target MSF-Net network model.

[0023] The beneficial effect of the present invention is that since the real high dynamic range image is often not available, this embodiment makes full use of the asymmetry between the training and inference (or test) stages, and proposes a new loss function (the sum of the pixel intensity loss value and the fusion quality loss value). In addition to the image set to be fused (i.e., the second image set) to be fused, other low dynamic range images of different exposures from the same high dynamic range scene are also used to define the loss function. Therefore, the fused high dynamic range image not only includes the image information of all low dynamic range images of different exposures in the second image set, but also can learn the image information of other low dynamic range images of different exposures in the same high dynamic range scene as the second image set, so that the fused high dynamic range image is closer to the actual high dynamic range scene. Among them, the new loss function is decoupled from the image set to be fused, so that the asymmetry between the training stage and the inference stage can be well utilized. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A flow chart of a MSF-Net network model training method provided by the present invention;

[0025] Figure 2 A flow chart of an image fusion method provided by the present invention;

[0026] Figure 3 An algorithm logic diagram of a target MSF-Net network model provided by the present invention;

[0027] Figure 4 A schematic diagram of the structure of a MSRRG provided by the present invention;

[0028] Figure 5 A structural schematic flow chart of an MSF-Net network model training system provided by the present invention. DETAILED DESCRIPTION

[0029] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0030] The following describes an MSF-Net network model training method, image fusion method and system according to an embodiment of the present invention in conjunction with the accompanying drawings.

[0031] like Figure 1 As shown, the embodiment of the present disclosure provides a MSF-Net network model training method, including:

[0032] Step S11, taking a pair of the first image set and the second image set as a sample, obtaining multiple samples as a training set, wherein the first image set includes multiple low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of multiple low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set.

[0033] It can be understood that all low dynamic range images in the first image set and the second image set are aligned.

[0034] Step S12, input the training set into the pre-trained MSF-Net network model for training. The training process is as follows:

[0035] Step S121, input each second image set into the pre-trained MSF-Net network model for image fusion, and obtain a fused high dynamic range image corresponding to each second image set.

[0036] Step S122, obtaining pixel intensity loss values ​​between each pair of high dynamic range images and the corresponding first image set, and obtaining fusion quality loss values ​​between each pair of high dynamic range images and the corresponding first image set.

[0037] It can be understood that the high dynamic range image is obtained by fusing all the images in the second image set, so one second image set corresponds to one fused high dynamic range image. Since one second image set corresponds to one first image set, there is also a corresponding relationship between a fused high dynamic range image and the first image set corresponding to its corresponding second image set. That is, the high dynamic range image is an image obtained by fusing all the images in the second image set corresponding to its corresponding first image set.

[0038] Step S123: The sum of the pixel intensity loss value and the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set is taken as the total loss value.

[0039] Step S124, adjust the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeat steps S121 to S124 until each total loss value converges, and / or the number of iterations reaches a preset number, then use the current pre-trained MSF-Net network model as the target MSF-Net network model.

[0040] It can be understood that a total loss value is the sum of a pixel intensity loss value and a fusion quality loss value between a high dynamic range image and a corresponding first image set.

[0041] A MSF-Net network model training method provided by the disclosed embodiment is adopted, and the beneficial effect of the present invention is as follows: since the real high dynamic range image is often not available, the present embodiment makes full use of the asymmetry between the training and inference (or test) stages, and proposes a new loss function (the sum of the pixel intensity loss value and the fusion quality loss value). In addition to the image set to be fused (i.e., the second image set) to be fused, other low dynamic range images of different exposures from the same high dynamic range scene are also used to define the loss function. Therefore, the fused high dynamic range image not only includes the image information of all low dynamic range images of different exposures in the second image set, but also can learn the image information of other low dynamic range images of different exposures in the same high dynamic range scene as the second image set, so that the fused high dynamic range image is closer to the actual high dynamic range scene. Among them, the new loss function is decoupled from the image set to be fused, so that the asymmetry between the training stage and the inference stage can be well utilized.

[0042] In addition, the MSF-Net network model (also known as the multi-scale exposure fusion network structure, Multi-Scale Fusion Network) of this embodiment can maintain the scene depth and local contrast of the image to be fused, and retain the information of the brightest and darkest areas, avoiding the appearance of halo.

[0043] Preferably, obtaining the pixel intensity loss value between each pair of high dynamic range images and the corresponding first image set includes: for each pair of high dynamic range images and the corresponding first image set, obtaining first contrast, saturation and exposure metrics of each image in the first image set, and determining the weight of each pixel point of each image in the first image set based on the first contrast, saturation and exposure metrics of each image. Determine the corresponding pixel intensity loss value based on the weight of each pixel point of each image in the high dynamic range image and the corresponding first image set.

[0044] The pixel intensity loss value is used to characterize the intensity distribution difference at the pixel level between the fused high dynamic range image and all the images in the corresponding first image set. Therefore, the weights of each pixel of each image in the first image set are obtained through the first contrast, saturation and exposure metrics, and based on the weights of each pixel of the high dynamic range image and the corresponding first image set, the intensity distribution difference at the pixel level between the high dynamic range image and the corresponding first image set can be evaluated from the three aspects of contrast, saturation and exposure metrics, which is conducive to the pre-trained MSF-Net network model to learn other high dynamic information in the same high dynamic range scene in addition to the image set to be fused (i.e., the second image set) when being trained.

[0045] In some embodiments, the weight of each pixel point of each image in the first image set is calculated by the following formula: Among them, W′ k (p) is the weight of P pixel in image k, S k 、E k They represent the first contrast, saturation and exposure metrics (also called good exposure metrics) corresponding to the kth weight value. is the first contrast of pixel P in image k, S k is the saturation of pixel P in image k, E k is the exposure metric of pixel P in image k. It can be understood that more reliable areas containing bright colors and details will be given greater weights (i.e., weights), so that the target MSF-Net network model pays more attention to and obtains more reliable information in the first image set. In this embodiment, P represents the coordinates of pixel P.

[0046] Furthermore, the weight of each pixel is normalized, and the calculation formula is as follows:

[0047] Among them, W m(k′) (p) is the normalized weight of the pixel point P of the k'th image in the first image set, θ 2 is the number of weights.

[0048] Preferably, determining the corresponding pixel intensity loss value based on the weights of each pixel of each image in the high dynamic range image and the corresponding first image set includes: determining the pixel value of each pixel of the high dynamic range image based on the high dynamic range image. Obtaining the pixel value of each pixel of each image in the corresponding first image set. Determining the corresponding pixel intensity loss value based on the pixel value of each pixel of the high dynamic range image, the pixel value of each pixel of each image in the corresponding first image set, and the weights of each pixel of each image in the corresponding first image set.

[0049] By using the pixel value of each point in the high dynamic range image, the pixel value of each point in each image in the corresponding first image set, and the weight of each pixel point in each image in the corresponding first image set, the difference in intensity distribution at the pixel level between the high dynamic range image and the corresponding images in the first image set can be determined more accurately.

[0050] In some embodiments, based on the pixel value of each pixel of the high dynamic range image, the pixel value of each pixel of each image in the corresponding first image set, and the weight of each pixel of each image in the corresponding first image set, determining the corresponding pixel intensity loss value includes calculating the pixel intensity loss value by the following formula:

[0051] Among them, Z F (p) is the pixel value of the high dynamic range image at pixel point P, Z m(k′) (p) is the pixel value of the kth image in the first image set at pixel point P.

[0052] Preferably, obtaining the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set includes: for each pair of high dynamic range images and the corresponding first image set, obtaining multiple first patch blocks of the high dynamic range images, and obtaining multiple groups of second patch blocks of the first image set. Among them, one first patch block corresponds to a group of second patch blocks, and a group of patch blocks are patch blocks at the same position of all images in the first image set. Each second patch block in the same group is processed to obtain a corresponding fused patch block; based on each first patch block and the fused patch block corresponding to each first patch block, the fusion quality loss value is determined.

[0053] It can be understood that, for different images in the first image set, the same serial number of the second patch block indicates that the position of the second patch with the same serial number on the image is the same. For example, the first image set includes 5 images, and each image includes the 1st second patch block, the 2nd second patch block and the 3rd second patch block. And the position of the 1st patch block on each image is the same, the position of the 2nd patch block on each image is the same, and the position of the 3rd patch block on each image is the same. Therefore, the set of the 1st second patch blocks of the 5 images is a group of second patch blocks in the first image set, which is the 1st group of second patch blocks. The set of the 2nd second patch blocks of the 5 images is a group of second patch blocks in the first image set, which is the 2nd group of second patch blocks. The set of the 3rd second patch blocks of the 5 images is a group of second patch blocks in the first image set, which is the 3rd group of second patch blocks. In addition, if the first patch block is in area A of the fused high dynamic range image, then a corresponding set of second patch blocks is in area A of the first image set, and the area A is a collection of areas A of all images in the first image set.

[0054] The patch block represents the local image obtained after the local extraction operation is performed on the image. The fusion quality loss value represents the quality difference between the fusion image actually obtained and the fusion image expected to be obtained from the first image set, that is, represents the quality of the fusion image actually obtained. In this embodiment, the MEF-SSIM (Multi-Exposure Fusion-Structural Similarity) index is used as the objective function of the fusion quality loss value, which can effectively measure the quality of the fusion image.

[0055] In this way, by extracting the local image of the fused high dynamic range image (i.e., the first patch block) and the local image of each image in the first image set (i.e., the second patch block), the local image of the same position of each image in the first image set is processed to obtain the corresponding fused patch block. The fused patch block is actually the expected local image of the first image set after the fusion of the local image at that position. Therefore, by obtaining the local image, it is convenient to evaluate the quality difference between the fused high dynamic range image and the fused image expected to be obtained in the first image set from the perspective of the local image, so as to improve the quality of the fused high dynamic range image, so that the fused high dynamic range image is closer to the fused image expected to be obtained in the first image set.

[0056] Preferably, the second patch blocks of the same group are processed to obtain the corresponding fused patch blocks, including: determining the strength, second contrast and structural measurement value of each second patch block based on the second patch blocks of the same group. An expected strength is determined based on the strength of each second patch block of the same group, an expected contrast is determined based on the second contrast of each second patch block, and an expected structural measurement value is determined based on the structural measurement value of each second patch block of the same group. Based on the expected strength, expected contrast and expected structural measurement value of the same group, the corresponding fused patch block is obtained.

[0057] In this way, by decomposing the second patch blocks into intensity, second contrast and structural measurement values, it is convenient to determine the expected intensity, expected contrast and expected structural measurement values ​​pointed to by the same group of second patch blocks in the first image set, thereby facilitating the determination of a fused patch block pointed to by the group of second patch blocks.

[0058] In some embodiments, R i (·) is an operator that extracts the ith second patch from the image, i.e., R i (Z m(k′) ) is from the first image set Ω m Image Z in m(k′) The i-th second patch extracted from is the i-th second patch of the m-th image in the first image set. The second patch block R is converted into i (Z m(k′) ) is decomposed into three independent components:

[0059] R i (Z m(k′) )=c m(k′),i S m(k′),i +l m(k′),i , where l m(k′)i 、c m(k′),i and m(k′),i Respectively represent the second patch block R i (Z m(k′)), the second contrast and the structural metric value, whose mathematical expressions are: in, The second patch R i (Z m(k′) ) is the mean value of the intensity.

[0060] Since the second patch block R has been obtained i (Z m(k′) ), so after determining the intensity of the second patch, the contrast and structure metrics of the second patch can be calculated. The calculation formula is as follows:

[0061]

[0062] Specifically, determining an expected contrast based on the second contrasts of the second patches in the same group includes: determining the largest second contrast among the second patches in the same group as the expected contrast corresponding to the group. The calculation formula is as follows: c i ' represents the expected contrast corresponding to the i-th group, and the i-th second patch block of all images in the first image set is determined as the i-th group. A higher contrast means a second patch block with higher quality, so the expected contrast is the highest second contrast in the second patch block of the group.

[0063] Specifically, determining an expected structure metric value based on the structure metric values ​​of each second patch block in the same group includes: a calculation formula is as follows:

[0064] s i ' Characterizes the expected structural metric value corresponding to the i-th group.

[0065] Specifically, determining an expected intensity based on the intensity of each second patch block in the same group includes: obtaining the global average intensity of each image in the first image set and the local average intensity of each second patch block. Based on the global average intensity of each image and the local average intensity of each second patch block, determining an expected intensity corresponding to each second patch block in the same group. The calculation formula is as follows:

[0066] Among them, l i ′ is the expected strength corresponding to the i-th group, μ m(k′) For image Z m(k′) The global average intensity of (the mth image in the first image set), R i (Z m(k′) ) is the image Z m(k′) The local average intensity of the second patch of the ith block. g , σ land τ are constants, which are set to 0.2, 0.5 and 0.5 respectively in this embodiment.

[0067] Specifically, based on the expected intensity, expected contrast, and expected structure metric values ​​of the same group, a corresponding fused patch block is obtained, including: calculating the fused patch block by the following formula, R i (Z′)=c i 's i ′+l i ′, where R i (Z′) is the fused patch corresponding to the i-th group.

[0068] Preferably, determining the fusion quality loss value based on each first patch block and the fusion patch block corresponding to each first patch block includes: determining the block fusion quality loss value between each pair of first patch blocks and the corresponding fusion patch block based on each first patch block and the fusion patch block corresponding to each first patch block. Determining the fusion quality loss value based on each block fusion quality loss value.

[0069] The fusion patch block is the expected local image after local extraction and fusion of all the images in the first image set at the same position. Therefore, by evaluating the difference between each pair of first patch blocks and the corresponding expected local image, the quality of each local image of the fused high dynamic range image can be determined, that is, the fusion quality loss value of each block is determined. Then determine the fusion quality loss value. It is convenient to train the pre-trained MSF-Net network model, and the local image of the fused image generated by the target MSF-Net network model is expected to be more in line with the expected local image after the fusion of the first image set. It can improve the high dynamic range image generated by the target MSF-Net network model to be closer to the actual high dynamic range scene.

[0070] Specifically, the block fusion quality loss value is calculated by the following formula:

[0071]

[0072] Among them, R i (Z F ) is the fused high dynamic range image Z F The corresponding first patch block, R i (Ω m ) is the first data set Ω m The i-th fused patch block in S(R i (Ω m ), R i (Z F )) is the block fusion quality loss value between the first patch block and a corresponding fused patch block, that is, the i-th block fusion quality loss value. Represents the first patch block R i (Z F ), Represents the fused patch block R i (Z′) and the first patch R i (Z F ), C 1 and C 2 Represents a constant.

[0073] Preferably, determining the fusion quality loss value based on the fusion quality loss value of each block includes: determining the fusion quality loss value based on the fusion quality loss value of each block, and the calculation formula is as follows: Among them, L S (Ω m , Z F ) is the fusion quality loss value, M is the number of block fusion quality loss values, S(R i (Ω m ), R i (Z F )) is the fusion quality loss value of the i-th block.

[0074] In this way, the fusion quality loss value is determined based on the fusion quality loss values ​​of each block, so that the decoupling between the first image set and the second image set can be achieved. The unsupervised fusion method proposed in this embodiment is improved from the perspective of MEF-SSIM.

[0075] Combination Figure 2 As shown, the embodiment of the present disclosure provides an image fusion method, which is applied to the MSF-Net network model training method shown above, including

[0076] Step S21 , obtaining a plurality of low dynamic range images to be fused with different exposures in the same high dynamic range scene.

[0077] Step S22, inputting a plurality of low dynamic range images to be fused into the trained target MSF-Net network model to obtain a fused target high dynamic range image.

[0078] An image fusion method provided by an embodiment of the present disclosure has the following beneficial effects: since real high dynamic range images are often not available, this embodiment makes full use of the asymmetry between the training and inference (or test) stages to propose a new loss function (the sum of the pixel intensity loss value and the fusion quality loss value). In addition to the image set to be fused (i.e., the second image set) to be fused, other low dynamic range images of different exposures from the same high dynamic range scene are also used to define the loss function. Therefore, the fused high dynamic range image not only includes the image information of all low dynamic range images of different exposures in the second image set, but also can learn the image information of other low dynamic range images of different exposures in the same high dynamic range scene as the second image set, so that the fused high dynamic range image is closer to the actual high dynamic range scene. Among them, the new loss function is decoupled from the image set to be fused, so that the asymmetry between the training stage and the inference stage can be well utilized.

[0079] In addition, the image fusion algorithm adopted in this embodiment is a multi-scale exposure fusion method based on unsupervised learning, based on the observation that multi-scale helps to maintain the scene depth of the fused low dynamic range image and improve the clarity of details. This embodiment decouples the loss function (pixel intensity loss value and fusion quality loss value) from the image set to be fused, so that exposure interpolation and exposure extrapolation can be easily implemented, so that the fused target high dynamic range image is closer to the high dynamic range scene, rather than a group of images to be fused (i.e., the image set to be fused). The MSF-Net network model (also known as the multi-scale exposure fusion network structure, Multi-Scale Fusion Network) of this embodiment can maintain the scene depth and local contrast of the image to be fused, and retains the information of the brightest and darkest areas, avoiding the appearance of halo.

[0080] Preferably, the target MSF-Net network model includes a feature extraction module, a sampling module and a feature fusion module; a plurality of low dynamic range images to be fused are input into the trained target MSF-Net network model to obtain a fused target high dynamic range image, including: using the feature extraction module to extract the features of each low dynamic range image to be fused respectively; using the sampling module to downsample each extracted feature to obtain feature pyramids of different scales; using the feature fusion module to fuse the feature pyramids of different scales to obtain a fused target high dynamic range image.

[0081] In this way, the image to be fused is processed by the feature extraction module, the sampling module and the feature fusion module, so that the fused target high dynamic range image is closer to the high dynamic range scene.

[0082] For ease of understanding, in this embodiment, Ω f is a set of low dynamic range images to be fused. In this embodiment, the feature extraction model F(·) of the feature extraction module is used to extract the features of the input image (i.e., the low dynamic range image to be fused) to obtain the corresponding feature set {F(Ω f )}. The feature extraction module contains two convolutional layers and one activation layer. In this way, the input sequence is transferred from the image domain to the feature domain, and features that are conducive to fusion are adaptively extracted. This method can effectively avoid artifacts in the fused image. Then, the sampling module is used to perform a series of downsampling operations to obtain feature pyramids of different scales. L represents the number of downsampling, and l represents the lth downsampling. That is, the intermediate image can be represented as The intermediate image is the image after feature extraction and downsampling of the image to be fused. Where N(·) is the network of the feature fusion module, v l The calculation formula is as follows: Among them, N -1 represents the feature extraction network of the previous scale, ↑ represents the upsampling operation, and C(·) represents the concatenation operation. Intermediate images of different scales are constructed through the network N(·), and then the intermediate images of different scales are fused to obtain the fused target high dynamic range image. The whole process is similar to the Gaussian pyramid, which can well preserve the scene depth and local contrast of the image to be fused, thereby ensuring the quality of the target high dynamic range image.

[0083] Combination Figure 3 As shown, Figure 3 An algorithm logic diagram of a target MSF-Net network model is provided. Figure 3 The image set in (Z 1 ,Z 2 ,...Z N ) is a set of low dynamic range images to be fused (hereinafter referred to as images to be fused). The weight function is the function for obtaining the weights of each pixel of each image to be fused as shown above. The multi-scale weighted L1 loss function is the function of the pixel intensity loss value shown above; the multi-scale MEF-SSIN loss function is the function of the fusion quality loss value shown above. The connection function is the C(·) concatenation operation shown above. UP is the upsampling operation. Scale is the scale. Z F is the target high dynamic range image. MSRRG (multi-scale recursive residual group). For ease of understanding, Figure 4A schematic diagram of the structure of MSRRG is shown. The structure of MSRRG in this embodiment is a residual network, which can reuse features and avoid possible gradient disappearance. In addition, the MSRRG in this embodiment uses a new multi-scale feature attention to suppress less useful features and only allows features with more information to be propagated, thereby effectively improving the quality of the fused target high dynamic range image. The MSRRG in this embodiment includes different scale convolutions (conv) and dual-attention blocks (DABs). Each DAB contains a channel attention block (ChannelAttention) and a spatial attention block (Spatial Attention), which can provide additional flexibility in processing different types of information. In addition, the proposed multi-scale structure is more conducive to retaining the details and scene depth of the fused image. Among them, the spatial attention block includes a global average pooling layer (Global Avg Pooling), a global maximum pooling layer (GlobalMax Pooling), a concatenation operation (Concat), a 3*3 convolution (Conv3*3) and a sigmoid function; and the global average pooling layer and the global maximum pooling layer are connected to the concatenation operation, the concatenation operation is connected to the 3*3 convolution, and the 3*3 convolution is connected to the sigmoid function. The channel attention block includes a global average pooling layer (Global Avg Pooling), two 1*1 convolutions (Conv1*1), an activation function (ReLU) and a sigmoid function; and the activation function is set between the two 1*1 convolutions, and the two 1*1 convolutions are also connected to the global average pooling layer and the sigmoid function respectively.

[0084] Combination Figure 5As shown, the disclosed embodiment provides an MSF-Net network model training system, including: an acquisition module and a training module. The training module includes a fusion unit, a first loss value unit, a second loss value unit and a target unit. Specifically, the acquisition module is used to take a pair of first image sets and second image sets as a sample, and obtain multiple samples as a training set; wherein the first image set includes multiple low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of multiple low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set. The training module is used to input the training set into the pre-trained MSF-Net network model for training. The fusion unit is used to input each second image set into the pre-trained MSF-Net network model for image fusion, and obtain the fused high dynamic range image corresponding to each second image set. The first loss value unit is used to obtain the pixel intensity loss value between each pair of high dynamic range images and the corresponding first image set, and obtain the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set. The second loss value unit is used to take the sum of the pixel intensity loss value and the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set as the total loss value. The target unit is used to adjust the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeatedly execute the steps corresponding to the fusion unit, the first loss value unit, the second loss value unit and the target unit until each total loss value converges and / or the number of iterations reaches a preset number, then the current pre-trained MSF-Net network model is used as the target MSF-Net network model.

[0085] A MSF-Net network model training system of the present embodiment is adopted, and the beneficial effect of the present invention is as follows: since the real high dynamic range image is often not available, the present embodiment makes full use of the asymmetry between the training and inference (or test) stages, and proposes a new loss function (the sum of the pixel intensity loss value and the fusion quality loss value). In addition to the image set to be fused (i.e., the second image set) to be fused, other low dynamic range images of different exposures from the same high dynamic range scene are also used to define the loss function. Therefore, the fused high dynamic range image not only includes the image information of all low dynamic range images of different exposures in the second image set, but also can learn the image information of other low dynamic range images of different exposures in the same high dynamic range scene as the second image set, so that the fused high dynamic range image is closer to the actual high dynamic range scene. Among them, the new loss function is decoupled from the image set to be fused, so that the asymmetry between the training stage and the inference stage can be well utilized.

[0086] Preferably, the first loss value unit is specifically used to obtain, for each pair of high dynamic range images and the corresponding first image set, first contrast, saturation and exposure measurement indicators of each image in the first image set, and determine the weight of each pixel point of each image in the first image set based on the first contrast, saturation and exposure measurement indicators of each image; and determine the corresponding pixel intensity loss value based on the weight of each pixel point of each image in the high dynamic range image and the corresponding first image set.

[0087] Preferably, the first loss value unit is specifically used to determine the pixel value of each pixel of the high dynamic range image based on the high dynamic range image; obtain the pixel value of each pixel of each image in the corresponding first image set; and determine the corresponding pixel intensity loss value based on the pixel value of each pixel of the high dynamic range image, the pixel value of each pixel of each image in the corresponding first image set, and the weight of each pixel of each image in the corresponding first image set.

[0088] Preferably, the first loss value unit is specifically used to obtain the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set, including: for each pair of high dynamic range images and the corresponding first image set, obtaining multiple first patch blocks of the high dynamic range images, and obtaining multiple groups of second patch blocks of the first image set; wherein, one first patch block corresponds to a group of second patch blocks, and a group of patch blocks are patch blocks at the same position of all images in the first image set; processing each second patch block in the same group to obtain a corresponding fused patch block; determining the fusion quality loss value based on each first patch block and the fused patch block corresponding to each first patch block.

[0089] Preferably, the first loss value unit is specifically used to determine the intensity, second contrast and structural measurement value of each second patch block based on the second patch blocks of the same group; determine an expected intensity based on the intensity of the second patch blocks of the same group, determine an expected contrast based on the second contrast of the second patch blocks, and determine an expected structural measurement value based on the structural measurement values ​​of the second patch blocks of the same group; for the expected intensity, expected contrast and expected structural measurement value of the same group, obtain the corresponding fused patch block.

[0090] Preferably, the first loss value unit is specifically used to determine the block fusion quality loss value between each pair of first patch blocks and the corresponding fusion patch block based on each first patch block and the fusion patch block corresponding to each first patch block; and determine the fusion quality loss value based on each block fusion quality loss value.

[0091] Preferably, the first loss value unit is specifically used to determine the fusion quality loss value based on the fusion quality loss values ​​of each block, and the calculation formula is as follows:

[0092]

[0093] Among them, L S (Ω m , Z F ) is the fusion quality loss value, M is the number of block fusion quality loss values, S(R i (Ω m ), R i (Z F )) is the fusion quality loss value of the i-th block.

[0094] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0095] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A MSF-Net network model training method, characterized in that: include: A pair of a first image set and a second image set is taken as a sample, and a plurality of samples are obtained as a training set; wherein the first image set includes a plurality of low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of a plurality of low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set; The training set is input into the pre-trained MSF-Net network model for training. The training process is as follows: S121, inputting each of the second image sets into the pre-trained MSF-Net network model for image fusion, to obtain a fused high dynamic range image corresponding to each of the second image sets; S122, obtaining a pixel intensity loss value between each pair of the high dynamic range images and the corresponding first image set, and obtaining a fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set; S123, taking the sum of the pixel intensity loss value and the fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set as the total loss value; S124, adjusting the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeating steps S121 to S124 until each of the total loss values ​​converges, and / or the number of iterations reaches a preset number, then using the current pre-trained MSF-Net network model as the target MSF-Net network model; The obtaining of the fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set comprises: For each pair of the high dynamic range image and the corresponding first image set, a plurality of first patch blocks of the high dynamic range image are obtained, and a plurality of groups of second patch blocks of the first image set are obtained; wherein one first patch block corresponds to a group of second patch blocks, and the group of second patch blocks are patch blocks at the same position of all images of the first image set; Processing each second patch block in the same group to obtain a corresponding fused patch block; Determining the fusion quality loss value based on each first patch block and a fusion patch block corresponding to each first patch block; The step of processing the second patch blocks of the same group to obtain corresponding fused patch blocks includes: Based on the second patches of the same group, determining the intensity, the second contrast and the structure metric value of each second patch; Determine an expected intensity based on the intensity of each second patch block in the same group, determine an expected contrast based on the second contrast of each second patch block, and determine an expected structure metric value based on the structure metric value of each second patch block in the same group; Based on the expected intensity, the expected contrast and the expected structure metric value of the same group, a corresponding fused patch block is obtained.

2. The method according to claim 1, characterized in that The obtaining of the pixel intensity loss value between each pair of the high dynamic range images and the corresponding first image set comprises: For each pair of the high dynamic range image and the corresponding first image set, obtaining first contrast, saturation and exposure metrics of each image in the first image set, and determining a weight of each pixel point of each image in the first image set based on the first contrast, saturation and exposure metrics of each image; Based on the weight of each pixel point of the high dynamic range image and each image in the corresponding first image set, a corresponding pixel intensity loss value is determined.

3. The method according to claim 2, characterized in that The determining of the corresponding pixel intensity loss value based on the weight of each pixel point of each image in the high dynamic range image and the corresponding first image set includes: Based on the high dynamic range image, determining a pixel value of each pixel point of the high dynamic range image; Obtaining the pixel value of each pixel point of each image in the corresponding first image set; Based on the pixel value of each pixel point of the high dynamic range image, the pixel value of each pixel point of each image in the corresponding first image set, and the weight of each pixel point of each image in the corresponding first image set, the corresponding pixel intensity loss value is determined.

4. The method according to claim 1, characterized in that: The determining the fusion quality loss value based on each first patch block and the fusion patch block corresponding to each first patch block includes: Based on each first patch block and the fused patch block corresponding to each first patch block, determining a block fusion quality loss value between each pair of first patch blocks and the corresponding fused patch block; The fusion quality loss value is determined based on the fusion quality loss values ​​of the individual blocks.

5. The method according to claim 4, characterized in that The determining the fusion quality loss value based on the fusion quality loss values ​​of each block includes: Based on the fusion quality loss values ​​of each block, the fusion quality loss value is determined, and the calculation formula is as follows: , in, is the fusion quality loss value, M is the number of block fusion quality loss values, For the The quality loss value of each block fusion.

6. An image fusion method, applied to the MSF-Net network model training method according to any one of claims 1 to 5, characterized in that: include Obtain multiple low dynamic range images to be fused at different exposures in the same high dynamic range scene; The multiple low dynamic range images to be fused are input into the trained target MSF-Net network model to obtain a fused target high dynamic range image.

7. The method according to claim 6, characterized in that The target MSF-Net network model includes a feature extraction module, a sampling module and a feature fusion module; the step of inputting the plurality of low dynamic range images to be fused into the trained target MSF-Net network model to obtain a fused target high dynamic range image includes: Using the feature extraction module to extract features of each low dynamic range image to be fused; Using the sampling module to downsample each extracted feature to obtain feature pyramids of different scales; The feature fusion module is used to fuse the feature pyramids of different scales to obtain the fused target high dynamic range image.

8. A MSF-Net network model training system, characterized in that: include: An acquisition module, configured to take a pair of a first image set and a second image set as a sample, and acquire multiple samples as a training set; wherein the first image set includes multiple low dynamic range images with different exposures in the same high dynamic range scene, and the second image set is a collection of multiple low dynamic range images in the first image set, and the number of images in the second image set is less than the number of images in the first image set; A training module, used for inputting the training set into a pre-trained MSF-Net network model for training, wherein the training module includes a fusion unit, a first loss value unit, a second loss value unit and a target unit; The fusion unit is used to input each of the second image sets into the pre-trained MSF-Net network model for image fusion, so as to obtain a fused high dynamic range image corresponding to each of the second image sets; The first loss value unit is used to obtain a pixel intensity loss value between each pair of the high dynamic range images and the corresponding first image set, and to obtain a fusion quality loss value between each pair of the high dynamic range images and the corresponding first image set; The second loss value unit is used to take the sum of the pixel intensity loss value and the fusion quality loss value between each pair of the high dynamic range image and the corresponding first image set as the total loss value; The target unit is used to adjust the network parameters in the pre-trained MSF-Net network model based on each total loss value, and repeatedly execute the steps corresponding to the fusion unit, the first loss value unit, the second loss value unit and the target unit until each of the total loss values ​​converges, and / or the number of iterations reaches a preset number, then the current pre-trained MSF-Net network model is used as the target MSF-Net network model; The first loss value unit is specifically used to obtain the fusion quality loss value between each pair of high dynamic range images and the corresponding first image set, including: for each pair of high dynamic range images and the corresponding first image set, obtaining multiple first patch blocks of the high dynamic range image, and obtaining multiple groups of second patch blocks of the first image set; wherein one first patch block corresponds to a group of second patch blocks, and a group of second patch blocks are patch blocks at the same position of all images of the first image set; processing each second patch block of the same group to obtain a corresponding fused patch block; determining the fusion quality loss value based on each first patch block and the fused patch block corresponding to each first patch block; The first loss value unit is specifically used to determine the strength, second contrast and structural measurement value of each second patch block based on the second patch blocks of the same group; determine an expected strength based on the strength of each second patch block of the same group, determine an expected contrast based on the second contrast of each second patch block, and determine an expected structural measurement value based on the structural measurement value of each second patch block of the same group; for the expected strength, expected contrast and expected structural measurement value of the same group, obtain the corresponding fused patch block.

Citation Information

Patent Citations

  • Food image segmentation method and system based on dynamic transformer

    CN114648535A

  • Multi-spectral image gradient fusion model establishment method and fusion method

    CN116108889A