Network structure capable of realizing tone mapping and contrast enhancement functions and training method

By designing a convolutional neural network that combines weight sharing and weight-sharing modules, the functions of tone mapping and contrast enhancement are realized, the problems of low efficiency and complex parameter adjustment in the prior art are solved, and stability and significant effects are achieved in different dynamic range scenarios.

CN120031079APending Publication Date: 2025-05-23HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311576291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems in terms of tone mapping and contrast enhancement, complex parameter adjustment and poor adaptability to different scenarios. Especially in scenarios with very large dynamic range, the tone mapping results are difficult to fit, the brightness of the LDR image is unstable, and flickering is prone to occur.

Method used

A network structure that can realize tone mapping and contrast enhancement functions is designed. Through the weight sharing module and the compressed parameter module with non-shared weights, combined with the convolutional neural network, the tone mapping degree and contrast enhancement effect of different dynamic range scenes are learned to reduce the complexity of parameter adjustment.

Benefits of technology

It realizes the stability of tone mapping output and the significance of contrast enhancement effect in various dynamic range scenarios, reduces hardware storage and computing pressure, and is suitable for fields such as image processing and intelligent monitoring video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031079A_ABST
    Figure CN120031079A_ABST
Patent Text Reader

Abstract

The invention provides a network structure capable of realizing tone mapping and contrast enhancement functions and a training method. The training method comprises the following steps: S1, preparing data; s2, network pre-processing, namely data processing, S2.1, design network pre-processing, and S2.2, logarithm domain conversion; s3, designing a convolutional neural network, S3.1, a basic feature extraction module, S3.2, a global condition vector module and S3.3, a scaling offset module; s4, designing a compression parameter module, S4.1, a TMO compression parameter module and S4.2, a CST compression parameter module, of which the weights are not shared; s5, network post-processing, S5.1, TMO network post-processing and TMO result image processing, and S5.2, CST network post-processing and convolution correction module and CST result image processing; and S6, network training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent monitoring video processing, and in particular relates to a network structure and a training method capable of realizing tone mapping and contrast enhancement functions. Background Art

[0002] With the development of artificial intelligence technology, image processing, intelligent surveillance video processing, etc. have also developed accordingly. In the existing technology, although HDR images have a larger dynamic range and can reflect the real scene more delicately, HDR images require larger storage space and transmission bandwidth than LDR images, and HDR is difficult to output and display. At present, the dynamic range of most graphics output devices such as monitors and printers is much smaller than that of ordinary HDR images. Image contrast enhancement is an indispensable task in the ISP process. If a network can perform both image tone mapping and image contrast enhancement, it will help reduce the storage and computing pressure of the hardware, which is a very necessary and meaningful research.

[0003] Current tone mapping methods are mainly divided into traditional tone mapping methods and tone mapping methods based on convolutional neural networks; traditional tone mapping methods are further divided into global and local methods; methods based on convolutional neural networks usually select relatively good LDR images as labels through a variety of traditional methods and HDR images as input.

[0004] Traditional contrast enhancement methods are mainly divided into two categories: global contrast enhancement and local contrast enhancement. For most scenes, local contrast enhancement has certain advantages over global contrast enhancement. The mainstream contrast enhancement algorithms mainly include AHE (Adaptive Histogram Equalization), ACE (Adaptive Contrast Enhancement) and CLAHE (Contrast-Limited Adaptive Histogram Equalization). When using supervised deep learning to achieve contrast enhancement, its labels are mainly made based on the above methods.

[0005] However, among traditional tone mapping methods, the global method is fast but has poor effect, while the local method is slow and has relatively good effect but has more parameters, and it is difficult to adjust the parameters for scenes with different dynamic ranges. The existing methods based on convolutional neural networks have poor fitting ability for scenes with very large dynamic ranges, and the brightness of LDR images is difficult to stabilize. When the dynamic range of the scene changes dramatically, the LDR image will flicker and other unstable conditions.

[0006] Existing de-contrast enhancement algorithms have poor adaptability to different scenarios, and fixed parameters are difficult to meet the needs of richer scenarios. Color anomalies or different degrees of contrast enhancement in different areas of the same image often occur in the enhancement results. The tone mapping and contrast enhancement algorithms in existing ISPs are often functionally divided into two modules, which have higher requirements on hardware, computing power, etc.

[0007] In addition, the commonly used terms in the prior art include:

[0008] HDR: high dynamic range, dynamic range refers to the brightness ratio between the brightest object and the darkest object in the scene. The larger the dynamic range, the richer the levels that can be expressed. The dynamic range in real scenes reaches 109:1, the dynamic range that the human visual system can perceive is about 105:1, and the dynamic range of general image sensors is about 102:1.

[0009] LDR: low dynamic range, low dynamic range.

[0010] Convolutional neural network: A convolutional neural network is a deep neural network with a convolutional structure. Images can be directly used as input to the network, avoiding the complex feature extraction and data reconstruction process in traditional algorithms. It has great advantages in the processing of two-dimensional images. For example, the network can extract image features including color, texture, shape and image topology by itself, and is robust and efficient in processing two-dimensional images.

[0011] Tone mapping: Tone mapping enables HDR images to be displayed on our low dynamic range displays and to be as consistent as possible with the visual experience of the human eye. Its essence is to convert HDR images into LDR images, and the process is dynamic range compression.

[0012] Contrast: Contrast refers to the measurement of the different brightness levels between the brightest white and the darkest black in the light and dark areas of an image. The larger the difference range, the greater the contrast, and the smaller the difference range, the smaller the contrast. The impact of image contrast on visual effects is very critical. Generally speaking, the greater the contrast, the clearer and more eye-catching the image, and the brighter and more colorful the color. The smaller the contrast, the grayer the image. High contrast is very helpful for image clarity, detail expression, and grayscale expression. There is no effective and fair standard to measure contrast today, so the best way to identify it is to rely on the user's visual experience. Summary of the invention

[0013] In order to solve the above problems, the purpose of this application is to: design a network structure and training method that can realize tone mapping and contrast enhancement functions, and let the network learn the tone mapping degree and contrast enhancement effect of different dynamic range scenes in a data-driven manner; ensure that the tone mapping output is stable and does not flicker in various dynamic range scenes, and the contrast enhancement effect is stable, avoiding the complex parameter adjustment of traditional image methods.

[0014] Specifically, the present invention provides a network structure capable of realizing tone mapping and contrast enhancement functions, wherein the network structure is a convolutional neural network capable of realizing both tone mapping and contrast enhancement functions, comprising:

[0015] A. Weight sharing modules, including:

[0016] The weight-sharing basic feature extraction module uses three consecutive convolutions with a kernel of 3x3, a stride of 1, and a pad of 1. Each convolution is followed by a ReLU activation function. The input image is passed through the basic feature extraction module to obtain the basic features.

[0017] The weight-sharing global conditional vector module uses three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is calculated to obtain a global conditional vector of size 32.

[0018] The weight-sharing scaling offset module uses 6 fully connected layers to form the scaling offset module. The global condition vectors generated by the weight-sharing global condition vector module are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters to obtain 6 groups of parameters, which are divided into a first group of scaling parameters and offset parameters, a second group of scaling parameters and offset parameters, and a third group of scaling parameters and offset parameters in turn; the basic features extracted by the weight-sharing basic feature extraction module are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain a first global feature; the first global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain a second global feature; the second global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain a third global feature;

[0019] B. Compression parameter module with unshared weights, including:

[0020] TMO compression parameter module, when the input image realizes the tone mapping function, the third global feature and the basic feature extracted by the weight-sharing basic feature extraction module are concat fused, and then pass through the convolution kernel of 3x3, the number of input channels is 32, the number of output channels is 1, the stride is 1, and the pad is 1, and then pass through the hyperbolic tangent tanh function to obtain the dynamic range compression parameter scale_hdr; CST compression parameter module, when the input image realizes the contrast enhancement function, the third global feature and the basic feature are concat fused, and then pass through the convolution kernel of 3x3, the number of input channels is 32, the number of output channels is 32, the stride is 1, and the pad is 1, and then pass through the tanh function to obtain the contrast enhancement parameter scale_cst.

[0021] The network implements two different functions based on different inputs, which are two different outputs, and are restricted by GT respectively during training; each function corresponds to its own data set, each data set consists of multiple data pairs, each data pair contains two data, input and GT, and the GT is the GT in the data pair corresponding to the two functions;

[0022] When TMO, or tone mapping, is implemented, TMO network post-processing is performed. When the input image is I HDRN That is, when realizing the tone mapping function, I HDRN_log The image is multiplied by the tone mapping parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6):

[0023]

[0024] Among them I LDR_OUT The final result image of tone mapping is the TMO result image;

[0025] The available python code is: HDR_OUT =2**(scale_hdr*I HDRN_log );

[0026] When CST, i.e., contrast enhancement function, is implemented, CST network post-processing is performed. When the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7):

[0027]

[0028] Among them I CST_out1 It is the contrast enhanced intermediate result image;

[0029] The available Python code is: CST_out1 =2**(scale_cst*I LDRN_log );

[0030] For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image.

[0031] The present application also relates to a training method for a network structure capable of realizing tone mapping and contrast enhancement functions, which is applicable to the above-mentioned network structure and includes the following steps:

[0032] S1, data preparation: prepare 900 sets of data pairs, each set of data contains HDR images of the same scene I HDR , LDR image after dynamic range compression of HDR image I LDR , and the CST image I after contrast enhancement of the LDR image CST , where I LDR and I CST As the GT for dynamic range compression and contrast enhancement tasks, the acquisition method uses traditional methods to adjust parameters and select the optimal one;

[0033] S2, data processing part, processing of image data before entering the network, including:

[0034] S2.1, design network pre-processing, first normalize the HDR image I through equations (1), (2), and (3) HDR 、LDR image I LDR , and the gt image I for the contrast enhancement task CST :

[0035]

[0036]

[0037]

[0038] Among them I HDR , ILDR, ICST are 16-bit wide images, I HDRN ,I LDRN ,I CSTN are their normalized images respectively;

[0039] S2.2, convert the network input image to the log2 domain. HDRNThat is, when the dynamic range compression function is realized, after processing as in formula (4), when the input image is I LDRN That is, when the contrast enhancement function is realized, it is processed as in formula (5);

[0040]

[0041] I HDRN_log = -1×log2(I′ HDRN ) Formula (4)

[0042]

[0043] I LDRN_log = -1×log2(I′ LDRN ) Formula (5)

[0044] Among them, I HDRN_log I after switching to the log domain HDRN Image, I LDRN_log I after switching to the log domain LDRN image;

[0045] S3, a convolutional neural network designed for tone mapping and contrast enhancement:

[0046] S3.1, design a weight-sharing basic feature extraction module, use three consecutive convolutions with a kernel of 3x3, stride of 1, and pad of 1 in series, each convolution is followed by a ReLU activation function, and the input image is passed through the basic feature extraction module to obtain the basic features;

[0047] S3.2, design a weight-sharing global conditional vector module, use three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1, and connect them in series. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is finally calculated to obtain a global conditional vector of size 32;

[0048] S3.3, design a weight-sharing scaling offset module, use 6 fully connected layers to form the scaling offset module, the global condition vectors generated in step S3.2 are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters, so as to obtain 6 groups of parameters, which are divided into the first group of scaling parameters and offset parameters, the second group of scaling parameters and offset parameters, and the third group of scaling parameters and offset parameters in turn; the basic features of step 3.1 are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain the first global features; the first global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain the second global features; the second global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain the third global features;

[0049] S4, designed compression parameter module with unshared weights,

[0050] S4.1, TMO compression parameter module, when the input image is I HDRN That is, when the tone mapping function is realized, the third global feature and the basic feature of step S3.1 are concat-fused and then convolved with a convolution kernel of 3x3, an input channel number of 32, an output channel number of 1, a stride of 1, and a pad of 1, and then the dynamic range compression parameter scale_hdr is obtained by the hyperbolic tangent tanh function;

[0051] S4.2, CST compression parameter module, when the input image is I LDRN That is, when realizing the contrast enhancement function, the third global feature and the basic feature are concat fused and then convolved with a convolution kernel of 3x3, 32 input channels, 32 output channels, 1 stride, and 1 pad, and then the contrast enhancement parameter scale_cst is obtained by the tanh function;

[0052] S5, design network post-processing, perform S5.1 and S5.2 respectively; that is, this network implements two different functions, when TMO, i.e., tone mapping function, is implemented, step S5.1 is performed; when CST, i.e., contrast enhancement function, is implemented, step S5.2 is performed, these two are two different outputs, and gt is used to restrict them respectively during training;

[0053] S5.1, TMO network post-processing, when the input image is I HDRN That is, when the dynamic range compression function is realized, I HDRN_log The image is multiplied by the dynamic range compression parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6):

[0054]

[0055] Among them I LDR_OUT The final result image of tone mapping is the TMO result image;

[0056] The available Python code is: HDR_OUT =2**(scale_hdr*I HDRN_log );

[0057] Then further perform step S6;

[0058] S5.2, CST network post-processing, when the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7):

[0059]

[0060] Among them I CST_out1 It is the contrast enhanced intermediate result image;

[0061] The available Python code is: CST_out1 =2**(scale_cst*I LDRN_log );

[0062] For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image;

[0063] Then further perform step S6;

[0064] S6. Network training: Use the convolutional neural network designed in steps S3 to S5 for training, and use Adam as the optimizer; the initial learning rate is 0.0001, the training cycle is 100, and the learning rate is reduced by 0.1 times every 30 training cycles.

[0065] S6.1, the tone mapping loss is a combination of L1 loss and vggloss, and the loss weight is 10:1;

[0066] In S6.2, the loss for contrast enhancement is a combination of L1loss and ssimloss, and the loss weight is 1:1; the total loss weight of the two tasks is 1:1.

[0067] Steps S1 and S2 are the data preparation and data processing parts. Before training the network, it is necessary to prepare data pairs (input, gt) corresponding to the two functions and process them; Steps S3 to S5 are the network structure part. The two functions have weight sharing, that is, step S3, and the weights are no longer shared, that is, steps S4 and S5; Step S6 is the training part. Step S6 includes the training strategy and loss usage for dual outputs. According to the different attributes and requirements of the two functions, different tasks correspond to different input, gt and different loss combinations for supervision; Among them, the network output I of step S5.1 LDR_OUT Tone-mapped gt(I HDRN ) and the loss combination of L1 loss and vggloss for supervision; the network output I of step S5.2 CST_OUT The contrast-enhanced gt(I CSIN ) and the loss combination of L1loss and ssimloss for supervision.

[0068] In step S2.1, the normalization refers to image normalization, which converts the pixel values ​​of an image into values ​​within a certain range to better perform the next step of processing. In this application, image I HDR ,I LDR and I CST The initial data type is uint16, whose theoretical maximum value is 65535. The goal is to normalize it to float32 data with a maximum value of 1, so I HDR ,I LDR and I CST Divide by 65535, then convert to float32, normalize uint16, and then convert to float32, no data loss will occur;

[0069] The available Python code is:

[0070] img_hdrn=(img_hdr / 65535).astype(np.float32)

[0071] img_ldrn=(img_ldr / 65535).astype(np.float32)

[0072] img_cstn=(img_cst / 65535).astype(np.float32);

[0073] In subsequent training, the tone-mapped gt is I HDRN , contrast enhanced gt is I CSTN .

[0074] The implementation of step S2.2 further includes: in the tasks of tone mapping and contrast enhancement, the transformation in the log domain is more in line with the visual characteristics of the human eye, so it is converted to the log domain, where the negative sign is due to the fact that the base of the log function used in this application is 2 and the value of the independent variable is 0-1, and the function value is negative at this time, so it is negated;

[0075] The available Python code is:

[0076] Tone Mapping: I HDRN_log = -1*torch.log2(I' HDRN )

[0077] Contrast Enhancement: I LDRN_log = -1*torch.log2(I' LDRN ).

[0078] Therefore, the advantages of this application are: through this method, the network learns the degree of tone mapping of scenes with different dynamic ranges in a data-driven way; it ensures that the tone mapping output is stable and does not flicker in various dynamic range scenes, avoiding the complex parameter adjustment of traditional image methods; it solves the problem that the tone mapping results of scenes with very large dynamic ranges are difficult to fit, and the tone mapping has good robustness; at the same time, the contrast enhancement effect is significant, and the idea of ​​weight sharing is used in the middle to reduce power consumption and computing costs for subsequent algorithm implementation and deployment. It is particularly suitable for multiple fields including image processing, intelligent surveillance video processing, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.

[0080] Figure 1 It is a schematic diagram of the main process flow chart of this application.

[0081] Figure 2 It is a schematic diagram of the process of an embodiment of the method. DETAILED DESCRIPTION

[0082] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail in conjunction with the accompanying drawings.

[0083] The present application provides a network structure that can realize tone mapping and contrast enhancement functions. The network structure is a convolutional neural network that realizes the two functions of tone mapping and contrast enhancement, including:

[0084] A. Weight sharing modules, including:

[0085] The weight-sharing basic feature extraction module uses three consecutive convolutions with a kernel of 3x3, a stride of 1, and a pad of 1. Each convolution is followed by a ReLU activation function. The input image is passed through the basic feature extraction module to obtain the basic features.

[0086] The weight-sharing global conditional vector module uses three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is calculated to obtain a global conditional vector of size 32.

[0087] The weight-sharing scaling offset module uses 6 fully connected layers to form the scaling offset module. The global condition vectors generated by the weight-sharing global condition vector module are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters to obtain 6 groups of parameters, which are divided into the first group of scaling parameters and offset parameters, the second group of scaling parameters and offset parameters, and the third group of scaling parameters and offset parameters in turn; the basic features extracted by the weight-sharing basic feature extraction module are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain the first global feature; the first global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain the second global feature; the second global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain the third global feature;

[0088] B. Compression parameter module with unshared weights, including:

[0089] TMO compression parameter module, when the input image realizes the tone mapping function, the third global feature and the basic feature extracted by the weight-sharing basic feature extraction module are concat fused, and then pass through the convolution kernel of 3x3, the number of input channels is 32, the number of output channels is 1, the stride is 1, and the pad is 1, and then pass through the hyperbolic tangent tanh function to obtain the dynamic range compression parameter scale_hdr; CST compression parameter module, when the input image realizes the contrast enhancement function, the third global feature and the basic feature are concat fused, and then pass through the convolution kernel of 3x3, the number of input channels is 32, the number of output channels is 32, the stride is 1, and the pad is 1, and then pass through the tanh function to obtain the contrast enhancement parameter scale_cst.

[0090] The network implements two different functions based on different inputs. These two functions are two different outputs, and they are restricted by gt respectively during training. Each function corresponds to its own data set, and each data set consists of multiple data pairs, each of which contains two data pairs, input and gt. The gt is the gt in the data pair corresponding to the two functions.

[0091] When TMO, or tone mapping, is implemented, TMO network post-processing is performed. When the input image is I HDRN That is, when the dynamic range compression function is realized, I HDRN_log The image is multiplied by the dynamic range compression parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6):

[0092]

[0093] Among them I LDR_OUT The final result image of tone mapping is the TMO result image;

[0094] The available python code is: HDR_OUT =2**(scale_hdr*I HDRN_log );

[0095] When CST, i.e., contrast enhancement function, is implemented, CST network post-processing is performed. When the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7):

[0096]

[0097] Among them I CST_out1 It is the contrast enhanced intermediate result image;

[0098] The available python code is: CST_out1 =2**(scale_cst*I LDRN_log ).

[0099] For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image.

[0100] The present application also relates to a training method for a network structure that can realize tone mapping and contrast enhancement functions, which is applicable to any of the above network structures. Figure 1 As shown, the training method includes, S1, data preparation; S2, network pre-processing, i.e., data processing, S2.1, design of network pre-processing, S2.2, logarithmic domain conversion; S3, design of convolutional neural network, S3.1, basic feature extraction module, S3.2, global conditional vector module, S3.3, scaling offset module; S4, design of compression parameter module with non-shared weights, S4.1, TMO compression parameter module, S4.2, CST compression parameter module; S5, network post-processing, S5.1, TMO network post-processing, TMO result image, S5.2, CST network post-processing, convolution correction module, CST result image.

[0101] like Figure 2 As shown, the main implementation steps of the method are as follows:

[0102] S1, data preparation: prepare 900 sets of data pairs, each set of data contains HDR images of the same scene I HDR , LDR image after dynamic range compression of HDR image I LDR , and the image I after contrast enhancement of the LDR image CST , where I LDR and I CST As the GT for dynamic range compression and contrast enhancement tasks, the acquisition method uses traditional methods to adjust parameters and select the optimal one;

[0103] S2, data processing part, processing of image data before entering the network, including:

[0104] S2.1, design network pre-processing, first normalize the HDR image I through equations (1), (2), and (3) HDR 、LDR image I LDR , and the gt image I for the contrast enhancement task CST :

[0105]

[0106]

[0107]

[0108] Among them I HDR ,I LDR ,I CST For 16-bit wide images, I HDRN ,I LDRN ,I CSTNare their normalized images respectively; the normalization refers to image normalization, which converts the pixel values ​​of an image into values ​​within a certain range in order to better perform the next step of processing. In this application, image I HDR ,I LDR and I CST The initial data type is uint16, whose theoretical maximum value is 65535. The goal is to normalize it to float32 data with a maximum value of 1, so I HDR ,I LDR and I CST Divide by 65535, then convert to float32, normalize uint16, and then convert to float32, no data loss will occur;

[0109] The available Python code is:

[0110] img_hdrn=(img_hdr / 65535).astype(np.float32)

[0111] img_ldrn=(img_ldr / 65535).astype(np.float32)

[0112] img_cstn=(img_cst / 65535).astype(np.float32);

[0113] In subsequent training, the tone-mapped gt is I HDRN , contrast enhanced gt is I CSTN ;

[0114] S2.2, convert the network input image to the log2 domain. HDRN That is, when the dynamic range compression function is realized, after processing as in formula (4), when the input image is I LDRN That is, when the contrast enhancement function is realized, it is processed as in formula (5);

[0115]

[0116] I HDRN_log = -1×log2(I′ HDRN ) Formula (4)

[0117]

[0118] I HDRN_log = -1×log2(I′ LDRN ) Formula (5)

[0119] Among them, I HDRN_log I after switching to the log domainHDRN Image, I LDRN_log I after switching to the log domain LDRN Image; In the task of dynamic range compression and contrast enhancement, the transformation in the log domain is more in line with the visual characteristics of the human eye, so it is converted to the log domain, where the negative sign is due to the fact that the base of the log function used in this application is 2 and the value of the independent variable is 0-1, and the function value is negative at this time, so it is negated; it can be expressed in Python code as:

[0120] Tone Mapping: I HDRN_log = -1*torch.log2(I' HDRN )

[0121] Contrast Enhancement: I LDRN_log = -1*torch.log2(I' LDRN );

[0122] S3, a convolutional neural network designed for tone mapping and contrast enhancement:

[0123] S3.1, design a weight-sharing basic feature extraction module, use three consecutive convolutions with a kernel of 3x3, a stride of 1, and a pad of 1, and connect them in series. Each convolution is followed by a ReLU activation function. The input image is passed through the basic feature extraction module to obtain basic features. The ReLU is an important activation function in deep learning. When the input is greater than 0, the input value is directly returned. When the input is 0 or less, the return value is 0. There is a relu interface in pytorch, which can be used by directly calling torch.nn.relu() (python3).

[0124] S3.2, design a weight-sharing global conditional vector module, use three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1, and connect them in series. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is finally calculated to obtain a global conditional vector of size 32;

[0125] S3.3, design a weight-sharing scaling offset module, use 6 fully connected layers to form the scaling offset module, the global condition vectors generated in step S3.2 are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters, so as to obtain 6 groups of parameters, which are divided into the first group of scaling parameters and offset parameters, the second group of scaling parameters and offset parameters, and the third group of scaling parameters and offset parameters in turn; the basic features of the step 3.1 are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain the first global features; the first global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain the second global features; the second global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain the third global features;

[0126] S4, designed compression parameter module with unshared weights,

[0127] S4.1, TMO compression parameter module, when the input image is I HDRN That is, when the tone mapping function is realized, the third global feature and the basic feature of step S3.1 are concat-fused and then convolved with a convolution kernel of 3x3, an input channel number of 32, an output channel number of 1, a stride of 1, and a pad of 1, and then the dynamic range compression parameter scale_hdr is obtained by the hyperbolic tangent tanh function;

[0128] S4.2, CST compression parameter module, when the input image is I LDRN That is, when realizing the contrast enhancement function, the third global feature and the basic feature are concat fused and then convolved with a convolution kernel of 3x3, 32 input channels, 32 output channels, 1 stride, and 1 pad, and then the contrast enhancement parameter scale_cst is obtained by the tanh function;

[0129] S5, design network post-processing, perform S5.1 and S5.2 respectively; that is, this network implements two different functions, when TMO, i.e., tone mapping function, is implemented, step S5.1 is performed; when CST, i.e., contrast enhancement function, is implemented, step S5.2 is performed, these two are two different outputs, and gt is used to restrict them respectively during training;

[0130] S5.1, TMO network post-processing, when the input image is I HDRN That is, when the dynamic range compression function is realized, I HDRN_log The image is multiplied by the dynamic range compression parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6):

[0131]

[0132] Among them I LDR_OUT The final result image of tone mapping is the TMO result image;

[0133] The available Python code is: HDR_OUT =2**(scale_hdr*I HDRN_log );

[0134] Then further perform step S6;

[0135] S5.2, CST network post-processing, when the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7):

[0136]

[0137] Among them I CST_out1 It is the contrast enhanced intermediate result image;

[0138] The available python code is: CST_out1 =2**(scale_cst*I LDRN_log ).

[0139] For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image;

[0140] Then further perform step S6;

[0141] S6. Network training: Use the convolutional neural network designed in steps S3 to S5 for training. The optimizer uses Adam. The Adam optimizer is an adaptive optimization algorithm used to update the weights of the neural network. It can adaptively adjust the learning rate. The initial learning rate is 0.0001, the training cycle is 100, and the learning rate is reduced by 0.1 times every 30 training cycles.

[0142] S6.1, the tone mapping loss is a combination of L1 loss and vggloss, and the loss weight is 10:1;

[0143] In S6.2, the loss for contrast enhancement is a combination of L1loss and ssimloss, and the loss weight is 1:1; the total loss weight of the two tasks is 1:1.

[0144] In the method, steps S1 and S2 are data preparation and data processing parts. Before training the network, it is necessary to prepare data pairs (input, gt) corresponding to the two functions and process them; steps S3 to S5 are network structure parts. The two functions have weight sharing, i.e., step S3, and the weights are no longer shared, i.e., steps S4 and S5; step S6 is the training part. Step S6 includes training strategies and loss usage for dual outputs. According to the different attributes and requirements of the two functions, different tasks correspond to different input, gt and different loss combinations for supervision; wherein, the network output I of step S5.1 is LDR_OUT Tone-mapped gt(I HDRN ) and the loss combination of L1 loss and vggloss for supervision; the network output I of step S5.2 CST_OUT The contrast-enhanced gt(I CSTN ) and the loss combination of L1loss and ssimloss for supervision.

[0145] In summary, the present invention realizes two functions of dynamic range compression and contrast enhancement with one network at the same time, that is, a set of networks is used to simultaneously complete the functions of tone mapping and contrast enhancement. That is, the innovation of this technical solution is to use the model obtained by current network training to realize two different functions (tone mapping and contrast enhancement):

[0146] ②The network always has two different output result images.

[0147] ② In the specific use and reasoning process, when the input is hdr, netout_tmo (the output of step S5.1) is taken as the final result image, and the output of S5.2 is not used; when the input is ldr, netout_cst (the output of step S5.2) is taken as the final result image, and the output of step S5.1 is not used.

[0148] ③ To achieve the above functions, when training this network, it is necessary to supervise the network input with different data pairs (input, gt) and loss for the two tasks.

[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A network structure that can achieve tone mapping and contrast enhancement, It is characterized in that The network structure is a convolutional neural network that realizes two functions of tone mapping and contrast enhancement, including: A. Weight sharing modules, including: The weight-sharing basic feature extraction module uses three consecutive convolutions with a kernel of 3x3, a stride of 1, and a pad of 1. Each convolution is followed by a ReLU activation function. The input image is passed through the basic feature extraction module to obtain the basic features. The weight-sharing global conditional vector module uses three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is calculated to obtain a global conditional vector of size 32. The weight-sharing scaling offset module uses 6 fully connected layers to form the scaling offset module. The global condition vectors generated by the weight-sharing global condition vector module are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters to obtain 6 groups of parameters, which are divided into the first group of scaling parameters and offset parameters, the second group of scaling parameters and offset parameters, and the third group of scaling parameters and offset parameters in turn; the basic features extracted by the weight-sharing basic feature extraction module are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain the first global feature; the first global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain the second global feature; the second global feature is subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain the third global feature; B. Compression parameter module with unshared weights, including: TMO compression parameter module, when the input image realizes the tone mapping function, the third global feature and the basic feature extracted by the weight-sharing basic feature extraction module are concat-fused and then convolved with a convolution kernel of 3x3, 32 input channels, 1 output channel, 1 stride, and 1 pad, and then the dynamic range compression parameter scale_hdr is obtained by the hyperbolic tangent tanh function; CST compression parameter module, when the input image realizes the contrast enhancement function, the third global feature and the basic feature are concat fused and then passed through a convolution with a convolution kernel of 3x3, 32 input channels, 32 output channels, stride 1, and pad 1, and then passed through the tanh function to obtain the contrast enhancement parameter scale_cst.

2. The network structure capable of realizing tone mapping and contrast enhancement functions according to claim 1, It is characterized in that The network implements two different functions based on different inputs, which are two different outputs, and are restricted by GT respectively during training; each function corresponds to its own data set, each data set consists of multiple data pairs, each data pair contains two data, input and GT, and the GT is the GT in the data pair corresponding to the two functions; When TMO, or tone mapping, is implemented, TMO network post-processing is performed. When the input image is I HDRN That is, when realizing the tone mapping function, I HDRN_log The image is multiplied by the tone mapping parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6): Among them I LDR_OUT The final result image of tone mapping is the TMO result image; The available python code is: HDR_OUT =2**(scale_hdr*I HDRN_log ); When CST, i.e., contrast enhancement function, is implemented, CST network post-processing is performed. When the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7): Among them I CST_out1 It is the contrast enhanced intermediate result image; The available Python code is: CST_out1 =2**(scale_cst*I LDRN_log ); For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image.

3. A training method for a network structure capable of realizing tone mapping and contrast enhancement functions, applicable to the network structure described in claim 1 or 2, It is characterized in that The following steps are involved: S1, data preparation: prepare 900 sets of data pairs, each set of data contains HDR images of the same scene I HDR , LDR image after dynamic range compression of HDR image I LDR , and the CST image I after contrast enhancement of the LDR image CST , where I LDR and I CST As the GT for dynamic range compression and contrast enhancement tasks, the acquisition method uses traditional methods to adjust parameters and select the optimal one; S2, data processing part, processing of image data before entering the network, including: S2.1, design network pre-processing, first normalize the HDR image I through equations (1), (2), and (3) HDR 、LDR image I LDR , and the gt image I for the contrast enhancement task CST : Among them I HDR ,I LDR ,I CST For 16-bit wide images, I HDRN ,I LDRN ,I CSTN are their normalized images respectively; S2.2, convert the network input image to the log2 domain. HDRN That is, when the dynamic range compression function is realized, after processing as in formula (4), when the input image is I LDRN That is, when the contrast enhancement function is realized, it is processed as in formula (5); I HDRN_log = -1 × log2(I′ HDRN ) Formula (4) I LDRN_log = -1 × log2(I′ LDRN ) Formula (5) Among them, I HDRN_log I after switching to the log domain HDRN Image, I LDRN_log I after switching to the log domain LDRN image; S3, a convolutional neural network designed for tone mapping and contrast enhancement: S3.1, design a weight-sharing basic feature extraction module, use three consecutive convolutions with a kernel of 3x3, stride of 1, and pad of 1 in series, each convolution is followed by a ReLU activation function, and the input image is passed through the basic feature extraction module to obtain the basic features; S3.2, design a weight-sharing global conditional vector module, use three consecutive convolutions with a kernel of 3x3, a stride of 2, and a pad of 1, and connect them in series. Each convolution is followed by a ReLU activation function. After the basic features pass through the global conditional vector module, the average value in the w and h directions is finally calculated to obtain a global conditional vector of size 32; S3.3, design a weight-sharing scaling offset module, use 6 fully connected layers to form the scaling offset module, the global condition vectors generated in step S3.2 are respectively input into the 6 fully connected layers, and are respectively used to generate 3 groups of scaling parameters and offset parameters, so as to obtain 6 groups of parameters, which are divided into the first group of scaling parameters and offset parameters, the second group of scaling parameters and offset parameters, and the third group of scaling parameters and offset parameters in turn; the basic features of the step 3.1 are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the first group of scaling parameters and offset parameters to obtain the first global features; the first global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the second group of scaling parameters and offset parameters to obtain the second global features; the second global features are subjected to a convolution with a convolution kernel of 3x3, a stride of 1, and a pad of 1, and then mapped through the third group of scaling parameters and offset parameters to obtain the third global features; S4, designed compression parameter module with unshared weights, S4.1, TMO compression parameter module, when the input image is I HDRN That is, when the tone mapping function is realized, the third global feature and the basic feature of step S3.1 are concat-fused and then convolved with a convolution kernel of 3x3, an input channel number of 32, an output channel number of 1, a stride of 1, and a pad of 1, and then the dynamic range compression parameter scale_hdr is obtained by the hyperbolic tangent tanh function; S4.2, CST compression parameter module, when the input image is I LDRN That is, when realizing the contrast enhancement function, the third global feature and the basic feature are concat fused and then convolved with a convolution kernel of 3x3, 32 input channels, 32 output channels, 1 stride, and 1 pad, and then the contrast enhancement parameter scale_cst is obtained by the tanh function; S5, design network post-processing, perform S5.1 and S5.2 respectively; that is, this network implements two different functions, when TMO, i.e., tone mapping function, is implemented, step S5.1 is performed; when CST, i.e., contrast enhancement function, is implemented, step S5.2 is performed, these two are two different outputs, and gt is used to restrict them respectively during training; S5.1, TMO network post-processing, when the input image is I HDRN That is, when the dynamic range compression function is realized, I HDRN_log The image is multiplied by the dynamic range compression parameter scale_hdr, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (6): Among them I LDR_OUT The final result image of tone mapping is the TMO result image; The available Python code is: HDR_OUT =2**(scale_hdr*I HDRN_log ); then further execute step S6; S5.2, CST network post-processing, when the input image is I LDRN That is, when the contrast enhancement function is realized, I LDRN_log The image is multiplied by the contrast enhancement parameter scale_cst, and then the target image is obtained by exponentially increasing it by 2, as shown in formula (7): Among them I CST_out1 It is the contrast enhanced intermediate result image; The available Python code is: CST_out1 =2**(scale_cst*I LDRN_log ); For the contrast enhancement task, we add a convolution operation with unshared weights, namely the convolution correction module. We design a convolution kernel of 3x3, with 32 input channels, 1 output channel, 1 stride, and 1 pad. CST_out1 , output I CST_OUT , which is the final result image of contrast enhancement, namely the CST result image; Then further perform step S6; S6. Network training: Use the convolutional neural network designed in steps S3 to S5 for training, and use Adam as the optimizer; the initial learning rate is 0.0001, the training cycle is 100, and the learning rate is reduced by 0.1 times every 30 training cycles. S6.1, the tone mapping loss is a combination of L1 loss and vggloss, and the loss weight is 10:1; In S6.2, the loss for contrast enhancement is a combination of L1loss and ssimloss, and the loss weight is 1:1; the total loss weight of the two tasks is 1:

1.

4. The method for training a network structure capable of realizing tone mapping and contrast enhancement functions according to claim 3, It is characterized in that Steps S1 and S2 are the data preparation and data processing parts. Before training the network, it is necessary to prepare data pairs (input, gt) corresponding to the two functions and process them; Steps S3 to S5 are the network structure part. The two functions have weight sharing, that is, step S3, and the weights are no longer shared, that is, steps S4 and S5; Step S6 is the training part. Step S6 includes the training strategy and loss usage for dual outputs. According to the different attributes and requirements of the two functions, different tasks correspond to different input, gt and different loss combinations for supervision; Among them, the network output I of step S5.1 LDR_OUT Tone-mapped gt(I HDRN ) and the loss combination of L1 loss and vggloss for supervision; the network output I of step S5.2 CST_OUT The contrast-enhanced gt(I CSTR ) and the loss combination of L1loss and ssimloss for supervision.

5. The method for training a network structure capable of realizing tone mapping and contrast enhancement functions according to claim 3, It is characterized in that In step S2.1, the normalization refers to image normalization, which converts the pixel values ​​of an image into values ​​within a certain range to better perform the next step of processing. In this application, image I HDR ,I LDR and I CST The initial data type is uint16, whose theoretical maximum value is 65535. The goal is to normalize it to float32 data with a maximum value of 1, so I HDR ,I LDR and I CST Divide by 65535, then convert to float32, normalize uint16, and then convert to float32, no data loss will occur; The available Python code is: img_hdrn=(img_hdr / 65535).astype(np.float32) img_ldrn=(img_ldr / 65535).astype(np.float32) img_cstn=(img_cst / 65535).astype(np.float32); In subsequent training, the tone-mapped gt is I HDRN , contrast enhanced gt is I CSTN .

6. The method for training a network structure capable of realizing tone mapping and contrast enhancement functions according to claim 5, It is characterized in that The implementation of step S2.2 further includes: in the tasks of tone mapping and contrast enhancement, the transformation in the log domain is more in line with the visual characteristics of the human eye, so it is converted to the log domain, where the negative sign is due to the fact that the base of the log function used in this application is 2 and the value of the independent variable is 0-1, and the function value is negative at this time, so it is negated; The available Python code is: Tone Mapping: I HDRN_log = -1*torch.log2(I' HDRN ) Contrast Enhancement: I LDRN_log = -1*torch.log2(I' LDRN ).