A Video Contrast Enhancement Method and System
A self-supervised deep learning method for video contrast enhancement addresses adaptability issues by using a neural network and unsupervised image quality models for adaptive grayscale mapping, ensuring friendly and efficient scene-independent performance.
Patent Information
- Application Number
- CN202111655763.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The existing video contrast enhancement algorithm has poor applicability in different scenarios and is unfriendly, relies on human parameter settings and the training results are affected.
The deep neural network learning model weight is adopted, combined with the unsupervised image quality evaluation model, and the adaptive grayscale mapping mechanism is designed, and video contrast enhancement is achieved through L1 loss and IQAloss supervised learning.
No human intervention is required, it can adapt to different image scenes, achieve real-time video contrast enhancement, and improve algorithm friendliness and effect.
Smart Images

Figure CN114372930B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video processing technology, and in particular to a method and system for enhancing video contrast. Background Art
[0002] Existing algorithms for enhancing video contrast mainly perform certain operations on the Y-component data of an image to enhance the contrast of the image. Generally, there are methods such as histogram equalization, gamma transformation, and multiplying the image data by a constant. These methods rely on manual parameter settings and cannot be well applied to various scenarios, and there is unfriendliness in the use of the algorithms.
[0003] Deep learning-based methods can be divided into two categories: unsupervised and supervised. Self-supervised methods belong to a type of unsupervised method. In the supervised methods based on deep learning, paired low-illumination images and normal-illumination images are often required for training. This type of method can generally well suppress noise in the enhancement result. A self-supervised low-illumination image enhancement method based on deep learning is disclosed in 202010097457.4, which includes the following steps: inputting the low-illumination image to be enhanced into an image enhancement network; the input of the image enhancement network is the low-illumination image S and its maximum-value channel image Smax. S is a matrix of M*N*3, where M is the number of rows, N is the number of columns, and 3 represents the three color channels of {r, g, b}. Smax is obtained by taking the maximum value of the three color channels and is a matrix of M*N*1. S and Smax are combined into a matrix of M*N*4 as the input of the network: the reflected image R output by the image enhancement network is the enhanced image. The structure of the image enhancement network is as follows: the input is respectively input into a first convolutional layer and a second convolutional layer. The first convolutional layer and the second convolutional layer are convolutional layers of 9*9 and 3*3 respectively; the first convolutional layer is connected to a third convolutional unit, and the third convolutional unit is a convolutional layer of 3*3 followed by a ReLU layer; the third convolutional unit is connected to a fourth convolutional unit, the fourth convolutional unit is connected to a fifth convolutional unit, the fifth convolutional unit is connected to a sixth convolutional unit. The fourth convolutional unit, the fifth convolutional unit, and the sixth convolutional unit are all convolutional layers of 3*3 followed by a ReLU layer; the output of the sixth convolutional unit and the output of the third convolutional unit are subjected to a Concat operation and then input into a seventh convolutional unit, and the seventh convolutional unit is a convolutional layer of 3*3 followed by a ReLU layer; the output of the seventh convolutional unit and the output of the second convolutional unit are subjected to a Concat operation and then input into an eighth convolutional layer, and the eighth convolutional layer is connected to a ninth convolutional layer. The eighth convolutional layer and the ninth convolutional layer are both convolutional layers of 3*3; the ninth convolutional layer is connected to a Sigmoid activation function layer; the Sigmoid activation function layer is connected to an output layer, and the output is the reflected image R and the illumination image I.
[0004] The image enhancement network is a trained image enhancement network, and the training process is as follows:
[0005] A1. Collect any n low-illumination images, where n >= 1, to construct a training dataset; A2. For each low-illumination image S in the training dataset, extract its corresponding maximum channel image Smax, and process Smax using histogram equalization to obtain the histogram-equalized maximum channel image SHe_max; A3. Using the histogram-equalized maximum channel image SHe_max as supervision, construct a loss function by combining the Retinex theory and the assumption of smooth illumination image I, and train the image enhancement network.
[0006] The result of this processing is directly affected by training and cannot be well applied to various scenarios, and there is inapplicability in the use of the algorithm. Summary of the Invention
[0007] The present invention provides a video contrast enhancement method and system to solve the technical problem that the existing technology cannot be well applied to various scenarios and there is inapplicability in the use of the algorithm.
[0008] A video contrast enhancement method includes:
[0009] S1: Obtain model weights weights
[1024]
[256] through deep neural network learning;
[0010] S2: Obtain a Map mapping matrix, which further includes: calculating the Map mapping matrix Map
[1024] through the model weights weights
[1024]
[256] , sorting the Map mapping matrix Map
[1024] from small to large, and then averaging every four bits to obtain the final Map
[256] ;
[0011]
[0012] histY[k] calculates the histogram on the Y component;
[0013] S3: Obtain the current YCbCr data of the image to be processed, and obtain the Y component data therefrom. According to the Map mapping matrix in step S2, obtain the current data y of the Y component of the image after video contrast enhancement:
[0014] y = Map[x] x ∈ [0, 255]
[0015] where x is the Y component data of the input video.
[0016] The present invention further includes:
[0017] S4: Adopt the effective supervised learning method of L1 loss and IQA loss;
[0018]
[0019] The expression of function F is as follows:
[0020]
[0021] N represents the number of pixels of the Y component of the image; T is a constant; iqa represents a no-reference image quality assessment model.
[0022] The present invention is a video contrast enhancement algorithm based on self-supervised learning, which uses an unsupervised image quality assessment model to guide the training of the video contrast enhancement algorithm model. The present invention can achieve real-time video contrast enhancement on the CPU. Inspired by the image histogram equalization algorithm, the present invention designs an adaptive gray mapping mechanism, uses the unsupervised image quality assessment model as a guide, and regresses to obtain the Map mapping matrix; according to different images, different Map mapping matrices will be obtained for video contrast enhancement, without human intervention, which is very friendly in the use of the algorithm. Description of the Drawings
[0023] Figure 1 It is a principle flow chart of a video contrast enhancement method. Detailed Embodiments
[0024] The present invention will be specifically described below with reference to the accompanying drawings.
[0025] The applicant found that histogram equalization expands the dynamic range of an image with a relatively small dynamic range of gray value distribution (such as an image with gray values concentrated on the right side of the histogram, and the image is too bright at this time), and the number of gray levels of the changed image may decrease. The known conditions in Table 1 are: the total number of gray levels is 8, and the distribution probability corresponding to each level of the original image is p s (s k ) After histogram equalization, gray level 0 is mapped to 1, gray level 1 is mapped to 3, gray level 2 is mapped to 5, gray levels 3 and 4 are mapped to 6, and gray levels 5, 6, and 7 are mapped to 7. It can be seen that the histogram equalization algorithm reduces the gray levels of the image from the 8 gray levels of 0, 1, 2, 3, 4, 5, 6, 7 to the 5 gray levels of 1, 3, 5, 6, 7. Numerically, it is not difficult to see that the contrast of some local parts of the image must be enhanced.
[0026] Table 1 Histogram Equalization Calculation Process
[0027]
[0028] The contrast enhancement algorithm is for a video in YUV420p format. Its luminance component is y. According to the formula, the adjusted luminance component y′ can be obtained as follows:
[0029] y′ = α·(y - 127) + 127 + β·255
[0030] where the value range of y is [0, 255]; α is the contrast enhancement factor.
[0031] How can we avoid designing various parameters in the video contrast enhancement algorithm? For this purpose, the applicant refers to histogram equalization to improve the video contrast enhancement method, as Figure 1 shown below.
[0032] A video contrast enhancement method includes:
[0033] S110: Obtain the model weights weights
[1024]
[256] through deep neural network learning;
[0034] S120: Obtain the Map mapping matrix, which further includes: Calculate the Map mapping matrix Map
[1024] through the model weights weights
[1024]
[256] , sort the Map mapping matrix Map
[1024] from small to large, and then average every four bits to obtain the final Map
[256] ;
[0035]
[0036] histY[k] is to calculate the histogram on the Y component;
[0037] S130: Obtain the current YCbCr data of the image to be processed, and obtain the Y component data from it. According to the Map mapping matrix in step S2, obtain the current data y of the Y component of the image after video contrast enhancement:
[0038] y = Map[x] x ∈ [0, 255]
[0039] where x is the Y component data of the input video.
[0040] The present invention may also include an effective supervised learning method using L1 loss and IQA loss;
[0041]
[0042] where the expression of the function F is as follows:
[0043]
[0044] N represents the number of pixels in the Y component of the image; T is a constant; iqa represents that in the model training of the no-reference image quality assessment model, T is set to 0.0784.
[0045] The no-reference image quality assessment model is a common model in the industry. A brief introduction is as follows. No-reference image quality assessment refers to directly calculating the visual quality of a distorted image in the absence of a reference image. According to whether the no-reference image quality assessment model needs the subjective scores of images for training when calculating the visual quality of images, the no-reference image quality assessment algorithms can be divided into supervised learning-based no-reference image quality assessment algorithms and unsupervised learning-based no-reference image quality assessment algorithms. The unsupervised learning-based no-reference image quality assessment algorithms mainly include traditional machine learning-based methods and deep learning-based methods. Traditional machine learning-based methods: In 2013, Mittal et al. proposed an unsupervised method based on their previous work to achieve no-reference image quality assessment NIQE. The NSS features of intact image patches were used to fit a multivariate Gaussian (MVG) model. Based on NIQE, Zhang et al. proposed the Integrated Local NIQE (IL-NIQE) algorithm by integrating structural statistical features, multi-scale direction and frequency statistics, and color statistical features. In 2013, Xue et al. proposed an image quality assessment method based on quality-aware clustering of image patches
[16] . The quality scores of each image patch were calculated through full-reference image quality assessment, and were divided into L major categories according to the range of all image patch quality scores. Based on the extracted structural features, the major categories of image patches were divided into K minor categories by the K-means clustering algorithm. Among them, each minor category corresponds to a clustering center, which is an image patch (structural features and visual quality scores). Given a distorted image, image patches were extracted, and the distances between the image patches and the clustering centers of each minor category in the L major categories were calculated. The visual scores of the clustering centers with the minimum distance in each major category were fused to calculate the scores of the image patches. Finally, using the average weight strategy, the visual quality score of the distorted image was obtained. Zhang et al. first classified the distortion types of images, then extracted NSS features for different distortion types, trained the corresponding regression models using SVR, estimated the distortion parameters and saved them. Finally, the visual quality score of the distorted image was calculated. Deep learning-based methods: Ma et al. designed a fully connected neural network image quality assessment model based on weight sharing. Its input is two images, and the method CORNIA is used to extract the visual features of the images. The image quality assessment model is trained by fusing the uncertainty of the visual differences between images and the quality of the visual quality of the images. In Lin et al., an image quality assessment method based on the generative adversarial network (GAN) was proposed. This method includes a generator network and a discriminator network. The goal of the generator network is to generate a pseudo-reference image, and the goal of the discriminator network is to distinguish the generated pseudo-reference image from the real reference image.
[0046] "Obtaining the current YCbCr data of the image to be processed" further includes converting the RGB image into YCbCr data.
[0047] By using an unsupervised image quality evaluation model to guide the training of a video contrast enhancement algorithm model, the present invention can achieve real-time video contrast enhancement on a CPU. Inspired by the image histogram equalization algorithm, the present invention designs an adaptive gray-scale mapping mechanism. Using the unsupervised image quality evaluation model as a guide, the Map mapping matrix is obtained by regression; different Map mapping matrices are obtained for different images for video contrast enhancement without human intervention.
[0048] Application Example
[0049] A video contrast enhancement method includes:
[0050] Step0: The first step still refers to the content of machine learning. The way to learn effectiveness is L1 loss. The following is the calculation formula of the L1 loss index:
[0051]
[0052] Where N represents the number of pixels in a single channel of the image; U gt represents the U component data of GT in the fiveK open-source dataset; V gt represents the V component data of GT in the fiveK open-source dataset; U' represents the U component in the predicted dataset output during the model training process; V' represents the V component in the predicted dataset output during the model training process.
[0053] This part is the outermost framework of model training, which is used to judge whether the training result reaches the best effect. Here, the predicted data of the model is used, and an L1Loss calculation is performed through the predicted data and the source data. When the change of L1Loss has been stable at a very small level, we consider that the training has achieved the best effect. The so-called predicted data here calculates the Y component, and a histogram histY
[256] is calculated on the Y component. The trained model is also a linear model, and the training is carried out by applying pytorch.nn.linear. Its working principle is to input a Y-component histogram data with a dimension of 256 dimensions, and through a linear transformation, the model weights weights
[1024]
[256] of the Y component are obtained. We mainly use pytorch.nn.linear, which is a model for linear regression. The input is histogram data. For the training of the UV component, we input the UV histogram data, and for the training of the Y component, we input the Y-component histogram data. linear will directly output a list of weights after linear regression by continuously learning these data. The dimension of this list is predefined by us and passed as a parameter to linear, and linear will output the weight data of the specified dimension according to the dimension parameters we defined.
[0054] Effective supervised learning method. Using L1loss and IQAloss.
[0055]
[0056] The expression of function F is as follows:
[0057]
[0058] T is a constant, which is set to 0.0784 in model training; iqa represents a no-reference image quality assessment model.
[0059] Step1: This step refers to model training. The process of machine learning is the so-called model training. This part mainly includes the model algorithm logic, input training data, and output prediction set.
[0060] The input data for model training is the open-source REDS deblur dataset. The trained model is a linear model, and the training is carried out by applying pytorch.nn.linear. Its working principle is to input a Y-component histogram data histUV[256×256] with a dimension of 256 dimensions, and through the following linear transformation formula, we can obtain the model weights weights
[1024]
[256] of the Y component
[0061] Linear transformation formula: y = xA + b
[0062] After obtaining the above weight data, it can be transposed during actual operations and then matrix multiplication can be performed with x.
[0063] Here, the Y component is calculated. A histogram histY
[256] is calculated for the Y component. The trained model is also a linear model, and it is trained by applying pytorch.nn.linear. Its working principle is to input histogram data of the Y component with a dimension of 256, and through linear transformation, obtain the model weights weights
[1024]
[256] of the Y component.
[0064] Through the above weights
[1024]
[256] , we can obtain a Map
[1024] . The purpose of this Map is to map a corresponding weight for each value of the Y component. Ultimately, there should be 256 weight data. Here, 1024 data are first trained simply to increase the quantity and diversity of the data so that the change of the trained weight coefficients is smoother. The subsequent processing still has one weight corresponding to one Y component.
[0065] This map weight is processed as follows: Sort Map
[1024] from smallest to largest and average every four elements to obtain the final Map
[256] . After such processing, there is one weight corresponding to one Y component. Through practice, it is found that such weights are smoother in actual enhancement processing.
[0066] Step2: First, calculate the histogram of the Y component of the current image data to obtain histY
[256] . Then, according to the weight data obtained in the above steps, perform the following processing to obtain the enhanced Y component.
[0067]
[0068] Then, the Y component data y of the image after video contrast enhancement:
[0069] y = Map[x] x ∈ [0, 255]
[0070] where x is the Y component data of the input video.
[0071] k all refer to an index value in the range of 0 - 255.
[0072] A video contrast enhancement system includes:
[0073] A model weight training model: used to learn the model weights weights
[1024]
[256] through a deep neural network;
[0074] Map Mapping Matrix Obtaining Unit: Used to obtain the Map mapping matrix, which further includes: calculating the Map mapping matrix Map
[1024] through the model weights weights
[1024]
[256] , sorting the Map mapping matrix Map
[1024] from small to large, and then averaging every four bits to obtain the final Map
[256] ;
[0075]
[0076] histY[k] calculates the histogram on the Y component;
[0077] Current Data y Calculation Unit of Image: Used to obtain the current YCbCr data of the image to be processed, obtain the Y component data from it, and obtain the subsequent video contrast according to the Map mapping matrix in step S2; The current data y of the enhanced image:
[0078] y = Map[x] x ∈ [0, 255]
[0079] Where x is the Y component data of the input video.
[0080] Supervised Learning Processing Unit: Used to adopt an effective supervised learning method of L1 loss and IQA loss;
[0081]
[0082] The expression of the function F is as follows:
[0083]
[0084] N represents the number of pixels in the Y component of the image; T is a constant; iqa represents a no-reference image quality assessment model.
[0085] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, provided that these changes fall within the scope of the claims of the present invention and their equivalent technologies, they still fall within the protection scope of the present invention.
Claims
1. A method for enhancing video contrast, characterized in that: including: S1: Obtain model weights through deep neural network learning , and perform training by applying pytorch.nn.linear. The input dimension is a 256-dimensional Y-component histogram data calculated on the Y component of GT in the fiveK open-source dataset. After the linear transformation y = xA + b, where A is the weight matrix of the pytorch.nn.linear layer and b is the bias vector, the model weights of the Y component are obtained ; S2: Obtain the Map mapping matrix, which further includes: through the model weights Calculate to obtain the Map mapping matrix , and sort the Map mapping matrix in ascending order, and then take the average of every four digits to obtain the final ; is in the Y-component histogram data; S3: Obtain the current YCbCr data of the image to be processed, and obtain the Y component data therefrom. According to the Map mapping matrix in step S2, obtain the current Y component data of the image after video contrast enhancement : Among them is the Y-component data of the input video.
2. The method according to claim 1, characterized in that, further including: S4: Adopt and effective supervised learning methods; Among them, the function has the following expression: N represents the number of pixels of the Y component of the image; is a constant; represents a no-reference image quality assessment model.
3. The method according to claim 2, characterized in that, During model training Set to 0.0784.
4. The method according to claim 1, characterized in that, further including: The way of learning effectiveness is , and the following is the calculation formula of the indicator: Where N represents the number of pixels in a single channel of the image; represents the U component data of GT in the fiveK open-source dataset; represents the V component data of GT in the fiveK open-source dataset; represents the U component in the predicted dataset output during the model training process; represents the V component in the predicted dataset output during the model training process; This part is the outermost framework of model training, used to determine whether the training results reach the best effect. Here, the predicted data of the model is used, and a calculation is performed through the predicted data and the open-source data When the change of has been stable within the preset change range, it is considered that the training has achieved the effect. For the Y-component data of GT in the fiveK open-source dataset, a histogram is calculated on the Y-component data The trained model is also a linear model, and it is trained by applying pytorch.nn.linear. Its working principle is to input the Y-component histogram data with a dimension of 256 dimensions, and through linear transformation, obtain the model weights of the Y-component .
5. The method according to claim 1 or 4, characterized in that, further including: Model training. The process of machine learning is the so-called model training. This part mainly includes the model algorithm logic, inputting training data, and outputting a prediction set. The input data for model training is the open-source REDS deblur dataset, and the trained model is a linear model, which is trained by applying pytorch.nn.linear. Its working principle is to input a Y-component histogram data with a dimension of 256 dimensions. After the following linear transformation formula, the model weights of the Y component ; Linear transformation formula: y = xA + b Having obtained the above weight data, it can be transposed again during actual operations, and then matrix multiplication can be performed with x. The Y component is calculated here, and the histogram is calculated on the Y component , and the trained model is also a linear model. It is trained by applying pytorch.nn.linear. Its working principle is to input the Y-component histogram data with a dimension of 256, and through linear transformation, obtain the model weights of the Y component .
6. A video contrast enhancement system, characterized in that: including: Model weight training model: used to obtain model weights through deep neural network learning ; It is trained by applying pytorch.nn.linear. The input dimension is a 256-dimensional Y-component histogram data calculated on the Y component of GT in the fiveK open-source dataset. After the linear transformation y = xA + b, where A is the weight matrix of the pytorch.nn.linear layer and b is the bias vector, the model weights of the Y component are obtained ; Map mapping matrix acquisition unit: used to obtain the Map mapping matrix, which further includes: through the model weights Calculate to obtain the Map mapping matrix , and sort the Map mapping matrix from small to large, and then average every four digits to obtain the final ; is the Y-component histogram data; Current data of the image Calculation unit: It is used to obtain the current YCbCr data of the image to be processed, obtain the Y component data therefrom, and obtain the current data of the image after video contrast enhancement according to the Map mapping matrix : Among them is the Y component data of the input video.
7. The system according to claim 6, wherein further including: Supervised learning processing unit: for adopting and an effective supervised learning method; Among them, the function has the following expression: N represents the number of pixels of the Y component of the image; is a constant; represents a no-reference image quality assessment model.
Citation Information
Patent Citations
A Deep Learning-Based Self-Supervised Low-Light Image Enhancement Method
CN111402145B
Image processing method and device
CN104599238A
Method, device and storage medium for enhancing image contrast
CN108513672A