Perception-preserving color image decoloring method

The color image is grayscaled through the deep learning network model, and the iterative perception enhancement curve is used to retain perceptual features, which solves the problem that grayscale images do not conform to human eye perception in the prior art, and achieves high-quality grayscale image generation.

CN120259448APending Publication Date: 2025-07-04HENAN UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411973134.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing color image decolorization method cannot effectively retain the perceptual characteristics of the image, resulting in the grayscale image not meeting the visual perception of the human eye.

Method used

Deep learning network model is used to decolorize color images, and color images are grayscaled through monochromatic networks and perception enhancement networks, and perceptual feature enhancement curves are used to enhance perception characteristics to generate grayscale images that conform to human eye perception.

Benefits of technology

The visual perception effect of grayscale images is improved, the contrast information, detail information and perception information of the image are effectively preserved, and the visual perception quality of the image is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259448A_ABST
    Figure CN120259448A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image decoloring, and particularly relates to a perception-preserving color image decoloring method. The method comprises the following steps: acquiring a color image and inputting the color image into a trained decoloration network model to obtain a decoloration image; wherein the decoloration network model comprises a monochromatization network (Gray-Net) and a perception enhancement network (PE-Net), and the processing process comprises the following steps: carrying out graying processing on an input color image by using the monochromatization network to obtain an initialized grey-scale map; and extracting perception features of an input color image by using the perception enhancement network to generate n feature parameter maps, taking the n feature parameter maps as parameters of a perception enhancement curve, and performing n times of iterative perception enhancement on the initialized grey-scale map by using the perception enhancement curve so as to output the initialized grey-scale map. The perception enhancement curve used in the method achieves perception feature enhancement in sequence through an iterative fitting mode, the visual perception effect of the result image is improved, and the finally obtained gray level image conforms to perception of human eyes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image decolorization, and particularly relates to a perceptual-preserving color image decolorization method. Background Art

[0002] Color image decolorization is the process of converting a color image into a grayscale image. This process can reduce data redundancy, lower storage requirements, and reduce computational complexity, so it is widely used in the image preprocessing stage of various fields, such as medical imaging, face recognition, and feature extraction. In these applications, the real-time performance of the decolorization algorithm is crucial for improving the performance of the entire image processing algorithm. In addition, image decolorization also plays an important role in other application scenarios, such as black-and-white printing for cost savings in daily life, ink painting rendering and black-and-white photography in the field of art aesthetics.

[0003] Existing decolorization methods are mainly divided into traditional decolorization methods and deep learning-based decolorization methods. Traditional decolorization methods are traditional grayscale methods, and traditional grayscale methods are further divided into local mapping methods and global mapping methods. Among them, the local mapping method has position correlation and can better retain the detail information of the image, but it ignores the global color consistency. Therefore, in the grayscale image, there will be situations such as the loss of the global contrast feature of the image and edge artifacts in the resulting image; the global mapping method has global consistency, but it ignores the local detail features of the image, which may lead to the image being too smooth and causing the loss of detail information. For example, the Chinese invention patent with the application publication number CN109903247A and the application publication date of June 18, 2019 discloses a high-precision grayscale method for color images based on the correlation of the Gaussian color space. The whole process of this method is as follows: first, convert the color image from the RGB color space to the LMN color space, and extract the L component as the luminance information. Then, obtain the standard deviations of the L, M, and N channels in the LMN color space and the correlation coefficients between the three channels. Next, use the second-order linear mapping method to obtain the chromaticity information by using the above standard deviations and correlation coefficients. Finally, add the luminance information and the chromaticity information and normalize them to obtain the final grayscale image. This method undoubtedly has the problems mentioned above. Moreover, in addition, traditional decolorization methods rarely pay attention to the retention of the perceptual features of the image, making the resulting grayscale image not conform to human visual perception.

[0004] Compared with traditional decolorization methods, deep learning-based decolorization methods begin to pay attention to the retention of the perceptual features of the image. How to more effectively retain the perceptual features, realize the retention of the details of the image information, and improve visual perception are urgent problems to be solved at present. Summary of the Invention

[0005] The purpose of the present invention is to provide a perceptual-preserving color image decolorization method to solve the technical problem that the existing decolorization methods do not conform to human visual perception.

[0006] To solve the above technical problems, the present invention provides a color image decolorization method for perceptual retention, and the method includes:

[0007] Obtain a color image and input it into a trained decolorization network model to obtain a decolorized image;

[0008] The processing process of the decolorization network model is as follows: perform grayscale processing on the input color image to obtain an initial grayscale image; extract the perceptual features of the input color image to generate n feature parameter maps, use the n feature parameter maps as the parameters of the perceptual enhancement curve, and perform n - time iterative perceptual enhancement on the initial grayscale image by using the perceptual enhancement curve and then output, and the perceptual enhancement curve is:

[0009] g j = g j-1 + β j · g j-1 · (1 - g j-1 )

[0010] In the formula, g j represents the output after the j - th iterative perceptual enhancement, g j-1 represents the output after the (j - 1) - th iterative perceptual enhancement, and g0 is the initial grayscale image; β j represents the j - th feature parameter map.

[0011] Further, use a perceptual enhancement network to extract the perceptual features of the input color image, and the perceptual enhancement network includes M convolutional layers connected in sequence, where M ≥ 2.

[0012] Further, the first M - 1 convolutional layers each include at least two convolutional kernels with the same size and stride and a Relu activation function located behind the convolutional kernels, and the M - th convolutional layer includes n convolutional kernels with the same size and stride and a Tanh activation function located behind the convolutional kernels.

[0013] Further, the process of performing grayscale processing on the input color image to obtain an initial grayscale image is as follows: input the color image into a monochromatization network to generate three weight parameter matrices, and perform weighted summation on the three weight parameter matrices and the RGB three channels of the color image, and the summation result is the initial grayscale image.

[0014] Further, the monochromatization network includes L convolutional layers. The first (L + 1) / 2 convolutional layers are connected in sequence. The output of the ((L + 1) / 2 + i) - th convolutional layer is jump - connected to the output of the ((L - 1) / 2 - i) - th convolutional layer and then input into the ((L + 3) / 2 + i) - th convolutional layer, where 0 ≤ i ≤ (L - 3) / 2. Then, the output of the (L - 1) - th convolutional layer is input into the L - th convolutional layer, and three weight parameter matrices are output by the L - th convolutional layer. L is an odd number and L ≥ 3.

[0015] Further, the first L-1 convolutional layers each include at least two convolutional kernels with the same size and stride, and a Relu activation function located behind the convolutional kernels. The L-th convolutional layer includes three convolutional kernels with the same size and stride, and a Sigmoid activation function located behind the convolutional kernels.

[0016] Further, the loss function used in training the decolorization network model is:

[0017] L sum = λ1L c + λ2L t + λ3L p

[0018]

[0019]

[0020] In the formula, L sum represents the total loss function; λ1, λ2, and λ3 represent hyperparameters used to balance different loss terms; L c represents the contrast preservation loss function, (x, y) is a pixel pair, P represents the set of global and local pixel pairs, Δg x,y represents the difference in gray values between x and y, δ x,y represents the Euclidean distance in the RGB space, σ represents the set variance; L t represents the smoothing loss function, N1 represents the number of iterations, N1 = n, β m represents the m-th generated feature parameter map, and respectively represent the horizontal and vertical gradients; L p represents the perceptual preservation loss function, N2 represents the total number of training samples, ‖·‖1 represents the L1 norm, x i represents the i-th input color image, g ni represents the perceptually enhanced gray-scale image generated corresponding to the i-th color image, and VGG(·) represents calculating the loss of the 12th convolutional layer in VGG19.

[0021] Further, M = 5.

[0022] Further, L = 7.

[0023] Further, the obtained color image is an image obtained by scaling the originally acquired color image.

[0024] The present invention is an improved invention, and its beneficial effects are as follows: When the present invention uses a deep learning network model for color image decolorization, the deep learning network model used is different from the deep learning network models in the prior art. This network model is called a decolorization network model. First, it grayscales the color image to obtain an initial grayscale image, and then uses a perceptual enhancement curve to perform perceptual enhancement processing on the initial grayscale image. The feature parameter map used in this perceptual enhancement curve contains the perceptual feature information in the original color image. The perceptual enhancement curve can automatically convert the initial grayscale image into a perceptually adjusted grayscale image, and realizes perceptual feature enhancement time after time through an iterative fitting method, improving the visual perception effect of the resulting image and making the finally obtained grayscale image conform to the perception of the human eye. Description of the Drawings

[0025] Figure 1 is the overall processing flow chart of the decolorization network model of the present invention;

[0026] Figure 2 is the structural diagram of the monochromatization network Gray-Net in the decolorization network model of the present invention;

[0027] Figure 3 is the structural diagram of the perceptual enhancement network PE-Net in the decolorization network model of the present invention;

[0028] Figure 4 is the processing flow chart of iteratively perceiving and enhancing the initial grayscale image using the perceptual enhancement curve of the present invention;

[0029] Figure 5 is a schematic diagram of the decolorization result of image decolorization using the method of the present invention. Detailed Embodiments

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0031] A method for perceptual-preserving color image decolorization according to the present invention has the following process:

[0032] Step 1: Obtain a color image and perform preprocessing on it.

[0033] Obtain the captured color image I. First, adjust its size (i.e., scale it). In this embodiment, it is adjusted to 256×256, and then the pixel values are normalized to between [0,1].

[0034] Step 2: Input the image processed in Step 1 into the decolorization network model, and extract important features through the decolorization network model to obtain the final decolorized image.

[0035] The decolorization network model is trained in an unsupervised manner and includes a monochromatization network Gray-Net and a perceptual enhancement network PE-Net.

[0036] The process of the decolorization network model for decolorizing the input color image is as Figure 1 shown. Specifically: 1) Input the color image obtained in Step 1 into the monochromatization network Gray-Net to obtain three weight parameter matrices α1, α2, α3, and then perform weighted superposition with the RGB three channels c1, c2, c3 of the color image, that is to obtain the initial grayscale image g0. 2) Input the color image obtained in Step 1 into the perceptual enhancement network PE-Net for perceptual feature extraction to generate n feature parameter maps β1, β2, … β j …, β n (j = 1, 2, …, n). Use the n feature parameter maps generated by the perceptual enhancement network as the n parameters for iteratively fitting the perceptual enhancement curve, and perform iterative enhancement on the initial grayscale image g0 to generate the perceptually enhanced feature map g n . Figure 1 In it, Paramenter represents a parameter, and Enhance-Curve represents the optimization result.

[0037] The monochromatization network Gray-Net uses a convolutional neural network ( Figure 1 the CNN-Net in it), specifically including multiple convolutional layers and adding skip connections. The specific structural connection form is: including L convolutional layers. The first (L + 1) / 2 convolutional layers are connected in sequence. The output of the (L + 1) / 2 + i-th convolutional layer is skip-connected (Skip-Connection) with the output of the (L - 1) / 2 - i-th convolutional layer and then input into the (L + 3) / 2 + i-th convolutional layer, 0 ≤ i ≤ (L - 3) / 2. Furthermore, the output of the L - 1-th convolutional layer is input into the L-th convolutional layer, and three weight parameter matrices are output by the L-th convolutional layer. L is an odd number and L ≥ 3. Moreover, the first L - 1 convolutional layers all include at least two convolutional kernels with the same size and stride and a Relu activation function behind the convolutional kernels. The L-th convolutional layer includes three convolutional kernels with the same size and stride and a Sigmoid activation function behind the convolutional kernels. The structure used in this embodiment is as Figure 2 shown. L = 7. Among them, the first six layers use 32 convolutional kernels with a size of 3×3 and a stride of 1, followed by a Relu activation function; the last layer consists of 3 convolutional kernels with a size of 3×3 and a stride of 1, followed by a Sigmoid activation function, so as to generate three weight parameter matrices. Among them, Figure 2Convolution in it represents convolution, Relu represents the Relu activation function, Sigmoid represents the Sigmoid activation function, and Skip-Connection represents the skip connection.

[0038] The perception enhancement network PE-Net adopts a convolutional neural network ( Figure 1 the CNN-Net in it), specifically including M convolutional layers connected in sequence, M≥2. The first M - 1 convolutional layers each include at least two convolutional kernels with the same size and stride, and a Relu activation function behind the convolutional kernels. The Mth convolutional layer includes n (the number is the same as the number of iterations) convolutional kernels with the same size and stride, and a Tanh activation function behind the convolutional kernels. The structure used in this embodiment is as Figure 3 shown, and it includes a total of five convolutional layers, M = 5. The first four layers (that is, Figure 3 the 4 Conv Blocks in it) are all composed of 32 convolutional kernels with a size of 3×3 and a stride of 1, followed by a Relu activation function; the last layer (that is, Figure 3 the Conv Block-1 in it) is composed of 5 convolutional kernels with a size of 3×3 and a stride of 1, followed by a Tanh activation function, thereby generating five feature parameter maps.

[0039] Moreover, the initial grayscale image g0 is subjected to n - times iterative perception enhancement through the iteratively fitted perception enhancement curve to improve the perception retention effect of the image. The number of iterations is set to 5 times, that is, n = 5. The curve can be expressed as:

[0040] E p (g; β) = g + β·g·(1 - g)

[0041] In the formula, g represents the input grayscale image; β represents the feature parameter map generated by the perception enhancement network; E p (g; β) represents the output grayscale image after enhancement processing for the given input g. The iterative enhancement process is as Figure 4 shown, and the formula is as follows:

[0042] E p1 = g1 = g0 + β1·g0·(1 - g0)

[0043] ……

[0044] E pj = g j = g j-1 + β j ·g j-1 ·(1 - g j-1 )

[0045] ……

[0046] Epn = g n = g n-1 + β n ·g n-1 ·(1 - g n-1 )

[0047] Wherein, g0 represents the initialized grayscale image generated by the monochromatization network; β j is a feature parameter map with the same size as the input image; g n represents the grayscale image after perceptual enhancement.

[0048] The total loss function for training the above decolorization network model consists of a contrast preservation loss function, a smoothness loss function, and a perceptual preservation loss function (i.e., Figure 1 the Loss-Function in). The total loss function is expressed as:

[0049] L sum = λ1L c + λ2L t + λ3L p

[0050] Wherein, λ1, λ2, and λ3 represent hyperparameters for balancing different loss terms, which are respectively set to 1, 200, 1 in this embodiment; L c represents the contrast preservation loss function, which is used to penalize the contrast difference between the color image and the initialized grayscale image in the monochromatization network; L t represents the smoothness loss function, which helps to reduce the generation of edge artifacts during the decolorization process; L p represents the perceptual preservation loss function, which is used to penalize the perceptual feature preservation difference between the color image and the resulting grayscale image in the perceptual enhancement network. The calculation formulas of these three loss functions are respectively:

[0051]

[0052]

[0053] Wherein, the grayscale values of pixels x and y are respectively represented by g x and g y , (x, y) is a pixel pair; P represents the set of global pixel pairs and local pixel pairs, and the difference in the grayscale values of x and y in P is represented as Δg x,y = g x - g y ; δ x,y represents the Euclidean distance in the RGB space and represents the contrast of the input color image; σ represents the set variance, which can take 0.05. N1 represents the number of iterations, and N1 = 5 in this embodiment; β mIndicates the generated feature parameter map, where the horizontal and vertical gradients are respectively denoted as and N2 represents the total number of training samples; ‖·‖1 represents the L1 norm; x i represents the i-th input color image; g ni represents the perceptually enhanced grayscale image corresponding to the i-th color image; VGG(·) represents calculating the loss of conv4-4 in VGG19 (i.e., the 12th convolutional layer in VGG19).

[0054] The method of the present invention is used to perform decolorization processing on color images, and the results before and after processing are as Figure 5 shown. It can be seen from this figure that the decolorized image finally obtained by the present invention has better improved the visual perception effect of the image.

[0055] Next, experiments are carried out to illustrate the real-time performance of the method of the present invention. The running times of different decolorization algorithms are shown in Table 1. The size of the color images used in the experiments is 256×256, and the experimental equipment is equipped with an Intel Core i9-13900k, an NVIDIA GeForce RTX4070 graphics card, and 64GB of RAM. Among them, the rgb2gray, Sowmya, Yu, and Liu methods are run on the Matlab platform, while the method of the present invention and the Cai method are run on the Pytorch platform. By running the rgb2gray code on the above two platforms respectively, it can be obtained that the running time ratio of rgb2gray on the Matlab and Pytorch platforms is approximately (1:1.4). According to this result, the running times of different decolorization algorithms on different platforms can be unified to the running time based on the Matlab platform and shown in Table 1. It can be seen from the results in the table that the running time of the method of the present invention is at a leading level in the running time comparison, which can ensure the real-time performance of decolorization.

[0056] Table 1 Comparison of running times of decolorization algorithms

[0057] Method rgb2gray Sowmya Yu Cai Liu The method of the present invention Running time 0.0026s 0.26s 0.025s 100s 8.1s 0.0035s

[0058] In summary, the present invention has the following characteristics: The present invention uses a decolorization network model to extract key features and perceptual features in the image, and enhances the perceptual features of the initialized grayscale image through an iteratively fitted perceptual enhancement curve, so that the finally obtained grayscale image conforms to the perception of the human eye; moreover, during the decolorization process, contrast loss, smooth loss, and perceptual preservation loss are used to train the network in an unsupervised manner, obtaining a high-quality grayscale image while ensuring the real-time performance of the algorithm, realizing the effective preservation of contrast information, detail information, and perceptual information in the image during the decolorization process, and improving the visual perception effect of the decolorized image; at the same time, the algorithm efficiency is at a leading level.

[0059] Specific embodiments are given above, but the present invention is not limited to the described embodiments. The basic idea of the present invention lies in the above basic solution. For those of ordinary skill in the art, according to the teachings of the present invention, it does not require creative labor to design various deformed models, formulas, and parameters. Changes, modifications, substitutions, and variations made to the embodiments without departing from the principles and spirit of the present invention still fall within the protection scope of the present invention.

Claims

1. A color image decolorization method for perception retention, characterized in that, The method includes: Obtaining a color image and inputting it into a trained decolorization network model to obtain a decolorized image; The processing process of the decolorization network model is: performing grayscale processing on the input color image to obtain an initial grayscale image; extracting the perceptual features of the input color image to generate n feature parameter maps, using the n feature parameter maps as the parameters of the perceptual enhancement curve, and performing n - time iterative perceptual enhancement on the initial grayscale image using the perceptual enhancement curve and then outputting, where the perceptual enhancement curve is: g j = g j-1 + β j · g j-1 · (1 - g j-1 ) where g j represents the output after the j-th iteration of iterative perception enhancement, and g j-1 represents the output after the (j - 1)-th iteration of iterative perception enhancement, and g0 is the initialized grayscale image; β j represents the j-th feature parameter map.

2. The color image decolorization method for perception retention according to claim 1, characterized in that, Using a perceptual enhancement network to extract the perceptual features of the input color image, and the perceptual enhancement network includes M convolutional layers connected in sequence, where M ≥ 2.

3. The color image decolorization method for perceptual retention according to claim 2, wherein, The first M - 1 convolutional layers each include at least two convolutional kernels with the same size and stride and a Relu activation function behind the convolutional kernels, and the M - th convolutional layer includes n convolutional kernels with the same size and stride and a Tanh activation function behind the convolutional kernels.

4. The method for decolorizing a color image with perceptual retention according to claim 1, characterized in that, The process of performing grayscale processing on the input color image to obtain an initial grayscale image is: inputting the color image into a monochromatization network to generate three weight parameter matrices, and performing weighted summation of the three weight parameter matrices with the RGB three channels of the color image, and the summation result is the initial grayscale image.

5. The method for decoloring a color image with perception retention according to claim 4, wherein The monochromatization network includes L convolutional layers. The first (L + 1) / 2 convolutional layers are connected in sequence. The output of the (L + 1) / 2 + i - th convolutional layer is jump - connected to the output of the (L - 1) / 2 - i - th convolutional layer and then input into the (L + 3) / 2 + i - th convolutional layer, where 0 ≤ i ≤ (L - 3) / 2. Then the output of the L - 1 - th convolutional layer is input into the L - th convolutional layer, and the L - th convolutional layer outputs three weight parameter matrices. L is an odd number and L ≥ 3.

6. The method for decolorizing a color image with perceptual retention according to claim 5, characterized in that, The first L - 1 convolutional layers each include at least two convolutional kernels with the same size and stride and a Relu activation function behind the convolutional kernels, and the L - th convolutional layer includes three convolutional kernels with the same size and stride and a Sigmoid activation function behind the convolutional kernels.

7. The color image decolorization method for perceptual retention according to claim 1, characterized in that The loss function used in training the decolorization network model is: L sum = λ1L c + λ2L t + λ3L p Where L sum represents the total loss function; λ1, λ2, and λ3 represent hyperparameters used to balance different loss terms; L c represents the contrast preservation loss function, (x, y) is a pixel pair, P represents the set of global and local pixel pairs, and Δg x,y represents the difference in grayscale values between x and y, and δ x,y represents the Euclidean distance in the RGB space, and σ represents the set variance; L t represents the smoothing loss function, N1 represents the number of iterations, N1 = n, and β m represents the m-th generated feature parameter map, and represent the horizontal and vertical gradients respectively; L p represents the perception preservation loss function, N2 represents the total number of training samples, ‖·‖1 represents the L1 norm, and x i represents the i-th input color image, and g ni represents the grayscale image corresponding to the i-th color image after the generated perception enhancement, and VGG(·) represents calculating the loss of the 12th convolutional layer in VGG19.

8. The color image decolorization method for perception retention according to claim 2, wherein M=5。 9. The color image decolorization method for perceptual retention according to claim 5, characterized in that, L=7。 10. The method for decolorizing a perceptually maintained color image according to any one of claims 1 to 9, characterized in that, The obtained color image is an image obtained by scaling the originally collected color image.

Citation Information

Patent Citations

  • A color image high-precision graying method based on Gaussian color space correlation

    CN109903247A