Automatic image coloring method based on semantic segmentation and generative adversarial network

By introducing semantic segmentation and generation adversarial networks into the automatic image coloring technology, combining color histogram information and shape prior modules, the problems of color uncertainty and color overflow are solved, and high-quality and high-saturation image coloring effects are achieved.

CN120070612APending Publication Date: 2025-05-30DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411839870.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as color uncertainty, color overflow, and color incompleteness in the process of automatic image coloring, resulting in poor color quality and authenticity.

Method used

Using an automatic image coloring method based on semantic segmentation and generation adversarial networks, a brand new storage network is designed to store and utilize color histogram information and category information, combined with the shape prior module to reduce color overflow and overflow and improve color saturation.

Benefits of technology

It significantly improves the color quality and color saturation, reduces color overflow and unnatural color transitions, and improves the generalization ability and training stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070612A_ABST
    Figure CN120070612A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic image coloring method based on semantic segmentation and a generative adversarial network, and relates to the technical field of computer vision and deep learning. According to the technical scheme, a model with ChromaGAN as a basic framework is adopted; designing a semantic segmentation network used for helping the model to obtain more semantic analysis information; adding a shape prior module into the segmentation network; an external storage network is used for guiding the coloring main network, color histogram information and category information of the image are extracted, and the information is stored in the storage network; and fusing the stored color histogram information and semantic information to improve the color saturation and finish coloring. The method has the beneficial effects that through an innovative technical scheme, key problems in an existing image coloring technology are solved, the coloring quality, the color saturation and the sense of reality are remarkably improved, meanwhile, the generalization ability and the training stability of the model are enhanced, the user interaction cost is reduced, and the user experience is improved. And a new thought and a new solution are provided for development and application of an image coloring technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and relates to a technology for converting grayscale images into color images and automating the process; more specifically, it relates to an automatic image coloring method based on semantic segmentation and generative adversarial networks. Background Art

[0002] In today's digital age, image coloring, as an important technology in the field of computer vision, has attracted extensive attention in both academia and industry. This technology aims to convert monochromatic black-and-white images into colorful color images to restore and reproduce the original appearance of black-and-white historical photos and old movies, as well as add colors to the line drawings of anime and comics. In this way, not only can the visual effect of the image be enhanced, but also a more intuitive and rich historical and cultural experience can be provided for people. However, despite the significant progress made in image coloring technology, it still faces a series of challenges. Among them, the problem of color uncertainty is one of the most critical challenges. Due to the diversity of the colors of actual objects, the same object may exhibit different colors in different environments. For example, a flower can be red, blue, yellow, etc. This subjectivity and diversity of colors bring great difficulties to automatic image coloring, resulting in possible results that do not match the colors of the original scene during the automation process, affecting the coloring quality and authenticity. In addition, problems such as uneven coloring effect with overflow phenomenon and low vividness are also relatively critical challenges.

[0003] To solve these problems, researchers have tried various methods. Traditional methods based on reference images and color scribbles that require user guidance need to introduce additional color images as guidance. The quality of the generated colors depends to a large extent on the reference images. However, obtaining these reference images requires a great deal of effort, and the finally generated images often lack color saturation. With the development of neural networks, automatic image coloring methods based on convolutional neural networks have been widely studied. Initially, some methods aimed to predict the color distribution of each pixel in the image using convolutional neural networks. However, due to the lack of a comprehensive understanding of the image semantics, these methods usually produce incoherent colors and color confusion phenomena. To obtain more semantic information, some recent methods have combined other tasks such as classification, detection, and segmentation to enhance global or object-level semantic representations. Unfortunately, they still cannot build long-range visual dependencies, and the actual coloring effect still has problems such as color overflow and unsaturated colors. To solve the problem of unsaturated colors in the coloring effect, some methods use additional color information as guidance, such as generating prior images and external color memories. However, due to the limited color information obtained, they still cannot generate vivid and saturated images. Summary of the Invention

[0004] To solve the technical problems existing in the above-mentioned prior art, the present invention provides an automatic image coloring method based on semantic segmentation and generative adversarial network. The method designs a brand-new storage network for storing the color information and classification information of the input image. Among them, the color information of the present invention is represented by a color histogram vector. Compared with other color vector extraction methods, the color histogram vector can capture more comprehensive color information and obtain a more saturated effect when guiding coloring. A semantic segmentation network combined with shape prior is designed to obtain accurate semantic information of the image, thereby reducing the problems of unnatural color transition and color spillover. On some challenging data sets, experiments show that the method of the present invention is superior to most current image coloring methods.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] An automatic image coloring method based on semantic segmentation and generative adversarial network, the steps are as follows:

[0007] S1. Adopt a model based on ChromaGAN as the basic framework;

[0008] S2. Design a semantic segmentation network in the model to help the model obtain more semantic parsing information;

[0009] S3. Add a shape prior module in the segmentation network to effectively model the long-distance context dependence of the network, so as to obtain rich texture information related to the global region;

[0010] S4. Use an external storage network to guide the main coloring network, extract the color histogram information and category information of the image, and store this information in the storage network;

[0011] S5. Use the stored color histogram information and semantic information for fusion to enhance the color saturation to complete coloring.

[0012] Further, in the segmentation network, the following steps are executed:

[0013] T1. Use SAM to segment the data set of the present invention;

[0014] T2. Introduce a shape prior module in the segmentation network part. By using explicit shape prior information as auxiliary information, use this information to guide the model to segment the target object, thereby improving the segmentation performance of the model;

[0015] T3. The segmentation network obtains the parsing features of the image through continuous upsampling under the supervision of the pre-segmented labels, and then fuses the upsampled parsing information with the main coloring network. The specific process is shown in Formula 3.2 and Formula 3.3.

[0016] S = U(VGG(I)) (3.2)

[0017]

[0018] Where I is the input image, VGG(I) represents the features extracted from the input image using the VGG network, and U represents the upsampling process to convert the extracted features into a higher-resolution segmentation map S. represents the loss function, which is used to calculate the difference between S and T.

[0019] Further, the shape prior module includes a self-updating block for helping the network effectively model long-range context dependencies to obtain rich texture information related to the global region, and a cross-updating block for injecting inductive bias into the network to obtain more detailed local shape information.

[0020] Further, the model also includes a classification network for obtaining high-level features of the image and class label information for image coloring; in the classification network, the feature information extracted from VGG is first processed by several convolutional blocks, and then split into two different fully connected layers. The fully connected layer located on the upper side of the network is used to compress the obtained information into a 256-dimensional vector for stacking with the coloring backbone network;

[0021] The fully connected layer on the right side of the network is used to output class labels. The feature vector compression and class label output processes are shown in Equations 3.4 and 3.5:

[0022] Fcompredssde = FC(ConvBlocks(VGG(I))) (3.4)

[0023] C102 = FCright(ConvBlocks(VGG(I))) (3.5)

[0024] Where VGG(I) represents the information extracted from the input image through the VGG network, ConvBlocks represents the convolutional blocks, FC represents the fully connected layer; C102 represents 102 classes, and if the ImageNet dataset is used, it is replaced with C1000 for representation. The class information finally extracted by the network is stored in the storage network.

[0025] Further, in the storage network: after a picture passes through the classification network, the class information is obtained, and the class information is expanded into a one-dimensional vector and stored in the memory network; at the same time, the RGBHisBlock is used to extract the color information of the image, and the extracted color histogram vector is stored in the corresponding memory slot; the entire memory network M of the model is expressed as:

[0026] M = (K1, V1), (K2, V2),..., (Km, Vm) (3.6)

[0027] Where M represents the size of the memory, K represents the category feature information, and V represents the color vector;

[0028] In the test phase of the model, according to the category label of the input image, the corresponding color histogram information is retrieved from the storage network, and the most matching color histogram feature is selected to color the input image.

[0029] Furthermore, the model also includes reducing the dimensionality of the color histogram vector and converting the vector into two dimensions through 2D convolution to adapt to the projection module of the colorization backbone network; the vector is first converted into two low-dimensional vectors through a fully connected layer, and these two vectors are then processed through 2D convolutional layers respectively to obtain the feature map suitable for the colorization backbone network.

[0030] Furthermore, the loss function is defined as:

[0031]

[0032] The first three losses in Equation 3.7 are all generator losses, and the last one is the discriminator loss; among them, the first loss function is used to calculate the color difference between the generated image and the gt image; compared with the (a gt , b gt ) channels of the gt image, the specific design is as follows:

[0033]

[0034] Among them, (L, a gt , b gt ) is the representation of the gt image in the CIEL ab color space, Pr is the distribution of the color image, and ||·|| 2 is the Euclidean distance; by calculating the and the Euclidean distance between (a gt , b gt ), the color distribution of the generated prediction image and the gt image in the Lab space is made closer; the L2 loss ensures that the network generates more realistic colors.

[0035] Furthermore, the second loss function in Equation 3.7 is used to obtain the classification information of the obtained image. Through this loss function, the model identifies and utilizes the specific category to which the image belongs, and performs more detailed and accurate colorization according to the category to which the image belongs. The specific design of the loss is as follows:

[0036]

[0037] Among them, P rgis the distribution of the input grayscale image; y v is the classification vector obtained by training through the VGG network and then classifying the images in the dataset; KL(·||·) is used to calculate the loss caused by the fitting between y v and ; after the input grayscale image is processed by the classification network, it guides the coloring network to select more suitable colors for coloring.

[0038] Furthermore, the third loss in Equation 3.7 is specifically used to constrain the performance of the segmentation network to ensure that the network can accurately identify and segment each target object in the image; through this constraint, the segmentation network improves its recognition accuracy of the object edges in the image and effectively distinguishes different objects or categories in complex scenarios. The specific design of the loss is as follows:

[0039]

[0040] where S is the image segmented by the segmentation network, and (a gt , b gt ) is the image segmented by using SAM; the model uses the image segmented by SAM as the GT to constrain the segmentation network, and obtains more accurate semantic segmentation information by minimizing the Euclidean distance between the image generated by the segmentation network and the GT, so as to reduce color bleeding in the final coloring.

[0041] Furthermore, the last loss function in Equation 3.7 is the discriminator loss, and the WGAN loss is selected as the loss function of the discriminator; WGAN constrains the network by minimizing the Wasserstein distance between the predicted distribution output by the generator and the real distribution (GT); this loss mechanism is used to improve the stability and final generation quality in the GAN training process. The specific loss design is as follows:

[0042]

[0043] where represents the model distribution in ; represents a linear uniform sampling, which is the sampling located between the data distribution P r and ; the model also uses the Kantorovich-Rubinstein duality and adds gradient penalty to constrain the L2 norm; during the training process, the hyperparameters λ seg are set to 0.003, λ cls are set to 0.003, and λ WGAN are set to 0.1.

[0044] Advantages of the present invention:

[0045] Compared with the prior art, the automatic image colorization method based on semantic segmentation and generative adversarial network of the present invention has the following technical features or beneficial effects:

[0046] (1) Significantly improve the colorization quality: By combining a semantic segmentation network and a generative adversarial network, the present invention can more accurately capture the semantic information in the image, thereby effectively reducing the problems of unnatural color transitions and color bleeding. At the same time, using the color histogram vector as external guidance information makes the colorization effect more saturated and vivid, significantly improving the overall quality of image colorization.

[0047] (2) Enhance color saturation and realism: The storage network designed in the present invention can store and retrieve color histogram information and category information related to the input image, and these information are effectively utilized during the colorization process, making the generated color images significantly improved in terms of color saturation and realism.

[0048] (3) Improve the model generalization ability: By introducing a shape prior module and a classification network, the present invention can more effectively model the long-distance context dependence and obtain rich texture information related to the global region. This enhances the model's generalization ability for complex scenes and different types of images, making the colorization effect more stable and reliable.

[0049] (4) Optimize the training process and stability: The loss function design adopted in the present invention comprehensively considers multiple aspects such as the color difference of the generated image, the accuracy of classification information, the performance of the segmentation network, and the discriminator loss, ensuring the stability and effectiveness of the training process. At the same time, by choosing the WGAN loss as the loss function of the discriminator, the stability and the final generated quality in the GAN training process are further improved.

[0050] (5) Reduce the user interaction cost: Compared with traditional methods based on reference images and color scribbles, etc., the present invention realizes a more automated image colorization process without the need for users to provide additional color information as guidance, thereby reducing the user interaction cost and improving the practicality of the colorization technology.

[0051] In summary, through the innovative technical solution, the present invention solves the key problems in the existing image colorization technology, significantly improves the colorization quality, color saturation and realism, while enhancing the model's generalization ability and training stability, and reducing the user interaction cost, providing new ideas and solutions for the development and application of image colorization technology. Brief Description of the Drawings

[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in conjunction with the drawings and specific embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0053] Figure 1 is the model structure diagram of the present invention;

[0054] Figure 2 is the structure diagram of the SPM module of the present invention;

[0055] Figure 3 is the storage network design diagram of the present invention;

[0056] Figure 4 is the color vector-guided backbone network coloring diagram of the present invention;

[0057] Figure 5 is the projection module diagram of the present invention;

[0058] Figure 6 is the comparison result diagram of the present invention on the OxfordFlower102 dataset;

[0059] Figure 7 is the comparison result diagram of the present invention on the MiniImageNets dataset;

[0060] Figure 8 is the ablation experiment diagram of the storage network of the present invention;

[0061] Figure 9 is the ablation experiment diagram of the segmentation network of the present invention. Specific Embodiments

[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The following combines the attached Figures 1-9 to further illustrate the image automatic coloring method based on semantic segmentation and generative adversarial networks, and clearly and completely describe the technical solutions in the embodiments of the present application; obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0063] Embodiment 1

[0064] The present invention proposes an automatic image coloring algorithm based on a color storage and semantic parsing module, and the backbone model of this algorithm is based on ChromaGAN. First, the present invention designs a semantic segmentation network in the model, which aims to help the model obtain more semantic parsing information. For example, separating each part of the image to ensure that the coloring boundaries are clearer and more accurate, avoiding color spillover. Different from previous networks, the present invention adds a Shape Prior Module (SPM) to the segmentation network. This method can enable the network to effectively model long-distance context dependencies, thereby obtaining rich texture information related to the global region, so the actual effect is relatively obvious and can avoid the problem of color mixing. Secondly, the present invention also uses an external storage network to guide the main coloring network. The method of the present invention extracts the color histogram information and category information of the image and stores this information in the storage network. The stored color histogram information and semantic information are fused. This method can effectively improve the saturation of colors and make the colored pictures have more vivid colors.

[0065] Specifically, the task of the model is to learn a mapping F: G ∈ C H×W×1 to its corresponding color channel values (a, b). Here, G represents the grayscale image, and (a, b) represents the chromaticity channel values in the Lab color space predicted by the model. By fusing the predicted color channels with the grayscale channels, the final output color image C H×W×3 of the model can be obtained. The core objective of the model of the present invention is to find an effective mapping relationship F(G H×W |C H×W×3 ) to accurately predict the Yab values, as Figure 1 shown.

[0066] PLab = Pab + L

[0067] Generally speaking, after the model prediction, the ab channel is added to the L channel to obtain the final complete predicted color map PLab.

[0068] The model of the present invention consists of three major parts. The first part is the semantic segmentation part. In the semantic segmentation network, the method of the present invention creatively introduces the SPM shape prior, which can help the model obtain a more realistic semantic parsing. Finally, the upsampled parsing information is fused with the colorization main network. In this way, the colorization main network can obtain the fine-grained semantic parsing information of the image. The second part is the classification network of the yellow part and the subsequent storage network. During the training process, the present invention extracts the classification information and color histogram information of the image and stores them in the storage network. When colorizing, the present invention selects the corresponding color histogram according to the classification information of the input image, and then the color histogram vector outputs color features through projection and fuses them with the semantic features of the colorization backbone network, which can effectively guide the model in terms of color to improve the color saturation of the colorization result. The third part is the core network responsible for colorization. It uses VGG16 as the encoder and outputs the predicted Yab values through convolutional neural networks and continuous upsampling. The present invention initializes the VGG16 encoder with the weights pre-trained on the ImageNet dataset, and this strategy ensures that the colorization network can recognize the feature information of different categories from the beginning of training. This method significantly improves the sensitivity of the model to category features, thereby enhancing its classification performance.

[0069] Segmentation network

[0070] The present invention provides richer and more comprehensive semantic information for the model by introducing a semantic segmentation network. Although previous studies have considered the importance of semantic segmentation, they are still insufficient in obtaining the comprehensiveness of semantic information. In recent years, with the development of large models, SAM has achieved great success in the field of segmentation. Therefore, the present invention first uses SAM to segment the dataset of the present invention. Of course, in order to make the segmented images meet the network requirements of the present invention, the present invention makes some fine-tuning to the existing SAM model. The fine-tuned model can pay more attention to the semantic information of the image when segmenting the image, such as accurately distinguishing different elements such as the sky, water body, and vegetation in a picture. Then the segmented data is used as the label of this model to supervise the semantic segmentation network part.

[0071] In the present invention, the advanced large model SAM is adopted for image segmentation because this approach has significant advantages compared to traditional segmentation techniques. Firstly, through deep learning and big data training, the SAM model can more accurately understand and capture complex semantic information in images, thus achieving a more refined and accurate segmentation effect. Secondly, the SAM model uses the attention mechanism to focus on the key information in the image, effectively improving the accuracy and efficiency of segmentation. Especially when dealing with images with rich details and complex backgrounds, it demonstrates excellent performance. In addition, compared with traditional methods, the SAM model can better retain the detail and texture information of the image during the segmentation process, avoiding problems such as over-smoothing or detail loss, and thus being closer to the real scenario in terms of visual effect. Most importantly, using the segmented images by the SAM model as labels for network learning can significantly improve the efficiency and reliability of the segmentation network.

[0072] Meanwhile, a Shape Prior Module (SPM) is introduced in the segmentation network part. By using explicit shape priors, that is, taking shape prior information as auxiliary information, these information are used to guide the model to better segment the target object, thereby improving the segmentation performance of the model of the present invention. The SPM module is as Figure 2 shown. Specifically, the SPM module consists of two parts: a Self-Update Block and a Cross-Update Block. The Self-Update Block helps the network effectively model long-distance context dependencies, thereby obtaining rich texture information related to the global region. The Cross-Update Block injects inductive biases into the network to obtain more detailed local shape information.

[0073] Generally speaking, the segmentation network obtains the parsed features of the image through continuous upsampling under the supervision of the pre-segmented labels, and then fuses the upsampled parsed information with the coloring main network. The specific process is shown in Equations 3.2 and 3.3. In this way, the coloring main network can obtain the fine-grained semantic parsing information of the image. For example, separating the main body and the background to ensure clear and accurate coloring boundaries.

[0074] S = U(VGG(I)) (3.2)

[0075]

[0076] where I is the input image, VGG(I) represents the features extracted from the input image using the VGG network, U represents the upsampling process to convert the extracted features into a higher-resolution segmentation map S, represents the loss function, which is used to calculate the difference between S and T. The specific loss formula will be introduced in the loss function part.

[0077] Through this segmentation step, the network can learn more semantic information. Finally, these semantic information are concatenated with the colorization main network. This process enables the network to obtain semantic information of various parts of the input image, even at the pixel level, thus making the colorization effect more accurate and concentrated.

[0078] Classification network

[0079] The bottom classification network aims to obtain high-level features of the image and class label information for image colorization. The feature information extracted from VGG in the classification network is first processed by several convolutional blocks and then split into two different fully connected layers. The fully connected layer located on the upper side of the network is used to compress the obtained information into a 256-dimensional vector for stacking with the colorization backbone network, which helps the model learn different class information and enables the model to colorize for different classes.

[0080] The fully connected layer on the right side of the network is used to output class labels. Here, the present invention designs 102 output heads or 1000 output heads for different datasets because the dataset Oxford Flower102 used in the present invention contains 102 common flower classes while ImageNet contains 1000 classes. The design of the classification network enables the colorization backbone network to select more correct colors for different classes.

[0081] The processes of feature vector compression and class label output are shown in Equations 3.4 and 3.5:

[0082] Fcompressde = FC(ConvBlocks(VGG(I))) (3.4)

[0083] C102 = FCright(ConvBlocks(VGG(I))) (3.5)

[0084] Where VGG(I) represents the input image passing through the VGG network to extract overall information, ConvBlocks represents convolutional blocks, and FC represents fully connected layers. C102 represents having 102 classes, and if the ImageNet dataset is used, it is replaced by C1000 for representation. The class information finally extracted by the network will be stored in the memory network.

[0085] Memory network

[0086] The present invention proposes a memory network, which aims to accommodate two key pieces of information: one is the image class label obtained from the classification network; the other is the color histogram information of the image. The color histogram information is extracted by the RGBHisBlock module proposed according to the research of HistoGAN, and this module is used to calculate the RGB color histogram features of a specified image. The memory network is designed asFigure 3 As shown in Figure 3 , the design of RGBHisBlock is based on the research results in the field of color constancy. The design uses the logarithmic chromaticity space, which has better invariance to illumination changes, to construct a differentiable histogram of color distribution. This feature is a two-dimensional histogram obtained by projecting the colors of an image into the logarithmic chromaticity space. The 2D histogram is parameterized by uv and can convey more color information of the image in a more compact form than the 3D histogram defined in the traditional RGB space. The logarithmic chromaticity space is defined by dividing the intensity of one channel by the other two channels, thus providing three possible definition methods. Instead of choosing to use only one of these spaces, the present invention constructs three different histograms using these three possible definition methods respectively and combines them to form a histogram feature H, which is represented as a tensor of h×h×3.

[0087] The memory module is updated during the training process. After a picture passes through the classification network, the class information is obtained and expanded into a one-dimensional vector and stored in the memory network. At the same time, the RGBHisBlock is used to extract the color information of the image, and the extracted color histogram vector is stored in the corresponding memory slot. Therefore, the entire memory network M of the model can be expressed as:

[0088] M = (K1, V1), (K2, V2),..., (Km, Vm) (3.6)

[0089] Where M represents the size of the memory, K represents the class feature information, and V represents the color vector.

[0090] In the test stage of the model, according to the class label of the input image, the corresponding color histogram information is retrieved from the storage network, and the most matching color histogram feature (Top-1 Color Feature) is selected to color the input image. The advantage of this method is that by using this color information to guide the model to color, not only can the color saturation of the output image be improved, but also the color performance of the image can be made more vivid and distinct. As shown in Figure 4 , this is the color vector-guided backbone network coloring proposed by the present invention. Figure 4 As shown in Figure 4 , this is the color vector-guided backbone network coloring proposed by the present invention.

[0091] In addition, the projection module mainly includes a fully connected layer and a 2D convolutional layer. The design purpose of this module is to reduce the dimension of the color histogram vector and convert the vector into two dimensions through 2D convolution to adapt to the coloring backbone network. Specifically, the vector is first converted into two low-dimensional vectors through the fully connected layer, and then these two vectors are processed by the 2D convolutional layer respectively to obtain the feature map adapted to the coloring backbone network. The design details of the projection module are shown in Figure 5 . Figure 5 As shown in Figure 5 .

[0092] Loss function

[0093] According to the description of the network structure diagram, the total loss function of the present invention is defined as:

[0094]

[0095] Among the first three losses in Formula 3.7, they are all generator losses, and the last one is the discriminator loss. Among them, the first loss function is used to calculate the color difference between the generated image and the gt image. The generator network of the present invention generates an image of the (a, b) channels. Therefore, the present invention needs to compare with the (a gt , b gt ) channels of the gt image. The specific design is as follows:

[0096]

[0097] Among them, (L, a gt , b gt ) is the representation of the gt image in the CIEL ab color space, Pr is the distribution of the color image, and ||·|| 2 is the Euclidean distance. The present invention makes the color distribution of the generated predicted image and the gt image closer in the Lab space by calculating the Euclidean distance between and (a gt , b gt ). The L2 loss can ensure that the network of the present invention generates more realistic colors.

[0098] The design of the second loss function in Formula 3.7 is to obtain the classification information of the image. This loss is crucial for improving the ability of the model to recognize and process images. Through this loss function, the model can effectively recognize and utilize the specific category to which the image belongs, and perform more detailed and accurate coloring according to the category to which the image belongs. The specific design of the loss is as follows:

[0099]

[0100] Among them, P rg is the distribution of the input grayscale image. y v is the classification vector obtained by training through the VGG network and then classifying the images in the dataset. KL(·||·) is used to calculate the loss caused by the fitting between y v and . After the input grayscale image is processed by the classification network, it can guide the coloring network to select more suitable colors for coloring.

[0101] The third loss in Equation 3.7 is specifically used to constrain the performance of the segmentation network, ensuring that the network can accurately identify and segment each target object in the image. Through this constraint, the segmentation network can not only improve its recognition accuracy of object edges in the image, but also effectively distinguish different objects or categories in complex scenarios, significantly enhancing the accuracy and robustness of segmentation. The specific design of the loss is as follows:

[0102]

[0103] Among them, S is the image after being segmented by the segmentation network, and (a gt , b gt ) is the image segmented by SAM. The model of the present invention uses the image segmented by SAM as the GT to constrain the segmentation network, and obtains more accurate semantic segmentation information by minimizing the Euclidean distance between the image generated by the segmentation network and the GT, so as to reduce color bleeding in the final coloring.

[0104] The last loss function in Equation 3.7 is the discriminator loss. The model of the present invention selects the WGAN loss as the loss function of the discriminator. WGAN constrains the network by minimizing the Wasserstein distance between the predicted distribution output by the generator and the real distribution (GT). This loss mechanism significantly improves the stability and final generation quality in the GAN training process. Specifically, the WGAN loss can effectively alleviate the problems of vanishing gradients and mode collapse that may occur during training, making the training process smoother. In addition, it helps the generated image colors to be closer to the real ones, enhancing the naturalness and realism of the generated image. By adopting the WGAN loss, the model can achieve higher stability and quality when generating images, further enhancing the application value and practicality of the model. The specific loss design is as follows:

[0105]

[0106] Among them, represents the model distribution in . represents a linear uniform sampling, which is the sampling located between the data distribution P r and . The model also uses the Kantorovich-Rubinstein duality and adds gradient penalty to constrain the L2 norm. During the training process, the hyperparameters λ seg are set to 0.003, λ cls are set to 0.003, and λ WGAN are set to 0.1.

[0107] The present invention proposes an image coloring method combining semantic information, which integrates semantic segmentation and an external storage network to solve the problems of color overflow and unsaturated color effects during image coloring. By combining the semantic segmentation and the color information stored in the external storage network, the model of the present invention can more accurately restore the real colors, and is more delicate and natural in color transition and boundary processing. The model of the present invention also shows higher performance in color saturation, making the finally colored images look more vivid and lively.

[0108] Generally speaking, the model proposed in this section of the present invention not only significantly improves the visual attractiveness of the colored images, but also enhances the structural integrity and visual consistency of the images, which is of great significance for the development of automatic image coloring technology. This improvement is particularly applicable to fields that require high-quality coloring, such as black-and-white historical photos and manuscript art works, providing an effective technical means for these applications.

[0109] Embodiment 2

[0110] This embodiment provides a specific application and experimental comparison.

[0111] Dataset

[0112] The present invention uses the Oxford Flower102 dataset and the subset ILSVRC2012 dataset of ImageNet (hereinafter referred to as ImageNet for short) to train the model proposed by the present invention.

[0113] The Oxford Flower102 dataset was created by the Visual Geometry Group at the University of Oxford and is a public dataset widely used in image classification and image recognition research. Its main content is various flowers, with more than 8,000 photos covering 102 different flower species.

[0114] ILSVRC2012 is a subset of the ImageNet Large Scale Visual Recognition Challenge (usually also simply referred to as ImageNet). This dataset is one of the most famous benchmarks in the field of computer vision and is widely used to train and test algorithms such as deep learning, image recognition, and object detection. It contains approximately 1.2 million images distributed in 1,000 different categories.

[0115] It should be noted that these data sets contain some black and white pictures. Given that the objective of the present invention is to color pictures, these grayscale pictures are excluded in the present invention. However, it should be emphasized that the number of these grayscale pictures is extremely small and has almost no impact on the coloring effect of the present invention. In addition, to meet the requirements of the network of the present invention, all pictures are simply cropped to ensure that the size of the processed pictures is uniformly 224×224 pixels.

[0116] In addition, it should be pointed out that in order to obtain the true segmentation labels required for the segmentation network part, the present invention uses the high-performance model SAM to comprehensively segment the data sets used, and the segmented images are used as the labels for constraining the segmentation network.

[0117] Training details

[0118] The model described in the present invention is constructed based on the ChromaGAN architecture. Before starting all training processes, the SAM model is used to segment the above-mentioned two data sets on 1080Ti and 3090 graphics cards respectively, so as to be used as the labels for the segmentation network part in the model. When training, the weights of the VGG-16 model pre-trained on the ImageNet data set are used to initialize the weights of the network. Then, the model is trained on the Oxford Flower102 data set on a computer equipped with an NVIDIA GTX 1080Ti graphics card, the training period is set to 300 epochs, and the batch size is set to 8. In addition, the model is also trained on the ImageNet data set on a computer equipped with an NVIDIA GTX 3090 graphics card, the training period is set to 30 epochs, and the batch size is set to 16. The model is optimized using the Adam optimizer, the learning rate is set to 2×10-5, and the momentum parameters β1 and β2 are set to 0.5 and 0.999 respectively.

[0119] Evaluation metrics

[0120] To comprehensively evaluate the performance of the model proposed in the present invention, the present invention selects three main evaluation metrics: LPIPS, PSNR, and SSIM.

[0121] Among them, LPIPS (Learned Perceptual Image Patch Similarity) is a method for measuring perceptual similarity proposed by Zhang et al., which is particularly suitable for challenging visual prediction tasks and other model evaluation scenarios. The core of LPIPS lies in evaluating the perceptual differences between images. The lower its value, the smaller the perceptual difference between the generated image and the original image, that is, the higher the similarity.

[0122] PSNR (Peak Signal-to-Noise Ratio) is the peak signal-to-noise ratio, which is a commonly used metric for measuring image quality and is widely applied especially in the field of image compression. The higher the value of PSNR, the closer the quality of the generated image is to the original image, that is, the smaller the image distortion.

[0123] SSIM (Structural Similarity Index Measure) is the structural similarity index measure, which is used to compare the similarity of two images. SSIM not only considers the pixel values of the images, but also factors such as the structure, texture, and brightness of the images. Therefore, it is a more comprehensive metric for measuring image similarity. Considering these three evaluation metrics comprehensively, the present invention can comprehensively evaluate the performance of the proposed model in the image colorization task from multiple dimensions, including aspects such as perceptual quality, signal-to-noise ratio, and structural similarity.

[0124] Comparative experiments

[0125] The model proposed in the present invention was trained on the Oxford Flower102, MiniImageNet, and ImageNet datasets respectively, and was compared with the baseline models (HistoryNet and HistoryNet), as well as nine currently mainstream methods (CIC, Zhang etal., UGColor, DeOldify (2019), InstColor, ColTran, CT2, Disentangled, DDColor). It should be noted that DeOldify (2019) is an old photo colorization project developed by ANTIC.J. et al. and publicly available on GitHub.

[0126] To ensure the fairness of the evaluation process, the present invention adopts the following strategy: for those codes that provide pre-trained models, these pre-trained models will be directly used for testing. For codes that do not provide pre-trained models, re-training will be carried out, and the number of training rounds and other related parameters will be consistent with the model settings in the present invention section.

[0127] (1) Quantitative comparison

[0128] As shown in the experimental results in Table 3.1, the colorization method proposed in the present invention was comprehensively evaluated on the Oxford Flower and MiniImageNet datasets. In the three key evaluation metrics - peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual image patch similarity (LPIPS), the method of the present invention significantly outperforms the existing baseline models.

[0129] In terms of the two metrics of PSNR and LPIPS, the performance of the method of the present invention is particularly outstanding, which further proves that there is higher consistency and similarity between the images processed by the colorization technology proposed by the present invention and the real images (Ground Truth, GT).

[0130] It is worth noting that the performance of all methods on the Mini-ImageNet dataset is generally better than that on the OxfordFlower dataset. The reason behind this phenomenon lies in the difference in the number of pictures in the two datasets. The Mini-ImageNet dataset has far more pictures than the Oxford Flower dataset, which provides the model with richer and more diverse learning materials, helping the model capture a wider range of color distributions and image features, thus improving the accuracy and naturalness of colorization.

[0131] Table 3.1 Quantitative comparison results of the present method and the basic model

[0132] Table 3.1 Results of quantitative comparison of the presentmethodology with the underlying model

[0133]

[0134] In the experimental results shown in Table 3.2, the model proposed by the present invention was comprehensively evaluated based on the ImageNet dataset, and was quantitatively compared with historical methods and the current state-of-the-art model DDColor.

[0135] The experimental results show that on the three evaluation metrics, the scores of all the models participating in the evaluation are better than those of the other two datasets. This phenomenon can be attributed to the significant advantages of the ImageNet dataset in terms of the number and variety of images, enabling the network to learn richer information.

[0136] In particular, in terms of the structural similarity index (SSIM), the method proposed by the present invention outperforms all other models,

[0137] including the current state-of-the-art model. This achievement benefits from the color histogram information stored in the network, which greatly improves the similarity between the images generated by the network and the ground truth (Ground Truth, GT) images. In terms of the peak signal-to-noise ratio (PSNR) evaluation metric, the model of the present invention is second only to the current state-of-the-art model.

[0138] In general, the proposed model not only significantly outperforms previous methods in multiple evaluation indicators, but also shows competitive performance compared with the latest advanced models. These excellent evaluation indicators strongly verify the effectiveness of the proposed method in achieving semantically consistent image colorization. In addition, the proposed method also shows its obvious advantages in generating images with high saturation and less color spillover.

[0139] Table 3.2 Quantitative comparison of the ImageNet dataset with the current SOTA and previous colorization algorithms

[0140] Table 3.2 Quantitative comparison with current SOTA and previous coloring algorithms in ImageNet dataset

[0141]

[0142] (2) Qualitative analysis

[0143] In this section, a qualitative comparison is made between the base model and the proposed model on the Oxford Flower102 and MiniImageNet datasets.

[0144] Figure 6 The colorization results of the model of the present invention and the basic model on the Oxford Flower102 dataset are shown. It can be seen that the color saturation of the picture output by the model of the present invention is better than that of other methods, which is due to the information of the color histogram in the storage network.

[0145] In the coloring task, the low saturation of the coloring effect is a common problem that bothers people. In order to solve the problem of low saturation, ChromaGAN learns coloring by combining the perception and semantic understanding of color and category distribution. Specifically, three loss functions that combine color, perceptual information and semantic category distribution are proposed. This method aims to use the loss function to constrain the network to generate more saturated pictures. However, since the design of its loss function mainly focuses on the accuracy of color and ignores saturation, the images it generates have low saturation, which can be seen in the second and fourth rows. The author of HistoryNet proposed a coloring network combined with semantic parsing and proposed a new historical figure dataset MHMD. The dataset proposed by this method contains a large number of low-saturation images, which makes the model more inclined to generate low-saturation images. At the same time, the author uses Deeplab-V3 in the semantic segmentation part. This method can only separate the characters and background in the picture, but cannot show the specific details of the characters, resulting in overflow in the coloring effect.

[0146] In view of this problem, the present invention adds a storage network on the basis of ChromaGAN, and uses the color histogram information in the storage network to guide the network, enabling the network to learn more color information, thereby enhancing the saturation of the generated images. At the same time, to avoid color overflow, the present invention draws on the semantic parsing method in HistoryNet. And an SPM module is introduced into the parsing network because using shape prior information can help the model better focus on the edge information of each category in the picture. Secondly, this method also uses the picture segmented by SAM as a constraint for the parsing network. This method aims to enable the model to generate a semantic map closer to the real parsing, so that the model can obtain more semantic information, and finally enable the model to generate better-quality color images.

[0147] As Figure 7 shown, the present invention compares and analyzes the model proposed in the present invention with the infrastructure model on the MiniImageNet dataset. The results show that in terms of color saturation and image quality, the model of the present invention is significantly better than the basic model. Especially in terms of controlling color overflow, the model of the present invention rarely exhibits color overflow. This benefits from the semantic parsing network proposed in the present invention. It can be seen that in the underwater photo in the third row and the sailboat photo in the fourth row, there are significant color overflow problems in the coloring effects of the three methods of ChromaGAN, HistoryNet, and Bright-ColoredNet. Through careful visual analysis and quantitative evaluation, it is finally confirmed that in terms of both color saturation and overall image quality, the model of the present invention significantly outperforms the basic model and previous methods.

[0148] In summary, whether on the Oxford Flower102 or MiniImageNet dataset, the coloring effect of the method of the present invention is better than that of other models, and it has good generalization ability.

[0149] Ablation experiment

[0150] (1) Storage network

[0151] The present invention conducts a detailed ablation experiment on each part of the model of the present invention on the Oxford Flower102 dataset to evaluate the contribution and effectiveness of each module. The present invention first conducts an ablation analysis on the storage network proposed in the model. The experimental results show that the storage network helps the model improve its ability to obtain color information. Through the color histogram information in the storage network, the model can generate images with higher saturation and more vivid colors. As Figure 8 shown, the present invention demonstrates the significant impact of the storage network on the final coloring effect. The figure clearly shows how the addition of the storage network makes the pictures output by the model more vivid and attractive in terms of color performance.

[0152] (2) Segmentation Network

[0153] The present invention conducts ablation experiments on the proposed segmentation network. Observing Figure 9 the area marked with a red circle, it can be seen that there are obvious differences between the results with and without the segmentation network added. The existence of this difference proves that the segmentation network plays an important role in obtaining more shape and detail information. The application of the segmentation network significantly enhances the network's ability to capture image details, enabling the model to accurately color different regions of the image. This not only improves the coloring accuracy but also effectively avoids the problem of color bleeding, thus ensuring natural colors and clear boundaries in the image.

[0154] As described above, only the specific preferred embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. An automatic image colorization method based on semantic segmentation and generative adversarial network, characterized in that: Here are the steps: S1, using the model based on ChromaGAN; S2. A semantic segmentation network is designed in the model to help the model obtain more semantic parsing information; S3, adding a shape prior module to the segmentation network for enabling the network to effectively model long-range contextual dependencies, thereby obtaining rich texture information related to the global area; S4, using the external storage network to guide the main coloring network, extracting the color histogram information and category information of the image, and storing this information in the storage network; S5. Utilize the stored color histogram information and semantic information for fusion to enhance the color saturation and complete the coloring.

2. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 1, characterized in that: In the segmentation network, the following steps are performed: T1. Use SAM to segment the data set of this article; T2. Introducing a shape prior module in the segmentation network part, by using explicit shape prior information as auxiliary information, using the information to guide the model to segment the target object, thereby improving the segmentation performance of the model; T3. The segmentation network obtains the analytical features of the image through continuous upsampling under the supervision of the pre-segmented labels, and then fuses the upsampled analytical information with the main coloring network. The specific process is shown in Formula 3.2 and Formula 3.

3. S=U(VGG(I)) (3.2) Where I is the input image, VGG(I) represents the features extracted from the input image using the VGG network, and U represents the upsampling process to convert the extracted features into a higher resolution segmentation map S. Represents the loss function, which is used to calculate the difference between S and T.

3. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 2, characterized in that: The shape prior module includes a self-update block for helping the network effectively model long-range contextual dependencies to obtain rich texture information related to the global area and a cross-update block for injecting inductive bias into the network to obtain more detailed local shape information.

4. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 1, characterized in that: The model also includes a classification network for obtaining high-level features of the image and class label information for image coloring; in the classification network, the feature information extracted from VGG is first processed by several convolution blocks and then shunted to two different fully connected layers, wherein the fully connected layer located on the upper side of the network is used to compress the obtained information into a 256-dimensional vector for stacking with the coloring backbone network; The fully connected layer on the right side of the network is used to output the category label. The feature vector compression and category label output process are shown in Formula 3.4 and Formula 3.5: Fcompressde=FC(ConvBlocks(VGG(I))) (3.4) C102=FCright(ConvBlocks(VGG(I))) (3.5) Among them, VGG(I) means that the input image is extracted from the VGG network, ConvBlocks means convolution block, FC means fully connected layer; C102 means there are 102 categories. If the ImageNet dataset is used, it is replaced by C1000. Finally, the category information extracted by the network is stored in the storage network.

5. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 1, characterized in that: In the storage network: after a picture passes through the classification network, the category information is obtained, and the category information is expanded into a one-dimensional vector and stored in the memory network; at the same time, the color information of the image is extracted using RGBHisBlock, and the extracted color histogram vector is stored in the corresponding memory slot; the entire memory network M of the model is expressed as: M=(K1,V1),(K2,V2),...,(Km,Vm) (3.6) Where M represents the size of the memory, K represents the category feature information, and V represents the color vector; During the testing phase of the model, the corresponding color histogram information is retrieved from the storage network according to the category label of the input image, and the most matching color histogram feature is selected to colorize the input image.

6. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 1, characterized in that: The model also includes reducing the dimension of the color histogram vector and converting the vector into two dimensions through 2D convolution to adapt to the projection module of the coloring backbone network; the vector is first converted into two low-dimensional vectors through a fully connected layer, and the two vectors are then processed through a 2D convolution layer respectively to obtain a feature map adapted to the coloring backbone network.

7. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 2, characterized in that: The loss function is defined as: The first three losses in formula 3.7 are all generator losses, and the last one is a discriminator loss. Among them, the first loss function is used to calculate the color difference between the generated image and the gt image; and the (a gt ,b gt ) channel for comparison, the specific design is as follows: Among them, (L,a gt ,b gt ) is the gt image in CIEL ab Representation in color space, Pr is the distribution of color images, ||·||2 is the Euclidean distance; by calculating and (a gt ,b gt ), so that the color distribution of the generated predicted image and the gt image in Lab space is closer; L2 loss ensures that the network generates more realistic colors.

8. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 7, characterized in that: The second loss function in formula 3.7 is used to obtain the classification information of the image. Through this loss function model, the specific category of the image is identified and used, and more detailed and accurate coloring is performed according to the category of the image. The specific design of the loss is as follows: Among them, P rg is the distribution of the input grayscale image; y v is the classification vector obtained by training the VGG network and classifying the images in the dataset; KL(·||·) is used to calculate y v and The loss caused by fitting between them; after the input grayscale image is processed by the classification network, it guides the coloring network to select a more suitable color for coloring.

9. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 7, characterized in that: The third loss in formula 3.7 is specifically used to constrain the performance of the segmentation network to ensure that the network can accurately identify and segment each target object in the image; through this constraint, the segmentation network improves its recognition accuracy of the edges of objects in the image and effectively distinguishes different objects or categories in complex scenes. The specific design of the loss is as follows: Among them, S is the image segmented by the segmentation network, (a gt ,b gt ) uses the image segmented by SAM; the model uses the image segmented by SAM as GT to constrain the segmentation network, and obtains more accurate semantic segmentation information by minimizing the Euclidean distance between the image generated by the segmentation network and GT, so that the final coloring reduces color overflow.

10. The method for automatic image colorization based on semantic segmentation and generative adversarial network according to claim 7, characterized in that: The last loss function in formula 3.7 is the discriminator loss, and WGAN loss is selected as the discriminator loss function; WGAN constrains the network by minimizing the Wasserstein distance between the predicted distribution of the generator output and the true distribution (GT); this loss mechanism is used to improve the stability and final generation quality during GAN training; the specific loss design is as follows: in, represent exist Model distribution in ; represents a straight line uniform sampling, which is located along the data distribution P r and The model also uses Kantorovich-Rubinstein duality and adds gradient penalty to constrain the L2 norm; during training, the hyperparameter λ seg Set to 0.003, λ cls Set to 0.003, λ WGAN Set to 0.1.