An image coloring method based on multi-modal content coding

By using multimodal content coding and conditional generative adversarial networks, the problem of neglecting cultural influences in intelligent color design is solved, and the cultural relevance of automatic color palette generation and image coloring is realized. The generated images have the style of Chinese youth subculture.

CN113888660BActive Publication Date: 2025-11-21TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111141755.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-11-21
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

Existing research on intelligent color design has largely ignored cultural influences, resulting in a lack of cultural relevance and consistency in color generation.

Method used

A multimodal content coding method is adopted, which combines image, text and category information. A color palette is generated through a conditional generative adversarial network, and the colors in the palette are generated step by step using a recursive generation approach. The image is then colored using a U-Net network to achieve the fusion of cultural information and automatic image coloring.

Benefits of technology

It achieves automatic generation of color palettes, fully considers cultural influences, improves the relevance and cultural communication effect of image coloring, and generates images with the style of Chinese youth subculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888660B_ABST
    Figure CN113888660B_ABST
Patent Text Reader

Abstract

The application relates to an image coloring method based on multi-modal content coding, which comprises the following steps: acquiring multi-modal data, including an image, text, a category and a color palette, wherein the text is a descriptive sentence of the image, the category is a style attribute of the image, and the color palette comprises multiple colors to be generated; respectively coding the multi-modal data and fusing the multi-modal data to obtain color palette generation fusion coding; generating the colors in the color palette based on the color palette generation fusion coding by using a color palette generation network; re-coding the multi-modal data based on the generated color palette and fusing the multi-modal data to obtain coloring fusion coding; and coloring the image based on the coloring fusion coding by using an image coloring network. Compared with the prior art, the intelligent color design process fully considers the influence of culture, and realizes automatic coloring of the image with a specific cultural style.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a picture coloring method and system, in particular to an image coloring method based on multi-modal content coding. BACKGROUND

[0002] Color is one of the basic components of graphic design, which is not only a visual element, but also carries the conveying effect of internal semantics and can arouse people's deep resonance beyond vision. The generation of these internal semantics and deep resonance is largely due to the cultural influence behind the color. At the same time, color has always been a very active research topic in the field of machine learning and artificial intelligence, and in recent years, many intelligent color design researches and tools have appeared. However, most of the previous intelligent color design researches ignore the cultural influence. SUMMARY

[0003] The application aims to overcome the defects of the prior art and provides an image coloring method based on multi-modal content coding.

[0004] The purpose of the application can be achieved by the following technical solutions:

[0005] An image coloring method based on multi-modal content coding, the method comprising:

[0006] Obtaining multi-modal data, including images, texts, categories and color plates, the texts being descriptive sentences of the images, the categories being style attributes of the images, and the color plates including multiple colors to be generated;

[0007] Encoding and fusing the multi-modal data to obtain color plate generation fusion coding, and generating the colors in the color plate based on the color plate generation fusion coding using a color plate generation network;

[0008] Re-encoding and fusing the multi-modal data based on the generated color plate to obtain coloring fusion coding, and coloring the images based on the coloring fusion coding using an image coloring network.

[0009] Preferably, the images are encoded using an image encoder, and the image encoder is a VGG16 model-based encoder.

[0010] Preferably, the texts are encoded using a text encoder, and the text encoder is a BERT model-based encoder.

[0011] Preferably, the categories are encoded using a category encoder, and the category encoder is a one-hot encoder.

[0012] Preferably, the color plate is encoded by combining the RGB values of each color in the color plate into a vector.

[0013] Preferably, each color in the color palette is generated one by one in a recursive form when generating the colors in the color palette, specifically: one color is generated at a time, the color palette is updated and the corresponding color palette generation fusion code is generated, the next color is generated based on the new fusion code, and the generation of all colors in the color palette is completed.

[0014] Preferably, after the multi-modal data is encoded respectively, a multi-layer perception is used for fusion to obtain the corresponding fusion code during the color palette generation and image coloring process.

[0015] Preferably, the color palette generation network adopts a conditional adversarial generation network, the color palette generation network includes a color palette generator and a color palette discriminator, and when generating the color palette, the color palette generation fusion code and noise are input into the color palette generator, the color code corresponding to the generated color is obtained through a full connection layer, and the color code and the color palette generation fusion code are input into the color palette discriminator to judge whether the output meets the expectation through a full connection layer.

[0016] Preferably, the image coloring network adopts a conditional adversarial generation network, the image coloring network includes a coloring generator and a coloring discriminator, and when coloring the image, the image to be colored is input into the coloring generator, the coloring generator generates a colored image, and the colored image is input into the coloring discriminator to judge the true or false.

[0017] Preferably, the coloring generator and the coloring discriminator adopt a U-Net network structure, when coloring the image, the coloring generator is subjected to multiple times of down-sampling, the intermediate result is combined with the coloring fusion code for up-sampling, the number of times of up-sampling is the same as that of down-sampling, and each time of up-sampling adds the down-sampling result at the corresponding position, and the up-sampling process finally obtains the colored image as the output result of the coloring generator, and the colored image is input into the coloring discriminator, and the coloring discriminator is subjected to multiple times of down-sampling to judge the true or false of the coloring result of the colored image.

[0018] Compared with the prior art, the present application has the following advantages:

[0019] (1) The present application introduces multi-modal content into the process of color generation and image coloring, first realizes automatic generation of a color palette based on an image and cultural information (including text, categories) of the image, and then performs automatic coloring based on the generated color palette, so that the intelligent color design process fully considers the influence of culture.

[0020] (2) The present application uses a recursive generation idea to improve the correlation of colors in the color palette, thereby improving the coloring effect of subsequent images. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 It is a whole idea framework diagram of the image coloring method based on multi-modal content coding.

[0022] Figure 2 Structure diagram of the color plate generation network of the present application;

[0023] Figure 3 Structure diagram of the image coloring network of the present application;

[0024] Figure 4 Flow chart of the color plate generation of the present application. DETAILED DESCRIPTION

[0025] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. Note that the following description of the embodiments is merely illustrative in nature, and the present application is not intended to limit the scope of application or its use, and the present application is not limited to the following embodiments.

[0026] EMBODIMENT

[0027] As shown in Figure 1 The present embodiment provides an image coloring method based on multi-modal content encoding, which comprises:

[0028] Obtaining multi-modal data, including images, texts, categories, and color plates, the texts being descriptive sentences for the images, the categories being style attributes of the images, and the color plates including multiple colors to be generated;

[0029] Encoding and fusing the multi-modal data respectively to obtain color plate generation fusion encoding, and generating the colors in the color plate based on the color plate generation fusion encoding using a color plate generation network;

[0030] Re-encoding and fusing the multi-modal data based on the generated color plate to obtain coloring fusion encoding, and coloring the images based on the coloring fusion encoding using an image coloring network.

[0031] The present embodiment takes Chinese youth subculture with unique color design features as an example, and uses the method of the present application to color the images, which can generate color plates with Chinese youth subculture style, and can convert any picture into a colored picture with Chinese youth subculture style.

[0032] Mainly contains three parts:

[0033] 1) Multi-modal content encoding: encoding the input data of multiple modalities (images, texts, categories, and color plates) into a vector form that can be recognized by the model.

[0034] 2) Color plate generation: a basic model is built based on conditional generative adversarial network, and in the generation process, a recursive idea is adopted to generate one color at a time, thereby obtaining the entire color plate.

[0035] 3) Gray image automatic coloring: also based on conditional generative adversarial network, according to the generated color palette, the input gray image is converted into a colored result with Chinese youth subculture style.

[0036] The color palette generation network and the image coloring network both adopt conditional GAN, but they are different in structure, and the specific differences are shown in Figure 2 and Figure 3 .

[0037] Specifically, the color palette generation network includes a color palette generator and a color palette discriminator. When generating a color palette, the color palette generation fusion code and noise are input into the color palette generator, and the generated color code is obtained through a full connection layer. The color code will be judged together with the color palette generation fusion code through a full connection layer to determine whether it meets the expected result.

[0038] The image coloring network includes a coloring generator and a coloring discriminator. The coloring generator and the coloring discriminator adopt the U-Net network structure. When coloring the image, the image to be colored is input into the coloring generator, and after multiple down-sampling, the intermediate result is obtained. The coloring fusion code is combined for up-sampling. The up-sampling times are the same as the down-sampling times, and each up-sampling adds the down-sampling result at the corresponding position. The up-sampling process finally obtains the colored image as the output result of the coloring generator. The colored image is input into the coloring discriminator, and the coloring discriminator judges the coloring result of the colored image through multiple down-sampling.

[0039] The following describes the specific implementation process of the method of the application in three steps of multi-modal encoding, generating a color palette, and image coloring.

[0040] 1. Multi-modal encoding

[0041] The multi-modal data in the training data set includes four types of data: image, text, category, and color palette. The image data is stored in the form of URL. The text data is a descriptive sentence for the image, which is extracted from the original web page information and stored in the form of Chinese characters. The category data includes 14 types such as punk, alternative, hip hop, rock, metal, core, techno, noise, indie music, indie design, indie animation, indie illustration, experimental art, and cyberpunk, which are summarized from the corresponding content classification in the original image web page. The color palette is a 5-color color palette group, which is obtained by image calculation and stored in the form of RGB color coding. In the use process, the multi-modal data input by the user is in the same form as the multi-modal data in the data set, which also includes four types of image, text, category, and color palette.

[0042] First, the data of different modalities is input into the corresponding encoder to obtain the encoding of each modality data, that is, the input of one modality corresponds to a fixed-length vector. Among them, the text encoder is based on the pre-trained BERT model, the image encoder is based on the pre-trained VGG16 model, the category encoder adopts simple one hot encoding, and the color palette encoder is actually a combination of the RGB values of 5 colors into a 15-dimensional vector. Then these fixed-length vectors are connected into a larger vector and input into a multi-layer perceptron (MLP) to obtain the final fusion encoding.

[0043] This part adopts the combination of pre-trained model and multi-layer perceptron (MLP) to complete the feature extraction and fusion of multi-modal content.

[0044] The multi-modal content includes text, image, category, and color palette, among which the first three are user input content, and the color palette is a model generated result. For text, a pre-trained BERT model is used to encode into a fixed-length vector; for image, a pre-trained VGG16 model is used to encode into a fixed-length vector; for category, one hot encoding is used to encode into a fixed-length vector; for color palette, 5 RGB values are directly combined into a 15-dimensional vector.

[0045] After obtaining the encoding result, the four fixed-length vectors are connected into one vector, and the combined vector is input into a multi-layer perception (MLP) to obtain a fusion encoding. In the color palette generation and image coloring stage, different fusion encodings are obtained, which are a color palette generation fusion encoding and a coloring fusion encoding.

[0046] 2. Generating a color palette

[0047] A color palette is a color combination composed of five colors. In order to automatically generate such a color combination, the present application makes some innovations based on the idea of a generative adversarial network.

[0048] The fusion encoding of the input content is used as a condition, and a conditional GAN (conditional generative adversarial network) is used to obtain the color palette required by us. Here, in order to make the generated color palette have better relevance, we do not generate five colors (corresponding to five groups of RGB values) at one time; instead, we use a recursive method to generate one color (corresponding to one group of RGB values) at a time, a total of five times.

[0049] In the process of generating each color, the fusion encoding of the multi-modal content (color palette generation fusion encoding) needs to be input into the conditional GAN as a condition to generate a new color palette encoding; then the new color palette encoding is used to update the fusion feature, and the new fusion feature is put into the conditional GAN for color generation. Since the result of the last time is used in each generation, this process is called recursive color palette generation.

[0050] A specific implementation is shown in Figure 4 During the training process, first set the initial color palette encoding to 0 (the values in the 15-dimensional vector are all 0), input it into the multi-layer perception as the initial color data, and then input it into the generator together with other features for fusion to generate the first color, and let the discriminator judge whether the color is true or false. Then, the RGB values of the first color in the training data set are added to the remaining 12 bits of 0 to form a new color palette encoding, which is input into the network to generate the second color and judge true or false. Then, the RGB values of the first and second colors in the training data set are added to the remaining 9 bits of 0 to form a new color palette encoding, which is used to generate the third color and judge. In this way, the fourth color is generated, and the input color palette condition is the encoding formed by adding the RGB values of the first three colors to the last three bits of 0. After the five colors are generated in turn, the final five-color color palette result is obtained. During the generation process, the same process is repeated five times to generate a single color, and the previously generated color is used as the input for the next color generation process, and 0 is used to fill the bits to form a 15-dimensional encoding.

[0051] To summarize the above, this part is based on conditional GAN, and adopts the idea of recursion to generate colors in the color palette.

[0052] Wherein, the condition of conditional GAN is the fusion encoding of the last step, and the model can generate a color (i.e. a set of R, G, B values) according to the fusion encoding.

[0053] The whole recursive generation process is as follows:

[0054] 1) Initialize all 5 colors in the color palette to 0;

[0055] 2) Generate fusion encoding according to the current color palette and user input;

[0056] 3) Take the fusion encoding as a condition, input GAN, and generate a color;

[0057] 4) Update the color palette according to the color;

[0058] 5) Repeat steps 2-4 until the generation of 5 color palettes is completed.

[0059] 3. Image coloring

[0060] Image coloring, i.e. converting a grayscale image into a color image, is a typical end-to-end generation process. The present application combines conditional GAN and U-Net to complete this task; wherein conditional GAN is the framework of the whole algorithm, and U-Net is the specific structure of the coloring generator and the coloring discriminator.

[0061] Specifically, we first use the multi-modal content encoding of the first step to encode the color palette and other attributes into a fixed-length fusion vector as the condition of conditional GAN, and then input the grayscale image to be colored as an image; based on the reconstruction loss of the image, the model is trained, and finally the model can generate a color image that meets the color palette.

[0062] Based on the method of the present application, an image coloring system is developed, including a user end and a model end, and the specific business process is as follows:

[0063] User end: user inputs a text -> uploads a picture to be colored -> selects a category -> gets color palette results -> modifies the color palette results (optional) -> gets coloring results

[0064] Model end: user inputs a string, selects a category, and uploads an image -> multi-modal content encoding model generates fusion encoding -> color palette generation model recursively generates color palette according to fusion encoding -> user adjusts color palette results -> coloring model colors the picture according to the color palette input and fusion encoding -> gets the final result.

[0065] The above-described embodiments are merely illustrative and are not intended to limit the scope of the present application. These embodiments can be carried out in other various modes, and various omissions, substitutions, and changes can be made thereto without departing from the scope of the technical thought of the present application.

Claims

1. A method for image coloring based on multi-modal content encoding, characterized in that, The method comprises: Obtaining multi-modal data, including images, texts, categories, and color plates, wherein the texts are descriptive sentences of the images, the categories are style attributes of the images, and the color plates include multiple colors to be generated; Encoding and fusing the multi-modal data to obtain color plate generation fusion encoding, and generating the colors in the color plate based on the color plate generation fusion encoding using a color plate generation network; Re-encoding and fusing the multi-modal data based on the generated color plate to obtain coloring fusion encoding, and coloring the images based on the coloring fusion encoding using an image coloring network; When generating the colors in the color plate, each color in the color plate is generated one by one in a recursive form, specifically: one color is generated at a time, the color plate and the corresponding color plate generation fusion encoding are updated, the next color is generated based on the new fusion encoding, and the generation of all colors in the color plate is completed. During the training process, an initial color plate encoding is set as 0, input into a multi-layer perception as initial color data, fused with other features, and then input into a generator to generate the first color, and a discriminator is used to determine whether the color is real or not. Then, the RGB values of the first color in the training data set are added to the remaining 12 zeros to form a new color plate encoding, which is input into the network to generate the second color and determine whether it is real or not. The RGB values of the first and second colors in the training data set are added to the remaining 9 zeros to form a new color plate encoding, which is used to generate the third color and determine whether it is real or not. The process is repeated to generate the fifth color, and the color plate condition input into the network is the encoding formed by the RGB values of the first four colors and the last three zeros. After the five colors are generated in turn, the final five-color color plate result is obtained. During the generation process, the process of generating a single color needs to be repeated five times for each complete five-color color plate, and the previously generated colors are used as the input of the next color generation process, and zeros are used to fill the positions to form a 15-dimensional encoding.

2. The method of claim 1, wherein, The images are encoded using an image encoder, and the image encoder is a VGG16 model-based encoder.

3. The method of claim 1, wherein, The texts are encoded using a text encoder, and the text encoder is a BERT model-based encoder.

4. The method of claim 1, wherein, The categories are encoded using a category encoder, and the category encoder is a one-hot encoder.

5. The method of claim 1, wherein, The encoding method of the color plate is to combine the RGB values of each color in the color plate into a vector.

6. The method of claim 1, wherein, During the color plate generation and image coloring process, the multi-modal data are encoded and fused using a multi-layer perception to obtain corresponding fusion encoding.

7. The method of claim 1, wherein, The color plate generation network adopts a conditional adversarial generation network, and the color plate generation network includes a color plate generator and a color plate discriminator. When generating the color plate, the color plate generation fusion encoding and noise are input into the color plate generator, the color encoding corresponding to the generated color is obtained through a full-connection layer, and the color encoding and the color plate generation fusion encoding are input into the color plate discriminator to determine whether the output meets the expectation through a full-connection layer.

8. The method of claim 1, wherein, The image coloring network adopts a conditional adversarial generative network, and comprises a coloring generator and a coloring discriminator.

9. The method of claim 8, wherein, The coloring generator and the coloring discriminator adopt a U-Net network structure, and when coloring the image, the coloring generator is subjected to multiple times of down-sampling, obtains an intermediate result, and is subjected to up-sampling combined with coloring fusion coding, the number of times of up-sampling is the same as that of down-sampling, and each time of up-sampling adds the down-sampling result at the corresponding position, and the up-sampling process finally obtains the coloring image as the output result of the coloring generator, the coloring image is input into the coloring discriminator, and the coloring discriminator judges the coloring result of the coloring image for multiple times of down-sampling.