Point interaction grey-scale map coloring method and system based on ViT coding and pixel shuffling
By using Vision Transformers and pixel shuffling technology in interactive grayscale image shading, the problem that user prompts are difficult to propagate to larger semantic areas of the image is solved, and efficient and accurate image shading effect is achieved.
Patent Information
- Application Number
- CN202510145375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
When the prior art realizes interactive grayscale image coloring, it is difficult to effectively propagate the larger semantic area of the image prompted by the user, resulting in poor shading effect, especially in areas with continuous grayscale values.
Using a Vision Transformers-based encoder, the self-attention mechanism enables the model to selectively propagate user prompts to the relevant areas of each single layer, and introduces pixel shuffling and local stabilization layers to achieve real-time shading and artifact reduction of images.
It achieves more accurate and efficient image shading, and can generate reasonable color images with less user interaction, improving color quality and speed.
Smart Images

Figure CN120070603A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, relates to interactive image coloring, and particularly relates to a point interaction grayscale image coloring method and system based on ViT encoding and pixel shuffling. Background Art
[0002] Unconditional image coloring technology has achieved remarkable achievements in fully automatically restoring the colors of grayscale photos. The interactive coloring method further expands this task, specifically allowing users to generate color images according to specific color conditions. In addition, recoloring existing images to achieve a new color theme has also become an effective means of photo editing. Among the various interaction types provided by users, such as reference images or color palettes, the interaction method based on point selection or scribbling is designed as a step-by-step coloring process. When the user specifies a color at a specific location in the image, the system will automatically fill the color of the image step by step according to these inputs.
[0003] The point interactive coloring method can help users generate color images with a minimum amount of interaction. Therefore, accurately identifying the regions associated with user cues is crucial for reducing the number of interactions. Early methods used manually designed filters to determine the regions that need to be colored according to user cues, but were only applicable to simple patterns in the image. Although deep learning-based models have been proposed, bringing new breakthroughs to the image coloring task, even in regions where the grayscale values are significantly continuous, existing methods often can only achieve partial coloring effects. This is mainly because the design efficiency of the convolutional layer is not sufficient to effectively spread the cues to distant relevant regions, making it more challenging to color larger semantic regions than smaller regions. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention proposes a point interaction grayscale image coloring method and system based on ViT encoding and pixel shuffling, where ViT is the abbreviation of Vision Transformers. The invention utilizes the self-attention mechanism in Vision Transformers to enable the model to selectively spread user cues to relevant regions in each single layer, introduces pixel shuffling to color the image in real time, and proposes a local stability layer to reduce the artifacts caused by pixel shuffling and improve the accuracy of coloring.
[0005] The point interaction grayscale image coloring method based on ViT encoding and pixel shuffling specifically includes the following steps:
[0006] Step 1, collect color images, and decompose the color images into grayscale images I g and user cues I hint where the user cues I hintIt includes position prompt information and color prompt information. The grayscale image I g and the user prompt I hint are cascaded at the channel level to form the sample X.
[0007] Step 2: Input the sample X obtained in Step 1 into the encoder of the Vision Transformer for multi-head self-attention calculation and layer normalization to obtain the output feature y p .
[0008] Step 3: For the feature y output by the encoder p , first pass through the local stability layer to limit the receptive field and reduce artifacts at the image patch boundaries, and then perform upsampling through the pixel shuffling method to obtain the color I ab pred , and finally merge it with the grayscale image I g to generate a color image.
[0009] Preferably, the local stability layer is one or a mixture of a linear layer, a convolutional layer, or local attention.
[0010] Step 4: Use the Huber loss function to compare the generated color image and the original color image in the CIELab color space, calculate the loss Lrecon, and use the AdamW optimizer for model training.
[0011] Step 5: Input the gray image to be colored into the trained model, add point-interactive user prompts, and generate a color image.
[0012] The point-interactive grayscale image coloring system based on ViT encoding and pixel shuffling includes an image reading module, a point-interactive module, a color selection module, and a color image generation module.
[0013] The image reading module is used to read the grayscale image I uploaded by the user g .
[0014] The point-interactive module obtains the interactive point position information by reading the positions clicked by the user on the grayscale image I g .
[0015] The color selection module is used to read the interactive point color information expected by the user and form the user prompt I with the interactive point position information obtained by the point-interactive module hint .
[0016] The color image generation module deploys the encoder of the trained Vision Transformer, and aims at the grayscale image I read by the image reading module g and the user prompt I given by the color selection module hint, extract the image feature y p , and generate the color information I using the local stability layer and pixel shuffling method ab pred , merge with the grayscale image I g , generate a color image and display it
[0017] The present invention has the following beneficial effects:
[0018] 1. The present invention innovatively proposes a point-interactive grayscale image coloring method based on Vision Transformer, enabling users to selectively color relevant regions. And it is convenient to operate, has a fast generation speed, and generates reasonable results through fewer user interactions.
[0019] 2. The present invention innovatively introduces pixel shuffling and local stability layer to effectively upsample the image at the lowest cost, thereby realizing real-time image coloring. Through lightweight pixel shuffling operations, the traditional decoder architecture can be abandoned, and a faster inference speed than existing baselines can be provided. A local stability layer is proposed to limit the receptive field of the last layer and reduce the artifacts caused by pixel shuffling. Description of the Drawings
[0020] Figure 1 is a schematic flow chart of a point-interactive grayscale image coloring method based on ViT provided by the present invention;
[0021] Figure 2 is a color image generated according to the grayscale image and user prompts in the embodiment;
[0022] Figure 3 is the PSNR comparison result of images generated by different methods in the embodiment;
[0023] Figure 4 is the LPIPS comparison result of images generated by different methods in the embodiment;
[0024] Figure 5 is the comparison result of color images generated by different methods in the embodiment;
[0025] Figure 6 is different color images generated by the present method for the same gray image in the embodiment. Detailed Embodiments
[0026] The following further explains the present invention with reference to the accompanying drawings; it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0027] As Figure 1As shown, the dot interaction grayscale image coloring method based on ViT encoding and pixel shuffling adopts the self-attention mechanism in Vision Transformers and utilizes its global receptive field, enabling the model to selectively propagate user prompts to relevant regions of each single layer. And it uses pixel shuffling to replace the decoder structure in Vision Transformers, and performs real-time image coloring through efficient upsampling. In addition, in order to reduce the artifacts caused by pixel shuffling with a large upsampling rate, a local stabilization layer is introduced before pixel shuffling, achieving more efficient and reasonable results with fewer user interactions. The specific steps are as follows:
[0028] Step 1. Prepare the grayscale image and user prompts, and preprocess them to form the input X for model training, which specifically includes the following steps:
[0029] Step 1.1. For color pictures from the large-scale dataset ImageNet, convert their color space from RGB to CIELab color space, where the L channel represents luminance, and the a and b channels represent chromaticity coordinates.
[0030] Step 1.2. Extract the information of the L channel to obtain a grayscale image I with the size of H×W×1 g . Extract the information of the a and b channels to form color prompt information I hint ’ with the size of H×W×2. Add a third channel as position prompt information to the color prompt information I hint ’, mark the prompt area with 1 and the non-prompt area with 0 to form user prompt I hint with the size of H×W×3. Here, H and W are the height and width of the original color picture.
[0031] Step 1.3. Perform channel concatenation on the grayscale image I g and the user prompt I hint to form a sample X with the size of H×W×4.
[0032] Step 2. Utilize Vision Transformer to achieve a global receptive field for propagating user prompts on the image to obtain the output feature y p , which specifically includes the following steps:
[0033] Step 2.1. Reshape the sample X with the size of H×W×4 into a series of block token sequences X 2 with the size of N×(P p ×4), taking an image patch with the size of P 2 ×4 as an input token. Here, P is the patch size, and N = H×W / P 2 is the number of input tokens.
[0034] Step 2.2: Input the token sequence Xp into the Transformer encoder for multi-head self-attention calculation and layer normalization:
[0035] z 0 = X p + Epos
[0036] z’ l = MSA(LN(z l-1 )) + z l-1
[0037] z l = MLP(LN(z’ l )) + z’ l
[0038] y p = LN(z l )
[0039] where Epos represents the sine position encoding, MSA(·) represents multi-head self-attention, LN(·) represents layer normalization, d represents the hidden dimension, l represents the number of layers, and y p represents the encoder output with a size of N×d.
[0040] Since multi-head self-attention does not use any position-related information, in addition to adding the position encoding Epos to the sequence Xp, a relative position bias is also added in the attention layer:
[0041]
[0042] Attention(Q, K, V) represents the attention calculation result of one head in multi-head self-attention, Q, K, and V represent the query matrix, key matrix, and value matrix obtained by mapping the input of multi-head self-attention, and B is the added relative position bias. Using the global receptive field of the self-attention mechanism, the color of the user prompt can be propagated to any spatial position in all layers.
[0043] Step 3: Based on the output feature y of the Transformer encoder p perform upsampling to generate color information and synthesize a color image with the grayscale image, which specifically includes the following steps:
[0044] Step 3.1: Reshape the output feature y p into a feature map y with a size of (H / P)×(W / P)×d.
[0045] Step 3.2: Process the feature map y through a local stabilization layer before pixel shuffling to limit the receptive field of the model and reduce artifacts at the boundaries of image patches.
[0046] Since the spatial resolution of the reshaped feature map y is reduced by P times compared to the original color image, in order to generate a full-resolution color image, it is necessary to upsample the feature map y. However, a large upsampling ratio will cause artifacts to appear at the boundaries of the image patches. Therefore, in order to promote the reasonable generation of colors, a local stability layer is used to limit the model to generate colors using neighboring features before pixel rearrangement. In this embodiment, a convolutional layer is used as the local stability layer.
[0047] Step 3.3: Upsample the output of the local stability layer through pixel shuffling technology to obtain the output I of the ab color channel. ab pred 。
[0048] Step 3.4: Connect the obtained final color I ab pred with the initial grayscale image I g to generate a color image.
[0049] Step 4: Use the Huber loss function to compare the generated color image and the original color image in the CIELab color space for model training, which specifically includes the following steps:
[0050] Step 4.1: Use the Huber loss function to compare the generated color image I pred and the original color image I GT to calculate the loss Lrecon. The loss function is as follows:
[0051]
[0052] Step 4.2: Use the AdamW optimizer for model training. During the training process, use the cosine annealing scheduler to manage the learning rate, set the number of iterative training times to 2.5 million times, and the batch size to 512.
[0053] Step 5: Save the model according to the optimal training result, input the grayscale image to be colored into the model, add point-interactive user prompts, and generate a color image.
[0054] Deploy the model saved in Step 5 to the interactive system to obtain a point-interactive grayscale image coloring system based on ViT encoding and pixel shuffling, including an image reading module, a point-interactive module, a color selection module, and a color image generation module.
[0055] The image reading module is used to read the grayscale image I uploaded by the user. g 。
[0056] The point-interactive module reads the user's operations on the grayscale image I gObtain the interactive point position information at the clicked position above.
[0057] The color selection module is used to read the color information of the interactive points expected by the user, and form the user prompt I with the interactive point position information obtained by the point interactive module hint .
[0058] The color image generation module is deployed with the encoder of the trained Vision Transformer. For the grayscale image I read by the image reading module g and the user prompt I given by the color selection module hint , extract the image feature y p , and generate the color information I using the local stability layer and the pixel shuffling method ab pred , merge with the grayscale image I g , generate a color image and display it.
[0059] Users can add different user prompts through the point interactive module and the color selection module, so as to generate different color images for a gray image.
[0060] Figure 2 On the left in is the gray image to be colored, and the colored squares in the image are the point interactive prompts added by the user. On the right is the color image generated based on the user prompt.
[0061] To illustrate the effectiveness of this method, on the three public datasets of ImageNet test, Oxford 102flowers, and CUB-200 respectively, use this method and other methods to generate color images, and select PSNR (Peak Signal-to-Noise Ratio) and LPIPS (Learned Perceptual Image Patch Similarity) as evaluation metrics. The results are as Figure 3 , Figure 4 shown. From Figure 3 it can be seen that the PSNR of this method on different datasets is ahead of other methods, indicating that the color images generated by this method have the best quality. From Figure 4 it can be seen that the LPIPS of this method on different datasets are lower than other methods, indicating that this method can achieve a better coloring effect with fewer interactive points.
[0062] Figure 5 shows the color images generated by different methods according to the same interactive point prompts. It can be seen that this method can achieve a better coloring effect when facing coloring tasks with different structures and complexities.
[0063] Figure 6 It shows different color images generated by the present method for the same gray image according to different user prompts.
[0064] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A point-interactive grayscale image coloring method based on ViT coding and pixel shuffling, characterized in that: The specific steps include: Step 1: Collect color images and decompose them into grayscale images I g and User Prompt I hint , where the user prompts I hint Including position hint information and color hint information; grayscale image I g and User Prompt I hint Perform channel cascading to form sample X; Step 2: Input the sample X obtained in step 1 into the encoder of Vision Transformer, perform multi-head self-attention calculation and layer normalization, and obtain the output feature y p ; Step 3: For the feature y output by the encoder p , first passes through a local stabilization layer to limit the receptive field and reduce artifacts at the image block boundaries, and then upsamples through a pixel shuffling method to obtain the color I ab pred , and finally with the grayscale image I g Merge to generate a color image; Step 4: Use the Huber loss function to compare the generated color image with the original color image in the CIELab color space, calculate the loss Lrecon, and use the AdamW optimizer for model training; Step 5: Input the gray image to be colored into the trained model, add some interactive user prompts, and generate a color image.
2. The point interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 1, characterized in that: The specific steps to form sample X are: Step 1.1: For the original color image, convert it from the RGB color space to the CIELab color space, where the L channel represents brightness and the a and b channels represent chromaticity coordinates; Step 1.2: Extract the information of L channel as grayscale image I g ; Extract the information of a and b channels as color prompt information I hint ', in the color prompt information I hint 'Add a third channel as the location prompt information and get the user prompt I hint ; step 1.
3. Grayscale image I g and User Prompt I hint Perform channel concatenation to form sample X.
3. The point-interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 1, characterized in that: Extract output feature y p The specific steps are: Step 2.1: Reshape the sample X into a series of block token sequences X p ; Step 2.2: Input the token sequence Xp into the Transformer encoder, perform layer normalization and multi-head self-attention calculation, and obtain the output feature y p .
4. The point-interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 3, characterized in that: is the block token sequence X p After adding the sinusoidal position code Epos, the layer is normalized and then sent to the multi-head self-attention mechanism.
5. The point-interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 3 or 4, characterized in that: In the multi-head self-attention calculation, add the relative position deviation B: Among them, Attention(Q,K,V) represents the attention calculation result of one head in the multi-head self-attention, and Q, K, and V represent the query matrix, key matrix, and value matrix obtained by the input mapping of the multi-head self-attention.
6. The point-interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 1, characterized in that: The local stabilization layer is a mixture of one or more of a linear layer, a convolutional layer, or a local attention layer.
7. The point-interactive grayscale image coloring method based on ViT coding and pixel shuffling as claimed in claim 1, characterized in that: The loss Lrecon is: Among them, I pred ,I GT Represent the generated color image and the original color image respectively.
8. A point-interactive grayscale image coloring system based on ViT coding and pixel shuffling, characterized by: It includes an image reading module, a point interactive module, a color selection module and a color image generation module; The image reading module is used to read the grayscale image to be colored uploaded by the user. g ; The point interactive module reads the user's grayscale image I g The clicked position on the screen obtains the interaction point location information; The color selection module is used to read the color information of the interaction point desired by the user, and the interaction point position information obtained by the point interactive module constitutes the user prompt I hint ; The color image generation module is equipped with a trained ViT encoder, and the grayscale image I read by the image reading module is g and user prompts given by the color selection module I hint , extract image features y p , and uses the local stabilization layer and pixel reorganization method to generate color information I ab pred , and the grayscale image I g Merge, generate color images and display them.