Image coloring method, and training method and device of image coloring network
Through the feature extraction and iterative denoising technology of the image colorization network, the coloring of animation line drawings is automatically completed, which solves the problem of traditional manual coloring being time-consuming and unstable, and realizes efficient and accurate color processing.
Patent Information
- Application Number
- CN202510852000.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional anime line drawing coloring relies on manual operation, which is time-consuming and has unstable visual effects, making it difficult to ensure consistency in color style.
An image colorization network is used for automatic colorization. By extracting features, adding noise, and iterative denoising, high-quality colorized images are generated in combination with color semantic features.
It improves the coloring efficiency and accuracy, achieves the stability and consistency of color style, and reduces the time consumption of manual intervention.
Smart Images

Figure CN120689463A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology and may involve fields such as artificial intelligence, image processing, and animation production. Specifically, the present application relates to an image coloring method, an image coloring network training method, and a device. Background Art
[0002] With the booming development of the animation industry, the demand for animation image production is increasing day by day. Among them, the coloring of animation line drawings is a key step in transforming black and white line drawings into vividly colored animation images, and plays a decisive role in the visual effect of the images.
[0003] Traditionally, coloring anime line art is done manually. Artists must carefully observe the line layout of the line art, understand the semantic information of the object or character outlined by each line, and use professional painting tools and software to color different areas one by one.
[0004] However, this traditional manual coloring method is time-consuming; coloring a complex line drawing often takes hours or even longer. Furthermore, the final visual effect is affected by factors such as the artist's personal state and artistic style preferences, making it difficult to ensure the consistency and stability of the color style of the work. Therefore, developing intelligent, automatic line drawing coloring technology to improve production efficiency while maintaining artistic quality has become a key development direction for the animation industry. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide an image colorization method, an image colorization network training method, and an apparatus that can effectively improve colorization efficiency and accuracy. To achieve this purpose, the technical solutions provided in the embodiments of the present application are as follows: In one aspect, an embodiment of the present application provides an image coloring method, the method comprising: Obtaining a line drawing to be colored and a coloring reference drawing corresponding to the line drawing to be colored; Based on the coloring reference image, the trained image coloring network performs the following coloring operations on the line drawing to be colored to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image to obtain color semantic features; Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; Iteratively denoising the noisy line drawing features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing features; The denoised colored image features are decoded to obtain a colored image corresponding to the line drawing to be colored.
[0006] On the other hand, an embodiment of the present application provides an image coloring method, the method comprising: Obtaining a line drawing to be colored, a coloring reference drawing of the line drawing to be colored, and a line drawing of the coloring reference drawing; Performing connected region division on the line drawing of the coloring reference image to obtain a plurality of connected domains; Based on the coloring reference image, the trained image coloring network performs the following coloring operations on the line drawing to be colored to obtain a colored image corresponding to the line drawing: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Adding noise to the line drawing feature to obtain a noisy line drawing feature; Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, determining a regional color feature of the connected domain according to a local color feature corresponding to the connected domain; Iteratively denoising the noisy line drawing image features based on the reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features, wherein the color semantic features include regional color features of each connected domain in the multiple connected domains; Decoding is performed based on the features of the denoised colored image to obtain a colored image corresponding to the line drawing to be colored.
[0007] Optionally, adding noise to the line drawing feature to obtain a noisy line drawing feature includes: Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; The method further comprises: Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; The iterative denoising of the noisy line drawing image features based on the noisy reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features includes: Based on the noisy reference image features corresponding to each denoising and the color semantic features, the noisy line drawing image features are iteratively denoised to obtain denoised colored image features corresponding to the noisy line drawing image features.
[0008] Optionally, the iterative denoising of the noisy line drawing image feature based on the reference image feature and the color semantic feature to obtain the denoised colored image feature corresponding to the noisy line drawing image feature includes: Based on the color semantic features and the noisy reference image features corresponding to the current denoising, predicting the noise of the number of denoising steps corresponding to the current denoising; Denoising the noisy line drawing features according to the predicted noise and the features of the noisy line drawing targeted by the current denoising, to obtain the features of the line drawing after the current denoising, which will be the features of the noisy line drawing targeted by the next denoising; The denoised colored image features are the denoised line drawing features obtained by the last denoising.
[0009] Optionally, the predicting the noise of the number of denoising steps corresponding to the current denoising based on the color semantic features and the features of the noisy reference image corresponding to the current denoising includes: Extract the features of the noisy reference image corresponding to the current denoising to obtain the first reference features; Splicing the noisy line drawing feature targeted by the current denoising and the first reference feature to obtain a spliced feature; According to the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature, weighted fusion is performed on the noisy line drawing feature and the first reference feature in the splicing feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, predict the noise of the number of denoising steps corresponding to the current denoising.
[0010] Optionally, extracting the feature of the noisy reference image corresponding to the current denoising to obtain the first reference feature includes any one of the following: Based on the self-attention mechanism, the feature of the noisy reference image corresponding to the current denoising is enhanced to obtain the first reference feature; The color semantic feature is weighted according to the correlation between the noisy reference image feature corresponding to the current denoising and the color semantic feature; and a first reference feature is obtained based on the weighted color semantic feature.
[0011] Optionally, the step of performing weighted fusion on the noisy line drawing feature and the first reference feature in the splicing feature based on the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature to obtain the first fused feature includes: Determining a correlation matrix between the features of the noisy line drawing targeted by the current denoising and the stitching features, wherein the correlation matrix includes correlations between each region in the line drawing to be colored and each region in the line drawing to be colored, and correlations between each region in the line drawing to be colored and each region in the coloring reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix to obtain a target probability matrix; Determining an optimal probability matrix based on the target probability matrix, the optimal probability matrix including the matching probabilities between each region in the line drawing to be colored and each region in the line drawing to be colored, and the matching probabilities between each region in the line drawing to be colored and each region in the coloring reference image, wherein the optimal probability matrix is a binary probability matrix; The noisy line drawing feature in the splicing feature and the first reference feature are weightedly fused according to the optimal probability matrix to obtain a first fused feature.
[0012] On the other hand, an embodiment of the present application provides a training method for an image colorization network, comprising: Acquire multiple samples, each of the samples including a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature. Color semantic feature extraction is performed on the colored reference image in the sample to obtain the color semantic feature. For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; According to the difference between the noise added to each sample and the predicted noise, the network parameters of the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0013] Optionally, the image colorization network includes a first encoding network, a network to be trained, and a decoding network, the first encoding network and the decoding network are pre-trained networks, and the network to be trained includes a color semantic extraction network and a denoising network; The feature extraction of the sample line drawing and the colored reference drawing in the sample is performed to obtain the line drawing features and the reference drawing features, including: Extracting features of the sample line drawing and the colored reference drawing in the sample using the first encoding network to obtain line drawing features and reference drawing features; The color semantic feature extraction of the colored reference image in the sample to obtain the color semantic feature includes: Extracting color semantic features from the colored reference image in the sample using the color semantic extraction network to obtain color semantic features; The step of predicting noise in the noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, comprises: Based on the noisy reference image features and color semantic features of the sample, predicting the noise in the noisy colored image features of the sample through the denoising network, and determining the denoised colored image features after denoising the noisy colored image features based on the predicted noise; The adjusting of network parameters of the image colorization network according to the difference between the noise added to each sample and the predicted noise includes: The network parameters of the network to be trained are adjusted according to the difference between the noise added to each sample and the predicted noise.
[0014] Optionally, the image colorization network includes a first encoding network, a network to be trained, and a decoding network, the first encoding network and the decoding network are pre-trained networks, and the network to be trained includes a second encoding network, a color semantic extraction network, and a denoising network; The feature extraction of the sample line drawing and the colored reference drawing in the sample is performed to obtain the line drawing features and the reference drawing features, including: Extracting features of the colored reference image in the sample through the first encoding network to obtain reference image features; The second encoding network is used to extract features of the sample line drawing in the sample to obtain line drawing features.
[0015] Optionally, predicting noise in the noisy colored image features of the sample based on the noisy reference image features and color semantic features of the sample includes: Extracting the noisy reference image features of the sample to obtain a first reference feature; Splicing the noisy colored image feature of the sample and the first reference feature to obtain a spliced feature; performing weighted fusion of the noisy colored image feature and the first reference feature in the spliced feature according to the correlation between the noisy colored image feature of the sample and the spliced feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, the noise in the noisy colored image feature of the sample is predicted.
[0016] Optionally, performing weighted fusion on the noisy colored image feature and the first reference feature in the spliced feature according to the correlation between the noisy colored image feature of the sample and the spliced feature to obtain the first fused feature includes: Determining a correlation matrix between the noisy colored image features in the sample and the stitching features, wherein the correlation matrix includes correlations between regions in the sample line drawing in the sample and regions in the sample line drawing, and correlations between regions in the sample line drawing and regions in the colored reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix to obtain a target probability matrix; Determining an optimal probability matrix based on the target probability matrix, the optimal probability matrix including the matching probabilities between each region in the sample line drawing and each region in the sample line drawing, and the matching probabilities between each region in the sample line drawing and each region in the coloring reference image, wherein the optimal probability matrix is a binary probability matrix; The noisy colored image feature and the first reference feature in the splicing feature are weightedly fused according to the optimal probability matrix to obtain a first fused feature.
[0017] On the other hand, an embodiment of the present application provides a method for training an image colorization network, the method comprising: Acquire multiple samples, each sample including a sample line drawing to be colored, a coloring reference drawing, a line drawing of the coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing; The line drawing of the colored reference image in each sample is divided into connected regions to obtain multiple connected domains; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature; For each sample, extract global color features from the colored reference image in the sample to obtain global color features; for each connected domain corresponding to the sample, determine the local color features corresponding to the connected domain in the global color features; for each connected domain, determine the regional color features of the connected domain based on the local color features corresponding to the connected domain; For each sample, add noise to the line drawing feature of the sample to obtain a noisy line drawing feature, and fuse the noisy line drawing feature corresponding to the sample with the standard image feature to obtain a noisy colored image feature; For each sample, based on the reference image features and color semantic features corresponding to the sample, predict the noise added to the noisy colored image features corresponding to the sample, and determine the denoised colored image features after denoising the noisy colored image features based on the predicted noise, so as to obtain the corresponding colored image by decoding the denoised colored image features; wherein the color semantic features include the regional color features of each connected domain corresponding to the sample; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0018] On the other hand, an embodiment of the present application provides an image coloring device, comprising: An image acquisition module is used to acquire a line drawing to be colored and a coloring reference image corresponding to the line drawing to be colored; An image coloring module is configured to perform the following coloring operations on the line drawing to be colored based on the coloring reference image using a trained image coloring network to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image to obtain color semantic features; Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; Iteratively denoising the noisy line drawing features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing features; The denoised colored image features are decoded to obtain a colored image corresponding to the line drawing to be colored.
[0019] Optionally, the image colorization module may also be used to: Obtaining a line drawing of the coloring reference image; Performing connected region division on the line drawing of the coloring reference image to obtain a plurality of connected domains; The image colorization module can be used to: Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, a regional color feature of the connected domain is determined according to a local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
[0020] Optionally, the image colorization module may be used to: Based on the color semantic features and the noisy reference image features corresponding to the current denoising, predicting the noise of the number of denoising steps corresponding to the current denoising; Denoising the noisy line drawing features according to the predicted noise and the features of the noisy line drawing targeted by the current denoising, to obtain the features of the line drawing after the current denoising, which will be the features of the noisy line drawing targeted by the next denoising; The denoised colored image features are the denoised line drawing features obtained by the last denoising.
[0021] Optionally, the image colorization module may be used to: Extract the features of the noisy reference image corresponding to the current denoising to obtain the first reference features; Splicing the noisy line drawing feature targeted by the current denoising and the first reference feature to obtain a spliced feature; According to the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature, weighted fusion is performed on the noisy line drawing feature and the first reference feature in the splicing feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, predict the noise of the number of denoising steps corresponding to the current denoising.
[0022] Optionally, the image colorization module may be configured to perform any of the following: Based on the self-attention mechanism, the feature of the noisy reference image corresponding to the current denoising is enhanced to obtain the first reference feature; The color semantic feature is weighted according to the correlation between the noisy reference image feature corresponding to the current denoising and the color semantic feature; and a first reference feature is obtained based on the weighted color semantic feature.
[0023] On the other hand, an embodiment of the present application provides an image coloring device, comprising: An image acquisition module is used to acquire a line drawing to be colored, a coloring reference image of the line drawing to be colored, and a line drawing of the coloring reference image; A region division module is used to divide the line drawing of the coloring reference image into connected regions to obtain multiple connected regions; An image coloring module is configured to perform the following coloring operations on the line drawing to be colored using a trained image coloring network based on the coloring reference image to obtain a colored image corresponding to the line drawing: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; adding noise to the line drawing feature to obtain a noisy line drawing feature; Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, determining a regional color feature of the connected domain according to a local color feature corresponding to the connected domain; Iteratively denoising the noisy line drawing image features based on the reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features, wherein the color semantic features include regional color features of each connected domain in the multiple connected domains; Decoding is performed based on the features of the denoised colored image to obtain a colored image corresponding to the line drawing to be colored.
[0024] Optionally, the image colorization module may be used to: Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; The image colorization module can also be used to: Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; The iterative denoising of the noisy line drawing image features based on the noisy reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features includes: Based on the noisy reference image features corresponding to each denoising and the color semantic features, the noisy line drawing image features are iteratively denoised to obtain denoised colored image features corresponding to the noisy line drawing image features.
[0025] On the other hand, an embodiment of the present application provides a training device for an image colorization network, comprising: A sample acquisition module is used to acquire multiple samples, each of which includes a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; A training module is configured to continuously perform the following training operations on the image colorization network to be trained based on the multiple samples until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature. Color semantic feature extraction is performed on the colored reference image in the sample to obtain the color semantic feature. For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; According to the difference between the noise added to each sample and the predicted noise, the network parameters of the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0026] On the other hand, an embodiment of the present application provides a training device for an image colorization network, comprising: A sample acquisition module is used to acquire multiple samples, each sample including a sample line drawing to be colored, a coloring reference drawing, a line drawing of the coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing; A region division module is used to divide the line drawing of the colored reference image in each sample into connected regions to obtain multiple connected regions; A training module is configured to continuously perform the following training operations on the image colorization network to be trained based on the multiple samples until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature; For each sample, extract global color features from the colored reference image in the sample to obtain global color features; for each connected domain corresponding to the sample, determine the local color features corresponding to the connected domain in the global color features; for each connected domain, determine the regional color features of the connected domain based on the local color features corresponding to the connected domain; For each sample, add noise to the line drawing feature of the sample to obtain a noisy line drawing feature, and fuse the noisy line drawing feature corresponding to the sample with the standard image feature to obtain a noisy colored image feature; For each sample, based on the reference image features and color semantic features corresponding to the sample, predict the noise added to the noisy colored image features corresponding to the sample, and determine the denoised colored image features after denoising the noisy colored image features based on the predicted noise, so as to obtain the corresponding colored image by decoding the denoised colored image features; wherein the color semantic features include the regional color features of each connected domain corresponding to the sample; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0027] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.
[0028] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.
[0029] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.
[0030] The beneficial effects of the technical solution provided by the embodiments of the present application are as follows: The image coloring method provided in the embodiment of the present application can add noise of the maximum number of noise addition steps to the extracted line drawing features when coloring the line drawing based on the coloring reference image, and add noise of the number of noise addition steps corresponding to each denoising to the reference image features. Based on the noisy reference image features and color semantic features corresponding to each denoising, the noisy line drawing features are iteratively denoised to obtain denoised coloring image features, and the coloring image corresponding to the line drawing is obtained by decoding the denoised coloring image features. In each iterative denoising process, this method uses the noisy reference image corresponding to the current denoising as a constraint to guide denoising, so that the reference image and the line drawing can be feature-aligned under the same noise interference, thereby more accurately achieving color migration and improving the accuracy of image coloring. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0032] Figure 1 A schematic diagram of the structure of an image colorization system provided in an embodiment of the present application; Figure 2 A flowchart of a training method for an image colorization network provided in an embodiment of the present application; Figure 3 Schematic diagram of the coloring reference image and line drawing provided in the embodiments of the present application; Figure 4 A schematic diagram of dividing a line drawing into regions provided in an embodiment of the present application; Figure 5 A schematic diagram of the process of color semantic feature extraction provided in an embodiment of the present application; Figure 6 A schematic diagram of the process of color semantic feature extraction provided in an embodiment of the present application; Figure 7 A denoising flow chart provided in an embodiment of the present application; Figure 8 A network structure diagram of an RU network and a DU network provided in an embodiment of the present application; Figure 9 A schematic diagram of the network structure of the image colorization network provided in an embodiment of the present application; Figure 10 A flowchart of a training method for an image colorization network provided in an embodiment of the present application; Figure 11 A schematic diagram of a process for coloring an image provided in an embodiment of the present application; Figure 12 A schematic diagram of a process for coloring an image provided in an embodiment of the present application; Figure 13A schematic diagram of an interface for coloring an animation line drawing provided in an embodiment of the present application; Figure 14 A schematic diagram of a process for coloring an image provided in an embodiment of the present application; Figure 15 A schematic structural diagram of an image coloring device provided in an embodiment of the present application; Figure 16 A schematic structural diagram of an image coloring device provided in an embodiment of the present application; Figure 17 A schematic diagram of the structure of a training device for an image colorization network provided in an embodiment of the present application; Figure 18 A schematic diagram of the structure of a training device for an image colorization network provided in an embodiment of the present application; Figure 19 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0034] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A including A1 or A2 or A3, and can also be implemented as parameter A including at least two of the three items A1, A2, and A3.
[0035] In order to better understand and illustrate the method provided in the embodiments of the present application, some technical terms involved in the embodiments of the present application are first explained and illustrated below.
[0036] Flooding method: A method for segmenting closed areas in computer vision. Its core principle is to start from a specified seed point and traverse the pixel area connected to the seed point until it encounters a boundary line (such as a black outline), thereby achieving automatic segmentation of the closed area.
[0037] The image coloring method provided in the embodiment of the present application can theoretically be applied to image coloring tasks in various scenes, including but not limited to coloring of animation line drawings, coloring of architectural drawings, etc.
[0038] Figure 1 This is a structural schematic diagram of an image colorization system provided in an embodiment of the present application, wherein the image colorization system includes a terminal 10, a server 20, and a network 30, wherein an application with an image colorization function runs in the terminal 10, the server 20 is a background server that provides image colorization function services, and a trained image colorization network is deployed in the server 20, and the terminal 10 and the server 20 can communicate through the network 30.
[0039] In one or more embodiments of the present application, a user may submit a line drawing to be colored and a coloring reference image through an image coloring application in the terminal 10. In response to the user's coloring trigger operation, the terminal 10 sends an image coloring request to the server 20. After receiving the image coloring request, the server 20 may input the line drawing and the coloring reference image in the image coloring request into the image coloring network, so that the image coloring network generates a coloring image corresponding to the line drawing with reference to the coloring reference image. The server 20 may send the generated coloring image corresponding to the line drawing to the terminal 10 and display the coloring image corresponding to the line drawing in the terminal 10 interface.
[0040] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server. The terminal can be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device (e.g., smart speaker), wearable electronic device (e.g., smart watch), vehicle-mounted terminal, smart home appliance (e.g., smart TV), AR / VR device, etc. The network can be a wired network or a wireless network. Wired networks can include local area networks, metropolitan area networks, and wide area networks, while wireless networks can include Bluetooth, Wi-Fi, and other networks that enable wireless communication.
[0041] The following describes several embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, learn from, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0042] An embodiment of the present application provides a method for training an image colorization network, which can be executed by any electronic device, for example, a server or a terminal.
[0043] Figure 2 This is a flow chart of the training method of the image colorization network provided in the embodiment of the present application. Figure 2 As shown, the method may include the following steps S110 to S120, wherein: S110: Acquire multiple samples, each sample including a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing.
[0044] Among them, the line drawing is an uncolored image composed of black and white lines, which mainly contains the outline, structure and other information of the object. The coloring reference image is a color image, which is usually a different posture or scene of the same object as the sample line drawing. It is used to provide guidance for coloring the objects in the sample line drawing. The standard coloring image is a color image after the objects in the line drawing to be colored are standardized.
[0045] like Figure 3 As shown, the image of a little bear holding a lollipop on the left is a coloring reference image. Five line drawings of the little bear in different states (reading, angry, cooking, sleeping, and standing) are provided above. Based on the color style of the little bear in the coloring reference image, each line drawing is colored separately to obtain the standard coloring images corresponding to each line drawing below.
[0046] S120: Based on the multiple samples, continuously perform a training operation on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network.
[0047] The training termination conditions and loss function of the image colorization network can be configured as needed. For example, the training termination conditions may include but are not limited to the number of training times reaching a preset number, the loss function convergence (such as the training loss is less than a preset value, or the training losses for multiple consecutive times are less than a preset value, etc.), the network test indicators meeting the preset indicators, etc. The network training loss represents the deviation between the noise predicted by the image colorization network during the denoising process and the actual added noise, and / or the deviation between the colored image generated by the image colorization network and the standard colored image of the line drawing. Through the training loss function of the model, the network is supervisedly trained using the gradient descent algorithm, so that the noise predicted by the network can continuously approach the actual added noise, and the generated colored image can continuously approach the standard colored image, thereby obtaining a trained image colorization network that meets the needs of actual applications.
[0048] The training operation includes the following steps S1201 to S1204: S1201: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain line drawing features and reference image features, and color semantic feature extraction is performed on the colored reference image in the sample to obtain color semantic features.
[0049] In this embodiment of the present application, the image colorization network may include a first encoding network, a color semantic extraction network, a denoising network (DU), and a decoding network. For each sample, the first encoding network may perform feature extraction on the sample line drawing and the colorized reference image, respectively, to obtain line drawing features and reference image features. The color semantic extraction network may then perform color semantic feature extraction on the colorized reference image, to obtain color semantic features.
[0050] Optionally, the image colorization network may also include a first encoding network, a second encoding network, a color semantic extraction network, a denoising network, and a decoding network, wherein the first encoding network is used to extract features from colored images, and the second encoding network is used to extract features from non-colored line drawings. For each sample, the first encoding network may be used to extract features from the colored reference image in the sample to obtain reference image features; the second encoding network may be used to extract features from the sample line drawing in the sample to obtain line drawing features; and the color semantic extraction network may be used to extract color semantic features from the colored reference image in the sample to obtain color semantic features.
[0051] Optionally, since the standard colored image is a colored image, the standard image features obtained above may also be obtained by encoding the standard colored image through the first encoding network.
[0052] In an embodiment of the present application, by using different encoding networks to perform feature extraction on the colored reference image and the non-colored sample line drawing respectively, the color, texture of the colored reference image and the line structure, spatial layout and other features of the line drawing can be captured in a targeted manner, thereby improving the accuracy of subsequent coloring.
[0053] The embodiments of this application do not restrict the network types of the encoding network (first encoding network, second encoding network) and the decoding network, and can be selected based on actual application requirements. As an alternative, the first encoding network in the embodiments of this application can use the encoder in a variational autoencoder (VAE), the second encoding network can use a sketch guide network (SG), and the decoding network can use a VAE decoder.
[0054] The embodiments of the present application do not impose any restrictions on the network structure of the color semantic extraction network. For example, a text encoder in a contrastive language-image pre-training (CLIP) network can be used to encode an image, and the encoded features can be used as the color semantic features of the image; or a color guide network (CG) can be used to extract color semantic features from the image.
[0055] S1202: For each sample, add the same noise to the reference image feature and the line drawing feature of the sample to obtain a noisy reference image feature and a noisy line drawing feature, and fuse the noisy line drawing feature and the standard image feature of the sample to obtain a noisy colored image feature.
[0056] S1203: For each sample, based on the noisy reference image features and color semantic features of the sample, predict the noise in the noisy colored image features of the sample, and based on the predicted noise, determine the denoised colored image features after the noisy colored image features are denoised, so as to obtain the corresponding colored image by decoding the denoised colored image features.
[0057] The image colorization method provided in the embodiment of the present application relies on the implementation principle of the diffusion model (Stable Diffusion, SD) and realizes colorization through two stages: forward diffusion (noise addition) and reverse denoising.
[0058] Specifically, in the forward stage, for each sample, the same noise is added to the reference image features and line drawing features of the sample respectively to obtain the noisy reference image features and the noisy line drawing features; in the reverse stage, based on the color semantic features of the sample and the noisy reference image features with the same noise added, the noise in the noisy colored image features of the sample is predicted through the denoising network, and the predicted noise is removed from the noisy colored image features to obtain the corresponding denoised colored image features, and then the denoised colored image features are decoded through the decoding network to obtain the colored image corresponding to the line drawing image.
[0059] In an embodiment of the present application, by adding the same noise as the line drawing feature to the reference image feature and performing guided denoising based on the noisy reference image feature with the same noise added, the noise in the noisy line drawing can be removed more accurately, thereby improving the accuracy of coloring the line drawing.
[0060] After obtaining the noisy line drawing features of the sample, the noisy line drawing features of the sample and the standard image features of the standard colored image may be concatenated to obtain the noisy colored image features.
[0061] In an embodiment of the present application, the noisy reference image features and color semantic features of the colored reference image may be used as conditions to guide the removal of noise in the noisy colored image features.
[0062] Specifically, the noisy reference image features, color semantic features, and noisy colored image features of the colored reference image are fed into a denoising network to predict the noise added to the noisy colored image features. The predicted noise is then removed from the noisy colored image features to obtain denoised colored image features. The denoised colored image features are then decoded by a decoding network to obtain the colored image corresponding to the line drawing.
[0063] Optionally, when adding noise to the reference image features and line drawing features of a sample, for each sample, a noise adding step number can be randomly selected from a preset noise adding step number range, where different noise adding step numbers represent different degrees of noise addition; and the noise corresponding to the selected noise adding step number is added to the reference image features and line drawing features of the sample respectively.
[0064] When denoising the noisy colored image features, for each sample, the noisy reference image features, color semantic features and the selected number of noisy steps of the sample are input into the denoising network. The noise corresponding to the number of noisy steps is predicted by the denoising network, and the noise corresponding to the number of noisy steps is removed from the noisy colored image features to obtain the denoised colored image features. The corresponding colored image is obtained by decoding the denoised colored image features.
[0065] As an implementation method, the addition of noise can be expressed as:
[0066] in, , t is the number of noise adding steps, Indicates the preset range of noise adding steps, that is, the selection range of noise adding intensity; Indicates the image features to be added with noise (reference image features / line drawing features), t is the number of noise addition steps, indicating the degree of noise addition, represents a predefined constant, Represents the image features after adding noise (noisy reference image features / noisy line drawing features), represents Gaussian noise.
[0067] The reference image features corresponding to sample A and line drawing features For example, adding noise, randomly select the number of noise adding steps t, and add noise to the reference image feature and line drawing features Add the real noise corresponding to t to obtain the noisy reference image features And the characteristics of noisy line drawings , the noisy line drawing features With standard coloring features Splice and get the noisy colored image features ; The noisy colored image features , noisy reference image features , color semantic features And the number of noise steps t is input into the denoising network, the noise corresponding to the number of noise steps t is predicted, and the noise coloring features are obtained from the noise Remove the predicted noise and obtain the denoised color image feature .
[0068] In the embodiment of the present application, when the image colorization network is trained, the added noise degree / intensity is randomly sampled and denoised through the denoising network, which can cover the distribution of noise of various added intensities, improve the adaptive processing capability of the denoising network, and make the image colorization network more flexible in the inference stage.
[0069] S1204: Adjust the network parameters of the image colorization network according to the difference between the noise added to each sample and the predicted noise.
[0070] In an embodiment of the present application, the first encoding network and the decoding network in the image colorization network can be pre-trained networks, and the other branch networks (semantic extraction network, denoising network and / or second encoding network) are networks to be trained, that is, the network parameters in the first encoding network and the decoding network are fixed network parameters obtained after pre-training. During training, only the network parameters in the network to be trained can be adjusted.
[0071] Therefore, the total training loss can be determined according to the difference between the noise added to each sample and the predicted noise, and the network parameters of the network to be trained in the image colorization network can be adjusted based on the total training loss.
[0072] Optionally, all branch networks in the image colorization network can be networks to be trained. A first training loss can be determined based on the difference between the noise added to each sample and the predicted noise. A second training loss can be determined based on the difference between the standard coloring image corresponding to each sample and the coloring image generated by the image colorization network. The total training loss can be determined based on the first training loss and the second training loss. Based on the total training loss, the network parameters of each branch network in the image colorization network can be adjusted.
[0073] based on Figure 2 The training method for an image colorization network shown in the figure adds the same noise to both the reference image features and the line drawing features when coloring a line drawing using a colorization reference image. The noise in the noisy line drawing features is predicted and removed using the color semantic features of the colorization reference image and the features of the noisy reference image with the same noise added as the color semantic features of the colorization reference image as conditions. During the denoising process, this method uses the color semantic features of the colorization reference image and the noisy reference image with the same noise added as the noisy line drawing features as conditions to guide denoising. This allows the reference image and line drawing features to be aligned despite the same noise interference, resulting in more precise color transfer and improved image colorization accuracy.
[0074] In an embodiment of the present application, each sample may also include a line drawing of a colored reference image in the sample. For each sample, the line drawing of the colored reference image in the sample may be divided into connected areas to obtain multiple connected domains.
[0075] Optionally, when dividing the line drawing of the coloring reference image into connected regions, a flooding method can be used. Starting from any key point in the line drawing (seed point), the method continuously traverses the surrounding pixels until it reaches the black outline, clustering pixels with the same attributes into independent connected regions, and dividing multiple connected regions. Each connected region represents a basic coloring unit in the image.
[0076] Take the image of the little bear standing as a reference for coloring. Figure 4 In the example, a line drawing of the standing bear is obtained as a reference for coloring, and each connected area in the line drawing is divided to obtain multiple closed connected domains, each of which can be represented by a different color.
[0077] For each sample, when performing color semantic feature extraction on the colored reference image in the sample, global color feature extraction can be performed on the colored reference image in the sample to obtain the global color feature of the colored reference image; for each connected domain in the multiple connected domains divided by the line drawing of the colored reference image, the local color feature corresponding to the connected domain in the global color feature is determined; for each connected domain, the regional color feature of the connected domain is determined based on the local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
[0078] The color semantic features of the colored reference image can be obtained by extracting the color semantic features of the colored reference image through a color semantic extraction network. The color semantic extraction network can include a global semantic feature extraction module (Semantic feature extraction) and a region-based feature pooling module (Mask embedding). The global semantic feature extraction module is used to extract global color features of the colored reference image in the sample, and the region-based feature pooling module can extract regional color features of each connected domain based on the global color features of the colored reference image and each connected domain divided by the line drawing based on the colored reference image.
[0079] like Figure 5 As shown in the figure, the reference image of the colored bear holding a lollipop is input into the global semantic feature extraction module to extract the global color semantic feature (Feature_Color). The extracted global color semantic feature has the same image size as the reference image, and both are ; Afterwards, the extracted global color semantic features and the connected domains segmented by the line drawing based on the reference image are input into the feature pooling module to obtain the regional color features of each connected domain.
[0080] In order to avoid the loss of semantic details, the above-mentioned global semantic feature extraction module can adopt a U-shaped network, that is, the image is first downsampled and then upsampled to keep the image resolution unchanged.
[0081] As an optional method, when extracting the color semantic features of the colored reference image, such as Figure 6 As shown, feature extraction can be performed on the colored reference image to obtain global color semantic features; the line drawing of the colored reference image can be segmented into regions to obtain multiple connected domains; from the global color features of the colored reference image, the local color features corresponding to each connected domain are intercepted according to the divided connected domains, and for each connected domain, the local color features corresponding to the connected domain are average pooled, and the average value of the feature values of each feature point in the local color feature is used as the regional color feature of the connected domain.
[0082] In the embodiment of the present application, when using the noisy reference image features and color semantic features as conditions to guide denoising, the noisy reference image features can be further processed to obtain detail features, which are then used as conditions. Specifically, feature extraction is performed on the noisy reference image features of the sample to obtain a first reference feature. This first reference feature is then used as a condition to input into the denoising network for denoising.
[0083] Optionally, the image colorization network further includes a reference feature guidance network (Reference Unet, RU), which can be used to extract features of the noisy reference image of the sample to obtain a first reference feature.
[0084] As an optional method, the first reference feature may be extracted by any of the following methods: The self-attention mechanism in the network is guided by the reference feature to enhance the feature of the noisy reference image of the sample and obtain the first reference feature; According to the correlation between the noisy reference image feature and the color semantic feature of the sample, the color semantic feature is weighted by the reference feature-guided network, and the first reference feature is obtained based on the weighted color semantic feature.
[0085] Among them, when the self-attention mechanism is used for feature enhancement, the noisy reference image feature can be input into the self-attention layer of the reference feature guidance network, and the noisy reference image feature can be multiplied by the network parameters W_q_RU, W_k_RU, and W_v_RU of the reference feature guidance network respectively to obtain the query feature Q and the key value KV, and the enhanced first reference feature can be calculated by the following formula:
[0086] Among them, the reference feature guidance network can also include a cross attention layer, which determines the query feature Q based on the noisy reference image feature, determines the key K and value V based on the color semantic feature, and calculates the weighted first reference feature through the above formula.
[0087] Optionally, the reference feature guided network can include multiple layers of self-attention layers and multiple layers of cross-attention layers. The self-attention layer can be used to perform self-attention enhancement on the input image features, and the enhanced image features are used as the query features Q of the next cross-attention layer. The color semantic features of the colored reference image are used as key values K and V, and are fused through cross-attention. After multiple layers of self-attention layers and cross-attention layers, the dynamic spatial features of the noisy reference image are finally output.
[0088] In the embodiment of the present application, after obtaining the processed first reference feature, denoising can be guided based on the first reference feature. Specifically, the noisy colored image feature of the sample and the first reference feature can be spliced to obtain a spliced feature; based on the correlation between the noisy colored image feature and the spliced feature of the sample, the noisy colored image feature and the first reference feature in the spliced feature are weightedly fused to obtain a first fused feature; based on the color semantic feature and the first fused feature, the noise in the noisy colored image feature of the sample is predicted.
[0089] In order to match the most relevant features (color features in the coloring reference image or structural and contour features in the line drawing) from the coloring reference image and the sample line drawing as reference conditions to guide denoising and coloring, in this embodiment of the application, the first fusion feature can be determined in the following manner: Determine a correlation matrix between the noisy colored image features and the stitching features in the sample, wherein the correlation matrix includes correlations between regions in the sample line drawing and regions in the sample line drawing, and correlations between regions in the sample line drawing and regions in the colored reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix, a target probability matrix is obtained; Based on the target probability matrix, an optimal probability matrix is determined, where the optimal probability matrix includes the matching probabilities between each region in the sample line drawing and each region in the sample line drawing, and the matching probabilities between each region in the sample line drawing and each region in the coloring reference image. The optimal probability matrix is a binary probability matrix; The noisy colored image feature and the first reference feature in the splicing feature are weightedly fused according to the optimal probability matrix to obtain the first fused feature.
[0090] Among them, the first probability matrix represents the correlation between each area in the sample line drawing and each area in the sample line drawing + colored reference image; the second probability matrix represents the correlation between each area in the sample line drawing + colored reference image and each area in the sample line drawing.
[0091] Optionally, when determining the optimal probability matrix, a Hungarian matching algorithm may be used to determine that each area of the sample line drawing corresponds to the most matching area in the sample line drawing + colored reference image, and a binary probability matrix may be determined based on the degree of matching.
[0092] In an embodiment of the present application, a denoising network in an image colorization network can be used to predict and denoise noise in the noisy colorization image features. The denoising network includes a self-attention layer and a cross-attention layer. The self-attention layer can be used to determine the first fusion feature, and the cross-attention layer can be used to introduce color semantic features.
[0093] Specifically, the first reference feature after the reference feature guide network processing is multiplied by the network parameters W_k and W_v respectively to obtain K_ref and V_ref; the noisy colored image is input into the denoising network and multiplied by the network parameters W_q_DU, W_k_ DU and matrix W_v_ DU in the denoising network respectively. The new features obtained are recorded as Q_tar, K_tar and V_tar. Then, K_tar and K_ref are concatenated to obtain the first concatenated feature as the new key K, and V_tar and V_ref are concatenated to obtain the second concatenated feature as the new value V. The first fusion feature is calculated by the following formula of the spatial self-attention mechanism:
[0094] Afterwards, the query feature Q is determined based on the first fused feature, and the key values K and V are determined based on the color semantic feature. Through the cross-attention mechanism, the correlation between the color semantic feature and the first fused feature is used to weight the color semantic feature to obtain the predicted noise in the noisy colored image feature of the sample.
[0095] Figure 7 A denoising flowchart is provided for an embodiment of the present application. After extracting the reference image features of the colored image through the VAE encoder, dynamic noise, that is, noise of random steps, can be added to the reference image features, and the noisy reference image features after adding noise are input into the reference feature guidance network. According to the first reference feature output by the reference feature guidance network, K_ref and V_ref corresponding to the noisy reference image features are determined; based on the noisy colored image features, the corresponding K_tar and V_tar are determined, K_tar and K_ref are spliced into a new K, and V_tar and V_ref are spliced into a new V. Based on Q_tar and the new [K_tar, K_ref] and [V_tar, V_ref], attention fusion based on binary probability is used to obtain the first fused feature as the new query feature Q_tar_new.
[0096] Among them, when performing attention fusion, Q_tar and [K_tar, K_ref] are multiplied to obtain the correlation matrix, and a two-way softmax strategy is adopted, that is, the softmax calculation is performed on the rows of the correlation matrix to obtain the probability P1, and then the softmax calculation is performed on the columns to obtain the probability P2. Calculate the target probability matrix P_final; Based on the target probability matrix, obtain the optimal matching relationship through the Hungarian matching algorithm, convert the target probability matrix into the optimal probability matrix (binary probability matrix P_final_binary), and perform weighted fusion on the splicing features [V_tar, V_ref] according to the binary probability matrix to obtain the first fusion feature.
[0097] It is understandable that in the embodiment of the present application, the attention fusion method based on binary probability can also be implemented separately. In the above training operation, it is also possible not to add noise to the reference image features, but to directly use the non-noised reference image features as conditions for denoising.
[0098] Figure 8 The network structure diagram of the RU network and DU network provided in the embodiment of the present application, wherein the network structure of RU and DU is the same, and the network structure includes a downsampling layer, an intermediate layer, and an upsampling layer. The downsampling layer is composed of the encoding module AD of the stable diffusion (SD) network, the SD encoding module A is the residual module (ResidualModule, Res), the SD encoding module BD is the residual-transformer module (Residual-Transformer Module, Res-Trans), the SD intermediate module contained in the intermediate layer is the Res-Trans module, the upsampling layer is composed of the SD decoding module AD, the SD decoding module D is the Res module, and the SD decoding module AC is the Res-Trans module. The Transformer module includes a self-attention mechanism (SA), a cross-attention mechanism (CA), and a feedforward network (FF).
[0099] The noisy reference image features (reference image features + noise addition steps) and color semantic features are input into the RU network. The noisy reference image features are enhanced through the self-attention layer in the RU network, and the color semantic features are introduced through cross-attention. The attention fusion result is continued to be input into the next attention layer. The output of each self-attention layer in the RU network is input into the corresponding layer in the DU network as the dynamic spatial feature (first reference feature). In the DU network, the noisy colored image features and the dynamic spatial features are spliced to obtain a new key value. The first fusion feature is calculated through the spatial self-attention mechanism, and the color semantic feature is introduced through cross-attention. The color semantic features and dynamic spatial features are combined to predict and remove noise.
[0100] Figure 9This is a schematic diagram of the network structure of the image colorization network provided in an embodiment of the present application. The image colorization network includes a VAE encoder (first encoding network), an SG network (second encoding network), a CG network (color semantic extraction network / color guidance network), an RU network (reference feature guidance network), a DU network (denoising network), and a VAE decoder (decoding network). The VAE encoder and VAE decoder are pre-trained networks, while the SG network, CG network, RU network, and DU network are networks to be trained.
[0101] The following takes a training operation of a training sample as an example to explain the training process in detail: A1: Input the colored reference image in the sample into the VAE encoder to extract the reference image features (hidden features) of the colored reference image.
[0102] For example, the image dimension of the color reference image of the given sample is , perform VAE encoding on the colored reference image, reduce the height and width of the image by 8 times respectively, and obtain Reference graph features of dimensions.
[0103] A2: Input the colored reference image in the sample into the CG network to extract the color semantic features of the colored reference image.
[0104] Specifically, global color semantic features are extracted from the colored reference image to obtain the global color semantic features of the colored reference image; a line drawing of the colored reference image is obtained, and the connected domains of the line drawing of the colored reference image are divided to obtain multiple connected domains; the local color features corresponding to each connected domain are determined from the global color semantic features, and the regional color features of each connected domain are determined based on the local color features of each connected domain, wherein the color semantic features include the regional color features of each connected domain.
[0105] A3: Input the sample line drawing in the sample into the SG network to obtain the line drawing features of the sample line drawing.
[0106] A4: Randomly select a noise step number t from the preset noise step number range [1, T], add noise corresponding to the noise step number t to the line drawing features in the sample, and then splice it with the standard colored image to obtain the noisy colored image.
[0107] A5: Add noise with a noise adding step number t (the same as the noise adding step number in the line drawing) to the reference image feature of the colored reference image to obtain a noisy reference image feature.
[0108] A6: Input the noisy reference image features and color semantic features into the RU network to obtain dynamic spatial features.
[0109] A7: The noisy colored image, dynamic spatial features, and color semantic features are input into the DU network. The noise corresponding to the number of noise addition steps t is predicted. Based on the predicted noise, the features of the denoised colored image are determined. The denoised colored image features are decoded through the decoding network to obtain the colored image corresponding to the sample line drawing.
[0110] A8: According to the difference between the predicted noise and the actual added noise, the network parameters in the SG network, CG network, RU network and DU network are adjusted.
[0111] An embodiment of the present application also provides a method for training an image colorization network, which can be executed by any electronic device, for example, by a server or a terminal.
[0112] Figure 10 This is a flow chart of the training method of the image colorization network provided in the embodiment of the present application. Figure 10 As shown, the method may include the following steps S210 to S230, wherein: Step S210: Acquire multiple samples, each sample including a sample line drawing to be colored, a coloring reference drawing, a line drawing of the coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing.
[0113] The sample line drawing and the coloring reference drawing in the same sample are usually different postures or scenes of the same object, and the standard coloring drawing is a color image after the objects in the sample line drawing are colorized in a standardized manner.
[0114] Step S220: performing connected region division on the line drawing of the colored reference image in each sample to obtain a plurality of connected regions.
[0115] In the embodiment of the present application, a flooding method can be used to divide connected regions. Specifically, starting from any key point in the line drawing (seed point), the surrounding pixels are continuously traversed until the black outline is reached. Pixels with the same attributes are clustered into independent connected domains, and multiple connected domains are divided. Each connected domain represents a basic colored unit in the image.
[0116] Step S230: Based on the multiple samples, the image colorization network to be trained is continuously trained to obtain a trained image colorization network.
[0117] The training operation includes the following steps S2301 to S2305.
[0118] Step S2301: For each sample, feature extraction is performed on the sample line drawing and the colored reference drawing in the sample to obtain line drawing features and reference drawing features.
[0119] Step S2302: For each sample, extract the global color features of the colored reference image in the sample to obtain the global color features; for each connected domain corresponding to the sample, determine the local color features corresponding to the connected domain in the global color features; for each connected domain, determine the regional color features of the connected domain based on the local color features corresponding to the connected domain.
[0120] Step S2303: For each sample, add noise to the line drawing feature of the sample to obtain a noisy line drawing feature, and fuse the noisy line drawing feature corresponding to the sample with the standard image feature to obtain a noisy colored image feature.
[0121] Step S2304: For each sample, based on the reference image features and color semantic features corresponding to the sample, predict the noise added to the noisy colored image features corresponding to the sample, and based on the predicted noise, determine the denoised colored image features after the noisy colored image features are denoised, so as to obtain the corresponding colored image by decoding the denoised colored image features; wherein the color semantic features include the regional color features of each connected domain corresponding to the sample.
[0122] Step S2305: According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0123] In an embodiment of the present application, in order to further improve the denoising effect, for each sample, the same noise as the line drawing image feature can be added to the reference image feature corresponding to the sample to obtain a noisy reference image feature; based on the noisy reference image feature and color semantic feature corresponding to the sample, the noise added to the noisy colored image feature corresponding to the sample is predicted.
[0124] During the denoising process, the color semantic features of the colored reference image and the noisy reference image with the same noise added to the noisy line drawing features are used as conditions to guide denoising. This allows the reference image and the line drawing to be aligned under the same noise interference, thereby achieving more accurate color transfer and improving the accuracy of image coloring. It can effectively handle the coloring of objects with large movements such as rotation.
[0125] based on Figure 10 The training method of the image colorization network shown in the figure extracts the global color features of the coloring reference image, and further extracts the regional color features of each connected domain from the global color features. The noisy coloring image features are denoised based on the regional color features of each connected domain in the coloring reference image, so that each connected domain in the sample line drawing can be more accurately matched to the color semantics of the corresponding area in the coloring reference image, thereby achieving accurate coloring of each area of the line drawing and significantly improving the coloring effect.
[0126] It should be noted that the specific training process of the image colorization network is described in detail in the above steps S110 to S120, and this application will not repeat it here.
[0127] The embodiment of the present application also provides an image coloring method. The method can be based on Figure 2 The image coloring network trained by the training method of the image coloring network shown realizes the coloring of line drawings. The method can be executed by any electronic device, for example, by a server or a terminal.
[0128] Figure 11 : is a flow chart of the image coloring method provided in the embodiment of the present application, such as Figure 11 As shown, the method may include the following steps S310-S320, wherein: Step S310: Obtain the line drawing to be colored and the coloring reference drawing corresponding to the line drawing to be colored.
[0129] Among them, the line drawing to be colored is an image draft that outlines the contour and structure of the object with lines and has not yet been filled with color. The coloring reference image corresponding to the line drawing to be colored is a color image that contains the same object or scene as the line drawing to be colored, but has different posture, perspective or details of the object. It is used to provide guidance for coloring the objects in the line drawing.
[0130] Step S320: Based on the coloring reference image, a coloring operation is performed on the line drawing to be colored by the trained image coloring network to obtain a colored image corresponding to the line drawing to be colored.
[0131] The coloring operation includes steps S3201 to S3206.
[0132] Step S3201: extract features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features.
[0133] In an embodiment of the present application, the image colorization network may include a first encoding network, a color semantic extraction network, a denoising network, and a decoding network. The first encoding network may be used to extract features from the line drawing to be colored and the reference image to be colored, respectively, to obtain line drawing features and reference image features.
[0134] Optionally, the image colorization network may also include a first encoding network, a second encoding network, a color semantic extraction network, a denoising network, and a decoding network, wherein the first encoding network is used to extract features from colored images, and the second encoding network is used to extract features from uncolored line drawings. The first encoding network can be used to extract features from the colorized reference image to obtain reference image features, while the second encoding network can be used to extract features from the line drawing to be colored to obtain line drawing features.
[0135] The embodiments of this application do not restrict the network types of the encoding network (first encoding network, second encoding network) and the decoding network, and can be selected based on actual application requirements. As an alternative, the first encoding network in the embodiments of this application can use the encoder in a variational autoencoder (VAE), the second encoding network can use a sketch guide network (SG), and the decoding network can use a VAE decoder.
[0136] Step S3202: extract color semantic features from the colored reference image to obtain color semantic features.
[0137] In an embodiment of the present application, a line drawing of a coloring reference image may be obtained, and the line drawing of the coloring reference image may be divided into connected regions to obtain a plurality of connected domains; When performing color semantic feature extraction, global color feature extraction can be performed on the colored reference image first to obtain the global color feature of the colored reference image; for each connected domain in the multiple connected domains divided by the line drawing of the colored reference image, the local color feature corresponding to the connected domain in the global color feature is determined; for each connected domain, the regional color feature of the connected domain is determined based on the local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
[0138] Step S3203: adding noise with a maximum number of noise adding steps within a preset noise adding step number range to the line drawing feature to obtain a noisy line drawing feature.
[0139] The image colorization method provided in the present embodiment leverages the inverse generation capability of the diffusion process. This method first adds high-intensity noise to the line drawing features, making the image closer to pure noise. Colorization is then performed based on the constraints of the colorization reference image, avoiding interference from implicit colors in the initial line drawing features. For example, assuming a preset noise addition step range of [1, T], noise corresponding to the maximum noise addition step number T can be added to the line drawing features.
[0140] Optionally, when adding noise to line drawing features, you can use either single addition or iterative addition modes. In single addition, the added noise includes the noise accumulated over the number of noise addition steps, with the noise corresponding to each noise addition step being the sum of the noise accumulated over that number of noise addition steps. In iterative addition, only the noise for the current number of noise addition steps is added each time.
[0141] Step S3204: adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, to obtain the noisy reference image features corresponding to each denoising process.
[0142] Step S3205: Based on the noisy reference image features and color semantic features corresponding to each denoising, iteratively denoise the noisy line drawing image features to obtain denoised colored image features corresponding to the noisy line drawing image features.
[0143] In the process of iterative denoising of the noisy line drawing features, the color semantic features and the noisy reference image features corresponding to the current denoising can be used as reference conditions for the current denoising to perform the current denoising on the noisy line drawing features.
[0144] The embodiments of the present application do not limit the denoising strategy adopted. For example, a denoising diffusion probabilistic model (DDPM) or a denoising diffusion implicit model (DDIM) may be used to iteratively remove noise. Among them, if the DDPM method is used for iterative denoising, the value of the denoising step number t gradually decreases from the maximum denoising step number to 0. Assuming that the maximum denoising step number T is 999, the value of the denoising step number t gradually decreases from 999 to 0. The denoising steps corresponding to each denoising are 999, 998, 997..., and the corresponding denoising steps in the noisy reference image features corresponding to each denoising are 999, 998, 997..., and 1000 denoising steps are required; if the DDIM method is used for iterative denoising, the value of the denoising step number t used for each denoising is discrete. Taking the maximum denoising step number T as 999 as an example, the possible values are 999, 889,..., 0, and a total of 30 denoising steps are required.
[0145] In an embodiment of the present application, for each denoising in the iterative denoising process, the noisy reference image features corresponding to the current denoising can be first determined; based on the color semantic features of the colored reference image and the noisy reference image features corresponding to the current denoising, the noise of the number of denoising steps corresponding to the current denoising is predicted; according to the predicted noise and the noisy line drawing features targeted by the current denoising, the noisy line drawing features are denoised to obtain the line drawing features after the current denoising, wherein the denoised line drawing features are the noisy line drawing features targeted by the next denoising, and the denoised line drawing features obtained by the last denoising are used as the denoised colored image features.
[0146] Optionally, for each denoising in the iterative denoising process, when predicting the noise of the number of denoising steps corresponding to the current denoising, feature extraction can be performed on the noisy reference image features corresponding to the current denoising to obtain a first reference feature; the noisy line drawing image features targeted by the current denoising and the first reference features are spliced to obtain a splicing feature; based on the correlation between the noisy line drawing image features targeted by the current denoising and the splicing features, the noisy line drawing image features and the first reference features in the splicing features are weightedly fused to obtain a first fused feature; based on the color semantic features and the first fused features, the noise of the number of denoising steps corresponding to the current denoising is predicted.
[0147] Assuming that the number of denoising steps corresponding to the current denoising is t, the noisy reference image feature corresponding to the current denoising can be expressed as , the noisy reference image features Input into the RU network to obtain the first reference feature , the first reference feature Multiply them with W_k and W_v respectively to get K_ref and V_ref, and use the noisy line drawing features targeted by the current denoising as the Input into the denoising network and multiply with the network parameters W_q_DU, W_k_DU and matrix W_v_DU in the denoising network respectively. The new features obtained are recorded as Q_tar, K_tar and V_tar. After that, K_tar and K_ref are spliced to obtain the first spliced feature as the new key K, and V_tar and V_ref are spliced to obtain the second spliced feature as the new value V. The first fusion feature is calculated by the following formula of the spatial self-attention mechanism:
[0148] Afterwards, the query feature Q is determined based on the first fused feature, and the key values K and V are determined based on the color semantic feature. Through the cross-attention mechanism, the correlation between the color semantic feature and the first fused feature is used to weight the color semantic feature to obtain the predicted noise in the noisy line drawing feature targeted by the current denoising, that is, the noise of the denoising steps corresponding to the current denoising.
[0149] The following uses the denoising of the noisy line drawing feature Z(T) as an example to illustrate iterative denoising. In the first denoising process, the number of denoising steps in the noisy line drawing feature Z(T) is the maximum number of steps T, so the number of denoising steps corresponding to the first denoising is T. The same number of steps can be added to the color reference image feature to obtain the noisy reference image feature corresponding to the first denoising. , and The color semantic features are input into the reference feature guided network to obtain the further extracted dynamic spatial features , with dynamic spatial features The noise w(T) in the noisy line drawing feature Z(T) is predicted by the denoising network based on the color semantic features. The predicted noise is then removed from the noisy line drawing feature Z(T) to obtain a new noisy line drawing feature Z(T-1). For the noisy line drawing feature Z(T-1), the number of noise adding steps corresponding to the current denoising is re-determined to be T-1, and the noisy reference image feature corresponding to the current denoising is the one with the same number of steps added. , and The color semantic features are input into the reference feature guided network to obtain the further extracted dynamic spatial features , re-using dynamic spatial features The noise w(T-1) in the noisy line drawing feature Z(T-1) is predicted through the denoising network based on the color semantic features. The predicted noise w(T-1) is then removed from the noisy line drawing feature Z(T-1) to obtain the new noisy line drawing feature Z(T-2). And so on, until the number of denoising steps t is equal to 0, it means that the denoising is completed, and the denoised colored image feature Z (0) that has been denoised and meets the conditions is obtained.
[0150] Step S3206: Decode the features of the denoised colored image to obtain a colored image corresponding to the line drawing to be colored.
[0151] Figure 12 A schematic diagram of a process for image coloring provided in an embodiment of the present application, wherein the image coloring network includes a VAE encoder, a color guidance network CG, a line drawing guidance network SG, a reference feature guidance network RU, a denoising network RU, and a VAE decoder. When coloring a line drawing to be colored, the corresponding colored reference image can be input into the VAE encoder for encoding to obtain reference image features, the colored reference image and the line drawing corresponding to the colored reference image are input into the color guidance network CG to extract the color semantic features of the colored reference image, the line drawing to be colored is input into the line drawing guidance network SG to extract the line drawing features, and Gaussian noise is added to the line drawing features to obtain noisy line drawing features. The feature dimension of the added Gaussian noise is the same as the feature dimension of the line drawing features.
[0152] Afterwards, noise is dynamically added to the reference image features to obtain noisy reference image features, where the noisy reference image features change with the number of noise addition steps t added; the noisy reference image features and color semantic features are input into the reference feature guidance network RU to obtain dynamic spatial features.
[0153] Finally, the dynamic spatial features, color semantic features, and noisy line drawing features are input into the denoising network DU to obtain the predicted noise. The noisy line drawing features are iteratively denoised based on the denoising strategy until the number of denoising steps t is 0. The denoised colored image features after noise removal are obtained, and then decoded through the VAE decoder to obtain the colored image of the line drawing.
[0154] based on Figure 11 The image colorization method shown here, when coloring a line drawing based on a coloring reference image, can add noise equal to the maximum number of noise addition steps to the extracted line drawing features, and add noise equal to the number of noise addition steps corresponding to each denoising step to the reference image features. Based on the noisy reference image features and color semantic features corresponding to each denoising step, the noisy line drawing features are iteratively denoised to obtain denoised coloring image features. The denoised coloring image features are then decoded to obtain the coloring image corresponding to the line drawing. During each iterative denoising process, this method uses a noisy reference image with the same number of noise addition steps as the noisy line drawing features as a condition to guide the denoising process. This allows the reference image and the line drawing to be aligned under the same noise interference, thereby achieving more precise color transfer and improving the accuracy of image colorization.
[0155] To better understand and illustrate the method and its use value provided by the embodiments of this application, the solution provided by this application is described below in conjunction with specific application scenarios. Taking the coloring of anime line drawings in anime scenes as an example, the trained image colorization network is used to automatically color the anime line drawings to improve the efficiency of animation production.
[0156] Users can import the animation video to be produced into the AI coloring tool, such as Figure 13 As shown, Figure 13 The top row shows several lines of art from an animation video, ready for coloring. Line art in closer frames has minimal movement. Users can select keyframes from the displayed line art and manually color them, using them as references for coloring the line art within the corresponding unit duration.
[0157] After that, drag the colored keyframe line drawing (colored reference image) into the specified reference area and click Auto Color. The terminal can send the colored reference image and the line drawing corresponding to the colored reference image to the server. Based on the line drawing to be colored and the colored reference image, the server uses the deployed trained image colorization network to generate the colored images (colored draft results) corresponding to each line drawing, and returns them to the terminal for display, as shown in the last row of the figure.
[0158] It should be noted that, for the specific implementation details of the image coloring process, please refer to the detailed description in the above steps S110 to S120, and this application will not go into details here.
[0159] The embodiment of the present application also provides an image coloring method. The method can be based on Figure 10 The image coloring network trained by the training method of the image coloring network shown realizes the coloring of the line drawing, which can be specifically executed by any electronic device, for example, it can be executed by a server or a terminal.
[0160] Figure 14 : is a flow chart of the image coloring method provided in the embodiment of the present application, such as Figure 14 As shown, the method may include the following steps S410 to S430, wherein: Step S410: Obtain the line drawing to be colored, the coloring reference drawing of the line drawing to be colored, and the line drawing of the coloring reference drawing.
[0161] Step S420: performing connected region division on the line drawing of the colored reference image to obtain a plurality of connected domains.
[0162] In the embodiment of the present application, a flooding method can be used to divide connected regions. Specifically, starting from any key point in the line drawing (seed point), the surrounding pixels are continuously traversed until the black outline is reached. Pixels with the same attributes are clustered into independent connected domains, and multiple connected domains are divided. Each connected domain represents a basic colored unit in the image.
[0163] Step S430: Based on the coloring reference image, a coloring operation is performed on the line drawing image to be colored through the trained image coloring network to obtain a coloring image corresponding to the line drawing image.
[0164] The coloring operation includes steps S4301 to S4307.
[0165] Step S4301: extract features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features.
[0166] Step S4302: adding noise to the line drawing feature to obtain a noisy line drawing feature.
[0167] Step S4302: extract global color features from the colored reference image to obtain global color features.
[0168] Step S4304: for each connected domain in the plurality of connected domains, determine the local color feature corresponding to the connected domain in the global color feature.
[0169] Step S4305: For each connected domain, determine the regional color feature of the connected domain according to the local color feature corresponding to the connected domain.
[0170] Step S4306: Based on the reference image features and the color semantic features, the noisy line drawing image features are iteratively denoised to obtain the denoised colored image features corresponding to the noisy line drawing image features, wherein the color semantic features include the regional color features of each connected domain in the multiple connected domains.
[0171] Step S4307: Decoding is performed based on the denoised colored image features to obtain a colored image corresponding to the line drawing to be colored.
[0172] The image coloring method provided in the embodiment of the present application is based on the inverse generation capability of the diffusion process. It first adds noise with relatively high intensity to the line drawing features to make the image close to pure noise. Then, the image is colored based on the conditional constraints of the coloring reference image to avoid interference from implicit colors in the initial line drawing features.
[0173] Therefore, when adding noise to the line drawing feature, we can add noise with the maximum number of noise steps within the preset noise step range to obtain a noisy line drawing feature. For example, assuming the preset noise step range is [1, T], we can add noise corresponding to the maximum number of noise steps T to the line drawing feature.
[0174] In the embodiment of the present application, noise corresponding to the number of noise addition steps corresponding to each denoising step during the iterative denoising process may be added to the reference image features to obtain the noisy reference image features corresponding to each denoising step. During the denoising process, the noisy line drawing features are iteratively denoised based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain the denoised colored image features corresponding to the noisy line drawing features.
[0175] It should be noted that, for the specific implementation details of the image coloring process, please refer to the detailed description in the above steps S110 to S120, and this application will not go into details here.
[0176] Based on Figure 11 The same principle as the image coloring method shown in FIG, the embodiment of the present application provides an image coloring device, such as Figure 15 As shown, the image coloring device 500 may include: an image acquisition module 510 and an image coloring module 520, wherein: The image acquisition module 510 is used to acquire a line drawing to be colored and a coloring reference image corresponding to the line drawing to be colored; The image coloring module 520 is configured to perform the following coloring operations on the line drawing to be colored using a trained image coloring network based on the coloring reference image to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image to obtain color semantic features; Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; Iteratively denoising the noisy line drawing features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing features; The denoised colored image features are decoded to obtain a colored image corresponding to the line drawing to be colored.
[0177] Optionally, the image colorization module 520 may also be used to: Obtaining a line drawing of the coloring reference image; Performing connected region division on the line drawing of the coloring reference image to obtain a plurality of connected domains; The image colorization module 520 can be used to: Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, a regional color feature of the connected domain is determined according to a local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
[0178] Optionally, the image colorization module 520 may be used to: Based on the color semantic features and the noisy reference image features corresponding to the current denoising, predicting the noise of the number of denoising steps corresponding to the current denoising; Denoising the noisy line drawing features according to the predicted noise and the features of the noisy line drawing targeted by the current denoising, to obtain the features of the line drawing after the current denoising, which will be the features of the noisy line drawing targeted by the next denoising; The denoised colored image features are the denoised line drawing features obtained by the last denoising.
[0179] Optionally, the image colorization module 520 may be used to: Extract the features of the noisy reference image corresponding to the current denoising to obtain the first reference features; Splicing the noisy line drawing feature targeted by the current denoising and the first reference feature to obtain a spliced feature; According to the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature, weighted fusion is performed on the noisy line drawing feature and the first reference feature in the splicing feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, predict the noise of the number of denoising steps corresponding to the current denoising.
[0180] Optionally, the image colorization module 520 may be configured to perform any of the following: Based on the self-attention mechanism, the feature of the noisy reference image corresponding to the current denoising is enhanced to obtain the first reference feature; The color semantic feature is weighted according to the correlation between the noisy reference image feature corresponding to the current denoising and the color semantic feature; and a first reference feature is obtained based on the weighted color semantic feature.
[0181] Optionally, the image colorization module 520 may be used to: Determining a correlation matrix between the features of the noisy line drawing targeted by the current denoising and the stitching features, wherein the correlation matrix includes correlations between each region in the line drawing to be colored and each region in the line drawing to be colored, and correlations between each region in the line drawing to be colored and each region in the coloring reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix to obtain a target probability matrix; Determining an optimal probability matrix based on the target probability matrix, the optimal probability matrix including the matching probabilities between each region in the line drawing to be colored and each region in the line drawing to be colored, and the matching probabilities between each region in the line drawing to be colored and each region in the coloring reference image, wherein the optimal probability matrix is a binary probability matrix; The noisy line drawing feature in the splicing feature and the first reference feature are weightedly fused according to the optimal probability matrix to obtain a first fused feature.
[0182] Optionally, the image coloring device further includes a training device, which can be used to: Acquire multiple samples, each of the samples including a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain line drawing features and reference image features, and color semantic feature extraction is performed on the colored reference image in the sample to obtain color semantic features; For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; Based on the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0183] Based on Figure 14 The same principle as the image coloring method shown in FIG, the embodiment of the present application provides an image coloring device, such as Figure 16 As shown, the image coloring device 600 may include: an image acquisition module 610, a region division module 620 and an image coloring module 630, wherein: An image acquisition module 610 is configured to acquire a line drawing to be colored, a coloring reference image of the line drawing to be colored, and a line drawing of the coloring reference image; A region division module 620 is configured to divide the line drawing of the coloring reference image into connected regions to obtain a plurality of connected regions; The image coloring module 630 is configured to perform the following coloring operations on the line drawing to be colored using a trained image coloring network based on the coloring reference image to obtain a colored image corresponding to the line drawing: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; adding noise to the line drawing feature to obtain a noisy line drawing feature; Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, determining a regional color feature of the connected domain according to a local color feature corresponding to the connected domain; Iteratively denoising the noisy line drawing image features based on the reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features, wherein the color semantic features include regional color features of each connected domain in the multiple connected domains; Decoding is performed based on the features of the denoised colored image to obtain a colored image corresponding to the line drawing to be colored.
[0184] Optionally, the image colorization module 630 may be used to: Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; The image colorization module 630 can be used to: Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; The iterative denoising of the noisy line drawing image features based on the noisy reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features includes: Based on the noisy reference image features corresponding to each denoising and the color semantic features, the noisy line drawing image features are iteratively denoised to obtain denoised colored image features corresponding to the noisy line drawing image features.
[0185] Based on Figure 2 The training method of the image coloring network shown in FIG. 1 is the same principle. The embodiment of the present application provides a training device for an image coloring network, such as Figure 17 As shown, the image colorization network training device 700 may include: a sample acquisition module 710 and a training module 720, wherein: The sample acquisition module 710 is configured to acquire a plurality of samples, each of which includes a sample line drawing, a coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing; The training module 720 is configured to continuously perform the following training operations on the image colorization network to be trained based on the multiple samples until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature. Color semantic feature extraction is performed on the colored reference image in the sample to obtain the color semantic feature. For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; According to the difference between the noise added to each sample and the predicted noise, the network parameters of the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0186] Optionally, the image colorization network includes a first encoding network, a network to be trained, and a decoding network, the first encoding network and the decoding network are pre-trained networks, and the network to be trained includes a color semantic extraction network and a denoising network; The training module 720 can be used to: Extracting features of the sample line drawing and the colored reference drawing in the sample using the first encoding network to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image in the sample using the color semantic extraction network to obtain color semantic features; Based on the noisy reference image features and color semantic features of the sample, predicting the noise in the noisy colored image features of the sample through the denoising network, and determining the denoised colored image features after denoising the noisy colored image features based on the predicted noise; The network parameters of the network to be trained are adjusted according to the difference between the noise added to each sample and the predicted noise.
[0187] Optionally, the image colorization network includes a first encoding network, a network to be trained, and a decoding network, the first encoding network and the decoding network are pre-trained networks, and the network to be trained includes a second encoding network, a color semantic extraction network, and a denoising network; The training module 720 can be used to: Extracting features of the colored reference image in the sample through the first encoding network to obtain reference image features; The second encoding network is used to extract features of the sample line drawing in the sample to obtain line drawing features.
[0188] Optionally, each of the samples further includes a line drawing of a coloring reference image in the sample; The training module 720 can be used to: The line drawing of the colored reference image in each sample is divided into connected regions to obtain multiple connected domains; The color semantic feature extraction of the colored reference image in the sample to obtain the color semantic feature includes: Performing global color feature extraction on the colored reference image in the sample to obtain the global color feature of the colored reference image; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, a regional color feature of the connected domain is determined according to a local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
[0189] Optionally, the training module 720 may be used to: Randomly select a noise adding step number within a preset noise adding step number range, where different noise adding step numbers represent different noise adding degrees; Add noise corresponding to the selected number of noise addition steps to the reference image features and line drawing features of the sample respectively; The training module 720 can be used to: Based on the noisy reference image features, color semantic features and the selected number of noise addition steps of the sample, the noise in the noisy colored image features of the sample is predicted.
[0190] Optionally, the training module 720 may be used to: Extracting the noisy reference image features of the sample to obtain a first reference feature; Splicing the noisy colored image feature of the sample and the first reference feature to obtain a spliced feature; performing weighted fusion of the noisy colored image feature and the first reference feature in the spliced feature according to the correlation between the noisy colored image feature of the sample and the spliced feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, the noise in the noisy colored image feature of the sample is predicted.
[0191] Optionally, the network to be trained further includes a reference feature guided network; The training module 720 may be configured to perform any of the following: Guide the self-attention mechanism in the network through the reference feature to enhance the noisy reference image feature of the sample to obtain a first reference feature; According to the correlation between the noisy reference image feature of the sample and the color semantic feature, the color semantic feature is weighted by the reference feature-guided network, and a first reference feature is obtained based on the weighted color semantic feature.
[0192] Optionally, the training module 720 may be used to: Determining a correlation matrix between the noisy colored image features in the sample and the stitching features, wherein the correlation matrix includes correlations between regions in the sample line drawing in the sample and regions in the sample line drawing, and correlations between regions in the sample line drawing and regions in the colored reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix to obtain a target probability matrix; Determining an optimal probability matrix based on the target probability matrix, the optimal probability matrix including the matching probabilities between each region in the sample line drawing and each region in the sample line drawing, and the matching probabilities between each region in the sample line drawing and each region in the coloring reference image, wherein the optimal probability matrix is a binary probability matrix; The noisy colored image feature and the first reference feature in the splicing feature are weightedly fused according to the optimal probability matrix to obtain a first fused feature.
[0193] Based on the same principle as the training method of the image coloring network shown in Figure 10, the present embodiment provides a training device for an image coloring network, such as Figure 18 As shown, the image colorization network training device 800 may include: a sample acquisition module 810, a region division module 820 and a training module 830, wherein: The sample acquisition module 810 is configured to acquire a plurality of samples, each sample including a sample line drawing to be colored, a coloring reference drawing, a line drawing of the coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing; A region division module 820 is used to divide the line drawing of the colored reference image in each sample into connected regions to obtain multiple connected regions; The training module 830 is configured to continuously perform the following training operations on the image colorization network to be trained based on the multiple samples until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature; For each sample, extract global color features from the colored reference image in the sample to obtain global color features; for each connected domain corresponding to the sample, determine the local color features corresponding to the connected domain in the global color features; for each connected domain, determine the regional color features of the connected domain based on the local color features corresponding to the connected domain; For each sample, add noise to the line drawing feature of the sample to obtain a noisy line drawing feature, and fuse the noisy line drawing feature corresponding to the sample with the standard image feature to obtain a noisy colored image feature; For each sample, based on the reference image features and color semantic features corresponding to the sample, predict the noise added to the noisy colored image features corresponding to the sample, and determine the denoised colored image features after denoising the noisy colored image features based on the predicted noise, so as to obtain the corresponding colored image by decoding the denoised colored image features; wherein the color semantic features include the regional color features of each connected domain corresponding to the sample; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
[0194] Optionally, the training module 830 may also be used to: For each sample, adding the noise to the reference image feature corresponding to the sample to obtain a noisy reference image feature; The training module 830 can be used to: Based on the noisy reference image features and color semantic features corresponding to the sample, the noise added to the noisy colored image features corresponding to the sample is predicted.
[0195] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or portion of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal. It can be implemented in whole or in part using software, hardware (such as processing circuits or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functionality of the module or unit.
[0196] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program stored in the memory, the method in any optional embodiment of the present application can be implemented.
[0197] Figure 19 FIG. 1 shows a schematic structural diagram of an electronic device to which an embodiment of the present invention is applicable. Figure 19 As shown, the electronic device may be a server or a terminal, and the electronic device may be used to implement the method provided in any embodiment of the present invention.
[0198] like Figure 19 As shown in FIG, the electronic device 2000 may mainly include at least one processor 2001 ( Figure 19), memory 2002, communication module 2003 and input / output interface 2004 and other components, optionally, the components can be connected and communicated through a bus 2005. It should be noted that, Figure 19 The structure of the electronic device 2000 shown in the figure is merely illustrative and does not constitute a limitation on the electronic devices to which the method provided in the embodiments of the present application is applicable.
[0199] Memory 2002 can be used to store an operating system and application programs, etc. Application programs can include computer programs that implement the methods described in the embodiments of the present invention when called by processor 2001, and can also include programs for implementing other functions or services. Memory 2002 can be, but is not limited to, ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices that can store information and computer programs, EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.
[0200] Processor 2001 is connected to memory 2002 via bus 2005 and implements corresponding functions by calling application programs stored in memory 2002. Processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 2001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0201] Electronic device 2000 can connect to a network via a communication module 2003 (which may include, but is not limited to, components such as a network interface) to communicate with other devices (such as a terminal or server) over the network to implement data exchange, such as sending data to or receiving data from other devices. Communication module 2003 may include a wired network interface and / or a wireless network interface, meaning that the communication module may include at least one of a wired communication module and a wireless communication module.
[0202] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 2004. The electronic device 2000 itself can have a display device, and can also be connected to other external display devices through the interface 2004. Optionally, a storage device, such as a hard disk, can also be connected through the interface 2004, so that data in the electronic device 2000 can be stored in the storage device, or data in the storage device can be read, and data in the storage device can also be stored in the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 2004 can be a component of the electronic device 2000, or it can be an external device connected to the electronic device 2000 when needed.
[0203] Bus 2005, used to connect the various components, may include a path for transmitting information between the components. Bus 2005 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Depending on their function, bus 2005 can be categorized as an address bus, a data bus, or a control bus.
[0204] Optionally, for the solution provided in the embodiment of the present invention, the memory 2002 can be used to store a computer program for executing the solution of the present invention, and be run by the processor 2001. When the processor 2001 runs the computer program, the actions of the method or device provided in the embodiment of the present invention are implemented.
[0205] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.
[0206] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.
[0207] It should be noted that the terms "first," "second," "third," "fourth," "1," "2," and so on (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described.
[0208] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0209] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. An image coloring method, characterized in that: The method comprises: Obtaining a line drawing to be colored and a coloring reference drawing corresponding to the line drawing to be colored; Based on the coloring reference image, the trained image coloring network performs the following coloring operations on the line drawing to be colored to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image to obtain color semantic features; Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; Iteratively denoising the noisy line drawing features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing features; The denoised colored image features are decoded to obtain a colored image corresponding to the line drawing to be colored.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining a line drawing of the coloring reference image; Performing connected region division on the line drawing of the coloring reference image to obtain a plurality of connected domains; The color semantic feature extraction of the colored reference image to obtain the color semantic feature includes: Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, a regional color feature of the connected domain is determined according to a local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
3. The method according to claim 1 or 2, characterized in that The iterative denoising of the noisy line drawing image features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features includes: Based on the color semantic features and the noisy reference image features corresponding to the current denoising, predicting the noise of the number of denoising steps corresponding to the current denoising; Denoising the noisy line drawing features according to the predicted noise and the features of the noisy line drawing targeted by the current denoising, to obtain the features of the line drawing after the current denoising, which will be the features of the noisy line drawing targeted by the next denoising; The denoised colored image features are the denoised line drawing features obtained by the last denoising.
4. The method according to claim 3, characterized in that The predicting the noise of the number of denoising steps corresponding to the current denoising based on the color semantic features and the noisy reference image features corresponding to the current denoising includes: Extract the features of the noisy reference image corresponding to the current denoising to obtain the first reference features; Splicing the noisy line drawing feature targeted by the current denoising and the first reference feature to obtain a spliced feature; According to the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature, weighted fusion is performed on the noisy line drawing feature and the first reference feature in the splicing feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, predict the noise of the number of denoising steps corresponding to the current denoising.
5. The method according to claim 4, characterized in that The extracting the feature of the noisy reference image corresponding to the current denoising to obtain the first reference feature includes any one of the following: Based on the self-attention mechanism, the feature of the noisy reference image corresponding to the current denoising is enhanced to obtain the first reference feature; weighting the color semantic feature according to the correlation between the noisy reference image feature corresponding to the current denoising and the color semantic feature; Based on the weighted color semantic features, a first reference feature is obtained.
6. The method according to claim 4, characterized in that The step of performing weighted fusion on the noisy line drawing feature and the first reference feature in the splicing feature according to the correlation between the noisy line drawing feature targeted by the current denoising and the splicing feature to obtain the first fused feature includes: Determining a correlation matrix between the features of the noisy line drawing targeted by the current denoising and the stitching features, wherein the correlation matrix includes correlations between each region in the line drawing to be colored and each region in the line drawing to be colored, and correlations between each region in the line drawing to be colored and each region in the coloring reference image; Normalizing the correlation matrix by rows to obtain a first probability matrix; normalizing the correlation matrix by columns to obtain a second probability matrix; Combining the first probability matrix and the second probability matrix to obtain a target probability matrix; Determining an optimal probability matrix based on the target probability matrix, the optimal probability matrix including the matching probabilities between each region in the line drawing to be colored and each region in the line drawing to be colored, and the matching probabilities between each region in the line drawing to be colored and each region in the coloring reference image, wherein the optimal probability matrix is a binary probability matrix; The noisy line drawing feature in the splicing feature and the first reference feature are weightedly fused according to the optimal probability matrix to obtain a first fused feature.
7. The method according to any one of claims 1 to 6, characterized in that The trained image colorization network is obtained by training in the following way: Acquire multiple samples, each of the samples including a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain line drawing features and reference image features, and color semantic feature extraction is performed on the colored reference image in the sample to obtain color semantic features; For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; Based on the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
8. An image coloring method, characterized in that: The method comprises: Obtaining a line drawing to be colored, a coloring reference drawing of the line drawing to be colored, and a line drawing of the coloring reference drawing; Performing connected region division on the line drawing of the coloring reference image to obtain a plurality of connected domains; Based on the coloring reference image, the trained image coloring network performs the following coloring operations on the line drawing to be colored to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Adding noise to the line drawing feature to obtain a noisy line drawing feature; Performing global color feature extraction on the colored reference image to obtain global color features; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, determining a regional color feature of the connected domain according to a local color feature corresponding to the connected domain; Iteratively denoising the noisy line drawing image features based on the reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features, wherein the color semantic features include regional color features of each connected domain in the multiple connected domains; Decoding is performed based on the features of the denoised colored image to obtain a colored image corresponding to the line drawing to be colored.
9. The method according to claim 8, characterized in that Adding noise to the line drawing feature to obtain the noisy line drawing feature includes: Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; The method further comprises: Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; The iterative denoising of the noisy line drawing image features based on the noisy reference image features and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing image features includes: Based on the noisy reference image features corresponding to each denoising and the color semantic features, the noisy line drawing image features are iteratively denoised to obtain denoised colored image features corresponding to the noisy line drawing image features.
10. A training method for an image colorization network, characterized in that: include: Acquire multiple samples, each of the samples including a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature. Color semantic feature extraction is performed on the colored reference image in the sample to obtain the color semantic feature. For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
11. The method according to claim 10, characterized in that Each of the samples also includes a line drawing of a coloring reference image in the sample; The method further comprises: The line drawing of the colored reference image in each sample is divided into connected regions to obtain multiple connected domains; The color semantic feature extraction of the colored reference image in the sample to obtain the color semantic feature includes: Performing global color feature extraction on the colored reference image in the sample to obtain the global color feature of the colored reference image; For each connected domain in the plurality of connected domains, determining a local color feature corresponding to the connected domain in the global color feature; For each of the connected domains, a regional color feature of the connected domain is determined according to a local color feature corresponding to the connected domain, wherein the color semantic feature includes the regional color feature of each connected domain.
12. The method according to claim 10, characterized in that For each sample, the same noise is added to the reference image feature and the line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature, including: Randomly select a noise adding step number within a preset noise adding step number range, where different noise adding step numbers represent different noise adding degrees; Adding noise corresponding to the selected number of noise addition steps to the reference image feature and the line drawing feature of the sample, respectively, to obtain a noisy reference image feature and a noisy line drawing feature; The step of predicting noise in the noisy colored image features of the sample based on the noisy reference image features and the color semantic features of the sample includes: Based on the noisy reference image features, color semantic features and the selected number of noise addition steps of the sample, the noise in the noisy colored image features of the sample is predicted.
13. The method according to claim 10, characterized in that The step of predicting noise in the noisy colored image features of the sample based on the noisy reference image features and the color semantic features of the sample includes: Extracting the noisy reference image features of the sample to obtain a first reference feature; Splicing the noisy colored image feature of the sample and the first reference feature to obtain a spliced feature; performing weighted fusion of the noisy colored image feature and the first reference feature in the spliced feature according to the correlation between the noisy colored image feature of the sample and the spliced feature to obtain a first fused feature; Based on the color semantic feature and the first fusion feature, the noise in the noisy colored image feature of the sample is predicted.
14. The method according to claim 13, wherein: The extracting the feature of the noisy reference image of the sample to obtain the first reference feature includes any one of the following: Based on the self-attention mechanism, the feature of the noisy reference image of the sample is enhanced to obtain the first reference feature; The color semantic feature is weighted according to the correlation between the noisy reference image feature of the sample and the color semantic feature, and a first reference feature is obtained based on the weighted color semantic feature.
15. A training method for an image colorization network, characterized in that: The method comprises: Acquire multiple samples, each sample including a sample line drawing to be colored, a coloring reference drawing, a line drawing of the coloring reference drawing, and standard image features obtained by extracting features from a standard coloring image corresponding to the sample line drawing; The line drawing of the colored reference image in each sample is divided into connected regions to obtain multiple connected domains; Based on the multiple samples, the following training operations are continuously performed on the image colorization network to be trained until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature; For each sample, extract global color features from the colored reference image in the sample to obtain global color features; for each connected domain corresponding to the sample, determine the local color features corresponding to the connected domain in the global color features; for each connected domain, determine the regional color features of the connected domain based on the local color features corresponding to the connected domain; For each sample, add noise to the line drawing feature of the sample to obtain a noisy line drawing feature, and fuse the noisy line drawing feature corresponding to the sample with the standard image feature to obtain a noisy colored image feature; For each sample, based on the reference image features and color semantic features corresponding to the sample, predict the noise added to the noisy colored image features corresponding to the sample, and determine the denoised colored image features after denoising the noisy colored image features based on the predicted noise, so as to obtain the corresponding colored image by decoding the denoised colored image features; wherein the color semantic features include the regional color features of each connected domain corresponding to the sample; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
16. The method according to claim 15, characterized in that The training operation further includes: For each sample, adding the noise to the reference image feature corresponding to the sample to obtain a noisy reference image feature; The step of predicting the noise added to the noisy colored image feature corresponding to the sample based on the reference image feature and the color semantic feature corresponding to the sample includes: Based on the noisy reference image features and color semantic features corresponding to the sample, the noise added to the noisy colored image features corresponding to the sample is predicted.
17. An image coloring device, characterized in that: The device comprises: An image acquisition module is used to acquire a line drawing to be colored and a coloring reference image corresponding to the line drawing to be colored; An image coloring module is configured to perform the following coloring operations on the line drawing to be colored based on the coloring reference image using a trained image coloring network to obtain a colored image corresponding to the line drawing to be colored: Extracting features of the line drawing to be colored and the reference drawing to be colored respectively to obtain line drawing features and reference drawing features; Extracting color semantic features from the colored reference image to obtain color semantic features; Adding noise with a maximum number of noise addition steps within a preset noise addition step number range to the line drawing feature to obtain a noisy line drawing feature; Adding noise corresponding to the number of denoising steps in each denoising process to the reference image features, respectively, to obtain the noisy reference image features corresponding to each denoising process; Iteratively denoising the noisy line drawing features based on the noisy reference image features corresponding to each denoising step and the color semantic features to obtain denoised colored image features corresponding to the noisy line drawing features; The denoised colored image features are decoded to obtain a colored image corresponding to the line drawing to be colored.
18. A training device for an image colorization network, characterized in that: include: A sample acquisition module is used to acquire multiple samples, each of which includes a sample line drawing, a coloring reference drawing, and a standard drawing feature obtained by extracting features from a standard coloring drawing corresponding to the sample line drawing; A training module is configured to continuously perform the following training operations on the image colorization network to be trained based on the multiple samples until a preset training end condition is met, thereby obtaining a trained image colorization network: For each sample, feature extraction is performed on the sample line drawing and the colored reference image in the sample to obtain the line drawing feature and the reference image feature. Color semantic feature extraction is performed on the colored reference image in the sample to obtain the color semantic feature. For each sample, add the same noise to the reference image feature and line drawing feature of the sample to obtain the noisy reference image feature and the noisy line drawing feature. Then, fuse the noisy line drawing feature and the standard image feature of the sample to obtain the noisy colored image feature. For each sample, predicting noise in a noisy colored image feature of the sample based on the noisy reference image feature and the color semantic feature of the sample, and determining a denoised colored image feature after denoising the noisy colored image feature based on the predicted noise, so as to obtain a corresponding colored image by decoding the denoised colored image feature; According to the difference between the noise added to each sample and the predicted noise, the network parameters in the image colorization network are adjusted to obtain the image colorization network based on which the next training operation is performed.
19. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 7, claims 8 to 9, claims 10 to 14, or claims 15 to 16.
20. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7, claims 8 to 9, claims 10 to 14, or claims 15 to 16.