Systems and methods for training an image colorization model

By using neighborhood color loss function in the image shading model, the problem of image shading model generating less saturation and brightness in the prior art is solved, and more efficient grayscale image shading is achieved to generate more saturated and brighter images.

CN113906468BActive Publication Date: 2025-05-30GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980097174.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-05-07
Publication Date
2025-05-30
Estimated Expiration
2039-05-07

AI Technical Summary

Technical Problem

Prior art When training an image shading model, the use of an absolute pixel color loss function results in the image generated by the model having less saturation and bright colors, and the inability to effectively handle a variety of different but acceptable solutions in grayscale images.

Method used

The neighborhood color loss function is used to train the image shading model, which evaluates the difference in color distance between pixels in the predicted color map and the corresponding pixels in the benchmark truth color map, avoiding penalizing acceptable shading solutions that are different from the benchmark truth.

Benefits of technology

By using the neighborhood color loss function, the shading model can generate more saturated and vivid images, effectively processing a variety of acceptable solutions in grayscale images, improving the effect of image shading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113906468B_ABST
    Figure CN113906468B_ABST
Patent Text Reader

Abstract

A method for training an image coloring model may include inputting a training input image into the coloring model and receiving a predicted color map as an output of the coloring model. A first color distance between a first pixel of the predicted color map and a second pixel of the predicted color map may be calculated. A second color distance between a third pixel included in a ground truth color map and a fourth pixel included in a ground truth coloring map may be calculated. The third pixel and the fourth pixel included in the ground truth color map may spatially correspond to the first pixel and the second pixel included in the predicted color map, respectively. The method may include adjusting parameters associated with the coloring model based on a neighborhood color loss function that evaluates a difference between the first color distance and the second color distance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to image processing using machine learning models, and more particularly, the present disclosure relates to systems and methods for training an image coloring model. Background Art

[0002] Typically, previous attempts to color grayscale images have resulted in dull, unsaturated colors or involved a significant amount of human interaction or supervision. In particular, some previous approaches use an absolute pixel color loss function to train a machine learning model, which directly compares the color values between the colored image and the ground truth image. This approach is undesirable because it can penalize the model for generating acceptable, realistic solutions that happen to not match the ground truth image. In other words, the absolute pixel color loss function does not consider the potential for multiple different but acceptable solutions for a given grayscale image. This hinders the training of models using such a loss function, resulting in images with less saturation and vivid colors. In particular, by using a loss function that penalizes the absolute color difference, the model may attempt to "split the difference" for an object with multiple different possible solutions by providing less saturated and vivid colors. For example, if the training set includes multiple images of apples, including some images where the apples are red and some images where the apples are green, the model may learn to provide a neutral color for the apples to minimize the overall loss across the entire training set. Additionally, the predicted neutral color may not accurately reflect the possible ground truth colors of the apples. Thus, a better method for training coloring models in the art would be welcome. Summary of the Invention

[0003] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned by practice of the embodiments.

[0004] One example aspect of the present disclosure relates to a method for training an image coloring model. The method may include inputting, by one or more computing devices, a training input image into a coloring model configured to receive the training input image and process the training input image to output a predicted color map depicting the predicted coloring of the training input image. The method may include receiving, by one or more computing devices, the predicted color map as an output of the coloring model. The method may include computing, by one or more computing devices, a first color distance between a first pixel included in the predicted color map and a second pixel included in the predicted color map. The method may include computing, by one or more computing devices, a second color distance between a third pixel included in a ground truth color map and a fourth pixel included in a ground truth coloring map. The third pixel and the fourth pixel included in the ground truth color map spatially correspond to the first pixel and the second pixel included in the predicted color map, respectively. The method may include evaluating, by one or more computing devices, a neighborhood color loss function that evaluates a difference between the first color distance and the second color distance. The method may include adjusting, by one or more computing devices, parameters associated with the coloring model based on the neighborhood color loss function.

[0005] Another example aspect of the present disclosure relates to a computing system including a coloring model configured to receive a training input image and, in response to receiving the training input image, output a predicted color map depicting the predicted coloring for the training input image. The computing system may include one or more processors and one or more non-transitory computer-readable media storing instructions jointly, which when executed by the one or more processors cause the computing system to perform operations. The operations may include: inputting the training input image into the coloring model; receiving the predicted color map as an output of the coloring model; computing a first color distance between a first pixel included in the predicted color map and a second pixel included in the predicted color map; computing a second color distance between a third pixel included in a ground truth color map and a fourth pixel included in a ground truth coloring map. The third pixel and the fourth pixel included in the ground truth color map may spatially correspond to the first pixel and the second pixel included in the predicted color map, respectively. The operations may include evaluating a neighborhood color loss function that evaluates a difference between the first color distance and the second color distance; and adjusting parameters associated with the coloring model based on the neighborhood color loss function.

[0006] Another example aspect of the present disclosure relates to a computing system including a coloring model configured to receive an input image and, in response to receiving the input image, output a predicted color map that describes the predicted coloring of the input image. The coloring model may have been trained based on a neighborhood color loss function that evaluates a difference between a first color distance and a second color distance. The first color distance may have been computed between a first pixel included in a training predicted color map output by the coloring model during training and a second pixel included in the training predicted color map. The second color distance may be between a third pixel included in a ground truth color map and a fourth pixel included in the ground truth color map. The third pixel and the fourth pixel included in the ground truth color map may spatially correspond to the first pixel and the second pixel included in the training predicted color map, respectively. The computing system may include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include inputting the input image into the coloring model; receiving the predicted color map as an output of the coloring model; and generating an output image based on the input image and the predicted color map.

[0007] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0008] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Example drawings are attached. A brief description of the drawings is as follows:

[0010] Figure 1A A block diagram of an example computing system for training an image coloring model in accordance with an example embodiment of the present disclosure is depicted.

[0011] Figure 1B A block diagram of an example computing device in accordance with an example embodiment of the present disclosure is depicted.

[0012] Figure 1C A block diagram of an example computing device in accordance with an example embodiment of the present disclosure is depicted.

[0013] Figure 2 A coloring model configured to generate a predicted color map based on an input image in accordance with an example embodiment of the present disclosure is depicted.

[0014] Figure 3A A coloring model including a polynomial prediction model and a refinement model in accordance with an example embodiment of the present disclosure is depicted.

[0015] Figure 3B depicts an example system for training a coloring model based on a neighborhood color loss function according to an example embodiment of the present disclosure.

[0016] Figure 4 depicts a training configuration for a coloring model including a polynomial prediction model and a refinement model based on a neighborhood color loss function and an additional loss function according to an example embodiment of the present disclosure.

[0017] Figure 5 depicts a training configuration for a discriminator model according to an example embodiment of the present disclosure.

[0018] Figure 6 depicts a flowchart of an example method for training a coloring model according to an example embodiment of the present disclosure.

[0019] Reference numerals repeated across multiple figures are intended to identify the same features in various implementations. Detailed Description

[0020] Overview

[0021] Generally, the present disclosure relates to systems and methods for coloring grayscale images, and more particularly to systems and methods for training a coloring model. In particular, aspects of the present disclosure relate to using a neighborhood color loss function to train a coloring model. The neighborhood color loss function can reward a coloring loss model for predicting pixel color values that have a correct magnitude distance from the corresponding color values of some or all of the other pixels in the image, rather than focusing on whether the coloring model has correctly predicted the actual ground truth color value itself. In this way, the coloring model can be enabled to produce vivid colors for objects that may have multiple correct recoloring solutions.

[0022] More specifically, some objects (e.g., vehicles, flowers, etc.) can be multiple colors. A grayscale image does not provide an indication of the original colors of these objects. Therefore, coloring a grayscale image is not a deterministic solution because there may be multiple acceptable or "correct" solutions for a single input image. Penalizing a coloring model for coloring such objects with colors different from the ground truth image is unproductive and may result in dull, unsaturated results. For example, penalizing a coloring model for predicting a vehicle in a grayscale image as red when the vehicle in the ground truth image is blue may be counterproductive. Therefore, to better train a coloring model, a neighborhood color loss function can be used to avoid penalizing (or at least offset the effect of the penalty) acceptable coloring solutions that differ from the ground truth coloring. Thus, a coloring model using a neighborhood color loss function can generate more saturated and vivid colored images.

[0023] The neighborhood color loss function can be calculated as the difference between the relative first color distance of the predicted image and the second color distance of the ground truth image. More specifically, the first color distance between a first and a second pixel included in the predicted color map output by the coloring model of the training input image can be calculated. The second color distance between a third pixel and a fourth pixel included in the ground truth color map corresponding to the ground truth image can be calculated. The third and fourth pixels in the ground truth color map spatially correspond to the first and second pixels in the predicted color map, respectively. The neighborhood color loss function can evaluate the difference between the first color distance and the second color distance. The parameters (e.g., weights) associated with the coloring model can be adjusted based on the neighborhood color loss function.

[0024] The neighborhood color loss function can be calculated for some or all of the pixels of the predicted color map. It should be understood that the neighborhood color loss function can be calculated iteratively for multiple "first pixels". For example, the neighborhood color loss function can be calculated for each pixel of the predicted color map.

[0025] Multiple suitable techniques can be used to select the first pixel of the predicted color map. The (multiple) first pixels can be randomly selected in the predicted color map (e.g., anywhere in the predicted color map or anywhere defined within certain portions of the predicted color map (e.g., within identified features or objects)). The (multiple) first pixels can be selected based on their position in the color map according to a shape or pattern (e.g., grid, square, circle, etc.). The (multiple) first pixels can be selected based on objects or features detected within the training input image. Some or all of the pixels of one or more detected features can be iteratively selected as the first pixel.

[0026] The neighborhood color loss function can be calculated for multiple second pixels of the predicted color map. Multiple suitable techniques (including random or systematic selection) can be used to select the second pixels. For example, the second pixels can be iteratively selected as each pixel of the predicted color map (except the first pixel) such that a "neighborhood color loss map" is generated for the first pixel with respect to the remaining pixels of the predicted color map. Alternatively, the second pixels can be selected based on their relative position with respect to the first pixel. For example, one or more pixels directly adjacent to the first pixel can be selected as the (multiple) second pixels. As another example, one or more pixels spaced apart from the first pixel by a set distance or arranged in a pattern (e.g., circle, square, grid, etc.) with respect to the first pixel can be selected.

[0027] In some implementations, the coloring model can employ multinomial classification to generate a predicted color map. More specifically, the coloring model can include a multinomial prediction model and a refinement model. The multinomial prediction model can be configured to receive a training input image and output a multinomial color distribution that describes multiple colorings of the training input image in a color space (such as a discretized color space). The discretized color space can include multiple color intervals (e.g., n×n color intervals) defined within a color space (e.g., CIELAB color space). For example, the color intervals can be selected to cover or include the colors that can be displayed on a display screen, such as the in-gamut colors in the RGB color space. The size of the color intervals can be selected considering the desired color resolution in the discretized color space and / or the desired simplicity of training the coloring model (e.g., reducing the required computational resources). For example, the color space can be discretized into 10x10 color intervals, resulting in approximately 310 discrete in-gamut colors (e.g., ab pairs in the CIELAB color space).

[0028] The multinomial prediction model can include an encoder model and a decoder model in an autoencoder configuration. The training input image can be input into the encoder, and the multinomial distribution can be received as the output of the decoder model. The multinomial prediction model can include at least one skip connection between the layers of the encoder model and the layers of the decoder model. Such a skip connection can transfer useful information from the hidden layers of the encoder model to the layers of the decoder model. This can facilitate backpropagation during the training of the multinomial prediction model.

[0029] The refinement model can be configured to receive the multinomial color distribution and output a predicted color map. An output image can be generated based on the predicted color map and the training input image (e.g., a luminance map that describes the luminance of the training input image). As described above, the multinomial color distribution can describe multiple colorings of the training input image. For example, the multinomial color distribution can include the corresponding color distribution for each pixel of the training input image. The refinement model can be configured to combine two or more colorings among the multiple colorings of the multinomial color distribution. The parameters of the refinement model can be adjusted based on a neighborhood color loss function such that the refinement model is trained to output a predicted color map that minimizes the neighborhood color loss function.

[0030] In some implementations, one or more additional loss functions can be used to train the coloring model. The additional loss functions can be defined with respect to the predicted color map, the multinomial color distribution, and / or the ground truth color map of the training input image. For example, a total loss function can be defined that includes the neighborhood color loss function and one or more additional loss functions. For example, an absolute color loss function can be used, which evaluates the difference (e.g., color distance) between the predicted color map and the ground truth color map.

[0031] As another example, a refined softmax loss function can be employed, which evaluates the difference between a polynomial color distribution and a predicted color map. The polynomial color distribution and the predicted color map can be encoded in different color spaces. For example, the polynomial color distribution can be encoded in a discretized color space, but the predicted color map may not necessarily be encoded in the discretized color space. Thus, to compare the predicted color map and the polynomial distribution, the method can include encoding the predicted color map into the discretized color space (e.g., encoding into multiple color intervals) to generate a discretized predicted color map. Thus, the refined softmax loss function can describe the difference between the polynomial color distribution in the discretized color space and the discretized predicted color map.

[0032] Encoding the predicted color map in the discretized color space can be achieved in a variety of ways. For example, the predicted color map can be "one hot" encoded in the discretized color space. Such encoding can include selecting a single color interval (e.g., the "closest" color interval) of the discretized color space for each pixel of the predicted color map. As another example, the predicted color map can be soft encoded in the discretized color space. Soft encoding can include representing each pixel as a distribution over two or more color intervals of the discretized color space.

[0033] As another example, the parameters of a polynomial prediction model can be adjusted based on a polynomial softmax loss function. The polynomial softmax loss function can evaluate the difference between a ground truth color map and a polynomial color distribution output by the polynomial prediction model. However, the ground truth color map and the polynomial color distribution may be encoded in different color spaces. Thus, the method can include discretizing the color space associated with the ground truth color map to generate a discretized color space including multiple color intervals. The ground truth color map can be encoded into the discretized color space to generate a discretized ground truth color map. The polynomial softmax loss function can describe the difference between the discretized ground truth color map and the polynomial color distribution.

[0034] In some implementations, the coloring model can be trained as part of a generative adversarial network (GAN). For example, the discriminator model can be configured to receive the output of the coloring model (e.g., a predicted color map and / or a colored image). In response to receiving the output of the coloring model, the discriminator model can output a discriminator loss that evaluates a score regarding the output of the coloring model compared to a ground truth image. For example, the discriminator loss can include a binary score that indicates whether the discriminator model identifies the predicted color map or output image as the ground truth image or corresponding to the ground truth image ("true"), or identifies the predicted color map or output image as a colored image ("false"). However, in some implementations, in addition to the binary score, the discriminator loss can include a probability distribution or confidence score relative to the above identification.

[0035] The discriminator model can be trained to better identify the ground truth image and distinguish the ground truth image from the colored image. During such training, a mixture of the ground truth image and the predicted image can be input into the discriminator model. The discriminator model can be trained based on a loss function that penalizes the model for misidentifying the predicted image.

[0036] The training of the discriminator model can be alternated with the training of the coloring model. For example, a first training phase can include training the coloring model, while a second training phase can include training the discriminator model. The first and second training phases can be alternated for a total number of training iterations, or until one or more performance criteria of the discriminator and / or coloring model are met. An example criterion for the coloring model is a threshold percentage of "fooling" the discriminator model's predicted color map. An example criterion for the discriminator model is a threshold percentage of correctly identifying the color map that is input into the discriminator as the predicted color map or the ground truth color map.

[0037] During the training of the coloring model, the predicted color map can be input into the discriminator model, and the discriminator loss can be received as the output of the discriminator model. Similar to the discriminator loss, the discriminator loss can include a binary indicator or distribution that describes whether the discriminator model classifies the predicted color map input into the discriminator model as a predicted coloring or a ground truth image. The coloring model can be trained based on the discriminator loss to generate a predicted color map that "fools" the discriminator model such that the discriminator model classifies the predicted color map as the ground truth image.

[0038] In some implementations, the coloring model can utilize object or feature detection or recognition to improve the training of the coloring model. For example, the coloring model can include a feature detection model that is configured to receive a training input image as input and output a plurality of feature representations that describe the locations of objects or features within the training input image. The feature detection model can be separate from and / or included in the polynomial prediction model. The feature representations can include bounding objects (e.g., boxes, points, pixel masks, etc.) that describe the locations of one or more features or objects identified or detected within the training input image.

[0039] In some embodiments, the feature representations can include labels (e.g., classes) of the identified objects. The feature representations can be input into a refinement model and / or a polynomial prediction model such that semantic relationships between objects and colors can be learned. For example, it can be learned that an octagonal sign is typically a red stop sign. However, in other embodiments, the feature representations can be without such labels. In such embodiments, the feature representations can simply identify the boundaries between various shapes or features depicted in the training input image.

[0040] Various transformations or encodings can be performed by the coloring model, the discriminator model, and / or on their inputs and / or outputs. The input image and / or the output image can be in the RGB color space. As described above, the coloring model can generally operate in the CIELAB color space. Additionally, in some embodiments, the discriminator model can operate in the RGB color space. Thus, conversions between the RGB color space and the CIELAB color space can be performed as needed. However, generally speaking, the various models and inputs / outputs described herein can operate or be represented in any number of different color spaces (including RGB, CIELAB, HSV, HSL, CMYK, etc.).

[0041] As an example, the systems and methods of the present disclosure can be included in or otherwise employed in the context of an application, a browser plugin, or other contexts. Thus, in some implementations, the models of the present disclosure can be included in or otherwise stored and implemented by a user computing device such as a laptop computer, a tablet computer, or a smart phone. As another example, the models can be included in or otherwise stored and implemented by a server computing device that communicates with the user computing device according to a client-server relationship. For example, the models can be implemented by the server computing device as part of a web service (e.g., an image coloring service).

[0042] The systems and methods of the present disclosure provide numerous technical effects and benefits. For example, compared to prior art systems and methods, the implementations described herein can provide perceptually more realistic coloring. For a computer-implemented image discriminator model configured to classify an image as "real" or "fake," the coloring may be substantially indistinguishable from the ground truth image. Additionally, compared to prior art methods, the systems and methods described herein may require less human supervision and / or manual input. Further, a model trained based on aspects of the present disclosure can replace multiple prior art machine learning models or more complex prior art machine learning models. Thus, aspects of the present disclosure can result in a machine learning model that consumes fewer computational resources than prior art models.

[0043] Referring now to the drawings, example embodiments of the present disclosure will be discussed in more detail.

[0044] Example Devices and Systems

[0045] Figure 1A A block diagram of an example computing system 100 for training an image coloring model in accordance with an example embodiment of the present disclosure is depicted. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0046] The user computing device 102 can be any type of computing device, such as a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0047] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple operatively connected processors. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0048] The user computing device 102 may store or include one or more coloring models 120. For example, the (multiple) coloring models 120 may be or may otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other multi-layer non-linear models. The neural network may include a recurrent neural network (e.g., a long short-term memory recurrent neural network), a feed-forward neural network, or other forms of neural networks. Refer to Figures 2 to 5 Discuss example coloring models 120.

[0049] In some implementations, one or more coloring models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single coloring model 120 (e.g., to perform parallel coloring operations across multiple instances of the coloring model 120).

[0050] More specifically, the coloring model 120 may be configured to color grayscale images such as old photos. In particular, aspects of the present disclosure relate to training the coloring model using a neighborhood color loss function. The neighborhood color loss function may reward the coloring loss model for predicting pixel color values that have the correct magnitude distance from the corresponding color values of some or all of the other pixels in the image, rather than focusing on whether the coloring model has correctly predicted the actual ground truth color value itself. In this way, the coloring model can be enabled to produce vivid colors for objects that may have multiple correct recoloring solutions.

[0051] Additionally or alternatively, one or more coloring models 140 may be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the coloring model 140 may be implemented by the server computing system 140 as part of a network service (e.g., an image coloring and / or storage service that may be provided, for example, as a feature of a photo storage application). Thus, one or more coloring models 120 may be stored and implemented at the user computing device 102 and / or one or more coloring models 140 may be stored and implemented at the server computing system 130.

[0052] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user may input communication.

[0053] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operably connected processors. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc. and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0054] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0055] As described above, the server computing system 130 can store or otherwise include one or more machine learning coloring models 140. For example, the model 140 can be or otherwise include various machine learning models, such as neural networks (e.g., deep recurrent neural networks) or other multi-layer non-linear models. Reference Figures 2 to 5 discusses example models 140.

[0056] The server computing system 130 can train the model 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be a part of the server computing system 130.

[0057] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operably connected processors. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.

[0058] The training computing system 150 may include a model trainer 160 that trains a machine learning model 140 stored at the server computing system 130 using various training or learning techniques (such as backpropagation of errors). In some implementations, performing backpropagation of errors may include performing truncated backpropagation through time. The model trainer 160 may perform a variety of generalization techniques (such as weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0059] In particular, the model trainer 160 may train the coloring model 140 based on a set of training data 142. The training data 142 may include, for example, ground truth images and training input images including grayscale versions of the ground truth images.

[0060] In some implementations, if the user has consented, training examples may be provided by the user computing device 102 (e.g., based on previous communications provided by the user of the user computing device 102). Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific communication data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.

[0061] The model trainer 160 includes computer logic for providing the required functionality. The model trainer 160 may be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.

[0062] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communications on the network 180 may be performed using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL) via any type of wired and / or wireless connection.

[0063] Figure 1AIllustrated is an example computing system that can be used to implement the present disclosure. Other computing systems can also be used. For example, in some implementations, the user computing device 102 can include a model trainer 160 and a training data set 162. In such implementations, the model 120 can be locally trained and used at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0064] Figure 1B Depicted is a block diagram of an example computing device 10 executing in accordance with an example embodiment of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0065] The computing device 10 includes a plurality of applications (e.g., Application 1 to N). Each application includes its own machine learning library and (a) machine learning model(s). For example, each application can include one machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0066] As Figure 1B shown, each application can communicate with a plurality of other components of the computing device (e.g., one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a common API). In certain implementations, the API used by each application is specific to that application.

[0067] Figure 1C Depicted is a block diagram of an example computing device 50 executing in accordance with an example embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0068] The computing device 50 includes a plurality of applications (e.g., Application 1 to N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0069] The central intelligence layer includes a plurality of machine learning models. For example, as Figure 1CAs shown, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 50.

[0070] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository of computing device 50. As Figure 1C shown, the central device data layer can communicate with multiple other components of the computing device (e.g., one or more sensors, context manager, device status component, and / or other components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a dedicated API).

[0071] Example Model Arrangement

[0072] Figure 2 A block diagram of an example coloring model 200 in accordance with an example embodiment of the present disclosure is depicted.

[0073] In some implementations, the coloring model 200 is trained to receive an input image 206 and process the input image 206 to output a predicted color map 204 that describes the predicted coloring of the input image 206.

[0074] Figure 3A A block diagram of an example coloring model 300 in accordance with an example embodiment of the present disclosure is depicted. In some implementations, the coloring model 300 can include a polynomial prediction model 302 and a refinement model 304. The polynomial prediction model 302 can be configured to receive a training input image 306 and, in response to receiving the training input image 306, output a polynomial color distribution 308. The polynomial color distribution 308 can describe multiple colorings of the training input image 306 (e.g., in a discretized color space). The refinement model 304 can be configured to receive the polynomial color distribution 308 and, in response to receiving the polynomial color distribution 308, output a predicted color map 310.

[0075] Figure 3BFIG. 0 depicts a block diagram of an example system 350 for training a coloring model 352 according to an example embodiment of the present disclosure. As described above, the coloring model 352 may include a polynomial prediction model 354 and a refinement model 356. The polynomial prediction model 354 may be configured to receive a training input image 358 and, in response to receiving the training input image 358, output a polynomial color distribution 360. The polynomial color distribution 360 may describe a plurality of colorings (e.g., in a discretized color space) of the training input image 358. The discretized color space may include a plurality of color intervals (e.g., n×n color intervals) defined within a color space (e.g., CIELAB color space). For example, the color intervals may be selected to cover or include colors that can be displayed on a display screen, such as colors within a color gamut in an RGB color space. The size of the color intervals may be selected considering a desired color resolution in the discretized color space and / or a desired simplicity of training the coloring model (e.g., reduced computational resources required for training the coloring model). For example, the color space may be discretized into 10×10 color intervals, resulting in approximately 310 discrete in-gamut colors (e.g., ab pairs in the CIELAB color space).

[0076] The polynomial prediction model 354 may include an encoder model and a decoder model in an autoencoder configuration. The training input image 358 may be input into the encoder, and the polynomial distribution 360 may be received as the output of the decoder model. The polynomial prediction model 354 may include at least one skip connection between layers of the encoder model and layers of the decoder model (e.g., as a trapezoidal model). Such a skip connection may transfer useful information from a hidden layer of the encoder model to a layer of the decoder model. This may facilitate backpropagation during training of the polynomial prediction model 354.

[0077] The refinement model 356 may be configured to receive the polynomial color distribution 360. In response to receiving the polynomial color distribution 360, the refinement model 356 may be configured to output a predicted color map 362.

[0078] The coloring model 352 may be trained using a neighborhood color loss function 364. Generally, the neighborhood color loss function 364 is configured to avoid penalizing (or at least reduce the impact of penalizing) a predicted color map 362 for the training input image 358 that is acceptable and different from a ground truth color map 366.

[0079] More specifically, the neighborhood color loss function 364 can be calculated as the difference between a relative first color distance of the predicted color map 362 and a second color distance of the ground truth color map 366. More specifically, a first color distance between a first and a second pixel included in the predicted color map 362 can be calculated. A second color distance between a third and a fourth pixel included in the ground truth color map 366 corresponding to the ground truth image can be calculated. The third and fourth pixels in the ground truth color map 366 can spatially correspond to the first and second pixels in the predicted color map 362, respectively. The neighborhood color loss function 364 can evaluate the difference between the first color distance and the second color distance. Parameters (e.g., weights) associated with the coloring model 352 can be adjusted based on the neighborhood color loss function 364.

[0080] The neighborhood color loss function 364 can be calculated for some or all pixels of the predicted color map 362. It should be understood that the neighborhood color loss function 364 can be calculated iteratively for multiple "first pixels". For example, the neighborhood color loss function 364 can be calculated for each pixel of the predicted color map 362.

[0081] Multiple suitable techniques can be used to select the first pixels of the predicted color map 362. The (multiple) first pixels can be randomly selected within the predicted color map 362 (e.g., anywhere within the predicted color map or within certain portions of the predicted color map, such as within identified features or objects). The (multiple) first pixels can be selected based on their positions within the predicted color map 362 according to a shape or pattern (e.g., a grid, a square, a circle, etc.). The (multiple) first pixels can be selected based on objects or features detected within the ground truth color map 366. Some or all pixels of one or more detected features can be iteratively selected as the first pixels.

[0082] The neighborhood color loss function 364 can be calculated for multiple second pixels of the predicted color map 362. Multiple suitable techniques (including random or systematic selection) can be used to select the second pixels. For example, the second pixels can be iteratively selected as each pixel of the predicted color map 362 (except for the first pixels) such that a "neighborhood color loss map" is generated for the first pixels relative to the remaining pixels of the predicted color map 362. Alternatively, the second pixels can be selected based on their relative positions with respect to the first pixels. For example, one or more pixels directly adjacent to the first pixel can be selected as the (multiple) second pixels. As another example, one or more pixels spaced apart from the first pixel by a set distance or arranged in a pattern (e.g., a circle, a square, a grid, etc.) relative to the first pixel can be selected.

[0083] As described above, the refinement model 356 can be configured to receive a polynomial color distribution 360 and output a predicted color map 362. An output image can be generated based on the predicted color map 362 and the training input image 358 (e.g., a luminance map describing the luminance of the training input image 358). As described above, the polynomial color distribution 360 can describe multiple colorings of the training input image 358. For example, the polynomial color distribution 360 can include a corresponding color distribution for each pixel of the training input image 358. The refinement model 356 can be configured to combine two or more colorings among the multiple colorings of the polynomial color distribution 360. The parameters (e.g., weights) of the refinement model 356 can be adjusted based on a neighborhood color loss function 364 such that the refinement model 356 is trained to output a predicted color map 362 that minimizes the neighborhood color loss function 364.

[0084] Figure 4 A block diagram of an example system 400 for training a coloring model 402 in accordance with an example embodiment of the present disclosure is depicted. As described above, the coloring model 402 can include a polynomial prediction model 404 and a refinement model 406. The polynomial prediction model 404 can be configured to receive a training input image 408 and, in response to receiving the training input image 408, output a polynomial color distribution 410. The polynomial color distribution 404 can describe multiple colorings of the training input image 408 (e.g., in a discretized color space). The refinement model 406 can be configured to receive the polynomial color distribution 410 and output a predicted color map 412. The coloring model 402 can be trained using a neighborhood color loss function 414, such as described above with respect to Figure 3B the neighborhood color loss function 364.

[0085] In some implementations, one or more additional loss functions can be used to train the coloring model 402. The additional loss functions can be defined with respect to the predicted color map 412, the polynomial color distribution 410, and / or a ground truth color map 411 of the training input image 408. For example, a total loss function can be defined that includes the neighborhood color loss function 414 and one or more additional loss functions. As an example, an absolute color loss function 416 can be employed that describes the difference (e.g., color distance) between the predicted color map 412 and the ground truth color map 411.

[0086] As another example of another additional loss function, a refined softmax loss function 418 that describes the difference between the polynomial color distribution 410 and the predicted color map 412 can be employed. The polynomial color distribution 410 and the predicted color map 412 can be encoded in different color spaces. For example, the polynomial color distribution 410 can be encoded in a discretized color space. To compare the predicted color map 412 with the polynomial distribution 410, the predicted color map 412 can be encoded into the discretized color space (e.g., encoded into multiple color intervals) to generate a discretized predicted color map. Thus, the refined softmax loss function 418 can describe the difference between the polynomial color distribution 410 in the discretized color space and the discretized predicted color map. However, it should be understood that the predicted color map 412 does not necessarily have to be encoded in the discretized color space.

[0087] Encoding the predicted color map 412 in the discretized color space can be achieved in various ways. For example, the predicted color map 412 can be "one-hot" encoded in the discretized color space. Such encoding can include selecting a single color interval (e.g., the "closest" color interval) of the discretized color space for each pixel of the predicted color map 412. As another example, the predicted color map 412 can be soft-encoded in the discretized color space. Soft-encoding can include representing each pixel as a distribution over two or more color intervals of the discretized color space.

[0088] As another example of another additional loss function, the parameters of the polynomial prediction model 404 can be adjusted based on a polynomial softmax loss function 420. The polynomial softmax loss function 420 can describe the difference between the ground truth color map 411 and the polynomial color distribution 410 output by the polynomial prediction model 404. However, the ground truth color map 411 and the polynomial color distribution 410 can be encoded in different color spaces. Thus, the color space associated with the ground truth color map 411 can be discretized to generate a discretized color space including multiple color intervals. The ground truth color map 411 can be encoded into the discretized color space to generate a discretized ground truth color map. The polynomial softmax loss function 420 can describe the difference between the discretized ground truth color map and the polynomial color distribution 410.

[0089] In some implementations, the coloring model 402 can be trained as part of a generative adversarial network (GAN). For example, the discriminator model 422 can be configured to receive the output of the coloring model 402 (e.g., the predicted color map 412 and / or the output image 417). In response to receiving the output of the coloring model 402, the discriminator model 422 can output a discriminator loss 424, which describes a score of the output of the coloring model 402 compared to the ground truth image. The discriminator loss 424 can describe a score of the predicted color map 412 compared to the ground truth color map 411.

[0090] For example, the discriminator loss 424 can include a binary score that indicates whether the discriminator model 422 identifies the predicted color map 412 as the ground truth color map 411 ("true") or as a predicted color map 412 ("false"). In other embodiments, in addition to the binary score, the discriminator loss 424 can include a probability distribution or confidence score relative to the above identification.

[0091] Figure 5 Illustrated is a training configuration 500 of a discriminator model 502 according to aspects of the present disclosure. The discriminator model 502 can be trained to better identify the ground truth image and distinguish the ground truth image from the colored image. During such training, a mixture of the ground truth image (or the ground truth color map 504) and the predicted image (or the predicted color map 506) can be input into the discriminator model 502. The discriminator model 502 can output a discriminator loss 508 in response to receiving the ground truth color map 504 or the predicted color map 506. The discriminator model 502 can be trained based on a loss function that penalizes the model for misidentifying the input image or color map 504, 506.

[0092] The training of the discriminator model 502 can be performed alternately with the training of the coloring model 402 (e.g., as described with reference to Figure 4 ). The first training phase can include training the coloring model 402 as described with reference to Figure 5 . The second training phase can include training the discriminator models 422, 502 as described with reference to Figure 4 . The first and second training phases can be alternated until the training of the discriminator model 502 is complete. For example, the discriminator model 502 can be trained for a total number of training iterations or until the discriminator model 502 and / or the coloring model 402 ( Figure 5 ). The first and second training phases can be alternated until the training of the discriminator model 502 is complete. For example, the discriminator model 502 can be trained for a total number of training iterations or until the discriminator model 502 and / or the coloring model 402 ( Figure 4) until one or more performance criteria are met. An example criterion for the coloring model is a threshold percentage of "fooling" the predicted color map of the discriminator model 502. An example criterion for the discriminator model 502 is a threshold percentage of correctly identifying the color map input to the discriminator model 502 as the predicted color map 506 or the ground truth color map 504.

[0093] During the training of the coloring model 402, the predicted color map 412 can be input into the discriminator model 422. The discriminator loss 422 can be received as the output of the discriminator model 422. As described above, the discriminator loss 422 can include a binary indicator or distribution that describes whether the discriminator 422 classifies the predicted color map input to the discriminator model 422 as a predicted coloring or a ground truth image. The coloring model 402 can be trained based on the discriminator loss 424 to produce a predicted color map 412 that "fools" the discriminator model 422 such that the discriminator model 422 classifies the predicted color map 412 as the ground truth color map 411 and / or as corresponding to the ground truth image (rather than identifying the predicted color map 412 as a generated coloring).

[0094] In some implementations, the coloring model 402 can utilize object or feature detection or recognition to improve the training of the coloring model 402. For example, the coloring model 402 can include a feature detection model 426 that is configured to receive the training input image 408 as input and output a plurality of feature representations or features 428 that describe the locations of objects or features within the training input image 408. The feature detection model can be separate from and / or included in the polynomial prediction model 404. The feature representation 428 can include bounding objects (e.g., boxes, points, pixel masks, etc.) that describe the locations of one or more features or objects identified or detected within the training input image 408.

[0095] In some embodiments, the feature representation 428 can include labels (e.g., classes) of the identified objects. The feature representation 428 can be input into the refinement model 406 and / or the polynomial prediction model 404 such that semantic relationships between objects and colors can be learned. For example, it can be learned that an octagonal sign is typically a red stop sign. However, in other embodiments, the feature representation 428 can not have such labels. In such embodiments, the feature representation 428 can simply identify the bounding objects (e.g., boxes, points, pixel masks, etc.) or locations of various shapes or features depicted in the training input image 408.

[0096] Various transformations or encodings can be performed by, and / or on the inputs and / or outputs of, the coloring model 402, the discriminator model 422. The input image 408 and / or the output image 417 can be in the RGB color space. As described above, the coloring model 402 can generally operate in the CIELAB color space. Additionally, in some embodiments, the discriminator model 422 can operate in the RGB color space. Thus, conversions between the RGB color space and the CIELAB color space can be performed as needed.

[0097] Example Method

[0098] Figure 6 Depicted is a flowchart of an example method 600 performed in accordance with an example embodiment of the present disclosure. Although Figure 6 Steps are depicted for purposes of illustration and discussion as being performed in a particular order, the method 600 of the present disclosure is not limited to the particular order or arrangement shown. The corresponding steps of the method 600 can be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of the present disclosure.

[0099] At 602, a computing system can input a training input image into a coloring model configured to receive the training input image, and in response to receiving the training input image, output a predicted color map that describes a predicted coloring for the training input image, such as as described above with reference to Figures 2 to 4 described.

[0100] At 604, the computing system can receive the predicted color map as the output of the coloring model, such as as described above with reference to Figures 2 to 4 described.

[0101] At 606, the computing system can compute a first color distance between a first pixel of the predicted color map and a second pixel of the predicted color map, such as as described above with reference to Figures 2 to 4 described.

[0102] At 608, the computing system can compute a second color distance between a third pixel included in a ground truth color map and a fourth pixel included in a ground truth coloring map, such as as described above with reference to Figures 2 to 4 described. The third pixel and the fourth pixel included in the ground truth color map can be spatially corresponding to the first pixel and the second pixel included in the predicted color map, respectively.

[0103] At 610, the computing system can evaluate a neighborhood color loss function that evaluates the difference between the first color distance and the second color distance, such as as described above with reference to Figures 2 to 4 described.

[0104] At 612, the computing system can adjust the parameters associated with the coloring model based on a neighborhood color loss function, such as described above with reference to Figures 2 to 4 as described.

[0105] Additional Disclosure

[0106] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from these systems. The inherent flexibility of computer-based systems allows for various possible configurations, combinations, and divisions of tasks and functionality among the components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. The database and applications can be implemented on a single system or distributed across multiple systems. The distributed components can run sequentially or in parallel.

[0107] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation and not limitation of the disclosure. Those skilled in the art can readily generate alterations, variations, and equivalents of these embodiments upon obtaining an understanding of the foregoing. Accordingly, the subject matter disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be apparent to a person of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used in another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover such alterations, variations, and equivalents.

Claims

1. A method for training an image coloring model, the method comprising: inputting, by one or more computing devices, a training input image into a coloring model, the coloring model being configured to receive the training input image and process the training input image to output a predicted color map that describes a predicted coloring of the training input image; receiving, by the one or more computing devices, the predicted color map as an output of the coloring model; computing, by the one or more computing devices, a first color distance between a first pixel included in the predicted color map and a second pixel included in the predicted color map; computing, by the one or more computing devices, a second color distance between a third pixel included in a ground truth color map and a fourth pixel included in the ground truth color map, wherein the third pixel and the fourth pixel included in the ground truth color map spatially correspond to the first pixel and the second pixel included in the predicted color map, respectively; evaluating, by the one or more computing devices, a neighborhood color loss function that evaluates a difference between the first color distance and the second color distance; and adjusting, by the one or more computing devices, parameters associated with the coloring model based on the neighborhood color loss function.

2. The method according to claim 1, further comprising randomly selecting, by the one or more computing devices, the first pixel of the predicted color map.

3. The method according to claim 1, further comprising selecting, by the one or more computing devices, at least one of the first pixel and the second pixel based on at least one of the following: the location of features detected within the training input image; and a predefined pattern defined in the predicted color map.

4. The method according to claim 1, further comprising iteratively generating, by the one or more computing devices, a corresponding neighborhood color loss map for a respective additional first pixel of the predicted color map.

5. The method according to claim 1, further comprising iteratively evaluating, by the one or more computing devices, the neighborhood color loss function for a plurality of additional second pixels with respect to the first pixel to generate a neighborhood color loss map for the first pixel of the predicted color map.

6. The method according to claim 1, wherein: the coloring model includes a polynomial prediction model and a refinement model, wherein the polynomial prediction model is configured to receive the training input image and, in response to receiving the training input image, output a polynomial color distribution that describes a plurality of colorings for the training input image in a discretized color space, and wherein the refinement model is configured to receive the polynomial color distribution and, in response to receiving the polynomial color distribution, output the predicted color map; inputting the training input image into the coloring model includes inputting the training input image into the polynomial prediction model; and the method further comprises: receiving the polynomial color distribution as an output of the polynomial prediction model; and Input the polynomial color distribution into the refinement model.

7. The method according to claim 6, wherein, Adjusting the parameters associated with the coloring model based on the neighborhood color loss function includes adjusting the parameters of the refinement model.

8. The method according to claim 6, further comprising: Encoding the predicted color map output by the refinement model into a discretized color space to generate a discretized predicted color map; Evaluating a refinement softmax loss function that describes the difference between the polynomial color distribution and the discretized predicted color map; and Adjusting the parameters of the refinement model based on the refinement softmax loss function.

9. The method according to claim 8, wherein, Encoding the predicted color map into the discretized color space to generate the discretized predicted color map includes one-hot encoding the corresponding color values of the predicted color map into the corresponding color intervals of multiple color intervals in the discretized color space.

10. The method according to claim 8, wherein, Projecting the predicted color map onto the discretized color space to generate the discretized predicted color map includes soft encoding the predicted color map with respect to multiple color intervals of the discretized color space.

11. The method according to claim 6, further comprising: Evaluating an absolute color loss function that describes the difference between the predicted color map and the ground truth color map; and Adjusting the parameters associated with the refinement model based on the absolute color loss function.

12. The method according to claim 6, further comprising: Encoding the ground truth color map into a discretized color space associated with the ground truth color map, the discretized color space including multiple color intervals to generate a discretized ground truth color map; Evaluating a polynomial softmax loss function that describes the difference between the discretized ground truth color map and the polynomial color distribution output by the polynomial prediction model; and Adjusting the parameters of the polynomial prediction model based on the polynomial softmax loss function.

13. The method according to any one of claims 6 to 12, wherein, The color space includes the CIELAB color space.

14. A computing system, comprising: One or more processors; A coloring model configured to receive a training input image and, in response to receiving the training input image, output a predicted color map describing the predicted coloring of the training input image; One or more non-transitory computer-readable media that jointly store instructions which, when executed by the one or more processors, cause the computing system to perform operations, the operations including: Inputting the training input image into the coloring model; Receiving the predicted color map as the output of the coloring model; Calculating a first color distance between a first pixel included in the predicted color map and a second pixel included in the predicted color map; Calculate a second color distance between a third pixel included in a ground truth color map and a fourth pixel included in the ground truth color map, wherein the third pixel and the fourth pixel included in the ground truth color map spatially correspond to the first pixel and the second pixel included in the predicted color map, respectively; Evaluate a neighborhood color loss function that evaluates a difference between the first color distance and the second color distance; and Adjust parameters associated with the coloring model based on the neighborhood color loss function.

15. The computing system according to claim 14, wherein, the operation further includes randomly selecting the first pixel of the predicted color map.

16. The computing system according to claim 14, wherein, the operation further includes selecting at least one of the first pixel and the second pixel based on at least one of: a location of a feature detected within the training input image; and a predefined pattern defined in the predicted color map.

17. The computing system according to any one of claims 14 to 16, wherein, the operation further includes iteratively generating a corresponding neighborhood color loss map for a respective additional first pixel of the predicted color map.

18. A computing system, comprising: one or more processors; a coloring model configured to receive an input image and, in response to receiving the input image, output a predicted color map depicting a predicted coloring of the input image, the coloring model having been trained based on a neighborhood color loss function that evaluates a difference between a first color distance and a second color distance, and wherein the first color distance has been calculated during training between a first pixel included in a training predicted color map output by the coloring model and a second pixel included in the training predicted color map, and wherein the second color distance is between a third pixel included in a ground truth color map and a fourth pixel in the ground truth color map, wherein the third pixel and the fourth pixel included in the ground truth color map spatially correspond to the first pixel and the second pixel included in the training predicted color map, respectively; one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations including: inputting the input image into the coloring model; receiving the predicted color map as an output of the coloring model; and generating an output image based on the input image and the predicted color map.