Image color processing method and device, equipment, medium and product
By introducing a combination of color mask loss function and perceptual loss function into the U-Net model for training, and combining it with VGG network feature extraction, the problems of color distortion and poor enhancement effect of the U-Net model in image color processing are solved, and better color processing effect is achieved.
Patent Information
- Application Number
- CN202510751455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-21
AI Technical Summary
The U-Net model suffers from color distortion and poor color enhancement effects when processing image colors.
By introducing a combination of color mask loss function and perceptual loss function into the U-Net model for training, and combining it with VGG network feature extraction, the color processing effect of images is optimized.
It improves the effect of image color processing, preserves the structure and detail information of the original image, and enhances the fidelity of the global color and some color range of the image.
Smart Images

Figure CN120823408A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image color processing method, device, equipment, medium and product. Background Art
[0002] With the continuous development and advancement of network and computer technologies, digital image processing technology has been widely used in fields such as autonomous driving, medical image analysis, and facial recognition. Current digital image processing technologies typically include traditional solutions that rely on predefined algorithms and rules, as well as intelligent and efficient image processing solutions that train neural network models to learn and capture image features.
[0003] The U-Net model, a neural network model, is a deep learning architecture used for image segmentation in medical and remote sensing images. It effectively extracts local features and performs multi-scale image fusion. However, when currently used for image color processing, the U-Net model can cause color distortion and poor color enhancement. Therefore, the current challenge is to improve the U-Net model's color processing performance. Summary of the Invention
[0004] The present application provides an image color processing method, apparatus, device, medium and product for improving the effect of U-Net model on image color processing.
[0005] On the one hand, the present application provides an image color processing method, comprising: obtaining feature data of an image to be processed; wherein the feature data includes multiple image channels and feature values corresponding to each image channel; inputting the feature data of the image to be processed into a U-Net model to obtain a model output image of the image to be processed output by the U-Net model; wherein the U-Net model is pre-trained based on an image data set and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and the area where at least one color range in the image to be processed is located; based on the model output image, a color processing result of the image to be processed is obtained.
[0006] In one possible implementation, the color range includes a black range and a white range; a color mask loss function is established based on the entire image area of the image to be processed and the area where at least one color range in the image to be processed is located, including: taking the area where pixels in the image to be processed whose grayscale values are greater than a first threshold are located as a white mask, and taking the area where pixels in the image to be processed whose grayscale values are less than a second threshold are located as a black mask; calculating a first loss function of the image to be processed and the model output image under the white mask, a second loss function of the image to be processed and the model output image under the black mask, and a third loss function of the image to be processed and the model output image under the entire image area; and determining the color mask loss function based on the first loss function, the second loss function, and the third loss function.
[0007] In one possible implementation, a color mask loss function is determined based on the first loss function, the second loss function, and the third loss function, including: taking the product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function as the color mask loss function.
[0008] In one possible implementation, the method further includes: performing feature extraction on the image to be processed based on the VGG network to obtain a first extracted feature before the first fully connected layer of the VGG network; performing feature extraction on the model output image based on the VGG network to obtain a second extracted feature before the first fully connected layer of the VGG network; and using the loss function of the first extracted feature and the second extracted feature as the perceptual loss function.
[0009] In one possible implementation, the U-Net model includes a first model, a second model, and a third model; the first model is used to perform convolutional layer feature extraction and at least one downsampling feature extraction on the feature data of the image to be processed to obtain the initial features corresponding to each channel; wherein the convolutional layer feature extraction includes at least one convolution and activation function activation after each convolution; the downsampling feature extraction includes maximum pooling processing and convolutional layer feature extraction; the second model is used to perform at least one upsampling feature extraction on the initial features corresponding to each channel to obtain the processing features corresponding to each channel; wherein the upsampling feature extraction includes: deconvolution processing, feature splicing, and convolutional layer feature extraction, and the magnification factor of the size of the eigenvalue matrix of the same channel in the deconvolution processing is equal to the reduction factor of the size of the eigenvalue matrix of the same channel in the maximum pooling processing; the third model is used to output the model output image of the image to be processed according to the processing features corresponding to each channel.
[0010] In one possible implementation, feature stitching includes adding, based on the same image channel, the feature values of the feature data in the result of the deconvolution processing and the feature values of the feature data in the result of the symmetrical convolution layer feature extraction in the first model.
[0011] In one possible implementation, before inputting the feature data of the image to be processed into the U-Net model, the method further includes: calculating a reduction ratio based on the target reduction size and the longer side length of the image to be processed, and scaling the image to be processed according to the reduction ratio to obtain a scaled image; determining whether the side length of the scaled image is a multiple of 2 to the power of n, and if not, enlarging the corresponding side length to a multiple of 2 to the power of n; wherein n is a positive integer greater than or equal to 3.
[0012] In one possible implementation, the method also includes: performing polynomial mapping on feature data of the image to be processed input into the U-Net model to obtain polynomial extracted features; determining the color processing result of the image to be processed based on the model output image, including: using the model output image as a training target, establishing an initial regression model of the polynomial extracted features and the feature values corresponding to each image channel in the model output image and performing model training to obtain a regression model; inputting the polynomial extracted features into the regression model, and using the output result of the regression model as the color processing result of the image to be processed.
[0013] On the other hand, the present application provides an image color processing device, including: an acquisition module for acquiring feature data of an image to be processed; wherein the feature data includes multiple image channels and feature values corresponding to each image channel; a processing module for inputting the feature data of the image to be processed into a U-Net model to obtain a model output image of the image to be processed output by the U-Net model; wherein the U-Net model is pre-trained based on an image data set and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and the area where at least one color range in the image to be processed is located; a determination module for obtaining a color processing result of the image to be processed based on the model output image.
[0014] On the other hand, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the above method.
[0015] On the other hand, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the above method.
[0016] On the other hand, the present application provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0017] The image color processing method, device, equipment, medium and product provided by the present application include: inputting feature data including multiple image channels and feature values corresponding to each image channel into a U-Net model output, and outputting an image based on the output model to obtain a color processing result of the processed image. The U-Net model is pre-trained based on an image data set and a loss function, wherein the perceptual loss function can extract features from the image to obtain the details and structural information of the image, and the color mask loss is established based on the entire image area and the area where at least one color range is located, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the effect of image color processing while retaining the structure and detail information of the original image. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0019] Figure 1 A flow chart of an image color processing method is shown in FIG.
[0020] Figure 2 A flow chart of an image color processing method is shown in FIG.
[0021] Figure 3 A flow chart of an image color processing method is shown in FIG.
[0022] Figure 4 A structural diagram of a U-Net model is shown as an example;
[0023] Figure 5 A flow chart of an image color processing method is shown in FIG.
[0024] Figure 6 A flow chart of an image color processing method is shown in FIG.
[0025] Figure 7 A flow chart of an image color processing method is shown in FIG.
[0026] Figure 8 exemplarily shows a structural diagram of an image color processing device;
[0027] Figure 9 Schematic diagram of the structure of an electronic device is shown in FIG.
[0028] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0029] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0030] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings. The terms "including" and "having" in the specification and claims of this application and the above-mentioned drawings, as well as any variations thereof, are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are clearly listed, but may include other components that are not clearly listed or are inherent to these products or devices. The term "module" used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that can perform the functions associated with the element.
[0031] Image processing technology refers to the various operations and processing performed on images to improve their quality, extract useful information, or prepare them for further analysis. Image processing includes image preprocessing, image transformation, image enhancement, image compression, and image recognition. Traditional image processing methods, such as those that rely on predefined algorithms and rules, can handle some simple and basic image processing tasks, but have difficulty handling complex and variable image data. With the development of deep learning, data-driven methods such as the U-Net model have shown excellent performance in image processing tasks.
[0032] The U-Net model is a convolutional neural network used primarily for image segmentation and generation tasks in biomedical and remote sensing fields. The U-Net model's primary advantage in image processing lies in its encoder-decoder architecture combined with skip connections, which efficiently extracts and reconstructs image features while preserving detailed image information. However, when using the U-Net model for image color processing, color fidelity is often compromised, such as color distortion or inconsistent color. Furthermore, when processing high-resolution images, the U-Net model can also suffer from a loss of color detail. Therefore, the current challenge is to improve the U-Net model's image color processing performance.
[0033] The technical content provided by this application is intended to solve the above-mentioned technical problems in related technologies. In the image color processing method, device, equipment, medium and product provided by this application, the method includes: inputting feature data including multiple image channels and feature values corresponding to each image channel into the U-Net model output, and outputting an image based on the output model to obtain a color processing result of the processed image. The U-Net model is pre-trained based on the image data set and the loss function, wherein the perceptual loss function can extract features from the image to obtain the details and structural information of the image, and the color mask loss is established based on the entire image area and the area where at least one color range is located, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the effect of image color processing while retaining the structure and detail information of the original image.
[0034] The technical solutions of the present application and the technical solutions of the present application are described in detail below with reference to specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in certain embodiments. In the description of the present application, unless otherwise clearly specified and limited, each term should be understood in a broad sense within the art. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0035] Example 1
[0036] Figure 1 A flow chart of an image color processing method is shown in FIG. , and the execution subject of this example may be an image color processing device, such as Figure 1 As shown, the method includes:
[0037] Step 101: Acquire feature data of an image to be processed; wherein the feature data includes multiple image channels and feature values corresponding to each image channel;
[0038] Step 102: Inputting feature data of the image to be processed into a U-Net model to obtain a model output image of the image to be processed output by the U-Net model; wherein the U-Net model is pre-trained based on the image dataset and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and an area where at least one color range in the image to be processed is located;
[0039] Step 103: Output the image based on the model to obtain a color processing result of the image to be processed.
[0040] In practical applications, the executor of this method can be an image color processing device, and there are many ways to implement it. For example, it can be implemented through a computer program, such as application software, etc.; or it can be implemented as a medium that stores relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip, etc.
[0041] In this example, the feature data of the image to be processed, including multiple image channels and the feature values corresponding to each image channel, are channels divided under different color standards and the feature values corresponding to each channel. For example, when the color standard is the RGB (Red, Green, Blue) standard, the feature data corresponds to the three channels of red, green and blue and the corresponding values of each pixel of the image to be processed under the three channels. Optionally, based on the current color standard, channels and the corresponding values of the channels can be added, such as adding a grayscale channel under the RGB channel and determining the corresponding value of each pixel of the image to be processed under the grayscale channel. Exemplarily, different color standards can be selected based on different color processing requirements, and the corresponding feature data can be determined; for example, the CMYK (Cyan, Magenta, Yellow, Key / Black) standard based on the ink color superposition during the printing process, or the HSV (Hue, Saturation, Value) standard based on the principle of human visual perception.
[0042] After obtaining the feature data corresponding to the object to be processed, it is input into the U-Net model. For example, a tensor can be created as the input function of the U-Net model. This tensor can include multiple dimensions, such as the number of images to be processed (i.e., the batch size), the height, width, and number of channels of the images to be processed, etc. In practical applications, by modifying the different dimensions of the tensor, the efficiency and quality of the U-Net model's image processing can be adjusted accordingly. The U-Net structure can be divided into two main parts: the encoder and the decoder, which form a U-shaped network architecture. The encoder consists of multiple convolutional and pooling layers, while the decoder gradually restores the spatial dimensions of the feature maps through upsampling. Skip connections connect the feature maps of corresponding encoder layers directly to the inputs of corresponding decoder layers. Specifically, upsampling can be achieved through deconvolution or transposed convolution with trainable parameters, or through upsampling layers without trainable parameters that utilize interpolation or repeated elements.
[0043] In this example, the U-Net model is pre-trained based on an image dataset and loss functions including a color mask loss and a perceptual loss function. The color mask loss function is established based on the entire image region of the image to be processed and the region containing at least one color range within the image to be processed. Depending on the model training objectives, high-resolution images of different themes, such as landscapes, people, or still lifes, can be selected as the image dataset. Optionally, selecting images of multiple themes as the image training set can increase the model's generalization capabilities. For example, when the U-Net model is used to train color processing methods such as color enhancement, color replacement, and color correction, an image dataset containing both the original and target images can be selected. For example, a perceptual loss function such as the Learned Perceptual Image Patch Similarity (LPIPS) perceptual loss function is used during training. This perceptual loss function utilizes features from multiple pre-trained convolutional neural networks and combines these features through learned weights to more accurately assess the perceptual quality of the image. Alternatively, other loss functions such as the VGG perceptual loss and style loss functions such as the Style Loss can be selected. The perceptual loss function compares the differences between the U-Net model's output image and the target image in the image dataset in high-level feature space, rather than simply at the pixel level. This allows the U-Net model to better capture image details and structural information. Furthermore, the U-Net model is trained using a color mask loss function. This color mask loss function, based on the entire image region of the processed image, can improve the overall color consistency between the U-Net model's output image and the processed image, avoiding color distortion or shift. Furthermore, a color mask loss function based on at least one color range in the processed image can improve the color fidelity of specific color regions and enhance the U-Net model's ability to recognize color features in color regions with key semantic color information. Exemplarily, the at least one color range can be determined by setting a threshold for the feature value or performing a cluster analysis on the pixels. Exemplarily, the color mask loss function can be implemented by calculating the mean squared error or mean absolute percentage error of the feature values for the same pixel in the processed image and the model output image.
[0044] After obtaining the model output image, the color processing result of the image to be processed can be obtained by denoising, resizing, and image format conversion. For example, in order to improve the performance of the U-Net model, image evaluation indicators can be introduced to evaluate and verify the model performance after each model training cycle. Specifically, the evaluation indicators may include: contrast, structural similarity index (SSIM) and vividness. The calculation formula for colorfulness under the RGB color standard is:
[0045]
[0046] The values of each pixel in the RGB three channels are R, G, and B, among which the difference between the red and green channels rg = |RG|, the difference between the blue and yellow channels yb = |0.5*(R+G)-B|, and the average value of the red and blue channels rb mean =(R+B) / 2.
[0047] The image color processing method provided in this example includes: inputting feature data including multiple image channels and feature values corresponding to each image channel into a U-Net model output, and outputting an image based on the output model, thereby obtaining a color processing result for the processed image. The U-Net model is pre-trained based on an image dataset and a loss function, wherein the perceptual loss function is capable of extracting features from the image to obtain image details and structural information. At the same time, the color mask loss is established based on the entire image area and the area containing at least one color range, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the image color processing effect while retaining the structure and detail information of the original image.
[0048] As another example, Figure 2 A flow chart of an image color processing method is shown in FIG. Figure 2 As shown, based on any example, the color range includes a black range and a white range. In step 102, the color mask loss function is established based on the entire image area of the image to be processed and the area where at least one color range in the image to be processed is located, including:
[0049] Step 201: The area where the pixels in the image to be processed have grayscale values greater than a first threshold are located is used as a white mask, and the area where the pixels in the image to be processed have grayscale values less than a second threshold are located is used as a black mask;
[0050] Step 202: Calculate a first loss function of the image to be processed and the model output image under the white mask, a second loss function of the image to be processed and the model output image under the black mask, and a third loss function of the image to be processed and the model output image over the entire image area;
[0051] Step 203: Determine a color mask loss function based on the first loss function, the second loss function, and the third loss function.
[0052] For example, the first threshold can be a fixed value, such as a grayscale value of 220, or a grayscale range, such as grayscale values 215 to 225; the second threshold can be a fixed value, such as a grayscale value of 20, or a grayscale range, such as grayscale values 15 to 25. Optionally, when the color standard is HSV, the first threshold is a brightness value. In practical applications, the white mask or black mask is composed of a binary matrix. After processing the image to be processed and the model output image using the mask, pixel pairs for which the loss function calculation is required are obtained. The error of each pixel pair in each color channel is calculated using an error calculation method, such as the mean square error algorithm or the root mean square error algorithm, to obtain the first loss function or the second loss function. Similarly, the third loss function can be obtained using the above algorithm. After obtaining the first, second, and third loss functions, each loss function can be multiplied by the same or different gain coefficients and then summed to obtain the color mask loss function. Alternatively, any one or two of the loss functions can be multiplied by the same or different gain coefficients and then summed to obtain the color mask loss function. The gain coefficient can be determined based on images with varying lighting conditions or environments. Optionally, the gain coefficient can also be adjusted based on the U-Net model's output image. This example solution considers the fidelity of the global color, dark areas, and bright areas of the processed photo. This allows the U-Net model to improve color adjustment performance in dark and bright areas while maintaining global color.
[0053] As another example, in step 203, determining the color mask loss function according to the first loss function, the second loss function, and the third loss function includes:
[0054] The product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function are used as the color mask loss function.
[0055] For example, the values of white weight and black weight can be set to the same default value such as 0.7. According to the changes in the proportion of dark areas and light areas in the image, the corresponding weight range can be increased or decreased. For example, when processing an image dataset with a large proportion of light areas, the white weight is increased and the black weight is decreased. For example, the weight adjustment range can be any value between 0.5 and 1.0. Color mask loss function L mask The formula for (x,y) is:
[0056] L(x,y)=L MSE (x,y)+L MSE_black (x,y)*W black +L MSE_white (x,y)*W w*ite
[0057] Among them, x is the image to be processed, y is the model output image, and the third loss function L MSE (x, y) is the mean square error calculation of x, y in the entire image area, and the second loss function L MSE_black (x,y) is the mean square error calculation of x,y in the black mask area, W black is the black weight, L MSE_white (x,y) is the mean square error calculation of x,y in the white mask area, W white is white weight.
[0058] This example uses the product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function as the color mask loss function. This improves the flexibility of the U-Net model in adjusting the colors of dark and light areas.
[0059] As another example, Figure 3 A flow chart of an image color processing method is shown in FIG. Figure 3 As shown, based on any example, the image color processing method further includes:
[0060] Step 301: Perform feature extraction on the image to be processed based on the VGG network to obtain the first extracted feature before the first fully connected layer of the VGG network;
[0061] Step 302: Perform feature extraction on the model output image based on the VGG network to obtain a second extracted feature before the first fully connected layer of the VGG network;
[0062] Step 303: Use the loss function of the first extracted feature and the second extracted feature as the perceptual loss function.
[0063] As a deep convolutional neural network architecture, the VGG network's intermediate layer features can capture high-level semantic information of the image. By calculating the difference between the image to be processed and the model output image in the intermediate layer feature space of the VGG network, the perceptual quality of the image can be better measured. At the same time, the VGG network has transfer learning capabilities, which enables the pre-trained VGG network to effectively extract image features when the perceptual loss function is established. In practical applications, VGG16 or VGG19 can be used for feature extraction. The extracted features before the first fully connected layer of the VGG network have the following beneficial effects: it inherits the rich feature information extracted by the previous multiple intermediate layers, and can capture the complex structure and object relationships in the image. At the same time, only using the features output by a single neural layer for the perceptual loss function calculation can reduce the amount of calculation. In this example, the perceptual loss function L perceptual The formula for (x,y) is:
[0064]
[0065] Among them, x is the image to be processed, y is the model output image, N is the number of extracted features of the VGG network before the first fully connected layer, φ i (x) represents the i-th feature extracted by the VGG network for the processed image, φ i (y) represents the i-th feature extracted by the VGG network for the model output image, and ‖·‖1 represents the L1 norm. Here, φ(x) corresponds to the first extracted feature in this example, and φ(y) corresponds to the second extracted feature in this example.
[0066] As yet another example, the U-Net model includes a first model, a second model, and a third model;
[0067] The first model is used to perform convolutional layer feature extraction and at least one downsampling feature extraction on the feature data of the image to be processed to obtain the initial features corresponding to each channel; wherein the convolutional layer feature extraction includes at least one convolution and activation function activation after each convolution; the downsampling feature extraction includes maximum pooling processing and convolutional layer feature extraction;
[0068] The second model is used to perform at least one upsampling feature extraction on the initial features corresponding to each channel to obtain processed features corresponding to each channel; wherein the upsampling feature extraction includes: deconvolution processing, feature splicing and convolution layer feature extraction, and the magnification factor of the size of the eigenvalue matrix of the same channel in the deconvolution processing is equal to the reduction factor of the size of the eigenvalue matrix of the same channel in the maximum pooling processing;
[0069] The third model is used to output a model output image of the image to be processed according to the processing features corresponding to each channel.
[0070] For example, before performing convolutional layer feature extraction, the eigenvalues can be normalized to reduce the subsequent computational effort of the model. Specifically, normalization involves maintaining the eigenvalues in the same dimension and keeping them between 0 and 1. For example, dividing all RGB values by 250 can achieve this effect. In the first model of this example, convolutional layer feature extraction includes at least one convolution, followed by activation function activation after each convolution. For example, the number of convolutions in a convolutional feature extraction process can be determined based on the size of the input eigenvalue matrix and the model's fitting capability. Using a 3×3 convolution kernel during convolution can effectively capture local features and improve computational efficiency; using 5×5 and 7×7 convolution kernels can capture features over a wider range. For example, by designing the convolution kernel size, stride, and padding, the size of the eigenvalue matrix before and after convolution can be kept constant. For example, a 3×3 convolution kernel, a stride of 1, and padding can keep the image size constant before and after convolution. The padding value is typically the default value of 0. In this example, activation functions such as the ReLU function after each convolution can alleviate the vanishing gradient problem. Leaky ReLU functions can mitigate the "neuron death" problem based on the ReLU function. Max pooling for downsampled feature extraction can reduce the size of feature maps while retaining important features. For example, a 2×2 pooling kernel with a stride of 2 can reduce the size of the input feature value matrix in downsampled feature extraction to half its original size.
[0071] In the second model of this example, the initial features corresponding to each channel are upsampled and extracted at least once to obtain processed features corresponding to each channel. The upsampled feature extraction includes deconvolution, feature concatenation, and convolutional layer feature extraction. The deconvolution process enlarges the size of the feature data for the same channel by the same factor as the size reduction factor for the feature data for the same channel during maximum pooling. Deconvolution can enlarge the size of the eigenvalue matrix of the input upsampled feature extraction by adjusting the deconvolution kernel size, stride, and padding. Exemplarily, the enlargement factor for the size of the eigenvalue matrix for the same channel during deconvolution is equal to the size reduction factor for the eigenvalue matrix for the same channel during maximum pooling, ensuring that the sizes of the two concatenated feature maps are the same. Exemplarily, feature concatenation can be performed by adding the eigenvalues of two 3×3 matrices to obtain a new 3×3 matrix; or by concatenating the two 3×3 matrices into a 3×3×2 matrix, and then converting the 3×3×2 matrix into a new 3×3 matrix through a single convolution.
[0072] The third model in this example outputs a model output image for the image to be processed based on the processing characteristics corresponding to each channel. The third model is an output model that can output a single image per channel through multiple convolutions based on output requirements, with the multiple channel output images serving as the model output image. Alternatively, a single convolution across multiple channels can output a single image, which serves as the model output image.
[0073] In this example, the U-Net model can extract downsampling features and sampling features of the image to be processed through the first model, the second model, and the third model, and finally output the model output image.
[0074] Figure 4 A schematic diagram of the structure of a U-Net model is shown in FIG. Figure 4 As shown, the image to be processed consists of four channels and an eigenvalue matrix for each channel. The first model uses two convolutions in the convolutional layer feature extraction, and the Leaky ReLU function is activated after each convolution, extracting 16 eigenvalue matrices from one channel and 64 eigenvalue matrices from four channels. The 64 eigenvalue matrices are then downsampled and extracted three times. Maximum pooling is performed in each downsampling feature extraction to reduce the feature matrix size by half. Furthermore, the number of convolution kernels in subsequent convolutional layer feature extraction is doubled, doubling the number of eigenvalue matrices. The second model uses three upsampling feature extractions on the 512 eigenvalue matrices obtained in the third downsampling feature extraction. Deconvolution is performed in each upsampling feature extraction to double the feature matrix size. Furthermore, the number of convolution kernels in subsequent convolutional layer feature extraction is reduced to half that of the previous convolutional layer feature extraction. The 64 eigenvalue matrices obtained in the third upsampling feature extraction are then output as an image using the third model, which serves as the model output image.
[0075] As yet another example, feature stitching includes:
[0076] Based on the same image channel, the characteristic values of the feature data in the deconvolution processing result and the characteristic values of the feature data in the result after the symmetrical convolution layer feature extraction in the first model are added.
[0077] Exemplarily, before feature splicing, if the eigenvalue matrix size of the feature data in the deconvolution processing result is the same as the eigenvalue matrix size of the feature data in the result after convolution layer feature extraction, then the two eigenvalues are added according to the coordinates of the matrices under the two feature data. The solution of this example adds the eigenvalues of the feature data in the deconvolution processing result and the eigenvalues of the feature data in the result after symmetrical convolution layer feature extraction in the first model based on the same image channel, and can perform feature fusion on the features after deconvolution processing and the features after convolution layer feature extraction in the first model, while reducing the computational complexity of feature fusion by adding eigenvalues.
[0078] As another example, Figure 5 A flow chart of an image color processing method is shown in FIG. Figure 5 As shown, based on any example, before step 102, the following steps are further included:
[0079] Step 401: Calculate a reduction ratio based on the target reduction size and the longer side length of the image to be processed, and scale the image to be processed according to the reduction ratio to obtain a scaled image;
[0080] Step 402: Determine whether the side length of the scaled image is a multiple of 2 to the power of n. If not, enlarge the corresponding side length to a multiple of 2 to the power of n; where n is a positive integer greater than or equal to 3.
[0081] In this example, the target size can be used to unify the image size and improve the computational efficiency of the model during batch processing. Among them, by scaling the image to be processed by reducing the ratio, the aspect ratio of the image to be processed can be maintained while reaching the target size, thereby avoiding image deformation. After obtaining the scaled image, the side length is enlarged to a multiple of 2 to the power of n, which can match the step size and pooling size in the downsampling and upsampling processes, thereby avoiding boundary effects in the convolution layer, pooling layer and deconvolution layer of the model. Exemplarily, when enlarging the side length of the image, linear interpolation or spline interpolation can be used. In the solution of this example, the image to be processed can reach the target size by reducing the ratio, and the image size can reach a multiple of 2 to the power of n by enlarging the side length, thereby improving the stability of the model and prediction accuracy.
[0082] As another example, Figure 6 The following is a flow chart showing an exemplary method for image color processing. Figure 6 It is shown that, based on any example, the image color processing method further includes:
[0083] Step 501: Perform polynomial mapping on the feature data of the image to be processed input into the U-Net model to obtain polynomial extraction features;
[0084] In step 103, based on the model output image, a color processing result of the image to be processed is determined, including:
[0085] Step 502: Using the model output image as a training target, establish an initial regression model based on the polynomial extracted features and the feature values corresponding to each image channel in the model output image, and perform model training to obtain a regression model;
[0086] Step 503: Input the polynomial extracted features into the regression model, and use the output of the regression model as the color processing result of the image to be processed.
[0087] In this example, the feature data of the image to be processed that is input to the U-Net model is subjected to polynomial mapping. Specifically, the feature values of each pixel of the image to be processed under different channels are memorized into polynomial mapping. For example, taking the polynomial mapping performed by the RGB standard as an example, the RGB value of a pixel before mapping is (r, g, b); the polynomial extracted features after polynomial mapping are (r, g, b, r*g, r*b, g*b, r 2 , g 2 , b 2 , r*g*b, 1); optionally, the polynomial mapping can also be other polynomial transformations. After obtaining the model output image, the model output image is used as the training target, an initial regression model of the polynomial extracted features and the eigenvalues corresponding to each image channel in the model output image is established and the model training is performed to obtain a regression model. The polynomial extracted features are input into the regression model, and the output result of the regression model is used as the color processing result of the image to be processed. Among them, by integrating the parameters obtained in the regression model, a mapping function is constructed, which can map the polynomial extracted features of the image to be processed to the three-channel feature space of the model output image. The scheme of this example can improve the color processing effect of the image while maintaining the consistency of the image content through the polynomial mapping of the image to be processed and the establishment of the regression model.
[0088] Figure 7 This schematic diagram illustrates a process flow for an image color processing method. After obtaining an image to be processed, it is scaled according to the reduction ratio and its side length is enlarged to obtain a resized image to be processed. This resized image to be processed is then input into a U-Net model and polynomial mapping is performed to obtain a model output image and polynomial extracted features. Using the model output image as a training target, an initial regression model is established between the polynomial extracted features and the eigenvalues corresponding to each image channel in the model output image, and the model is trained to obtain a regression model. After obtaining the regression model, the unresized image to be processed is input into the regression model to obtain the color processing result of the image to be processed.
[0089] The image color processing method provided in this embodiment includes: inputting feature data including multiple image channels and feature values corresponding to each image channel into a U-Net model output, and outputting an image based on the output model, thereby obtaining a color processing result of the processed image. The U-Net model is pre-trained based on an image dataset and a loss function, wherein the perceptual loss function is capable of extracting features from the image to obtain image details and structural information, while the color mask loss is established based on the entire image area and the area where at least one color range is located, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the image color processing effect while retaining the structure and detail information of the original image.
[0090] Example 2
[0091] Figure 8 FIG. 4 is a schematic diagram showing the structure of the image color processing device provided in an embodiment of the present application. Figure 8 As shown, the device includes:
[0092] An acquisition module 81 is configured to acquire feature data of an image to be processed; wherein the feature data includes a plurality of image channels and a feature value corresponding to each image channel;
[0093] a processing module 82 for inputting feature data of the image to be processed into a U-Net model to obtain a model output image of the image to be processed outputted by the U-Net model; wherein the U-Net model is pre-trained based on an image dataset and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and an area containing at least one color range in the image to be processed;
[0094] The determination module 83 is used to output the image based on the model and obtain the color processing result of the image to be processed.
[0095] In actual applications, there are many ways to implement the image color processing device. For example, it can be implemented through a computer program, such as application software; or it can be implemented as a medium that stores relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip.
[0096] In this example, the feature data of the image to be processed, including multiple image channels and the feature values corresponding to each image channel, are channels divided under different color standards and the feature values corresponding to each channel. For example, when the color standard is the RGB (Red, Green, Blue) standard, the feature data corresponds to the three channels of red, green and blue and the corresponding values of each pixel of the image to be processed under the three channels. Optionally, based on the current color standard, channels and the corresponding values of the channels can be added, such as adding a grayscale channel under the RGB channel and determining the corresponding value of each pixel of the image to be processed under the grayscale channel. Exemplarily, different color standards can be selected based on different color processing requirements, and the corresponding feature data can be determined; for example, the CMYK (Cyan, Magenta, Yellow, Key / Black) standard based on the ink color superposition during the printing process, or the HSV (Hue, Saturation, Value) standard based on the principle of human visual perception.
[0097] After obtaining the feature data corresponding to the object to be processed, it is input into the U-Net model. For example, a tensor can be created as the input function of the U-Net model. This tensor can include multiple dimensions, such as the number of images to be processed (i.e., the batch size), the height, width, and number of channels of the images to be processed, etc. In practical applications, by modifying the different dimensions of the tensor, the efficiency and quality of the U-Net model's image processing can be adjusted accordingly. The U-Net structure can be divided into two main parts: the encoder and the decoder, which form a U-shaped network architecture. The encoder consists of multiple convolutional and pooling layers, while the decoder gradually restores the spatial dimensions of the feature maps through upsampling. Skip connections connect the feature maps of corresponding encoder layers directly to the inputs of corresponding decoder layers. Specifically, upsampling can be achieved through deconvolution or transposed convolution with trainable parameters, or through upsampling layers without trainable parameters that utilize interpolation or repeated elements.
[0098] In this example, the U-Net model is pre-trained based on an image dataset and loss functions including a color mask loss and a perceptual loss function. The color mask loss function is established based on the entire image region of the image to be processed and the region containing at least one color range within the image to be processed. Depending on the model training objectives, high-resolution images of different themes, such as landscapes, people, or still lifes, can be selected as the image dataset. Optionally, selecting images of multiple themes as the image training set can increase the model's generalization capabilities. For example, when the U-Net model is used to train color processing methods such as color enhancement, color replacement, and color correction, an image dataset containing both the original and target images can be selected. For example, a perceptual loss function such as the Learned Perceptual Image Patch Similarity (LPIPS) perceptual loss function is used during training. This perceptual loss function utilizes features from multiple pre-trained convolutional neural networks and combines these features through learned weights to more accurately assess the perceptual quality of the image. Alternatively, other loss functions such as the VGG perceptual loss and style loss functions such as the Style Loss can be selected. The perceptual loss function compares the differences between the U-Net model's output image and the target image in the image dataset in high-level feature space, rather than simply at the pixel level. This allows the U-Net model to better capture image details and structural information. Furthermore, the U-Net model is trained using a color mask loss function. This color mask loss function, based on the entire image region of the processed image, can improve the overall color consistency between the U-Net model's output image and the processed image, avoiding color distortion or shift. Furthermore, a color mask loss function based on at least one color range in the processed image can improve the color fidelity of specific color regions and enhance the U-Net model's ability to recognize color features in color regions with key semantic color information. Exemplarily, the at least one color range can be determined by setting a threshold for the feature value or performing a cluster analysis on the pixels. Exemplarily, the color mask loss function can be implemented by calculating the mean squared error or mean absolute percentage error of the feature values for the same pixel in the processed image and the model output image.
[0099] After obtaining the model output image, the color processing result of the image to be processed can be obtained by denoising, resizing, and image format conversion. For example, in order to improve the performance of the U-Net model, image evaluation indicators can be introduced to evaluate and verify the model performance after each model training cycle. Specifically, the evaluation indicators may include: contrast, structural similarity index (SSIM) and vividness. The calculation formula for colorfulness under the RGB color standard is:
[0100]
[0101] The values of each pixel in the RGB three channels are R, G, and B, among which the difference between the red and green channels rg = |RG|, the difference between the blue and yellow channels yb = |0.5*(R+G)-B|, and the average value of the red and blue channels rb mean =(R+B) / 2.
[0102] In the image color processing device provided in this example, feature data including multiple image channels and feature values corresponding to each image channel are input into a U-Net model output, and an image is output based on the output model to obtain a color processing result of the processed image. The U-Net model is pre-trained based on an image dataset and a loss function, wherein the perceptual loss function can extract features from the image to obtain image details and structural information, while the color mask loss is established based on the entire image area and the area where at least one color range is located, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the effect of image color processing while retaining the structure and detail information of the original image.
[0103] In one example, the color range includes a black range and a white range. The processing module 82 is specifically configured to:
[0104] The area where the pixels in the image to be processed have a grayscale value greater than a first threshold value are located is used as a white mask, and the area where the pixels in the image to be processed have a grayscale value less than a second threshold value are located is used as a black mask;
[0105] Calculate a first loss function of the image to be processed and the model output image under the white mask, a second loss function of the image to be processed and the model output image under the black mask, and a third loss function of the image to be processed and the model output image under the entire image area;
[0106] A color mask loss function is determined according to the first loss function, the second loss function, and the third loss function.
[0107] For example, the first threshold can be a fixed value, such as a grayscale value of 220, or a grayscale range, such as grayscale values 215 to 225; the second threshold can be a fixed value, such as a grayscale value of 20, or a grayscale range, such as grayscale values 15 to 25. Optionally, when the color standard is HSV, the first threshold is a brightness value. In practical applications, the white mask or black mask is composed of a binary matrix. After processing the image to be processed and the model output image using the mask, pixel pairs for which the loss function calculation is required are obtained. The error of each pixel pair in each color channel is calculated using an error calculation method, such as the mean square error algorithm or the root mean square error algorithm, to obtain the first loss function or the second loss function. Similarly, the third loss function can be obtained using the above algorithm. After obtaining the first, second, and third loss functions, each loss function can be multiplied by the same or different gain coefficients and then summed to obtain the color mask loss function. Alternatively, any one or two of the loss functions can be multiplied by the same or different gain coefficients and then summed to obtain the color mask loss function. The gain coefficient can be determined based on images with varying lighting conditions or environments. Optionally, the gain coefficient can also be adjusted based on the U-Net model's output image. This example solution considers the fidelity of the global color, dark areas, and bright areas of the processed photo. This allows the U-Net model to improve color adjustment performance in dark and bright areas while maintaining global color.
[0108] In one example, the processing module 82 is specifically configured to:
[0109] The product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function are used as the color mask loss function.
[0110] For example, the values of white weight and black weight can be set to the same default value such as 0.7. According to the changes in the proportion of dark areas and light areas in the image, the corresponding weight range can be increased or decreased. For example, when processing an image dataset with a large proportion of light areas, the white weight is increased and the black weight is decreased. For example, the weight adjustment range can be any value between 0.5 and 1.0. Color mask loss function L mask The formula for (x,y) is:
[0111] L(x,y)=L MSE (x,y)+L MSE_black (x,y)*W black +L MSE_white (x,y)*W white
[0112] Among them, x is the image to be processed, y is the model output image, and the third loss function LMSE (x, y) is the mean square error calculation of x, y in the entire image area, and the second loss function L MSE_black (x,y) is the mean square error calculation of x,y in the black mask area, W black is the black weight, L MSE_white (x,y) is the mean square error calculation of x,y in the white mask area, W white is white weight.
[0113] This example uses the product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function as the color mask loss function. This improves the flexibility of the U-Net model in adjusting the colors of dark and light areas.
[0114] In one example, the processing module 82 is further configured to:
[0115] Perform feature extraction on the processed image based on the VGG network to obtain the first extracted feature before the first fully connected layer of the VGG network;
[0116] Perform feature extraction on the model output image based on the VGG network to obtain the second extracted feature before the first fully connected layer of the VGG network;
[0117] The loss function of the first extracted feature and the second extracted feature is used as the perceptual loss function.
[0118] As a deep convolutional neural network architecture, the VGG network's intermediate layer features can capture high-level semantic information of the image. By calculating the difference between the image to be processed and the model output image in the intermediate layer feature space of the VGG network, the perceptual quality of the image can be better measured. At the same time, the VGG network has transfer learning capabilities, which enables the pre-trained VGG network to effectively extract image features when the perceptual loss function is established. In practical applications, VGG16 or VGG19 can be used for feature extraction. The extracted features before the first fully connected layer of the VGG network have the following beneficial effects: it inherits the rich feature information extracted by the previous multiple intermediate layers, and can capture the complex structure and object relationships in the image. At the same time, only using the features output by a single neural layer for the perceptual loss function calculation can reduce the amount of calculation. In this example, the perceptual loss function L perceptual The formula for (x,y) is:
[0119]
[0120] Among them, x is the image to be processed, y is the model output image, N is the number of extracted features of the VGG network before the first fully connected layer, φ i(x) represents the i-th feature extracted by the VGG network for the processed image, φ i (y) represents the i-th feature extracted by the VGG network for the model output image, and ‖·‖1 represents the L1 norm. Here, φ(x) corresponds to the first extracted feature in this example, and φ(y) corresponds to the second extracted feature in this example.
[0121] In one example, the U-Net model includes a first model, a second model, and a third model;
[0122] The first model is used to perform convolutional layer feature extraction and at least one downsampling feature extraction on the feature data of the image to be processed to obtain the initial features corresponding to each channel; wherein the convolutional layer feature extraction includes at least one convolution and activation function activation after each convolution; the downsampling feature extraction includes maximum pooling processing and convolutional layer feature extraction;
[0123] The second model is used to perform at least one upsampling feature extraction on the initial features corresponding to each channel to obtain processed features corresponding to each channel; wherein the upsampling feature extraction includes: deconvolution processing, feature splicing and convolution layer feature extraction, and the magnification factor of the size of the eigenvalue matrix of the same channel in the deconvolution processing is equal to the reduction factor of the size of the eigenvalue matrix of the same channel in the maximum pooling processing;
[0124] The third model is used to output a model output image of the image to be processed according to the processing features corresponding to each channel.
[0125] For example, before performing convolutional layer feature extraction, the eigenvalues can be normalized to reduce the subsequent computational effort of the model. Specifically, normalization involves maintaining the eigenvalues in the same dimension and keeping them between 0 and 1. For example, dividing all RGB values by 250 can achieve this effect. In the first model of this example, convolutional layer feature extraction includes at least one convolution, followed by activation function activation after each convolution. For example, the number of convolutions in a convolutional feature extraction process can be determined based on the size of the input eigenvalue matrix and the model's fitting capability. Using a 3×3 convolution kernel during convolution can effectively capture local features and improve computational efficiency; using 5×5 and 7×7 convolution kernels can capture features over a wider range. For example, by designing the convolution kernel size, stride, and padding, the size of the eigenvalue matrix before and after convolution can be kept constant. For example, a 3×3 convolution kernel, a stride of 1, and padding can keep the image size constant before and after convolution. The padding value is typically the default value of 0. In this example, activation functions such as the ReLU function after each convolution can alleviate the vanishing gradient problem. Leaky ReLU functions can mitigate the "neuron death" problem based on the ReLU function. Max pooling for downsampled feature extraction can reduce the size of feature maps while retaining important features. For example, a 2×2 pooling kernel with a stride of 2 can reduce the size of the input feature value matrix in downsampled feature extraction to half its original size.
[0126] In the second model of this example, the initial features corresponding to each channel are upsampled and extracted at least once to obtain processed features corresponding to each channel. The upsampled feature extraction includes deconvolution, feature concatenation, and convolutional layer feature extraction. The deconvolution process enlarges the size of the feature data for the same channel by the same factor as the size reduction factor for the feature data for the same channel during maximum pooling. Deconvolution can enlarge the size of the eigenvalue matrix of the input upsampled feature extraction by adjusting the deconvolution kernel size, stride, and padding. Exemplarily, the enlargement factor for the size of the eigenvalue matrix for the same channel during deconvolution is equal to the size reduction factor for the eigenvalue matrix for the same channel during maximum pooling, ensuring that the sizes of the two concatenated feature maps are the same. Exemplarily, feature concatenation can be performed by adding the eigenvalues of two 3×3 matrices to obtain a new 3×3 matrix; or by concatenating the two 3×3 matrices into a 3×3×2 matrix, and then converting the 3×3×2 matrix into a new 3×3 matrix through a single convolution.
[0127] The third model in this example outputs a model output image for the image to be processed based on the processing characteristics corresponding to each channel. The third model is an output model that can output a single image per channel through multiple convolutions based on output requirements, with the multiple channel output images serving as the model output image. Alternatively, a single convolution across multiple channels can output a single image, which serves as the model output image.
[0128] In this example, the U-Net model can extract downsampling features and sampling features of the image to be processed through the first model, the second model, and the third model, and finally output the model output image.
[0129] In one example, feature stitching includes:
[0130] Based on the same image channel, the characteristic values of the feature data in the deconvolution processing result and the characteristic values of the feature data in the result after the symmetrical convolution layer feature extraction in the first model are added.
[0131] Exemplarily, before feature splicing, if the eigenvalue matrix size of the feature data in the deconvolution processing result is the same as the eigenvalue matrix size of the feature data in the result after convolution layer feature extraction, then the two eigenvalues are added according to the coordinates of the matrices under the two feature data. The solution of this example adds the eigenvalues of the feature data in the deconvolution processing result and the eigenvalues of the feature data in the result after symmetrical convolution layer feature extraction in the first model based on the same image channel, and can perform feature fusion on the features after deconvolution processing and the features after convolution layer feature extraction in the first model, while reducing the computational complexity of feature fusion by adding eigenvalues.
[0132] In one example, the processing module 82 is further configured to:
[0133] Calculating a reduction ratio according to a target reduction size and a longer side length of the image to be processed, and scaling the image to be processed according to the reduction ratio to obtain a scaled image;
[0134] Determine whether the side length of the scaled image is a multiple of 2 to the power of n. If not, enlarge the corresponding side length to a multiple of 2 to the power of n; where n is a positive integer greater than or equal to 3.
[0135] In this example, the target size can be used to unify the image size and improve the computational efficiency of the model during batch processing. Among them, by scaling the image to be processed by reducing the ratio, the aspect ratio of the image to be processed can be maintained while reaching the target size, thereby avoiding image deformation. After obtaining the scaled image, the side length is enlarged to a multiple of 2 to the power of n, which can match the step size and pooling size in the downsampling and upsampling processes, thereby avoiding boundary effects in the convolution layer, pooling layer and deconvolution layer of the model. Exemplarily, when enlarging the side length of the image, linear interpolation or spline interpolation can be used. In the solution of this example, the image to be processed can reach the target size by reducing the ratio, and the image size can reach a multiple of 2 to the power of n by enlarging the side length, thereby improving the stability of the model and prediction accuracy.
[0136] In one example, the processing module 82 is further configured to:
[0137] Perform polynomial mapping on the feature data of the image to be processed input to the U-Net model to obtain polynomial extraction features;
[0138] The determination module 83 is further configured to:
[0139] Taking the model output image as the training target, an initial regression model is established between the polynomial extracted features and the feature values corresponding to each image channel in the model output image, and the model training is performed to obtain a regression model;
[0140] The polynomial extracted features are input into the regression model, and the output result of the regression model is used as the color processing result of the image to be processed.
[0141] In this example, the feature data of the image to be processed that is input to the U-Net model is subjected to polynomial mapping. Specifically, the feature values of each pixel of the image to be processed under different channels are memorized into polynomial mapping. For example, taking the polynomial mapping performed by the RGB standard as an example, the RGB value of a pixel before mapping is (r, g, b); the polynomial extracted features after polynomial mapping are (r, g, b, r*g, r*b, g*b, r 2 , g 2 , b 2, r*g*b, 1); optionally, the polynomial mapping can also be other polynomial transformations. After obtaining the model output image, the model output image is used as the training target, an initial regression model of the polynomial extracted features and the eigenvalues corresponding to each image channel in the model output image is established and the model training is performed to obtain a regression model. The polynomial extracted features are input into the regression model, and the output result of the regression model is used as the color processing result of the image to be processed. Among them, by integrating the parameters obtained in the regression model, a mapping function is constructed, which can map the polynomial extracted features of the image to be processed to the three-channel feature space of the model output image. The scheme of this example can improve the color processing effect of the image while maintaining the consistency of the image content through the polynomial mapping of the image to be processed and the establishment of the regression model.
[0142] The image color processing device provided in this embodiment includes: inputting feature data including multiple image channels and feature values corresponding to each image channel into a U-Net model output, and outputting an image based on the output model, thereby obtaining a color processing result of the processed image. The U-Net model is pre-trained based on an image dataset and a loss function, wherein the perceptual loss function can extract features from the image to obtain image details and structural information, while the color mask loss is established based on the entire image area and the area where at least one color range is located, taking into account the global color and partial color range of the image, so that the color result of the image to be processed can improve the image color processing effect while retaining the structure and detail information of the original image.
[0143] Example 3
[0144] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, the electronic device comprising:
[0145] The electronic device includes a processor 291 and a memory 292; a communication interface 293, and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via bus 294. Communication interface 293 can be used for information transmission. The processor 291 can invoke logic instructions in memory 292 to execute the method described above.
[0146] In addition, the logic instructions in the memory 292 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0147] Memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present application. Processor 291 executes the software programs, instructions, and modules stored in memory 292 to execute functional applications and data processing, thereby implementing the methods in the above-mentioned method examples.
[0148] Memory 292 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Memory 292 may also include high-speed random access memory and non-volatile memory.
[0149] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method in any embodiment.
[0150] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method in any embodiment when the computer program is executed by a processor.
[0151] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0152] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for image color processing, characterized in that: include: Acquire feature data of an image to be processed; wherein the feature data includes multiple image channels and feature values corresponding to each image channel; Inputting feature data of the image to be processed into a U-Net model to obtain a model output image of the image to be processed output by the U-Net model; wherein the U-Net model is pre-trained based on an image dataset and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and an area where at least one color range in the image to be processed is located; Based on the model output image, a color processing result of the image to be processed is obtained.
2. The method according to claim 1, characterized in that The color range includes a black range and a white range; the color mask loss function is established based on the entire image area of the image to be processed and the area where at least one color range in the image to be processed is located, including: Using the area where the pixels in the image to be processed have grayscale values greater than a first threshold as a white mask, and using the area where the pixels in the image to be processed have grayscale values less than a second threshold as a black mask; Calculating a first loss function of the image to be processed and the model output image under the white mask, a second loss function of the image to be processed and the model output image under the black mask, and a third loss function of the image to be processed and the model output image under the entire image area; Determine the color mask loss function according to the first loss function, the second loss function, and the third loss function.
3. The method according to claim 2, characterized in that The determining the color mask loss function according to the first loss function, the second loss function, and the third loss function includes: The product of the first loss function and the white weight, the product of the second loss function and the black weight, and the sum of the third loss function are used as the color mask loss function.
4. The method according to claim 1, wherein The method further comprises: Performing feature extraction on the image to be processed based on a VGG network to obtain a first extracted feature before a first fully connected layer of the VGG network; Performing feature extraction on the model output image based on a VGG network to obtain a second extracted feature before a first fully connected layer of the VGG network; The loss function of the first extracted feature and the second extracted feature is used as the perceptual loss function.
5. The method according to claim 1, wherein The U-Net model includes a first model, a second model and a third model; The first model is used to perform convolutional layer feature extraction and at least one downsampling feature extraction on the feature data of the image to be processed to obtain initial features corresponding to each channel; wherein the convolutional layer feature extraction includes at least one convolution and activation function activation after each convolution; the downsampling feature extraction includes maximum pooling processing and the convolutional layer feature extraction; The second model is used to perform at least one upsampling feature extraction on the initial features corresponding to each channel to obtain processed features corresponding to each channel; wherein the upsampling feature extraction includes: deconvolution processing, feature splicing and the convolution layer feature extraction, and the magnification factor of the size of the eigenvalue matrix of the same channel in the deconvolution processing is equal to the reduction factor of the size of the eigenvalue matrix of the same channel in the maximum pooling processing; The third model is used to output the model output image of the image to be processed according to the processing features corresponding to each channel.
6. The method according to claim 5, characterized in that The feature splicing includes: Based on the same image channel, the characteristic values of the feature data in the deconvolution processing result and the characteristic values of the feature data in the result after the symmetrical convolution layer feature extraction in the first model are added.
7. The method according to claim 1, characterized in that Before inputting the feature data of the image to be processed into the U-Net model, the method further includes: Calculating a reduction ratio according to a target reduction size and a longer side length of the image to be processed, and scaling the image to be processed according to the reduction ratio to obtain a scaled image; Determine whether the side length of the scaled image is a multiple of 2 to the nth power. If not, enlarge the corresponding side length to a multiple of 2 to the nth power; wherein n is a positive integer greater than or equal to 3.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Performing polynomial mapping on the feature data of the image to be processed input into the U-Net model to obtain polynomial extraction features; The step of determining a color processing result of the image to be processed based on the model output image includes: Taking the model output image as a training target, establishing an initial regression model between the polynomial extracted features and the feature values corresponding to each image channel in the model output image and performing model training to obtain a regression model; The polynomial extracted features are input into the regression model, and the output result of the regression model is used as the color processing result of the image to be processed.
9. An image color processing device, characterized in that: include: An acquisition module, configured to acquire feature data of an image to be processed; wherein the feature data includes a plurality of image channels and a feature value corresponding to each image channel; a processing module, configured to input the feature data of the image to be processed into a U-Net model, and obtain a model output image of the image to be processed output by the U-Net model; wherein the U-Net model is pre-trained based on an image dataset and a loss function including a color mask loss function and a perceptual loss function; the color mask loss function is established based on the entire image area of the image to be processed and an area containing at least one color range in the image to be processed; The determination module is used to output an image based on the model to obtain a color processing result of the image to be processed.
10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.
12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.