Image Color Enhancement Method, Device, Storage Medium, and Electronic Device

The image color enhancement is achieved by extracting three-dimensional grid features through deep neural networks, which solves the problems of limited effects and high computational volume in the existing technology, and achieves efficient image color enhancement that adapts to diverse scenarios.

CN114359100BActive Publication Date: 2025-07-08GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111679517.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-08
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In the prior art, the image color enhancement effect is limited, and it is difficult to adapt to diverse actual scenarios, and the deep neural network has a large amount of computing and high cost.

Method used

The pre-trained deep neural network is used to extract features based on three-dimensional grids, generate an information matrix for image color enhancement, divide the airspace and pixel value domains through the three-dimensional grid, reduce the amount of direct output image calculations, and adapt to diverse scenarios.

Benefits of technology

It improves the image color enhancement effect, breaks through the limitations of artificially defined scenes, adapts to diverse scenes, reduces the computing volume of deep neural networks, and realizes a lightweight network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359100B_ABST
    Figure CN114359100B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image color enhancement method, apparatus, storage medium, and electronic device, relating to the technical field of image and video processing. The image color enhancement method includes: obtaining an image to be processed; extracting three-dimensional grid-based features from the image to be processed through a pre-trained deep neural network, and generating an information matrix according to the extracted features, where the three-dimensional grid is obtained by dividing a three-dimensional space formed by the spatial domain and pixel value domain of the image to be processed; and performing color enhancement processing on the image to be processed by using the information matrix to obtain a color-enhanced image corresponding to the image to be processed. The present disclosure improves the effect of image color enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image and video processing, and particularly to an image color enhancement method, an image color enhancement device, a computer-readable storage medium, and an electronic device. Background Art

[0002] Image color enhancement refers to beautifying the colors of an image (or video frame) according to the picture scene or a set image style to better meet the aesthetic needs of users. For example, enhancing the colors of an image of a sunset scene to render the picture with a more ambient atmosphere.

[0003] In the related art, the effect of image color enhancement needs to be improved.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present disclosure provides an image color enhancement method, an image color enhancement device, a computer-readable storage medium, and an electronic device, thereby at least to some extent improving the effect of image color enhancement.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, there is provided an image color enhancement method, including: obtaining an image to be processed; extracting three-dimensional grid-based features from the image to be processed through a pre-trained deep neural network, and generating an information matrix according to the extracted features, where the three-dimensional grid is obtained by partitioning a three-dimensional space formed by the spatial domain and the pixel value domain of the image to be processed; and performing color enhancement processing on the image to be processed by using the information matrix to obtain a color-enhanced image corresponding to the image to be processed.

[0008] According to a second aspect of the present disclosure, there is provided an image color enhancement device, including: an image acquisition module configured to obtain an image to be processed; an information matrix generation module configured to extract three-dimensional grid-based features from the image to be processed through a pre-trained deep neural network, and generate an information matrix according to the extracted features, where the three-dimensional grid is obtained by partitioning a three-dimensional space formed by the spatial domain and the pixel value domain of the image to be processed; and a color enhancement processing module configured to perform color enhancement processing on the image to be processed by using the information matrix to obtain a color-enhanced image corresponding to the image to be processed.

[0009] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the image color enhancement method of the first aspect and its possible implementation manners as described above.

[0010] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the image color enhancement method of the first aspect and its possible implementation manners by executing the executable instructions.

[0011] The technical solution of the present disclosure has the following beneficial effects:

[0012] Based on the image color enhancement method of the present disclosure, on the one hand, through the processing of the image to be processed by a deep neural network, an information matrix for color enhancement processing is obtained, so that the information matrix is adapted to the scene, style, etc. of the image to be processed, which is beneficial to improving the effect of image color enhancement, and this solution can break through the limitation of artificially defined scenes and adapt to diverse actual scenes. On the other hand, the deep neural network in this solution is used to output the information matrix and does not directly output the color-enhanced image, thereby reducing the computational amount of the deep neural network, which is beneficial to realizing a lightweight network and reducing the implementation cost of the solution.

[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0015] Figure 1 A schematic diagram showing a system architecture in this exemplary embodiment;

[0016] Figure 2 A flowchart showing an image color enhancement method in this exemplary embodiment;

[0017] Figure 3 A schematic diagram showing the structure of a deep neural network in this exemplary embodiment;

[0018] Figure 4 A flowchart showing a process of obtaining an information matrix in this exemplary embodiment;

[0019] Figure 5Shows a schematic diagram of processing through a fusion layer in this exemplary embodiment;

[0020] Figure 6 Shows a sub - flowchart of an image color enhancement method in this exemplary embodiment;

[0021] Figure 7 Shows a schematic diagram of an image color enhancement method in this exemplary embodiment;

[0022] Figure 8 Shows a flowchart of training a deep neural network in this exemplary embodiment;

[0023] Figure 9 Shows a schematic diagram of training a deep neural network in this exemplary embodiment;

[0024] Figure 10 Shows another flowchart of training a deep neural network in this exemplary embodiment;

[0025] Figure 11 Shows another schematic diagram of training a deep neural network in this exemplary embodiment;

[0026] Figure 12 Shows a schematic flowchart of an image color enhancement method in this exemplary embodiment;

[0027] Figure 13 Shows a schematic structural diagram of an image color enhancement device in this exemplary embodiment;

[0028] Figure 14 Shows a schematic structural diagram of an electronic device in this exemplary embodiment. Detailed implementation manners

[0029] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that one or more of the specific details may be omitted, or other methods, components, devices, steps, etc. may be used. In other cases, well - known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0030] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] In a solution of the related art, a LUT (Look Up Table, color lookup table) is used to perform image color enhancement. The basic process of this solution is as follows: perform scene recognition on the image; select the corresponding LUT according to the result of the scene recognition; use this LUT to perform a look-up mapping on each pixel value in the image to complete the image color enhancement. However, since the artificially defined scene categories are relatively limited (usually dozens), and each scene only corresponds to a fixed LUT, it is difficult for this solution to adapt to diverse actual scenes, affecting the effect of image color enhancement.

[0032] In view of the above one or more problems, the exemplary embodiments of the present disclosure provide an image color enhancement method for performing color enhancement processing on an image or a video frame. The following combines Figure 1 An exemplary description is given of the system architecture and application scenarios of the operating environment of this exemplary embodiment.

[0033] Figure 1 A schematic diagram of the system architecture is shown. The system architecture 100 may include a terminal 110 and a server 120. Among them, the terminal 110 may be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc. The server 120 generally refers to a background system that provides services related to image color enhancement in this exemplary embodiment, and may be a single server or a cluster formed by multiple servers. A connection may be formed between the terminal 110 and the server 120 through a wired or wireless communication link for data interaction.

[0034] In one embodiment, the terminal 110 can capture or obtain an image or video to be processed in other ways and upload it to the server 120. For example, the user opens an image processing related App (Application, an application program, and the image processing related App includes a beauty App, etc.) on the terminal 110, selects the image or video to be processed from the album, and uploads it to the server 120 for color enhancement. Or the user opens the color enhancement function in a video processing related App (such as a live broadcast App, an App with a video call function, etc.) on the terminal 110 and uploads the real-time captured video to the server 120 for beauty. The server 120 executes the above image color enhancement method to obtain the color-enhanced image or video and returns it to the terminal 110.

[0035] In one embodiment, the server 120 can perform the training of the deep neural network, send the trained deep neural network to the terminal 110 for deployment. For example, the relevant data of the deep neural network is packaged in the update package of the above image processing related App, so that the terminal 110 can obtain the deep neural network by updating the App and deploy it locally. Further, after the terminal 110 captures or obtains an image or video to be processed in other ways, it can call the deep neural network to implement the color enhancement processing of the image or video by executing the above image color enhancement method.

[0036] In one embodiment, the terminal 110 can perform the training of the deep neural network. For example, it obtains the basic architecture of the deep neural network from the server 120 and trains it through the local data set, or obtains the data set from the server 120 and trains the locally constructed deep neural network, or trains the deep neural network without relying on the server 120 at all. Further, the terminal 110 can call the deep neural network to implement the color enhancement processing of the image or video by executing the above image color enhancement method.

[0037] As can be seen from the above, the execution subject of the image color enhancement method in this exemplary embodiment can be the above terminal 110 or server 120, and the present disclosure does not limit this.

[0038] The following combines Figure 2 to illustrate the image color enhancement method in this exemplary embodiment. Figure 2 It shows an exemplary process of the image color enhancement method, which may include:

[0039] Step S210, obtain the image to be processed;

[0040] Step S220: Extract 3D grid-based features from the image to be processed through a pre-trained deep neural network, and generate an information matrix based on the extracted features. The 3D grid is obtained by dividing the 3D space formed by the spatial domain and the pixel value domain of the image to be processed.

[0041] Step S230: Use the information matrix to perform color enhancement processing on the image to be processed, and obtain a color-enhanced image corresponding to the image to be processed.

[0042] Based on the above method, on the one hand, through the processing of the image to be processed by the deep neural network, an information matrix for color enhancement processing is obtained, making the information matrix adaptable to the scene, style, etc. of the image to be processed, which is beneficial to improving the effect of image color enhancement, and this solution can break through the limitation of artificially defined scenes and adapt to diverse actual scenes. On the other hand, the deep neural network in this solution is used to output the information matrix and does not directly output the color-enhanced image, thereby reducing the computational amount of the deep neural network, being beneficial to realizing a lightweight network, and reducing the implementation cost of the solution.

[0043] Next, a specific description is given for Figure 2 each step in

[0044] Referring to Figure 2 , in step S210, the image to be processed is obtained.

[0045] The image to be processed is an image that needs to be subjected to color enhancement processing. It should be understood that image color enhancement can be a link in image processing. In addition, other aspects of image processing can also be performed, such as image deblurring, denoising, portrait beautification, etc. The present disclosure does not limit the sequence of image color enhancement and other image processing. For example, if image color enhancement is the first link in image processing, the image to be processed can be the original image; if image color enhancement is the last link in image processing, the image to be processed can be an image after image deblurring, denoising, portrait beautification, etc.

[0046] In one embodiment, the image to be processed can be a frame image in an image sequence. An image sequence refers to a sequence formed by multiple consecutive frame images, which can be a video or a series of consecutively captured images, etc. The image sequence can be an object that needs to be subjected to color enhancement processing. Taking a video as an example, it can be a currently live-captured or live-received video stream, or a complete video that has been captured or received, such as a video stored locally. The present disclosure does not limit parameters such as the frame rate and image resolution of the video. For example, the video frame rate can be 30fps (frames per second), 60fps, 120fps, etc., and the image resolution can be 720P, 1080P, 4K, etc. and corresponding different aspect ratios. Color enhancement processing can be performed on each frame image in the video, or a part of the images can be selected from the video for color enhancement processing, and the images that need to be subjected to color enhancement processing are used as the original beauty image or the image to be processed as described above. For example, when receiving a video stream in real time, each received frame image can be used as the image to be processed.

[0047] Continuing to refer to Figure 2 , in step S220, based on a three-dimensional grid, features of the image to be processed are extracted by a pre-trained deep neural network, and an information matrix is generated according to the extracted features. The three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the image to be processed.

[0048] The deep neural network (DNN) is used to output the information matrix. The information matrix is a parameter matrix for performing color enhancement processing on the image to be processed. That is to say, the deep neural network is used to indirectly implement image color enhancement processing.

[0049] The spatial domain of the image to be processed, i.e., the two-dimensional space where the image plane of the image to be processed is located, has two dimensions. The first dimension can be, for example, the width direction of the image, and the second dimension can be, for example, the height direction of the image. The pixel value range refers to the numerical range of the pixel values of the image to be processed. For example, it can be [0, 255], or if the pixel values are normalized, the pixel value range is [0, 1]. Taking the pixel value range as the third dimension, a three-dimensional space is formed with the above-mentioned first dimension and second dimension. In this exemplary embodiment, the three-dimensional space can be pre-divided, including dividing the spatial domain and dividing the pixel value range, to obtain a three-dimensional grid. The two-dimensional projection of the three-dimensional grid on the spatial domain is called the spatial domain grid; the one-dimensional projection of the three-dimensional grid on the pixel value range is called the value range partition. Exemplarily, a region of 16 pixels * 16 pixels can be used as the spatial domain grid, and [0, 1 / 8), [1 / 8, 1 / 4), [1 / 4, 3 / 8), etc. (dividing [0, 1] into 8 partitions evenly) can be used as the value range partition, so as to obtain a three-dimensional grid. Thus, features based on the three-dimensional grid can be extracted from the image to be processed, and an information matrix can be generated according to the extracted features.

[0050] In one embodiment, before inputting the image to be processed into the deep neural network, the image to be processed can be upsampled or downsampled according to the size of the input image required by the deep neural network. For example, if the size of the input image required by the deep neural network is 256 * 256, the image to be processed can be compressed to this size and then input into the deep neural network for processing.

[0051] In one embodiment, the structure of the deep neural network can refer to Figure 3 As shown, it includes four main parts: a basic feature extraction sub-network, a local feature extraction sub-network, a global feature extraction sub-network, and an output sub-network. Each part can include one or more intermediate layers. The local feature extraction sub-network and the global feature extraction sub-network are two parallel parts between the basic feature extraction sub-network and the output sub-network.

[0052] Refer to Figure 4 As shown, the above-mentioned steps of extracting features based on the three-dimensional grid from the image to be processed by the pre-trained deep neural network and generating an information matrix according to the extracted features can include the following steps S410 to S440:

[0053] Step S410, downsample the image to be processed according to the size of the spatial domain grid through the basic feature extraction sub-network to obtain basic features.

[0054] Through the downsampling process, the image to be processed can be converted into features on the scale of the spatial domain grid, and this feature is the basic feature. The form of the basic feature in this disclosure is not limited. For example, it can be a basic feature vector or a basic feature image.

[0055] In one embodiment, the downsampling process may include downsampling convolution processing, which refers to reducing the image size through convolution to achieve the downsampling effect. For example, a convolution layer with a stride greater than 1 can be used to implement the downsampling convolution processing.

[0056] Combined with Figure 3 For example, the dimension of the input image is (B, W, H, C), where B represents the number of images, which can be any positive integer, indicating that B images to be processed are taken as a batch and input into the deep neural network for processing; W represents the image width, H represents the image height, and C represents the number of image channels. When the image to be processed is an RGB image, C is 3. The size of the spatial grid is 16 pixels * 16 pixels. The basic feature extraction sub-network may include 4 3*3 convolution layers with a stride of 2 (3*3 represents the convolution kernel size, which is only exemplary and can also be replaced with other sizes). After the image to be processed is processed by it, both the height and width are reduced to 1 / 16; of course, in the present disclosure, convolution layers with other numbers and strides can also be set to achieve the same downsampling effect. For example, the above 4 3*3 convolution layers with a stride of 2 can be replaced with two 5*5 convolution layers with a stride of 4, etc. In addition, the basic feature extraction sub-network may further include one or more 3*3 convolution layers with a stride of 1 (3*3 represents the convolution kernel size, which is only exemplary and can also be replaced with other sizes) for further extracting features from the image after downsampling convolution without changing the scale of the features, to obtain the basic features; of course, setting the convolution layer with a stride of 1 is not necessary. The basic feature extraction sub-network can output the basic feature image corresponding to the image to be processed, and its dimension is (B, W / 16, H / 16, k1), where k1 represents the number of channels of the basic feature image and is related to the number of convolution kernels of the last convolution layer in the basic feature extraction sub-network, which is not limited in the present disclosure. For example Figure 3 as shown in

[0057] As can be seen from the above, the processing process of the basic feature extraction sub-network is to gradually extract features within the range of each spatial grid of the image to be processed, represent features of different dimensions in different channels, and finally obtain the basic features, which can be the features of the image to be processed at the scale of the spatial grid.

[0058] Step S420, extracting local features within the spatial grid of the basic features through the local feature extraction sub-network.

[0059] Based on the basic features extracted by the basic feature extraction sub-network, the local feature extraction sub-network can further extract deeper features within the spatial grid range to obtain local features.

[0060] Combined with Figure 3For example, the local feature extraction sub-network may include one or more 3*3 convolutional layers with a stride of 1 (3*3 represents the convolutional kernel size, which is only exemplary and can be replaced with other sizes) and a 1*1 convolutional layer with a stride of 1. The 3*3 convolutional layer is used to further extract local features from the basic features without changing the scale of the features, and the 1*1 convolutional layer is used to adjust the number of channels of the extracted local features. The local feature extraction sub-network can output a local feature image corresponding to the image to be processed. The dimension of the local feature image is (B, W / 16, H / 16, k2), where k2 represents the number of channels of the local features and is related to the number of convolutional kernels of the last convolutional layer in the local feature extraction sub-network. This disclosure does not make a limitation. For example Figure 3 shows that k2 is 64.

[0061] Step S430, extracting global features from the basic features through the global feature extraction sub-network.

[0062] Based on the basic feature image extracted by the basic feature extraction sub-network, the global feature extraction sub-network can further extract global features within the entire range of the image to be processed.

[0063] Combined with Figure 3 For example, the global feature extraction sub-network may include one or more convolutional layers (or a combination of convolutional layers and pooling layers) and one or more fully connected layers. The convolutional layers are used to further extract local features from the basic features, and the fully connected layers are used to fuse the local features to obtain global features. The global feature extraction sub-network can output global features corresponding to the image to be processed. The global features can be global feature vectors, and their dimension is (B, k2), that is, the dimension of the global feature vector is the same as the number of channels of the local feature image, so as to facilitate subsequent fusion.

[0064] Step S440, processing the local features and global features through the output sub-network according to the number of value range partitions to obtain an information matrix.

[0065] The local features and global features reflect the features of the image to be processed at different scales. The output sub-network realizes the fusion of the two aspects of features through further processing of the local features and global features, and obtains an information matrix that matches the image to be processed.

[0066] In one implementation, the output sub-network may include a fusion layer and a dimension conversion layer. The above process of processing the local features and global features through the output sub-network according to the number of value range partitions to obtain an information matrix may include the following steps:

[0067] Fusing the local features and global features into a comprehensive feature through the fusion layer;

[0068] The dimensionality transformation layer performs dimensionality transformation on the comprehensive features according to the number of value range partitions to obtain an information matrix.

[0069] Among them, the processing of the fusion layer can refer to Figure 5 As shown, taking the fusion process of the local features and the global features corresponding to a single image (i.e., the case of B = 1) as an example, the dimension of the local features is (W / 16, H / 16, 64), and the dimension of the global features is (1, 1, 64). In each channel, the value of the global feature is added to the whole of the local features. As Figure 5 shown, the value a1 of the global feature is added to each pixel point of the local features in the first channel, and the value a2 of the global feature is added to each pixel point of the local features in the second channel. Then, through the calculation of the activation function, the comprehensive features are output. Figure 5 It is shown that the activation function adopts ReLu (rectified linear unit). p1 in the first channel of the local features is added to the value a1 of the global feature in the first channel, and then through ReLu activation, the value ReLu(p1 + a1) of the comprehensive features is obtained; p2 in the second channel of the local features is added to the value a2 of the global feature in the second channel, and then through ReLu activation, the value ReLu(p2 + a2) of the comprehensive features is obtained. Of course, the specific form of the activation function in the present disclosure is not limited, and the result of directly adding the local features and the global features without passing through the calculation of the activation function can also be used as the comprehensive features.

[0070] As can be seen from the above, the number of channels of the comprehensive features is the same as that of the local features or the global features. The dimensionality transformation layer can further convert its number of channels into the number of value range partitions to correspond to the three-dimensional grid, and obtain the information matrix G. The dimensionality transformation layer can be implemented by one or more convolutional layers. For example, dimensionality transformation can be performed through a 1*1 convolutional layer with a stride of 1, and the number of convolutional kernels of this convolutional layer can be set to the number of value range partitions G_c = G_z * G_n. G_z is the number of value range partitions. For example, when dividing the three-dimensional grid, if the pixel value range is equally divided into 8 parts, then G_z is 8; G_n is the dimension of the sub-information matrix gi corresponding to each three-dimensional grid (i.e., the number of elements of the sub-information matrix gi), and i represents the ordinal number of the three-dimensional grid. After the comprehensive features pass through the convolution of this convolutional layer, the information matrix G is obtained, and its dimension is (B, W / 16, H / 16, G_c).

[0071] It should be understood that Figure 3 and Figure 5 the fusion method of the local features and the global features shown in

[0072] In one embodiment, the information matrix G obtained in step S440 can be regarded as a set of sub-information matrices gi. For each image to be processed, the deep neural network can output its corresponding information matrix G, including W / 16*H / 16*G_z sub-information matrices gi, and W / 16*H / 16*G_z is exactly the number of three-dimensional grids, that is, the information matrix G includes the sub-information matrix gi corresponding to each three-dimensional grid.

[0073] The above has described how to obtain the information matrix. Continuing to refer to Figure 2 , in step S230, the information matrix is used to perform color enhancement processing on the image to be processed, and a color-enhanced image corresponding to the image to be processed is obtained.

[0074] Generally, the pixel values of the image to be processed can be multiplied by the information matrix to implement numerical conversion of the pixel values and obtain a color-enhanced image.

[0075] In one embodiment, the information matrix may include a reference information matrix corresponding to each three-dimensional grid, and this reference information matrix is equivalent to the above-mentioned sub-information matrix gi. Referring to Figure 6 as shown, the above-mentioned use of the information matrix to process the image to be processed to obtain a color-enhanced image corresponding to the image to be processed may include the following steps S610 and S620:

[0076] Step S610: Interpolate the reference information matrix based on the image to be processed to obtain a color-enhanced information matrix corresponding to each pixel point of the image to be processed.

[0077] The reference information matrix can be the reference information for color enhancement processing of all pixel points within the three-dimensional grid, and can be regarded as a summary of the information required for color enhancement processing of all pixel points within the three-dimensional grid. The color-enhanced information matrix is the specific information for color enhancement processing of each pixel point. The reference information matrix can further correspond to the reference point of the three-dimensional grid. For example, this reference point can be the center point of the three-dimensional grid. Since each pixel point of the image to be processed is distributed at different positions within its respective three-dimensional grid and has an offset relative to the reference point within the three-dimensional grid, the reference information matrix can be interpolated to obtain a color-enhanced information matrix corresponding to each pixel point of the image to be processed.

[0078] In one embodiment, interpolation can be performed on one or more reference information matrices according to the offset of each pixel point of the image to be processed relative to the center point of one or more three-dimensional grids, so as to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed. Exemplarily, assume that the width of the image to be processed is 128, the height is also 128, and the size of the spatial domain grid is 16 pixels * 16 pixels. Then, both the first dimension and the second dimension of the three-dimensional space are equally divided into 8 parts; the pixel value range [0, 1] is also equally divided into 8 value range partitions, so that the three-dimensional space is divided into 8 * 8 * 8 three-dimensional grids. Represent the three-dimensional grid located at the upper left corner of the image to be processed with a pixel value in the range [0, 1 / 8) as {0, 0, 0}, and the center point coordinates of this three-dimensional grid are (8, 8, 1 / 16); obtain the pixel points within this three-dimensional grid in the image to be processed, calculate the offset of each pixel point from the center point for each pixel point, including the offset amounts in the first dimension, the second dimension, and the third dimension, and perform trilinear interpolation based on the reference information matrix of the {0, 0, 0} three-dimensional grid and the reference information matrices of its adjacent three-dimensional grids {1, 0, 0}, {0, 1, 0}, and {0, 0, 1} according to the offset amounts, so as to obtain the color enhancement information matrix corresponding to each pixel point in the {0, 0, 0} three-dimensional grid. It should be noted that if the three-dimensional grid is not on the boundary, trilinear interpolation can be performed based on the reference information matrix of this three-dimensional grid and the reference information matrices of its adjacent 6 three-dimensional grids, so as to obtain the color enhancement information matrix corresponding to each pixel point in this three-dimensional grid.

[0079] It should be understood that the present disclosure does not limit the specific interpolation algorithm. For example, a non-linear interpolation algorithm can also be used.

[0080] As can be seen from the above, when performing interpolation, it is necessary to calculate the offset between the pixel value of the pixel point and the pixel value of the reference point, that is, the offset between the pixel point and the reference point in the third dimension. When the image to be processed is a single-channel image, the pixel value of the image to be processed can be directly used for calculation. When the image to be processed is a multi-channel image, it is difficult to calculate based on the pixel values of multiple channels and the pixel value of the reference point. Based on this, in one embodiment, the above-mentioned interpolation of the reference information matrix based on the image to be processed to obtain the color enhancement information matrix corresponding to each pixel point of the image to be processed may include the following steps:

[0081] When the image to be processed is a multi-channel image, convert the image to be processed into a single-channel reference value image;

[0082] Perform interpolation on the reference information matrix based on the reference value image to obtain the color enhancement information matrix corresponding to each pixel point of the image to be processed.

[0083] Among them, the reference value image is an image that represents the multi-channels of the image to be processed through a single channel. When the image to be processed is an RGB image, the reference value image can be its corresponding grayscale image, and the grayscale can adopt normalized values with a value range of [0, 1].

[0084] In one implementation, the following formula can be used to convert the image to be processed into a single-channel reference value image:

[0085]

[0086] where R, G, and B are the normalized pixel values of each pixel in the image to be processed; n represents dividing the value ranges of R, G, and B into n partitions, and j represents the ordinal number of the partition; a rj , a gj , a bj are the conversion coefficients for each partition of R, G, and B respectively, which can be determined according to experience or actual requirements; shift rj , shift gj , shift bj are the conversion thresholds set in each partition of R, G, and B respectively, indicating that only pixel values greater than this conversion threshold are converted, and the conversion threshold can be set according to experience or actual requirements; guidemap r , guidemap g , guidemap b are the single-channel images of R, G, and B after partition conversion respectively; g r , g g , g b are the fusion coefficients of R, G, and B respectively, which can be empirical coefficients; guidemap bias is the offset added after fusion, which can also be determined according to experience; guidemap z is the reference value image with a value range of [0, 1].

[0087] In one implementation, the above a rj , a gj , a bj , shift rj , shift gj , shift bj , g r , g g , g b , guidemap bias and other parameters can also be obtained through preset model training. By setting the initial values of the model, the value range of the finally obtained reference value image satisfies [0, 1].

[0088] The reference information matrix in the information matrix G can be interpolated based on the reference value image to obtain the color enhancement information matrix corresponding to each pixel of the image to be processed. These color enhancement information matrices can be regarded as a set with a dimension of (B, W, H, G_n).

[0089] Step S620: Process each pixel of the image to be processed according to the color enhancement information matrix corresponding to each pixel of the image to be processed, to obtain a color-enhanced image.

[0090] The pixel value of each pixel can be multiplied by the corresponding color enhancement information matrix to obtain the processed pixel value, thereby forming a color-enhanced image. Exemplarily, the pixel value of pixel i is represented as a pixel value vector [r, g, b], and its corresponding color enhancement information matrix is:

[0091]

[0092] Then there is the following relationship:

[0093]

[0094] Among them, [r′ g′ b′] represents the pixel value after beauty treatment.

[0095] In one implementation manner, the above-mentioned process of processing each pixel of the image to be processed according to the color enhancement information matrix corresponding to each pixel of the image to be processed to obtain a color-enhanced image may include:

[0096] Add a new channel to the image to be processed according to the dimension of the color enhancement information matrix, and set the new channel to a preset value;

[0097] Multiply the pixel value vector of each pixel of the image to be processed by the color enhancement information matrix corresponding to each pixel respectively to obtain a color-enhanced image; the pixel value vector of each pixel is a vector formed by the values of each channel of each pixel.

[0098] Among them, the dimension of the color enhancement information matrix represents the number of rows and columns of the color enhancement information matrix. It can be seen from formula (2) that the pixel value vector of each pixel needs to be cross-multiplied with the color enhancement information matrix, indicating that the dimension of the pixel value vector needs to be the same as the number of rows of the color enhancement information matrix. And the dimension of the pixel value vector is equal to the number of channels of the image to be processed. Therefore, if the number of channels of the image to be processed is not equal to (generally less than) the number of rows of the color enhancement information matrix, a new channel can be added to the image to be processed. For the added new channel, a preset value can be filled, such as 1. Thus, it is equivalent to converting the pixel value vector of each pixel in the image to be processed into a homogeneous vector.

[0099] Exemplarily, assume that the color enhancement information matrix corresponding to pixel point i is as follows:

[0100]

[0101] That is, the number of rows of this color enhancement information matrix is 4. The image to be processed is an RGB image with 3 channels. Therefore, a new channel needs to be added and the new channel is uniformly filled with the value 1. Then the pixel value vector of pixel point i is [r, g, b, 1], thus satisfying the following relationship:

[0102]

[0103] Thus, through the processing of the information matrix, a color enhanced image is obtained, and its dimension is (B, W, H, C). In formulas (2) and (3), C = 3. The dimension of the color enhanced image is the same as that of the image to be processed, indicating that the color enhancement processing process of this exemplary embodiment does not change the image dimension.

[0104] If the pixel values of the image to be processed are normalized before being input into the deep neural network, then after obtaining the color enhanced image, its pixel values can be denormalized. For example, the pixel values in the range of [0, 1] can be uniformly multiplied by 255 to obtain pixel values in the range of [0, 255].

[0105] Figure 7 Shows a schematic flow of the image color enhancement method. Input the image to be processed with the dimension of (B, W, H, 3) into the deep neural network, and output the information matrix G with the dimension of (B, W / 16, H / 16, G_z * G_n), which includes the reference information matrix corresponding to each three-dimensional grid. Convert the image to be processed into a single-channel reference value image with the dimension of (B, W, H, 1). Interpolate the reference information matrix based on the reference value image to obtain the color enhancement information matrix with the dimension of (B, W, H, G_n), which includes the color enhancement information matrix corresponding to each pixel point. Finally, multiply the image to be processed by the color enhancement information matrix to obtain the color enhanced image with the dimension of (B, W, H, 3), which is the same as the dimension of the image to be processed.

[0106] In one embodiment, the image to be processed can be a frame image in an image sequence. The image color enhancement method may further include the following steps:

[0107] Use the information matrix to perform color enhancement processing on the subsequence of the image sequence that includes the image to be processed.

[0108] That is to say, the information matrix of the image to be processed can be reused in the color enhancement processing of other images in the subsequence, so that when performing color enhancement processing on the image sequence, it is not necessary to determine the information matrix for each frame of the image, thereby reducing the computational amount and improving the efficiency. The present disclosure does not limit the position and length of the subsequence. For example, it may include the previous frame or multiple frames of the image to be processed, or may also include the next frame or multiple frames of the image to be processed. Exemplarily, the subsequence may include the image to be processed and the subsequent continuous N - 1 frames of images, that is, the subsequence may be N consecutive frames of images starting from the image to be processed, and N may be a positive integer not less than 2.

[0109] For example, the image sequence is a video to be processed, and color enhancement processing needs to be performed on each frame of the video to be processed. In the video to be processed, every N frames are determined as an image to be processed. For example, the 1st frame, the (1 + N)th frame, and the (1 + 2N)th frame are used as the images to be processed. The information matrix of the image to be processed is determined by using a deep neural network, and color enhancement processing is performed on the image to be processed and the subsequent N - 1 frames of images based on this information matrix, thereby reducing the computational amount related to the deep neural network and facilitating the realization of real-time color enhancement processing of the video to be processed.

[0110] In one implementation manner, the length of the subsequence can be determined according to the severity of the change in the pictures in the image sequence. Generally, the lower the severity of the change in the pictures, the longer the length of the subsequence, so that more image frames can reuse the information matrix of the image to be processed.

[0111] In one implementation manner, the degree of change between two adjacent frames can be determined according to the difference between two adjacent frames or optical flow information, etc. When the degree of change is low (such as lower than a set threshold), the information matrix of the previous frame can be reused for the next frame of the image. In other words, when the degree of change is high, the next frame of the image is used as the new image to be processed, and the information matrix of the previous frame is not reused, but is input into the deep neural network to obtain a new information matrix.

[0112] In one implementation manner, the method for image color enhancement may further include the training process of the deep neural network. The present disclosure does not limit the specific training method. The following provides three specific examples:

[0113] ① As shown in Figure 8 the training process may include the following steps S810 to S830:

[0114] Step S810, input the sample image to be processed into the deep neural network to be trained to output the sample information matrix;

[0115] Step S820, process the sample image to be processed by using the sample information matrix to obtain the color - enhanced sample image corresponding to the sample image to be processed;

[0116] Step S830: Update the parameters of the deep neural network based on the difference between the annotated image corresponding to the sample image to be processed and the color-enhanced sample image.

[0117] Among them, the sample image to be processed can be an image that has not undergone color enhancement processing, and the annotated image corresponding to the sample image to be processed can be an image obtained by artificially color-enhancing the sample image to be processed. As shown in Figure 9 After obtaining the color-enhanced sample image, calculate the value of the first loss function based on the difference between the color-enhanced sample image and the annotated image, and then perform backpropagation update on the parameters of the deep neural network. The present disclosure does not limit the specific form of the first loss function. For example, L1 or L2 loss can be used.

[0118] By Figure 8 this training method, the deep neural network can indirectly achieve an effect similar to artificial color enhancement processing.

[0119] ② As shown in Figure 10 the training process may include the following steps S1010 to S1030:

[0120] Step S1010: Input the sample image to be processed into the deep neural network to be trained, use the first sample information matrix output by the deep neural network to perform color enhancement processing on the sample image to be processed to obtain a color-enhanced sample image, and perform transformation on the color-enhanced sample image through transformation parameters to obtain a first transformed sample image;

[0121] Step S1020: Perform transformation on the sample image to be processed through transformation parameters, input the transformed sample image to be processed into the deep neural network, and use the second sample information matrix output by the deep neural network to perform color enhancement processing on the transformed sample image to be processed to obtain a second transformed sample image;

[0122] Step S1030: Update the parameters of the deep neural network based on the difference between the first transformed sample image and the second transformed sample image.

[0123] As shown in Figure 11As shown, two types of processing are performed on the sample image to be processed: The first type of processing is similar to the processing flow of the image to be processed described above. The sample image to be processed is input into the deep neural network. For the sake of distinction, the information matrix output by the deep neural network is denoted as the first sample information matrix. Then, the first sample information matrix is used to perform color enhancement processing on the sample image to be processed, obtaining a color-enhanced sample image. Furthermore, the color-enhanced sample image is transformed by the pre-generated transformation parameters, obtaining a first transformed sample image. The second type of processing is to first transform the sample image to be processed by the transformation parameters, then input the transformed sample image to be processed into the deep neural network, obtaining a second sample information matrix. Finally, the second sample information matrix is used to perform color enhancement processing on the transformed sample image to be processed, obtaining a second transformed sample image. Based on the difference between the first transformed sample image and the second transformed sample image, the value of the second loss function is calculated, and the parameters of the deep neural network are updated by backpropagation accordingly. The present disclosure does not limit the specific form of the second loss function. For example, the L1 or L2 loss, etc., can be adopted.

[0124] Among them, the transformation of the image can include perspective transformation or affine transformation, etc. Specifically, one or more of the following transformations such as translation, rotation, scaling, and shearing can be performed on the image. In one implementation manner, the numerical range of the transformation parameters can be determined in advance, and then the transformation parameters are randomly generated within this range. For example, the preset first numerical interval, second numerical interval, and third numerical interval are obtained; the translation parameter is randomly generated within the first numerical interval, the rotation parameter is randomly generated within the second numerical interval, and the scaling parameter is randomly generated within the third numerical interval. This exemplary implementation manner can determine the three numerical intervals according to experience and the actual scenario. Exemplarily, the first numerical interval can be [-3, 3], with the unit being pixel, representing the number of pixels for translation; the second numerical interval can be [-5, 5], with the unit being degree, representing the degree of rotation; the third numerical interval can be [0.97, 1.03], with the unit being multiple, representing the magnification of scaling. Furthermore, random numbers are generated within the three numerical intervals respectively, obtaining the translation parameter, rotation parameter, and scaling parameter, that is, obtaining the transformation parameters in steps S1010 and S1020. This can avoid the difficulty of convergence in the training process caused by too large transformation parameters.

[0125] Generally, when performing color enhancement processing on an image sequence (such as the video to be processed), if there are changes in the image content between different frame images, especially between adjacent frame images, it may lead to inconsistent color enhancement effects for different frames, presenting a flickering phenomenon of the picture and affecting the visual experience. The difference between the first transformed sample image and the second transformed sample image reflects the anti-flickering effect of the deep neural network. By Figure 10The training method can endow the deep neural network with a certain degree of invariance to image transformation, that is, with the ability to resist flicker, so as to ensure the consistency of the effect of color enhancement processing on the image sequence.

[0126] ③ The training process may include the following steps:

[0127] Input the sample image to be processed into the deep neural network to be trained to output a sample information matrix;

[0128] Process the sample image to be processed by using the sample information matrix to obtain a color-enhanced sample image corresponding to the sample image to be processed;

[0129] Input the color-enhanced real image and the color-enhanced sample image into the discriminative network respectively, and update the parameters of the discriminative network and the parameters of the deep neural network according to the output results of the discriminative network; the discriminative network is used to discriminate whether the image is a target-style and real image.

[0130] Among them, the color-enhanced real image and the sample image to be processed do not need to correspond one by one, and it is only necessary that the categories or styles of the scenes are the same. Regarding the deep neural network as a generative network, using the discriminative function of the discriminative network to assist in training the generative network, so that the generative network can generate target-style and real images. Thus, it is ensured that the real sense and style consistency of the image after the deep neural network indirectly performs color enhancement processing.

[0131] In this exemplary embodiment, any one or more of the above training methods may be combined. For example, by combining the above training methods ① and ②, calculate the first loss function value based on the difference between the color-enhanced sample image and the annotated image, calculate the second loss function value based on the difference between the first transformed sample image and the second transformed sample image, calculate the total loss function value according to the first loss function value and the second loss function value, and update the parameters of the deep neural network through the total loss function value.

[0132] Figure 12 The schematic flow of the image color enhancement method is shown, including:

[0133] Step S1201, extract one frame from the video to be processed every N frames as the image to be processed.

[0134] Step S1202, input the image to be processed into the deep neural network to obtain a reference information matrix corresponding to the three-dimensional grid in the three-dimensional space of the image spatial domain - pixel threshold.

[0135] Step S1203, convert the multi-channel image to be processed into a single-channel reference value image.

[0136] Step S1204: Interpolate the reference information matrix based on the reference value image to obtain a color enhancement information matrix corresponding to each pixel point in the image to be processed.

[0137] Step S1205: Use the above color enhancement information matrix to perform color enhancement processing on the image to be processed and the subsequent N - 1 frames of images to obtain N consecutive frames of color-enhanced images.

[0138] Step S1206: Obtain a color-enhanced video by performing color enhancement processing on each frame in the video to be processed.

[0139] An exemplary embodiment of the present disclosure also provides an image color enhancement device. Referring to Figure 13 as shown, the image color enhancement device 1300 may include:

[0140] An image acquisition module 1310, configured to acquire an image to be processed;

[0141] An information matrix generation module 1320, configured to extract three-dimensional grid-based features from the image to be processed through a pre-trained deep neural network, and generate an information matrix according to the extracted features, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the image to be processed;

[0142] A color enhancement processing module 1330, configured to perform color enhancement processing on the image to be processed using the information matrix to obtain a color-enhanced image corresponding to the image to be processed.

[0143] In one embodiment, the deep neural network includes a basic feature extraction sub-network, a local feature extraction sub-network, a global feature extraction sub-network, and an output sub-network; the above-mentioned extraction of three-dimensional grid-based features from the image to be processed through a pre-trained deep neural network and generation of an information matrix according to the extracted features may include:

[0144] Perform downsampling processing on the image to be processed according to the size of the spatial domain grid through the basic feature extraction sub-network to obtain basic features, where the spatial domain grid is the two-dimensional projection of the three-dimensional grid in the spatial domain;

[0145] Extract local features within the spatial domain grid from the basic features through the local feature extraction sub-network;

[0146] Extract global features from the basic features through the global feature extraction sub-network;

[0147] Process the local features and global features according to the number of value domain partitions through the output sub-network to obtain an information matrix, where the value domain partition is the one-dimensional projection of the three-dimensional grid in the pixel value domain.

[0148] In one implementation, the output sub-network includes a fusion layer and a dimensionality conversion layer; the above-mentioned local features and global features are processed by the output sub-network according to the number of value range partitions to obtain an information matrix, including:

[0149] The local features and global features are fused into a comprehensive feature by the fusion layer;

[0150] The dimensionality of the comprehensive feature is converted according to the number of value range partitions by the dimensionality conversion layer to obtain an information matrix.

[0151] In one implementation, the information matrix includes a reference information matrix corresponding to each three-dimensional grid; the above-mentioned information matrix is used to perform color enhancement processing on the image to be processed to obtain a color-enhanced image corresponding to the image to be processed, including:

[0152] Interpolate the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed;

[0153] Process each pixel point of the image to be processed respectively according to the color enhancement information matrix corresponding to each pixel point of the image to be processed to obtain a color-enhanced image.

[0154] In one implementation, the above-mentioned interpolation of the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed includes:

[0155] When the image to be processed is a multi-channel image, convert the image to be processed into a single-channel reference value image;

[0156] Interpolate the reference information matrix based on the reference value image to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed.

[0157] In one implementation, the above-mentioned reference information matrix corresponds to the center point of the three-dimensional grid; the interpolation of the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed includes:

[0158] Interpolate one or more reference information matrices according to the offset of each pixel point of the image to be processed relative to the center point of one or more three-dimensional grids to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed.

[0159] In one implementation, the above-mentioned process of processing each pixel point of the image to be processed respectively according to the color enhancement information matrix corresponding to each pixel point of the image to be processed to obtain a color-enhanced image includes:

[0160] Add a new channel to the image to be processed according to the dimension of the color enhancement information matrix, and set the new channel to a preset value;

[0161] Multiply the pixel value vector of each pixel point of the image to be processed by the color enhancement information matrix corresponding to each pixel point to obtain a color enhanced image; the pixel value vector of each pixel point is a vector formed by the numerical values of each channel of each pixel point.

[0162] In one embodiment, the image to be processed is a frame image in an image sequence. The color enhancement processing module 1330 is further configured to:

[0163] Perform color enhancement processing on the subsequence including the image to be processed in the image sequence by using the information matrix.

[0164] In one embodiment, the image color enhancement device 1300 further includes a deep neural network training module, which is configured to:

[0165] Input the image to be processed as a sample into the deep neural network to be trained to output a sample information matrix;

[0166] Process the image to be processed as a sample by using the sample information matrix to obtain a color enhanced sample image corresponding to the image to be processed as a sample;

[0167] Update the parameters of the deep neural network based on the difference between the labeled image corresponding to the image to be processed as a sample and the color enhanced sample image.

[0168] In one embodiment, the image color enhancement device 1300 further includes a deep neural network training module, which is configured to:

[0169] Input the image to be processed as a sample into the deep neural network to be trained, perform color enhancement processing on the image to be processed as a sample by using the first sample information matrix output by the deep neural network to obtain a color enhanced sample image, and perform transformation on the color enhanced sample image by using transformation parameters to obtain a first transformed sample image;

[0170] Perform transformation on the image to be processed as a sample by using transformation parameters, input the transformed image to be processed as a sample into the deep neural network, and perform color enhancement processing on the transformed image to be processed as a sample by using the second sample information matrix output by the deep neural network to obtain a second transformed sample image;

[0171] Update the parameters of the deep neural network based on the difference between the first transformed sample image and the second transformed sample image.

[0172] The specific details of each part in the above device have been described in detail in the embodiments of the method part. The details not disclosed can be referred to the content of the embodiments in the method part, and thus will not be elaborated here.

[0173] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium, which can be implemented in the form of a program product. The program product includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. In an alternative embodiment, the program product can be implemented as a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0174] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0175] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0176] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0177] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., connected through the Internet using an Internet service provider).

[0178] Exemplary embodiments of the present disclosure also provide an electronic device. The electronic device may be the above-mentioned terminal 110 or server 120. Generally, the electronic device may include a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the above-mentioned image color enhancement method by executing the executable instructions.

[0179] The following takes Figure 14 the mobile terminal 1400 as an example to exemplarily illustrate the structure of the electronic device. Those skilled in the art should understand that, except for components specifically for mobile purposes, Figure 14 the structure in

[0180] As Figure 14 shown, the mobile terminal 1400 may specifically include: a processor 1401, a memory 1402, a bus 1403, a mobile communication module 1404, an antenna 1, a wireless communication module 1405, an antenna 2, a display screen 1406, a camera module 1407, an audio module 1408, a power module 1409, and a sensor module 1410.

[0181] The processor 1401 may include one or more processing units. For example, the processor 1401 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc. The image color enhancement method in this exemplary embodiment may be executed by the AP, the GPU, or the DSP. In addition, the NPU may execute the processing related to the deep neural network. For example, the NPU may load the parameters of the deep neural network and execute the algorithm instructions related to the deep neural network.

[0182] The encoder may encode (i.e., compress) an image or a video to reduce the data size for easy storage or transmission. The decoder may decode (i.e., decompress) the encoded data of the image or the video to restore the image or video data. The mobile terminal 1400 may support one or more encoders and decoders. For example: image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), etc., and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG14, H.1463, H.1464, HEVC (High Efficiency Video Coding), etc.

[0183] The processor 1401 may be connected to the memory 1402 or other components through the bus 1403.

[0184] The memory 1402 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 1401 executes various functional applications and data processing of the mobile terminal 1400 by running the instructions stored in the memory 1402. The memory 1402 may also store application data, such as storing files such as images and videos.

[0185] The communication function of the mobile terminal 1400 can be implemented by a mobile communication module 1404, an antenna 1, a wireless communication module 1405, an antenna 2, a modulation and demodulation processor, a baseband processor, etc. The antenna 1 and the antenna 2 are used for transmitting and receiving electromagnetic wave signals. The mobile communication module 1404 can provide mobile communication solutions such as 3G, 4G, 5G, etc. applied to the mobile terminal 1400. The wireless communication module 1405 can provide wireless communication solutions such as wireless local area network, Bluetooth, near field communication, etc. applied to the mobile terminal 1400.

[0186] The display screen 1406 is used to implement the display function, such as displaying a user interface, an image, a video, etc. The camera module 1407 is used to implement the shooting function, such as shooting an image, a video, etc. The audio module 1408 is used to implement the audio function, such as playing audio, collecting voice, etc. The power module 1409 is used to implement the power management function, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 1410 may include one or more sensors for implementing corresponding sensing and detection functions.

[0187] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0188] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here. After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily think of other embodiments of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0189] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only defined by the appended claims.

Claims

1. An image color enhancement method, characterized in that, Including: Obtain an image to be processed; Extract 3D grid-based features from the image to be processed through a pre-trained deep neural network and output an information matrix, where the 3D grid is obtained by dividing the 3D space formed by the spatial domain and pixel value domain of the image to be processed; Use the information matrix to perform color enhancement processing on the image to be processed to obtain a color-enhanced image corresponding to the image to be processed; Wherein, the information matrix includes a reference information matrix corresponding to each 3D grid, and the reference information matrix includes reference information for color enhancement processing of all pixel points within the 3D grid; the using the information matrix to perform color enhancement processing on the image to be processed to obtain a color-enhanced image corresponding to the image to be processed includes: interpolating the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed; respectively processing each pixel point of the image to be processed according to the color enhancement information matrix corresponding to each pixel point of the image to be processed to obtain the color-enhanced image.

2. The method according to claim 1, wherein The deep neural network includes a basic feature extraction sub-network, a local feature extraction sub-network, a global feature extraction sub-network, and an output sub-network; the extracting 3D grid-based features from the image to be processed through a pre-trained deep neural network and outputting an information matrix includes: Performing downsampling processing on the image to be processed according to the size of the spatial grid through the basic feature extraction sub-network to obtain basic features, where the spatial grid is the 2D projection of the 3D grid in the spatial domain; Extracting local features within the spatial grid from the basic features through the local feature extraction sub-network; Extracting global features from the basic features through the global feature extraction sub-network; Processing the local features and the global features according to the number of value domain partitions through the output sub-network to obtain the information matrix, where the value domain partition is the 1D projection of the 3D grid in the pixel value domain.

3. The method according to claim 2, characterized in that The output sub-network includes a fusion layer and a dimension conversion layer; the processing the local features and the global features according to the number of value domain partitions through the output sub-network to obtain the information matrix includes: Fusing the local features and the global features into a comprehensive feature through the fusion layer; Performing dimension conversion on the comprehensive feature according to the number of value domain partitions through the dimension conversion layer to obtain the information matrix.

4. The method according to claim 1, characterized in that, Use the following formula to convert the image to be processed into a single-channel reference value image: Wherein, R, G, and B are the normalized pixel values of each pixel in the image to be processed; n represents dividing the value ranges of R, G, and B into n partitions, and j represents the ordinal number of the partition; a rj , a gj , a bj are the conversion coefficients for each partition of R, G, and B respectively, which can be determined according to experience or actual requirements; shift rj , shift gj , shift bj are the conversion thresholds set in each partition of R, G, and B respectively; guidemap r , guidemap g , guidemap b are the single-channel images of R, G, and B after partition conversion respectively; g r , g g , g b are the fusion coefficients of R, G, and B respectively; guidemap bias is the offset added after fusion; guidemap z is the reference value image, and its value range is [0, 1].

5. The method according to claim 1, wherein The interpolating the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed includes: When the image to be processed is a multi-channel image, convert the image to be processed into a single-channel reference value image; Interpolate the reference information matrix based on the reference value image to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed.

6. The method according to claim 1, wherein The reference information matrix corresponds to the center point of the three-dimensional grid; interpolating the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed includes: Interpolating one or more of the reference information matrices according to the offset of each pixel point of the image to be processed relative to the center point of one or more of the three-dimensional grids to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed.

7. The method according to claim 1, characterized in that, Processing each pixel point of the image to be processed respectively according to the color enhancement information matrix corresponding to each pixel point of the image to be processed to obtain the color-enhanced image includes: Adding a new channel to the image to be processed according to the dimension of the color enhancement information matrix and setting the new channel to a preset value; Multiplying the pixel value vector of each pixel point of the image to be processed by the color enhancement information matrix corresponding to each pixel point respectively to obtain the color-enhanced image; the pixel value vector of each pixel point is a vector formed by the values of each channel of each pixel point.

8. The method according to claim 1, characterized in that The image to be processed is a frame image in an image sequence; the method further includes: Performing color enhancement processing on a subsequence of the image sequence including the image to be processed by using the information matrix.

9. The method according to claim 1, wherein The method further includes: Inputting a sample image to be processed into the deep neural network to be trained to output a sample information matrix; Processing the sample image to be processed by using the sample information matrix to obtain a color-enhanced sample image corresponding to the sample image to be processed; Updating the parameters of the deep neural network based on the difference between the labeled image corresponding to the sample image to be processed and the color-enhanced sample image.

10. The method according to claim 1, characterized in that, The method further includes: Inputting a sample image to be processed into the deep neural network to be trained, performing color enhancement processing on the sample image to be processed by using the first sample information matrix output by the deep neural network to obtain a color-enhanced sample image, and performing transformation on the color-enhanced sample image by using transformation parameters to obtain a first transformed sample image; wherein, performing transformation on the color-enhanced sample image includes affine transformation or perspective transformation; Performing transformation on the sample image to be processed by using the transformation parameters, inputting the transformed sample image to be processed into the deep neural network, and performing color enhancement processing on the transformed sample image to be processed by using the second sample information matrix output by the deep neural network to obtain a second transformed sample image; Updating the parameters of the deep neural network based on the difference between the first transformed sample image and the second transformed sample image.

11. An image color enhancement device, characterized in that, Including: An image acquisition module configured to acquire an image to be processed; An information matrix generation module configured to extract three-dimensional grid-based features from the image to be processed by a pre-trained deep neural network and output an information matrix, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and pixel value domain of the image to be processed; A color enhancement processing module, configured to perform color enhancement processing on the image to be processed by using the information matrix, so as to obtain a color-enhanced image corresponding to the image to be processed; Wherein, the information matrix includes a reference information matrix corresponding to each of the three-dimensional grids, and the reference information matrix includes reference information for color enhancement processing of all pixel points within the three-dimensional grid; the performing color enhancement processing on the image to be processed by using the information matrix to obtain a color-enhanced image corresponding to the image to be processed includes: interpolating the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed; processing each pixel point of the image to be processed respectively according to the color enhancement information matrix corresponding to each pixel point of the image to be processed to obtain the color-enhanced image; Wherein, the interpolating the reference information matrix based on the image to be processed to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed includes: when the image to be processed is a multi-channel image, converting the image to be processed into a single-channel reference value image; interpolating the reference information matrix based on the reference value image to obtain a color enhancement information matrix corresponding to each pixel point of the image to be processed.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.

13. An electronic device, characterized in that, Comprising: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 10 by executing the executable instructions.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN112950497A

  • Image enhancement method, model training method and equipment

    CN113066017A