Image Processing Method, Apparatus, Electronic Device, and Storage Medium
By adjusting image brightness using convolutional layer groups in pre-trained neural networks, the problem of local supersaturation in image brightness enhancement is solved, brightness equalization and contrast retention is achieved, and computing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202110956618.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-08-19
AI Technical Summary
The image brightness enhancement method in the prior art easily causes local supersaturation of the picture and may reduce the contrast of useful information, especially in dark light conditions, imaging quality is affected by low signal-to-noise ratio and low brightness.
A pre-trained neural network is adopted, including at least one convolutional layer group, consisting of a descending channel layer, a core convolutional layer and a rising channel layer. The image brightness is adjusted through convolutional operations to make it within the standard brightness range, avoiding local supersaturation problems, and retaining contrast.
Image brightness equalization is achieved, local oversaturation is avoided, prediction accuracy and computing speed are improved, and computing load and memory footprint are reduced.
Smart Images

Figure CN113706410B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] People use images extensively in daily life. However, under low-light conditions, due to the influence of low signal-to-noise ratio and low brightness, the imaging quality of images will be greatly affected. Therefore, it is necessary to enhance the images to improve their visual effects. The image brightness enhancement methods in related technologies are based on histogram equalization. However, such methods are prone to cause local over-saturation of pictures and may reduce the contrast of useful information. Summary of the Invention
[0003] To overcome the problems existing in the related technologies, embodiments of the present disclosure provide an image processing method, apparatus, electronic device, and storage medium to solve the defects in the related technologies.
[0004] According to a first aspect of the embodiments of the present disclosure, there is provided an image processing method applied to a terminal device, including:
[0005] Inputting the image to be processed into a pre-trained neural network, where the neural network outputs at least one level of curve parameter maps, and the neural network includes at least one convolutional layer group, and the convolutional layer group includes a down-channel layer, a core convolutional layer, and an up-channel layer connected in sequence;
[0006] Determining a target image according to the image to be processed and the at least one level of curve parameter maps, where the brightness of the target image is within a standard brightness range.
[0007] In one embodiment, the width and height of the convolutional kernel of the down-channel layer are both 1, and the number of output channels of the down-channel layer is less than the number of input channels of the down-channel layer;
[0008] The width and height of the convolutional kernel of the up-channel layer are both 1, and the number of output channels of the up-channel layer is greater than the number of input channels of the down-channel layer.
[0009] In one embodiment, the width and height of the convolutional kernel of the core convolutional layer are both greater than 1, the number of input channels of the core convolutional layer is equal to the number of output channels of the down-channel layer, and the number of output channels of the core convolutional layer is equal to the number of input and output channels of the down-channel layer.
[0010] In one embodiment, the number of convolutional kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolutional kernel is used to perform convolution on the input of one channel of the core convolutional layer.
[0011] In one embodiment, the neural network further includes an initial convolutional layer disposed before the at least one convolutional layer group. The input of the initial convolutional layer is the image to be processed. The width and height of the convolutional kernel of the initial convolutional layer are both greater than 1, and the number of output channels of the initial convolutional layer is the number of input channels of the downsampling layer in the first convolutional layer group of the at least one convolutional layer group.
[0012] In one embodiment, the convolutional stride of the initial convolutional layer is greater than 1; the convolutional stride of each convolutional layer group in the at least one convolutional layer group is 1.
[0013] In one embodiment, in the at least one convolutional layer group, the number of output channels of the upsampling layer in the last convolutional layer group is greater than the number of input channels of the downsampling layer, and the number of output channels of the upsampling layer in other convolutional layer groups is equal to the number of input channels of the downsampling layer;
[0014] The neural network further includes an upsampling layer disposed after the last convolutional layer group. The upsampling layer is used to generate the at least one-level curve parameter map according to the output of the last convolutional layer group, wherein the height and width of the curve parameter map are equal to the height and width of the image to be processed.
[0015] In one embodiment, the neural network further includes a splicing layer disposed between the at least one convolutional layer group; the splicing layer is used to take the outputs of at least two previous convolutional layer groups as inputs and output the fully connected result to the subsequent convolutional layer group.
[0016] In one embodiment, determining the target image according to the image to be processed and the at least one-level curve parameter map includes:
[0017] Determining the (i + 1)-th level image according to the i-th level image and the i-th level curve parameter map, where i ≥ 1, and the image to be processed is the 0-th level image;
[0018] Determining the N-th level image as the target image, where N is the number of levels of the at least one-level curve parameter map.
[0019] In one embodiment, it further includes:
[0020] Inputting multiple images in the image training set into the neural network, and the neural network outputs the at least one-level curve parameter map corresponding to each image, wherein the brightness values of the multiple images are not equal;
[0021] Determining the corresponding predicted image according to each image and the corresponding at least one-level curve parameter map;
[0022] Determine the smoothing loss value according to the at least one-level curve parameter diagram corresponding to each image, and determine the brightness loss value according to the predicted image corresponding to each image and the preset reference brightness value;
[0023] Adjust the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image.
[0024] According to the second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus applied to a terminal device, including:
[0025] A prediction module, configured to input the to-be-processed image into a pre-trained neural network, and the neural network outputs at least one-level curve parameter diagram, wherein the neural network includes at least one convolutional layer group, and the convolutional layer group includes a down-channel layer, a core convolutional layer, and an up-channel layer connected in sequence;
[0026] A determination module, configured to determine a target image according to the to-be-processed image and the at least one-level curve parameter diagram, wherein the brightness of the target image is within a standard brightness range.
[0027] In one embodiment, the width and height of the convolution kernel of the down-channel layer are both 1, and the number of output channels of the down-channel layer is less than the number of input channels of the down-channel layer;
[0028] The width and height of the convolution kernel of the up-channel layer are both 1, and the number of output channels of the up-channel layer is greater than the number of input channels of the down-channel layer.
[0029] In one embodiment, the width and height of the convolution kernel of the core convolutional layer are both greater than 1, the number of input channels of the core convolutional layer is equal to the number of output channels of the down-channel layer, and the number of output channels of the core convolutional layer is equal to the number of input and output channels of the down-channel layer.
[0030] In one embodiment, the number of convolution kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolution kernel is used to perform convolution on the input of one channel of the core convolutional layer.
[0031] In one embodiment, the neural network further includes an initial convolutional layer disposed before the at least one convolutional layer group, the input of the initial convolutional layer is the to-be-processed image, the width and height of the convolution kernel of the initial convolutional layer are both greater than 1, and the number of output channels of the initial convolutional layer is the number of input channels of the down-channel layer of the first convolutional layer group in the at least one convolutional layer group.
[0032] In one embodiment, the convolution stride of the initial convolutional layer is greater than 1; the convolution stride of each convolutional layer group in the at least one convolutional layer group is 1.
[0033] In one embodiment, in the at least one convolutional layer group, the number of output channels of the upsampling layer in the last convolutional layer group is greater than the number of input channels of the downsampling layer, and the number of output channels of the upsampling layer in other convolutional layer groups is equal to the number of input channels of the downsampling layer;
[0034] The neural network further includes an upsampling layer disposed after the last convolutional layer group, and the upsampling layer is configured to generate the at least one-level curve parameter map according to the output of the last convolutional layer group, wherein the height and width of the curve parameter map are equal to the height and width of the image to be processed.
[0035] In one embodiment, the neural network further includes a splicing layer disposed between the at least one convolutional layer group; the splicing layer takes the outputs of at least two previous convolutional layer groups as inputs and outputs the fully connected result to the subsequent convolutional layer group.
[0036] In one embodiment, the determining module is specifically configured to:
[0037] Determine the i-th level image according to the (i - 1)-th level image and the i-th level curve parameter map, where i ≥ 1, and the image to be processed is the 0-th level image;
[0038] Determine the N-th level image as the target image, where N is the number of levels of the at least one-level curve parameter map.
[0039] In one embodiment, it further includes a training module, configured to:
[0040] Input a plurality of images in an image training set into the neural network, and the neural network outputs the at least one-level curve parameter map corresponding to each image, where the brightness values of the plurality of images are not equal;
[0041] Determine the corresponding predicted image according to each image and the corresponding at least one-level curve parameter map;
[0042] Determine a smoothing loss value according to the at least one-level curve parameter map corresponding to each image, and determine a brightness loss value according to the predicted image corresponding to each image and a preset reference brightness value;
[0043] Adjust the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image.
[0044] According to the third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory and a processor, where the memory is configured to store computer instructions that can be run on the processor, and the processor is configured to perform the image processing method according to the first aspect when executing the computer instructions.
[0045] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0046] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0047] The present disclosure inputs the image to be processed into a pre-trained neural network. The neural network outputs at least one-level curve parameter map, and further determines the target image according to the image to be processed and the at least one-level curve parameter map, so that the brightness of the target image is within the standard brightness range. The brightness of the target image is relatively uniform, avoiding the local over-saturation problem of the image enhanced by the histogram equalization method, and retaining the contrast of the image to be processed. Moreover, since the neural network includes a convolutional layer group composed of a down-channel layer, a core convolutional layer, and an up-channel layer, compared with the traditional convolutional layer, the depth of the network can be increased, the number of convolutional kernels in the core convolutional layer can be reduced, the prediction accuracy and operation speed of the curve parameter map are improved, the operation load is reduced, and the occupation of operation memory is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0049] Figure 1 is a flowchart of an image processing method shown in an exemplary embodiment of the present disclosure;
[0050] Figure 2 is a schematic connection diagram of an image acquisition device and an electronic device shown in an exemplary embodiment of the present disclosure;
[0051] Figure 3 is a schematic diagram of the process of determining a target image according to a curve parameter map shown in an exemplary embodiment of the present disclosure;
[0052] Figure 4 is a schematic structural diagram of a neural network shown in an exemplary embodiment of the present disclosure;
[0053] Figure 5 is a schematic structural diagram of a volume and module shown in an exemplary embodiment of the present disclosure;
[0054] Figure 6 is a schematic structural diagram of an image processing apparatus shown in an exemplary embodiment of the present disclosure;
[0055] Figure 7 is a block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0056] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0057] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. The singular forms "a", "said", and "the" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0058] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0059] Based on this, in a first aspect, at least one embodiment of the present disclosure provides an image processing method. Please refer to the attached Figure 1 , which shows the flow of the method, including step S101 and step S102.
[0060] Among them, the sound processing method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA) handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the method can be executed by a server, and the server can be a local server, a cloud server, etc.
[0061] In step S101, the image to be processed is input into a pre-trained neural network, and the neural network outputs at least one-level curve parameter map. Among them, the neural network includes at least one convolutional layer group, and the convolutional layer group includes a down-channel layer, a core convolutional layer, and an up-channel layer connected in sequence.
[0062] Among them, please refer to the attached Figure 2 , the image to be processed can be an image captured by an image acquisition device 201 such as a webcam, or a frame of a video captured by the above image acquisition device 201. The electronic device 202 that executes this method can process the image to be processed and view the processing result. The brightness within a partial or entire range of the image to be processed is not within the standard brightness range, so this method needs to be used for processing to make the brightness within its entire range within the standard brightness range. When the brightness within the entire range of the image captured by the image acquisition device 201 is within the standard brightness range, the electronic device 202 that executes this method can directly display the captured image.
[0063] It should be noted that the above standard brightness range is generated according to an image standard or generated according to user settings. In other words, when the brightness within the entire range of an image is within the standard brightness range, the image can be clearly and comfortably viewed by the user. For example, the image captured by the above image acquisition device under relatively dim ambient light cannot be seen clearly, and the brightness can be enhanced through this method, so as to perform real-time processing on the image or video captured by the image acquisition device, enhancing the brightness and improving the visual effect.
[0064] As the input of the neural network, the image to be processed includes data of at least one channel. When the image to be processed is a grayscale image, it can include at least one channel of data, that is, the data composed of the brightness of each pixel of the image to be processed; when the image to be processed is a color image, it can include at least three channels of data, that is, the data of the red (R), green (G), and blue (B) channels of the image to be processed, and each channel also includes the color components of each pixel of the image to be processed.
[0065] The downsampling layer, the core convolutional layer, and the upsampling layer of the convolutional layer group are connected in sequence, and data is transmitted among the three layers in sequence, that is, the output of the downsampling layer is used as the input of the core convolutional layer, and the output of the core convolutional layer is used as the input of the upsampling layer. The convolutional layer group can control the number of input and output channels before and after the core convolutional layer, that is, reduce the number of input channels and increase the number of output channels. Compared with the traditional convolutional layer, this convolutional layer group can reduce the number of channels and the number of convolutional kernels of the convolutional kernel of the core convolutional layer.
[0066] In step S102, according to the image to be processed and the at least one-level curve parameter map, a target image is determined, where the brightness of the target image is within the standard brightness range.
[0067] The present disclosure inputs a to-be-processed image into a pre-trained neural network. The neural network outputs at least one-level curve parameter maps, and further determines a target image according to the to-be-processed image and the at least one-level curve parameter maps, so that the brightness of the target image is within a standard brightness range. The brightness of the target image is relatively balanced, avoiding the local over-saturation problem of the image enhanced by the histogram equalization method, and retaining the contrast of the to-be-processed image. Moreover, since the neural network includes a convolutional layer group composed of a downsampling channel layer, a core convolutional layer, and an upsampling channel layer, the depth of the network can be increased relative to the traditional convolutional layer, while the number of convolutional kernels in the core convolutional layer can be reduced, improving the prediction accuracy and operation speed of the curve parameter maps, reducing the operation load, and reducing the occupation of operation memory.
[0068] In one embodiment of the present disclosure, the width and height of the convolutional kernels of the downsampling channel layer are both 1, and the number of output channels of the downsampling channel layer is less than the number of input channels of the downsampling channel layer; the width and height of the convolutional kernels of the upsampling channel layer are both 1, and the number of output channels of the upsampling channel layer is greater than the number of input channels of the downsampling channel layer.
[0069] Among them, the width and height of the convolutional kernels of the downsampling channel layer are both 1, which can reduce the amount of computation of the convolutional operation, so that the downsampling channel layer is only used to reduce the number of channels of its input data. The number of channels of the convolutional layer of the downsampling channel layer is the same as the number of input channels of the downsampling channel layer. Therefore, in the convolutional operation, multiple channels of each convolutional kernel can be corresponding to multiple channels of the input data one by one for operation, and then the sum of the operation results of all channels of the convolutional kernel at one position is used as the value of the corresponding position of the convolutional result. The number of convolutional kernels of the downsampling channel layer is the same as the number of output channels of the downsampling channel layer. Therefore, the convolutional result of each convolutional kernel can be used as one channel of the output data.
[0070] Among them, the width and height of the convolutional kernels of the upsampling channel layer are both 1, which can reduce the amount of computation of the convolutional operation, so that the upsampling channel layer is only used to increase the number of channels of the output data of the core convolutional layer. The number of channels of the convolutional layer of the upsampling channel layer is the same as the number of input channels of the upsampling channel layer (i.e., the number of output channels of the core convolutional layer). Therefore, in the convolutional operation, multiple channels of each convolutional kernel can be corresponding to multiple channels of the input data one by one for operation, and then the sum of the operation results of all channels of the convolutional kernel at one position is used as the value of the corresponding position of the convolutional result. The number of convolutional kernels of the upsampling channel layer is the same as the number of output channels of the upsampling channel layer. Therefore, the convolutional result of each convolutional kernel can be used as one channel of the output data.
[0071] Correspondingly, the width and height of the convolutional kernels of the core convolutional layer are both greater than 1, the number of input channels of the core convolutional layer is equal to the number of output channels of the downsampling channel layer, and the number of output channels of the core convolutional layer is equal to the number of input and output channels of the downsampling channel layer.
[0072] Among them, the convolution kernel of the core convolutional layer can be a 3×3 convolution kernel, which can extract the features of the input data more accurately. The number of convolution kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolution kernel is used to perform convolution on the input of one channel of the core convolutional layer. That is to say, the core convolutional layer has only one convolution kernel, and multiple channels of this convolution kernel correspond one by one to multiple channels of the input data of the core convolutional layer. Each channel of the convolution kernel performs convolution with the corresponding channel of the data to obtain one channel of the output data. Through this setting of the convolution kernel, the computing load can be further reduced and the computing efficiency can be improved. Moreover, the number of channels of the output data of this core convolutional layer can be increased in the subsequent channel-increasing layer.
[0073] In some embodiments of the present disclosure, the neural network further includes an initial convolutional layer disposed before the at least one convolutional layer group. The input of the initial convolutional layer is the image to be processed. The width and height of the convolution kernel of the initial convolutional layer are both greater than 1, and the number of output channels of the initial convolutional layer is the number of input channels of the down-channel layer of the first convolutional layer group in the at least one convolutional layer group.
[0074] Among them, the initial convolutional layer is a traditional convolutional layer, which can perform the initial convolution operation on the image to be processed, so as to increase the number of channels of data with fewer channels. For example, the data of a grayscale image with 1 channel or the data of a color image with 3 channels can be increased to 24 channels or 32 channels, etc.
[0075] The convolution kernel of the initial convolutional layer can be a 3×3 convolution kernel, and the number of channels of the convolution kernel can be the same as the number of channels of the image data to be processed. Therefore, in the convolution operation, multiple channels of each convolution kernel can correspond one by one to multiple channels of the image data for operation, and then the sum of the operation results of all channels of the convolution kernel at one position is used as the value of the corresponding position of the convolution result. For example, when the image to be processed is grayscale image data with 1 channel, the number of channels of the convolution kernel can be 1; when the image to be processed is color image data with 3 channels, the number of channels of the convolution kernel can be 3. The number of convolution kernels can be the same as the number of channels of the output data. For example, if you want to increase the number of channels of the image data to 24, the number of convolution kernels can be 24, and the convolution result of each convolution kernel is used as one channel of the output data; if you want to increase the number of channels of the image data to 32, the number of convolution kernels can be 32, and the convolution result of each convolution kernel is used as one channel of the output data.
[0076] In addition, the output of the initial convolutional layer can be used as the input data of the first convolutional layer group. The down-channel layer of the convolutional layer group can reduce the number of channels of the input data by half. For example, reduce the input data with 24 channels to 12 channels, and reduce the input data with 32 channels to 16 channels.
[0077] Optionally, the convolution stride of the initial convolution layer is greater than 1; the convolution stride of each convolution layer group in the at least one convolution layer group is 1. The size of the convolution stride can affect the width and height of the convolution operation result, that is, the larger the convolution stride, the smaller the width and height of the convolution operation result, and the smaller the convolution stride, the larger the width and height of the convolution operation result. Therefore, the convolution stride of the initial convolution layer being greater than 1 can reduce the width and height of the output data of the initial convolution layer relative to the width and height of the input data, and further reduce the width and height of the data input to the convolution layer group relative to the width and height of the image data to be processed, thereby further reducing the operation load of the convolution layer group and improving the operation efficiency. And the convolution stride of the convolution layer group being 1 can keep the width and height of the data input to the convolution layer group unchanged. The initial convolution layer reduces the width and height of the convolution operation result through a convolution stride greater than 1, that is, it can reduce the resolution of the input image, so that the subsequent operations can be reduced. And this way of reducing the operation load by reducing the image resolution is simpler to adjust and more obvious in effect compared with various ways of reducing the operation load in the related art.
[0078] Based on the above, the initial convolution layer reduces the width and height of the image data to be processed through the convolution stride. In the at least one convolution layer group, the number of output channels of the upsampling layer in the last convolution layer group is greater than the number of input channels of the downsampling layer, and the number of output channels of the upsampling layer in other convolution layer groups is equal to the number of input channels of the downsampling layer; the neural network further includes an upsampling layer arranged after the last convolution layer group, and the upsampling layer is used to generate the at least one-level curve parameter map according to the output of the last convolution layer group, wherein the height and width of the curve parameter map are equal to the height and width of the image data to be processed.
[0079] The number of output channels of the last convolution layer group can be n*r 2, where n is the number of channels of the input data of the convolutional layer group, and r is the convolution stride of the initial convolutional layer. For example, if the number of channels of the input data of the convolutional layer group is 24 and the convolution stride of the initial convolutional layer is 2, the number of output channels of the last convolutional layer group can be 96. This is because the number of input channels of other convolutional layer groups except the last one is equal to their number of output channels, that is, the number of channels of the input data remains unchanged. Therefore, the number of input channels of the last convolutional layer group is the number of output channels of the initial convolutional layer. Since the convolution stride of the initial convolutional layer reduces the size of the output data in both the width and height directions, increasing the number of output channels of the last convolutional layer group can make the overall size of the output data equal to the overall size of the output data when the convolution stride of the initial convolutional layer is 1. Therefore, the upsampling layer can generate the at least one-level curve parameter map according to the output of the last convolutional layer group. For example, if the number of channels of the output data of the last convolutional layer group is 96, the upsampling layer can generate data with 24 channels, and then generate the at least one-level curve parameter map according to the data of these 24 channels. It should be noted that the number of channels of the curve parameter map is the same as the number of channels of the image to be processed. That is, if the image to be processed is three-channel color image data, the curve parameter map also includes three channels. For example, if the output data of the upsampling layer includes data with 24 channels, then the data of every 3 channels form one-level curve parameter map.
[0080] In some embodiments of the present disclosure, the neural network further includes a concatenation layer (cancat) disposed between the at least one convolutional layer group; the concatenation layer is configured to use the outputs of at least two previous convolutional layer groups as inputs and output the fully connected result to the subsequent convolutional layer group. The concatenation layer can concatenate the outputs of multiple convolutional layer groups in the channel dimension.
[0081] In some embodiments of the present disclosure, the target image can be determined according to the image to be processed and the at least one-level curve parameter map in the following manner: First, determine the i-th level image according to the (i - 1)-th level image and the i-th level curve parameter map, where i ≥ 1, and the image to be processed is the 0-th level image; then determine the N-th level image as the target image, where N is the number of levels of the at least one-level curve parameter map.
[0082] Please refer to the appendix Figure 3 , the image 301 to be processed and the first-level curve parameter map 302 determine the first-level image 303, and then the first-level image 303 and the second-level curve parameter map 304 determine the second-level image 305, and so on, until the last-level image 307 is determined as the target image according to the last-level curve parameter map 306. The i-th level image can be determined according to the following formula: LE i (x) = LE i-1 (x) + A i (x)LE i-1(x)(1 - LE i-1 (x)), where LE i (x) is the i-th level image, and LE i-1 (x) is the (i - 1)-th level image, and A i (x) is the i-th level curve parameter map. Optionally, N is 8.
[0083] In some embodiments of the present disclosure, the method further includes the training process of the following neural network: First, input multiple images in the image training set into the neural network, and the neural network outputs at least one level of curve parameter map corresponding to each image, where the brightness values of the multiple images are not equal; Next, determine the corresponding predicted image according to each image and the corresponding at least one level of curve parameter map; Next, determine the smoothing loss value according to the at least one level of curve parameter map corresponding to each image, and determine the brightness loss value according to the predicted image corresponding to each image and the preset reference brightness value; Finally, adjust the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image.
[0084] Among them, the images in the training set can be divided into multiple batches, and the multiple images input each time are one batch, and there may be overlaps between different batches of images. The above training process is only the training process of a certain batch of images in the training set, and each batch of images in the training set can be used to train the neural network in the above manner, that is, input into the neural network batch by batch and train the neural network in the above manner until convergence.
[0085] Among them, before the multiple images in the training set are input into the neural network, they can be adjusted to the same size, such as 256 * 256; and normalized, that is, each pixel value is divided by 255.
[0086] Among them, when adjusting the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image, the gradient descent method can be used for adjustment.
[0087] Please refer to the appendix Figure 4, an embodiment of the present disclosure exemplarily shows the structure of a neural network. Among them, the convolution kernel of the initial convolution layer 401 is a 3*3 convolution kernel, and the number of output channels is 24, which is used to input the image to be processed; then it is successively connected with the first-level convolution layer group 402, the second-level convolution layer group 403, the first-level splicing layer 404, the third-level convolution layer group 405, the second-level splicing layer 406, the fourth-level convolution layer group 407 and the upsampling layer 408; the outputs of the first-level convolution layer group 402 and the second-level convolution layer group 403 are both input to the first-level splicing layer 404, and the outputs of the initial convolution layer 401 and the third-level convolution layer group 405 are both input to the second-level splicing layer 406; the output of the initial convolution layer 401 is 24 channels, the outputs of the first-level convolution layer group 402, the second-level convolution layer group 403, and the third-level convolution layer group 405 are 24 channels, the output of the fourth-level convolution layer group 407 is 96 channels, and the upsampling layer 408 outputs an 8-level curve parameter map.
[0088] In addition, please refer to the attached Figure 5 , which shows the difference between the convolution layer group in the above embodiment and the traditional convolution layer. It can be seen that the convolution kernel of the down-channel layer 501 of the convolution layer group is a 1*1 convolution kernel with 24 channels, the convolution step is 1, and the number of convolution kernels is 12. Therefore, the output channels are 12. The convolution kernel of the core convolution layer 502 is a 3*3 convolution kernel with 12 channels, the convolution step is 1, the output channels are 12, and the number of convolution kernels is 1. Its convolution operation is that one channel of the convolution kernel is used to convolve one channel of the input data to obtain one channel of the output data. The convolution kernel of the up-channel layer 503 is a 1*1 convolution kernel with 12 channels, the convolution step is 1, and the output channels are 24; while the convolution kernel of the traditional convolution layer is 3*3 and the number of channels is 24, the convolution step is 1, the output channels are 24, and the number of convolution kernels is 24. All channels of each convolution kernel are used to perform convolution operations with the input data to obtain one channel of the output data.
[0089] Compared with the traditional convolution layer and the convolution method, this embodiment can reduce the network calculation amount and reduce the time consumption. At the same time, this embodiment increases the network depth. Therefore, this solution reduces the number of 3×3 convolution layers and reduces the number of convolution kernels in each convolution layer, which can accelerate the network operation and reduce the memory occupancy.
[0090] According to the second aspect of the embodiments of the present disclosure, an image processing device is provided, which is applied to a terminal device. Please refer to the attached Figure 6 , which shows the structural schematic diagram of the device, including:
[0091] A prediction module 601 for inputting the image to be processed into a pre-trained neural network, where the neural network outputs at least one level of curve parameter maps. The neural network includes at least one convolutional layer group, and the convolutional layer group includes a downsampling layer, a core convolutional layer, and an upsampling layer connected in sequence.
[0092] A determination module 602 for determining a target image according to the image to be processed and the at least one level of curve parameter maps, where the brightness of the target image is within a standard brightness range.
[0093] In some embodiments of the present disclosure, the width and height of the convolution kernel of the downsampling layer are both 1, and the number of output channels of the downsampling layer is less than the number of input channels of the downsampling layer.
[0094] The width and height of the convolution kernel of the upsampling layer are both 1, and the number of output channels of the upsampling layer is greater than the number of input channels of the downsampling layer.
[0095] In some embodiments of the present disclosure, the width and height of the convolution kernel of the core convolutional layer are both greater than 1. The number of input channels of the core convolutional layer is equal to the number of output channels of the downsampling layer, and the number of output channels of the core convolutional layer is equal to the number of input and output channels of the downsampling layer.
[0096] In some embodiments of the present disclosure, the number of convolution kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolution kernel is used to perform convolution on the input of one channel of the core convolutional layer.
[0097] In some embodiments of the present disclosure, the neural network further includes an initial convolutional layer disposed before the at least one convolutional layer group. The input of the initial convolutional layer is the image to be processed. The width and height of the convolution kernel of the initial convolutional layer are both greater than 1, and the number of output channels of the initial convolutional layer is the number of input channels of the downsampling layer of the first convolutional layer group in the at least one convolutional layer group.
[0098] In some embodiments of the present disclosure, the convolution stride of the initial convolutional layer is greater than 1; the convolution stride of each convolutional layer group in the at least one convolutional layer group is 1.
[0099] In some embodiments of the present disclosure, in the at least one convolutional layer group, the number of output channels of the upsampling layer in the last convolutional layer group is greater than the number of input channels of the downsampling layer, and the number of output channels of the upsampling layer in other convolutional layer groups is equal to the number of input channels of the downsampling layer.
[0100] The neural network further includes an upsampling layer disposed after the last convolutional layer group, and the upsampling layer is configured to generate the at least one level of curve parameter maps according to the output of the last convolutional layer group, wherein the height and width of the curve parameter maps are equal to the height and width of the image to be processed.
[0101] In some embodiments of the present disclosure, the neural network further includes a splicing layer disposed between the at least one convolutional layer group; the splicing layer takes the outputs of at least two previous convolutional layer groups as inputs and outputs the fully connected result to the subsequent convolutional layer group.
[0102] In some embodiments of the present disclosure, the determining module is specifically configured to:
[0103] Determine the i-th level image according to the (i - 1)-th level image and the i-th level curve parameter map, where i ≥ 1, and the image to be processed is the 0-th level image;
[0104] Determine the N-th level image as the target image, where N is the number of levels of the at least one level of curve parameter maps.
[0105] In some embodiments of the present disclosure, it further includes a training module, which is configured to:
[0106] Input a plurality of images in the image training set into the neural network, and the neural network outputs at least one level of curve parameter maps corresponding to each image, wherein the brightness values of the plurality of images are not equal;
[0107] Determine the corresponding predicted image according to each image and the corresponding at least one level of curve parameter maps;
[0108] Determine a smooth loss value according to the at least one level of curve parameter maps corresponding to each image, and determine a brightness loss value according to the predicted image corresponding to each image and a preset reference brightness value;
[0109] Adjust the network parameters of the neural network according to the smooth loss value and the brightness loss value of each image.
[0110] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments of the method in the first aspect, and will not be elaborated herein.
[0111] According to the third aspect of the embodiments of the present disclosure, please refer to the appendix Figure 7 , which exemplarily shows a block diagram of an electronic device. For example, the device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0112] Refer toFigure 7 , the apparatus 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0113] The processing component 702 generally controls the overall operation of the apparatus 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing element 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-described method. In addition, the processing component 702 may include one or more modules to facilitate interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate interaction between the multimedia component 708 and the processing component 702.
[0114] The memory 704 is configured to store various types of data to support the operation of the device 700. Examples of such data include instructions for any application or method operating on the apparatus 700, contact data, phone book data, messages, pictures, videos, and the like. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0115] The power component 706 provides power to the various components of the apparatus 700. The power component 706 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the apparatus 700.
[0116] The multimedia component 708 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the device 700 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0117] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the device 700 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0118] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0119] The sensor component 714 includes one or more sensors for providing an assessment of the various aspects of the state of the device 700. For example, the sensor component 714 can detect the on / off state of the device 700, the relative positioning of components, such as the display and the keypad of the device 700. The sensor component 714 can also detect a change in the position of the device 700 or a component of the device 700, the presence or absence of user contact with the device 700, the orientation or acceleration / deceleration of the device 700, and the temperature change of the device 700. The sensor component 714 can also include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can also include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0120] The communication component 716 is configured to facilitate communication between the device 700 and other devices in a wired or wireless manner. The device 700 can access a communication standard-based wireless network, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0121] In an exemplary embodiment, the device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the power supply method of the above-mentioned electronic device.
[0122] In a fourth aspect, in an exemplary embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium including instructions, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the device 700 to complete the power supply method of the above-mentioned electronic device. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0123] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0124] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, Applied to a terminal device, including: Inputting an image to be processed into a pre-trained neural network, the neural network outputting at least one level of curve parameter maps, wherein the neural network includes a plurality of convolutional layer groups, and each convolutional layer group includes a down-channel layer, a core convolutional layer, and an up-channel layer connected in sequence; Determining a target image according to the image to be processed and the at least one level of curve parameter maps, wherein the brightness of the target image is within a standard brightness range; The width and height of the convolutional kernel of the down-channel layer are both 1, and the number of output channels of the down-channel layer is less than the number of input channels of the down-channel layer; the width and height of the convolutional kernel of the up-channel layer are both 1, and the number of output channels of the up-channel layer is greater than the number of input channels of the up-channel layer; The width and height of the convolutional kernel of the core convolutional layer are both greater than 1, the number of input channels of the core convolutional layer is equal to the number of output channels of the down-channel layer, and the number of output channels of the core convolutional layer is equal to the number of input channels of the up-channel layer; The number of convolutional kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolutional kernel is used to perform convolution on the input of one channel of the core convolutional layer; The neural network includes an initial convolutional layer, a first-level convolutional layer group, a second-level convolutional layer group, a first-level splicing layer, a third-level convolutional layer group, a second-level splicing layer, a fourth-level convolutional layer group, and an upsampling layer connected in sequence; the outputs of the first-level convolutional layer group and the second-level convolutional layer group are both input to the first-level splicing layer, and the outputs of the initial convolutional layer and the third-level convolutional layer group are both input to the second-level splicing layer; Wherein, the input of the initial convolutional layer is the image to be processed, and the width and height of the convolutional kernel of the initial convolutional layer are both greater than 1; both the first-level splicing layer and the second-level splicing layer are used to take the outputs of at least two previous convolutional layer groups as inputs and output the fully-connected result to the subsequent convolutional layer groups; the number of output channels of the up-channel layer in the fourth-level convolutional layer group is greater than the number of input channels of the down-channel layer, the number of output channels of the up-channel layer in other convolutional layer groups is equal to the number of input channels of the down-channel layer, and the upsampling layer is used to generate the at least one level of curve parameter maps according to the output of the fourth-level convolutional layer group.
2. The image processing method according to claim 1, wherein The convolutional stride of the initial convolutional layer is greater than 1; the convolutional stride of each convolutional layer group in the plurality of convolutional layer groups is 1.
3. The image processing method according to claim 2, wherein The height and width of the curve parameter maps are equal to the height and width of the image to be processed.
4. The image processing method according to claim 1, wherein The neural network further includes a splicing layer disposed between the plurality of convolutional layer groups; the splicing layer is used to take the outputs of at least two previous convolutional layer groups as inputs and output the fully-connected result to the subsequent convolutional layer groups.
5. The image processing method according to claim 1, characterized in that The determining the target image according to the image to be processed and the at least one level of curve parameter maps includes: Determining the i-th level image according to the (i-1)-th level image and the i-th level curve parameter map, where i≥1, and the image to be processed is the 0-th level image; Determining the N-th level image as the target image, where N is the number of levels of the at least one level of curve parameter maps.
6. The image processing method according to claim 1, wherein Further including: Input multiple images in an image training set into the neural network, and the neural network outputs at least one-level curve parameter maps corresponding to each image, where the brightness values of the multiple images are not equal; Determine corresponding predicted images according to each image and the corresponding at least one-level curve parameter map; Determine a smoothing loss value according to the at least one-level curve parameter map corresponding to each image, and determine a brightness loss value according to the predicted image corresponding to each image and a preset reference brightness value; Adjust the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image.
7. An image processing apparatus, characterized in that, Applied to a terminal device, including: A prediction module, configured to input an image to be processed into a pre-trained neural network, and the neural network outputs at least one-level curve parameter maps, where the neural network includes a plurality of convolutional layer groups, and each convolutional layer group includes a down-channel layer, a core convolutional layer, and an up-channel layer connected in sequence; A determination module, configured to determine a target image according to the image to be processed and the at least one-level curve parameter maps, where the brightness of the target image is within a standard brightness range; The width and height of the convolution kernel of the down-channel layer are both 1, and the number of output channels of the down-channel layer is less than the number of input channels of the down-channel layer; The width and height of the convolution kernel of the up-channel layer are both 1, and the number of output channels of the up-channel layer is greater than the number of input channels of the up-channel layer; The width and height of the convolution kernel of the core convolutional layer are both greater than 1, the number of input channels of the core convolutional layer is equal to the number of output channels of the down-channel layer, and the number of output channels of the core convolutional layer is equal to the number of input channels of the up-channel layer; The number of convolution kernels of the core convolutional layer is the same as the number of input channels of the core convolutional layer, and each convolution kernel is used to perform convolution on the input of one channel of the core convolutional layer; The neural network includes an initial convolutional layer, a first-level convolutional layer group, a second-level convolutional layer group, a first-level splicing layer, a third-level convolutional layer group, a second-level splicing layer, a fourth-level convolutional layer group, and an upsampling layer connected in sequence; the outputs of the first-level convolutional layer group and the second-level convolutional layer group are both input to the first-level splicing layer, and the outputs of the initial convolutional layer and the third-level convolutional layer group are both input to the second-level splicing layer; Wherein, the input of the initial convolutional layer is the image to be processed, and the width and height of the convolution kernel of the initial convolutional layer are both greater than 1; the first-level splicing layer and the second-level splicing layer are both configured to use the outputs of at least two previous convolutional layer groups as inputs and output the fully connected results to the subsequent convolutional layer groups; the number of output channels of the up-channel layer in the fourth-level convolutional layer group is greater than the number of input channels of the down-channel layer, and the number of output channels of the up-channel layer in other convolutional layer groups is equal to the number of input channels of the down-channel layer, and the upsampling layer is configured to generate the at least one-level curve parameter maps according to the output of the fourth-level convolutional layer group.
8. The image processing apparatus according to claim 7, wherein The convolution stride of the initial convolutional layer is greater than 1; the convolution stride of each convolutional layer group in the plurality of convolutional layer groups is 1.
9. The image processing apparatus according to claim 8, wherein The height and width of the curve parameter maps are equal to the height and width of the image to be processed.
10. The image processing apparatus according to claim 7, wherein The neural network further includes a splicing layer disposed between the multiple convolutional layer groups; the splicing layer takes the outputs of at least two of the previous convolutional layer groups as inputs and outputs the fully connected result to the subsequent convolutional layer groups.
11. The image processing apparatus according to claim 7, wherein The determining module is specifically configured to: determine the i-th level image according to the (i - 1)-th level image and the i-th level curve parameter map, where i ≥ 1, and the image to be processed is the 0-th level image; determine the N-th level image as the target image, where N is the number of levels of the at least one level of curve parameter map.
12. The image processing apparatus according to claim 7, wherein It further includes a training module for: inputting multiple images in an image training set into the neural network, and the neural network outputs at least one level of curve parameter map corresponding to each image, where the brightness values of the multiple images are different; determine a corresponding predicted image according to each image and the corresponding at least one level of curve parameter map; determine a smoothing loss value according to the at least one level of curve parameter map corresponding to each image, and determine a brightness loss value according to the predicted image corresponding to each image and a preset reference brightness value; adjust the network parameters of the neural network according to the smoothing loss value and the brightness loss value of each image.
13. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory is used to store computer instructions that can run on the processor, and the processor is used to perform the image processing method according to any one of claims 1 to 6 when executing the computer instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
A medical image processing apparatus and method
CN109583576A
Unsupervised automatic correction method for image exposure
CN111640068A