Image processing method and device
By using an image processing method based on convolutional neural networks, the reflection and illumination components are extracted and synthesized, and then magnified and their resolution enhanced. This solves the problems of color distortion and information loss in low-light images, and improves the image's illumination and resolution.
Patent Information
- Application Number
- CN202110302555.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-03-22
AI Technical Summary
Existing illumination compensation methods are prone to color distortion and loss of important information when enhancing low-light images.
An image processing method based on convolutional neural networks is adopted. By extracting the reflection component and the illumination component, intermediate image information is synthesized, and upsampling layer and convolutional neural network are used for magnification and resolution enhancement, outputting an image with enhanced illumination and resolution.
While improving image illumination, it also increases image resolution and avoids color distortion and information loss.
Smart Images

Figure CN115115527B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method and apparatus. Background Technology
[0002] Limited by weather and lighting conditions during image acquisition, as well as the performance of the image acquisition equipment, the information retention in images often falls short of subsequent processing requirements. To address this issue, image enhancement methods are typically used to enhance images, highlighting important information while weakening less important information.
[0003] In the field of image processing, low-light images are typically enhanced using illumination compensation methods to improve image contrast and detail in dark areas, thereby addressing the problem of insufficient illumination. However, while most illumination compensation methods can improve the problem of insufficient illumination, they are prone to causing color distortion and loss of some important information in the image. Summary of the Invention
[0004] This application provides an image processing method and apparatus to solve the problem that existing illumination compensation methods are prone to causing image color distortion and loss of some important information in the image.
[0005] In a first aspect, this application provides an image processing method for an image processing model, the method comprising:
[0006] The reflection component and illumination component are extracted from the initial image information using the first convolutional neural network in the image processing model.
[0007] The reflection component and the illumination component are synthesized using the second convolutional neural network in the image processing model to output intermediate image information;
[0008] The intermediate image features are extracted from the intermediate image information using the third convolutional neural network in the image processing model, and the intermediate image features are amplified using the upsampling layer in the image processing model.
[0009] The resolution enhancement process is performed by using the intermediate image features magnified by the fourth convolutional neural network in the image processing model to obtain an image with enhanced illumination and resolution.
[0010] Secondly, this application also provides an image processing apparatus, the apparatus comprising:
[0011] The low-light enhancement module is used to extract the reflection component and the illumination component from the initial image information; and to perform synthesis processing on the reflection component and the illumination component to output intermediate image information.
[0012] The resolution enhancement module is used to extract intermediate image features from the intermediate image information; enlarge the intermediate image features; and perform resolution enhancement processing on the enlarged intermediate image features to obtain an illuminance and resolution enhanced image.
[0013] As can be seen from the above technical solutions, the embodiments of this application provide an image processing method and apparatus. This method first uses a first convolutional neural network to extract the reflection component and illumination component from the initial image information. Then, it uses a second convolutional neural network to synthesize the reflection component and illumination component into intermediate image information, thereby completing the illumination enhancement processing of the initial image. Next, it uses a third convolutional neural network to extract intermediate image features from the intermediate image information, then uses an upsampling layer to amplify the intermediate image features. Finally, it uses a fourth convolutional neural network to perform resolution enhancement processing on the amplified intermediate image features, obtaining an image with enhanced illumination and resolution. The image processing method provided by the embodiments of this application can improve image illumination while simultaneously increasing image resolution and avoiding image color distortion and information loss. Attached Figure Description
[0014] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram illustrating the structure of the low-light enhancement module and the resolution enhancement module according to an exemplary embodiment of this application;
[0016] Figure 2 This application presents a flowchart of an image processing method based on an example.
[0017] Figure 3 This application provides an exemplary schematic diagram of a second convolutional neural network.
[0018] Figure 4 This application is based on an exemplary illustration of a low-light initial image and an intermediate image.
[0019] Figure 5 This application is illustrated with an intermediate image and an illumination and resolution enhanced image based on an exemplary diagram.
[0020] Figure 6 This is a flowchart illustrating an image processing method according to an exemplary embodiment of this application;
[0021] Figure 7 This is a schematic diagram of an image processing apparatus according to an exemplary embodiment of this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.
[0023] This application provides an image processing method. The method first constructs a neural network-based image processing model for low-light images, and then uses the trained image processing model to process the low-light images. The purpose is to improve the illumination of the low-light images, increase the resolution of the images, avoid color distortion and information loss, and output an image with enhanced illumination and resolution.
[0024] Figure 1 This is a schematic diagram illustrating the structure of an image processing model according to an exemplary embodiment of this application. Figure 1 As shown, the image processing model includes a first convolutional neural network, a second convolutional neural network, a third convolutional neural network, at least one upsampling layer, and a fourth convolutional neural network connected in sequence. The first convolutional neural network extracts the reflection and illumination components from the input initial image (low-light image). The second convolutional neural network synthesizes the reflection and illumination components output by the first convolutional neural network into intermediate image information. Notably, the illumination of the intermediate image is enhanced compared to the initial image information. The third convolutional neural network extracts intermediate image features from the intermediate image information output by the second convolutional neural network. At least one upsampling layer performs at least one upsampling process on the intermediate image features output by the third convolutional neural network. The fourth convolutional neural network enhances the resolution of the upsampling intermediate image features, finally outputting an illumination and resolution-enhanced image. It is noteworthy that, compared to the intermediate image information, the final output illumination and resolution-enhanced image is enlarged in size and has enhanced resolution.
[0025] Figure 2 This application provides a flowchart of an image processing method based on an example, such as... Figure 2 As shown, the method may include:
[0026] S210 uses the first convolutional neural network in the image processing model to extract the reflection component and illumination component from the initial image information.
[0027] In this embodiment, the initial image information is a vectorized representation of the initial image. The initial image, as the processing object of the image processing method of this application, should be a low-light image whose brightness parameters meet preset conditions. For example, for an initial image of size a×b and number of channels c, its corresponding initial image information can be represented as a vector a×b×c. Similarly, for an RGB image of size 64×64, its corresponding initial image information can be represented as a vector of 64×64×3, where the number of channels in the RGB image is 3.
[0028] In this application, a first convolutional neural network (CNN) is used to segment initial image information into reflection and illumination components. The first CNN may include n sequentially connected convolutional layers. Extracting the reflection and illumination components from the initial image information using the first CNN involves performing n convolutions on the initial image information. The input to the first convolutional layer is the initial image information, the output of the ith convolutional layer is the input to the (i+1)th convolutional layer, and the output of the nth convolutional layer consists of four feature maps, where i is 1, 2, ..., n-1, and 4 is the number of channels in the nth convolutional layer. As a possible implementation, the first to third feature maps can be used as the reflection component extracted from the initial image information, and the fourth feature map can be used as the illumination component extracted from the initial image information. Furthermore, to ensure that the input and output images are the same size, the padding pixel size of each convolutional layer is set to half of (kernel size - 1).
[0029] In one example, the initial image information is a 64×64×3 vector. The first convolutional neural network includes three sequentially connected convolutional layers (i.e., n=3). The first and second convolutional layers each have 64 channels, and the third convolutional layer has 4 channels. In this example, the output of the first convolutional neural network is four feature maps. The first three feature maps represent the reflection component of the initial image information, specifically a 64×64×3 vector. The fourth feature map represents the illumination component of the initial image information, specifically a 64×64×1 vector.
[0030] S220, the reflection component and illumination component are synthesized using the second convolutional neural network in the image processing model to output intermediate image information.
[0031] Intermediate image information is the vectorized representation of the intermediate image. In this application, a second convolutional neural network is used to synthesize the reflection and illumination components output by the first convolutional neural network into intermediate image information. Specifically, the second convolutional neural network includes multiple convolutional layers based on a ResNet structure. When processing the reflection and illumination components using this second convolutional neural network, the reflection and illumination components are used as the input to the first convolutional layer. The input to each of the remaining layers is the sum of the input and output of the previous layer. The results of every three convolutional layers are stacked as the input to the next convolutional layer, and finally, three consecutive convolutional layers are connected. Among the last three convolutional layers, the first convolutional layer reduces the number of channels to one-third of the stacked input, and the other two convolutional layers form a small bottleneck result with the first convolutional layer, with 1 and 3 channels respectively. The output of the last convolutional layer is the intermediate image information.
[0032] Figure 3 This application exemplarily illustrates a second convolutional neural network, such as... Figure 3 As shown, the second convolutional neural network includes six convolutional layers based on the ResNet structure, namely convolutional layers 1-6. When processing the reflection and illumination components using this second convolutional neural network, the reflection and illumination components are used as inputs to convolutional layer 1; the input and output of convolutional layer 1 are summed as inputs to convolutional layer 2; the input and output of convolutional layer 2 are summed as inputs to convolutional layer 3; the outputs of convolutional layer 1, convolutional layer 2, and convolutional layer 3 are stacked as inputs to convolutional layer 4, and the number of channels in convolutional layer 4 is one-third of the number of channels in the aforementioned stacked result; the output of convolutional layer 4 is used as input to convolutional layer 5, and the output of convolutional layer 5 is used as input to convolutional layer 6. Convolutional layer 6 outputs intermediate image information, wherein convolutional layer 5 has 1 channel and convolutional layer 6 has 3 channels.
[0033] It should be understood that this application does not limit the number of convolutional layers included in the second convolutional neural network. For example, the second convolutional neural network may include... Figure 3 The six convolutional layers shown can also include more convolutional layers, such as nine, which will not be elaborated here.
[0034] Figure 4 This application is based on an exemplary illustration of a low-light initial image and an intermediate image, provided by [the relevant authority / organization]. Figure 4 As can be seen, after the initial low-light image A is processed by the first convolutional neural network and the second convolutional neural network, the intermediate image B is obtained. Compared with image A, the illumination of image B is significantly enhanced.
[0035] S230: The third convolutional neural network in the image processing model is used to extract intermediate image features from the intermediate image information, and the intermediate image features are amplified using an upsampling layer.
[0036] S230 utilizes the intermediate image features magnified by the fourth convolutional neural network in the image processing model to perform resolution enhancement processing, resulting in an image with enhanced illumination and resolution.
[0037] In this embodiment, the third convolutional neural network includes several convolutional layers, with the last convolutional layer connected to one or more upsampling layers. The last convolutional layer has 32 channels, while the remaining convolutional layers have 64 channels each. The upsampling layers are used to amplify intermediate image features, and the number of upsampling layers determines the magnification factor of the initial image. For example, when the last convolutional layer is connected to one upsampling layer, the processed image is magnified by a factor of two relative to the initial image; when the last convolutional layer is connected to two consecutive upsampling layers, the processed image is magnified by a factor of four relative to the initial image. It should be understood that those skilled in the art can determine the number of upsampling layers based on the image processing objective, which will not be elaborated upon herein.
[0038] The fourth convolutional neural network is used to enhance the resolution of the magnified intermediate image.
[0039] Figure 5 This application is based on the schematic diagram of the intermediate image and the illumination and resolution enhanced image shown by example. Figure 5 As can be seen, after the intermediate image B is processed by the third convolutional neural network, the upsampling layer and the fourth convolutional neural network, an image C with enhanced illumination and resolution is obtained. Compared with image B, the resolution of image C is enhanced, the size is enlarged and the illumination is further optimized.
[0040] As can be seen from the above embodiments, this application provides an image processing method. This method first uses a first convolutional neural network in the image processing model to extract the reflection component and illumination component from the initial image information. Then, it uses a second convolutional neural network in the image processing model to synthesize the reflection component and illumination component into intermediate image information, thereby completing the illumination enhancement processing of the initial image. Next, it uses a third convolutional neural network in the image processing model to extract intermediate image features from the intermediate image information. Then, it uses an upsampling layer in the image processing model to amplify the intermediate image features. Finally, it uses a fourth convolutional neural network in the image processing model to perform resolution enhancement processing on the amplified intermediate image features, thereby completing the amplification and resolution enhancement processing of the intermediate image. The image processing method provided by this application can improve image illumination while increasing image resolution and avoiding image color distortion and information loss.
[0041] Figure 6 This is a flowchart illustrating an image processing method according to an exemplary embodiment of this application, which specifically shows the process of... Figure 1 The training process of the image processing model shown is the training process of the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network. For example... Figure 6 As shown, the method may include:
[0042] S610, Obtain the training dataset, which includes several sets of corresponding input images, intermediate target images and target images.
[0043] In this context, the intermediate target image has the same size as the input image, and the illumination of the intermediate target image is higher than that of the input image; the target image has the same illumination as the intermediate target image, and the size of the target image is larger than that of the intermediate target image.
[0044] In one possible implementation of S610, several sets of original images are first acquired. Each set of original images includes a corresponding low-light image and a normal-light image. The corresponding low-light image and normal-light image refer to images acquired in the same shooting scene under different lighting conditions, and the two images are of the same size. For example, keeping the position and parameters of the image acquisition device unchanged, a low-light image is acquired under low-light conditions, and a normal-light image is acquired under normal-light conditions. Then, training data corresponding to each set of original images is generated. In each set of training data, the input image is a low-light image scaled down by a factor of k, the intermediate target image is a normal-light image scaled down by a factor of k, and the target image is the corresponding normal-light image, where k is a positive number. It should be understood that the scale reduction of a factor of k here refers to the simultaneous reduction of both the width and height of the image by a factor of k.
[0045] In one example, a set of original images includes a first low-light image and a first normal-light image. The training data corresponding to this set of original images includes a first input image, a first intermediate target image, and a first target image. According to S110, the first input image is a scaled-down version of the first low-light image, the first intermediate target image is a scaled-down version of the first normal-light image, and the first target image is the first normal-light image.
[0046] It should be noted that those skilled in the art can design specific values for k according to their needs. For example, when it is necessary to use the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network to enlarge the initial image to be processed by 4 times, then the input image in each set of training data is the image corresponding to the low-light image reduced by 4 times, and the intermediate target image is the image corresponding to the normal lighting image reduced by 4 times.
[0047] S620, the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer and the fourth convolutional neural network are jointly trained using the training dataset.
[0048] During training, the input image of each training data set is fed into the first convolutional neural network, and the output of the previous neural network in the image processing model is fed into the next adjacent neural network. Specifically, the input image from each training data set is used as the input to the first convolutional neural network, the input reflection component and input illumination component output from the first convolutional neural network are used as the input to the second convolutional neural network, the intermediate training image output from the second convolutional neural network is used as the input to the third convolutional neural network, the output of the third convolutional neural network is used as the input to the upsampling layer, and the output of the upsampling layer is used as the input to the fourth convolutional neural network.
[0049] In this process, while the input image is fed into the first convolutional neural network, the corresponding intermediate target image is fed into the auxiliary convolutional neural network. The auxiliary convolutional neural network extracts the target reflection component and the target illumination component from the intermediate target image. The auxiliary convolutional neural network has the same structure as the first convolutional neural network, and its parameters are optimized synchronously with those of the first convolutional neural network, so that the parameters of the auxiliary convolutional neural network and the parameters of the first convolutional neural network always remain the same.
[0050] After each training round, the training loss is calculated according to a preset loss function. This training loss includes the first local loss generated at the first convolutional neural network, the second local loss generated at the second convolutional neural network, and the third local loss generated at the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network. Therefore, the parameters of the first, second, third, upsampling layers, and fourth convolutional neural networks are optimized based on the training loss until preset validation conditions are met.
[0051] In one possible implementation, the default loss function is the MSE function. In this implementation, the training loss can be calculated using the following formula:
[0052]
[0053] Where N is the number of training data sets in this round of training;
[0054] This represents the local loss generated by the i-th set of training data in the first convolutional neural network, i.e., the first local loss;
[0055] This represents the local loss generated by the i-th set of training data in the second convolutional neural network, i.e., the second local loss;
[0056] This represents the local loss generated by the i-th set of training data in the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network, i.e., the third local loss;
[0057] ω1, ω2, and ω3 are the weights corresponding to each local loss.
[0058] In this application, the output of the first convolutional neural network is the input reflection component and the input illumination component extracted from the input image. The image reconstructed based on the input reflection component and the input illumination component is called the first reconstructed image, and the image reconstructed based on the target reflection component and the target illumination component is called the second reconstructed image, wherein:
[0059] First restored image = Input reflection component × Input illumination component;
[0060] The second reconstructed image = target reflection component × target illumination component.
[0061] In this application, the first partial loss This can include: a loss of the first reconstructed image relative to the corresponding input image; a loss of the second reconstructed image relative to the corresponding intermediate target image; and a loss of the input reflection vector relative to the corresponding target reflection vector. In this way, through continuous training, the input reflection component output by the first convolutional neural network can continuously approach the corresponding target reflection component. Simultaneously, the input reflection component and input illumination component output by the first convolutional neural network can reconstruct the corresponding input image as closely as possible, and the target reflection component and target illumination component can reconstruct the corresponding intermediate target image as closely as possible.
[0062] In addition, the second local loss is calculated based on the output of the second convolutional neural network and the corresponding intermediate target image, and the third local loss is calculated based on the output of the fourth convolutional neural network and the corresponding target image.
[0063] In this embodiment of the application, the weights ω1, ω2, and ω3 corresponding to each local loss can be preset according to the gradients of the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, and the fourth convolutional neural network. Generally, the smaller the gradient of the neural network, the smaller the corresponding weight.
[0064] In one implementation, ω1 is less than ω2 and ω3.
[0065] During training, when the result of a certain iteration is better than that of the previous iteration, the parameters used in that iteration are saved as the current optimal parameters. The training process synchronously outputs the loss corresponding to the training set and the loss corresponding to the validation set. Training stops when the loss on the validation set meets the preset validation conditions.
[0066] As can be seen from S610-S620 above, the image processing method provided in this application constructs a neural network-based image processing model for low-light images. The image processing model includes a first convolutional neural network, a second convolutional neural network, a third convolutional neural network, at least one upsampling layer, and a fourth convolutional neural network connected in sequence. By jointly training the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network, the image processing model can improve the illumination of low-light images and improve the resolution of intermediate images, while avoiding color distortion and information loss.
[0067] Based on the image processing method provided in the above embodiments, this application also provides an image processing apparatus, such as... Figure 7 As shown, the device may include:
[0068] The low-light enhancement module 710 is used to extract the reflection component and the illumination component from the initial image information; and to synthesize the reflection component and the illumination component to output intermediate image information. The resolution enhancement module 720 is used to extract intermediate image features from the intermediate image information; and to enlarge and enhance the resolution of the intermediate image features to obtain an enhanced image with higher illumination and resolution.
[0069] In some embodiments, the low-light enhancement module 710 includes a first enhancement unit and a second enhancement unit. The first enhancement unit is used to extract a reflection component and an illumination component from the initial image information; the second enhancement unit is used to synthesize the reflection component and the illumination component to output intermediate image information. The resolution enhancement module 720 includes a third enhancement unit and a fourth enhancement unit. The third enhancement unit is used to extract intermediate image features from the intermediate image information and magnify the intermediate image features; the fourth enhancement unit is used to perform resolution enhancement processing on the magnified intermediate image features to obtain an enhanced image with both illumination and resolution.
[0070] In some embodiments, the first enhancement unit is specifically a convolutional neural network including multiple convolutional layers, the second enhancement unit is specifically a convolutional neural network including multiple convolutional layers, the third enhancement unit is specifically a convolutional neural network including multiple convolutional layers and at least one upsampling layer, and the fourth enhancement unit is specifically a convolutional neural network including multiple convolutional layers.
[0071] In some embodiments, the first enhancement unit includes n convolutional layers. The first enhancement unit is specifically used to: perform n convolutional processing on the initial image information using the n convolutional layers, wherein the output of the i-th convolutional layer is the input of the (i+1)-th convolutional layer, and the output of the n-th convolutional layer is 4 feature maps, where i is 1, 2, ..., n-1; and use the 1st to 3rd feature maps as the reflection component and the 4th feature map as the illumination component.
[0072] In some embodiments, the apparatus further includes a training module, which includes a data preparation unit for acquiring a training dataset. The training dataset includes several sets of corresponding input images, intermediate target images, and target images, wherein the intermediate target images have the same size as the input images and the illuminance of the intermediate target images is higher than that of the input images, and the target images have the same illuminance as the intermediate target images and the size of the target images is larger than that of the intermediate target images; and the training unit is used to jointly train the first enhancement unit, the second enhancement unit, the third enhancement unit, and the fourth enhancement unit using the training dataset.
[0073] In some embodiments, the data preparation unit is specifically used to: acquire several sets of original images, each set of original images including a corresponding low-light image and a normal-light image; generate training data corresponding to each set of original images, wherein the input image in each set of training data is an image of the corresponding low-light image reduced by a factor of k, the intermediate target image is an image of the corresponding normal-light image reduced by a factor of k, and the target image is the corresponding normal-light image, where k is a positive number.
[0074] In some embodiments, the training unit is specifically configured to: input the input image of each set of training data into the first convolutional neural network; input the output of the previous neural network in the image processing model into the next adjacent neural network; calculate the training loss according to a preset loss function, the training loss including a first local loss calculated based on the output of the first convolutional neural network and the target reflection component and target illumination component extracted from the corresponding intermediate target image, a second local loss calculated based on the output of the second convolutional neural network and the corresponding intermediate target image, and a third local loss calculated based on the output of the fourth convolutional neural network and the corresponding target image; optimize the parameters of the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network according to the training loss until a preset verification condition is met. In some embodiments, the training unit is further configured to: input the corresponding intermediate target image into an auxiliary convolutional neural network to extract the target reflection component and the target illumination component from the first intermediate target image using the auxiliary convolutional neural network, wherein the auxiliary convolutional neural network has the same structure as the first enhancement unit; and, while optimizing the parameters of the first enhancement unit according to the training loss, simultaneously optimize the parameters of the auxiliary convolutional neural network so that the parameters of the auxiliary convolutional neural network are the same as those of the first enhancement unit.
[0075] In some embodiments, the output of the first convolutional neural network is the input reflection component and the input illumination component extracted from the input image; the local loss corresponding to the first enhancement unit includes: the loss of the first restored image relative to the corresponding input image, the first restored image being restored based on the input reflection component and the input illumination component; the loss of the second restored image relative to the corresponding intermediate target image, the second restored image being restored based on the target reflection component and the target illumination component; and the loss of the input reflection vector relative to the target reflection vector.
[0076] In some embodiments, the training loss is calculated according to the following formula:
[0077]
[0078] Where N is the number of training data sets;
[0079] This represents the local loss generated by the i-th set of training data in the first augmentation unit;
[0080] This represents the local loss generated by the i-th set of training data in the second augmentation unit;
[0081] This represents the local loss generated by the i-th training data in the third and fourth augmentation units;
[0082] ω1, ω2, and ω3 are the weights corresponding to each local loss.
[0083] In some embodiments, ω1 is less than ω2 and ω3.
[0084] In a specific implementation, the present invention also provides a computer storage medium, wherein the computer storage medium may store a program, which, when executed, may include some or all of the steps of the various embodiments of the image processing method provided by the present invention. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0085] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0086] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0087] The embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that, For image processing models, the method includes: The reflection component and illumination component are extracted from the initial image information using the first convolutional neural network in the image processing model. The reflection component and the illumination component are synthesized using the second convolutional neural network in the image processing model to output intermediate image information; The intermediate image features are extracted from the intermediate image information using the third convolutional neural network in the image processing model, and the intermediate image features are amplified using the upsampling layer in the image processing model. The fourth convolutional neural network in the image processing model is used to perform resolution enhancement processing on the features of the magnified intermediate image to obtain an image with enhanced illumination and resolution. The image processing model was trained according to the following steps: Obtain a training dataset, which includes several sets of corresponding input images, intermediate target images, and target images. The intermediate target images are the same size as the input images and have higher illumination than the input images. The target images have the same illumination as the intermediate target images and have larger size than the intermediate target images. The first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network in the image processing model are jointly trained using the training dataset. The first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network are connected sequentially; Jointly training the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network in the image processing model using the training dataset includes: The input image of each set of training data is input into the first convolutional neural network, and the output of the previous neural network in the image processing model is input into the next adjacent neural network. The training loss is calculated according to a preset loss function. The training loss includes a first local loss calculated based on the output of the first convolutional neural network and the target reflection component and target illumination component extracted from the corresponding intermediate target image, a second local loss calculated based on the output of the second convolutional neural network and the corresponding intermediate target image, and a third local loss calculated based on the output of the fourth convolutional neural network and the corresponding target image. The parameters of the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network are optimized based on the training loss until the preset verification conditions are met.
2. The method according to claim 1, characterized in that, The first convolutional neural network includes n convolutional layers. It extracts reflection and illumination components from initial image information, including: The initial image information is processed by n convolutional layers, where the output of the i-th convolutional layer is the input of the (i+1)-th convolutional layer, and the output of the n-th convolutional layer is four feature maps, where i is 1, 2, ..., n-1. The first to third feature maps are used as the reflection component, and the fourth feature map is used as the illumination component.
3. The method according to claim 1, characterized in that, The acquisition of the training dataset includes: Acquire several sets of original images, each set of original images including a corresponding low-light image and a normal-light image; Training data corresponding to each set of original images is generated. The input image in each set of training data is the image of the corresponding low-light image reduced by a factor of k. The intermediate target image in each set of training data is the image of the corresponding normal-light image reduced by a factor of k. The target image in each set of training data is the corresponding normal-light image. k is a positive number.
4. The method according to claim 1, characterized in that, While inputting the input images of a set of training data into the first convolutional neural network, the method further includes: The intermediate target image in the training data is input into an auxiliary convolutional neural network to extract the target reflection component and the target illumination component from the intermediate target image using the auxiliary convolutional neural network, wherein the auxiliary convolutional neural network has the same structure as the first convolutional neural network; Furthermore, while optimizing the parameters of the first convolutional neural network based on the training loss, the parameters of the auxiliary convolutional neural network are simultaneously optimized so that the parameters of the auxiliary convolutional neural network are the same as those of the first convolutional neural network.
5. The method according to claim 1, characterized in that, The output of the first convolutional neural network is the input reflection component and input illumination component extracted from the input image; a first local loss is calculated based on the output of the first convolutional neural network and the target reflection component and target illumination component extracted from the corresponding intermediate target image, including: The first restored image is obtained by restoring the first restored image based on the input reflection component and the input illumination component, and the loss of the first restored image relative to the corresponding input image is calculated. The second restored image is obtained by reconstructing the target reflection component and the target illumination component, and the loss of the second restored image relative to the corresponding intermediate target image is calculated. And, calculate the loss of the input reflection component relative to the target reflection component.
6. The method according to claim 1, characterized in that, The training loss is calculated using the following formula: Where N is the number of training data sets; This represents the first local loss generated by the i-th set of training data in the first convolutional neural network; This represents the second local loss generated by the i-th set of training data in the second convolutional neural network; This represents the third local loss generated by the i-th set of training data in the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network; ω1, ω2, and ω3 are the weights corresponding to each local loss.
7. The method according to claim 6, characterized in that, The ω1 is less than the ω2 and ω3.
8. An image processing apparatus, characterized in that, For an image processing model, the apparatus includes: The low-light enhancement module is used to extract the reflection component and the illumination component from the initial image information; and to perform synthesis processing on the reflection component and the illumination component to output intermediate image information. The low-light enhancement module includes a first enhancement unit and a second enhancement unit. The first enhancement unit is used to extract the reflection component and the illumination component from the initial image information. The second enhancement unit is used to perform synthesis processing on the reflection component and the illumination component to output intermediate image information. The resolution enhancement module is used to extract intermediate image features from the intermediate image information; enlarge the intermediate image features; and perform resolution enhancement processing on the enlarged intermediate image features to obtain an illumination and resolution-enhanced image. The resolution enhancement module includes a third enhancement unit and a fourth enhancement unit. The third enhancement unit is used to extract intermediate image features from the intermediate image information and magnify the intermediate image features. The fourth enhancement unit is used to perform resolution enhancement processing on the magnified intermediate image features to obtain an illuminance and resolution enhanced image. The device further includes a training module, which includes a data preparation unit for acquiring a training dataset. The training dataset includes several sets of corresponding input images, intermediate target images, and target images. The intermediate target images have the same size as the input images and their illumination is higher than that of the input images. The target images have the same illumination as the intermediate target images and their size is larger than that of the intermediate target images. The training unit is used to jointly train the first enhancement unit, the second enhancement unit, the third enhancement unit, and the fourth enhancement unit using the training dataset. The first enhancement unit is specifically a convolutional neural network including multiple convolutional layers; the second enhancement unit is specifically a convolutional neural network including multiple convolutional layers; the third enhancement unit is specifically a convolutional neural network including multiple convolutional layers and at least one upsampling layer; and the fourth enhancement unit is specifically a convolutional neural network including multiple convolutional layers. The training unit is specifically used for: inputting the input image of each set of training data into the first convolutional neural network; inputting the output of the previous neural network in the image processing model into the next adjacent neural network; calculating the training loss according to a preset loss function, the training loss including a first local loss calculated based on the output of the first convolutional neural network and the target reflection component and target illumination component extracted from the corresponding intermediate target image, a second local loss calculated based on the output of the second convolutional neural network and the corresponding intermediate target image, and a third local loss calculated based on the output of the fourth convolutional neural network and the corresponding target image; optimizing the parameters of the first convolutional neural network, the second convolutional neural network, the third convolutional neural network, the upsampling layer, and the fourth convolutional neural network according to the training loss until a preset verification condition is met.
Citation Information
Patent Citations
DEC_SE-based low-illumination image super-resolution reconstruction method
CN111784582A