An image processing method, apparatus, terminal device, and storage medium

The image is processed through the series optimization model, local enhancement model and compensation model, and the problems of information loss and artifact color deviation in the prior art are solved, and the quality of image optimization is improved.

CN113781320BActive Publication Date: 2025-07-22SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110882192.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-02
Publication Date
2025-07-22
Estimated Expiration
2041-08-02

AI Technical Summary

Technical Problem

Existing image processing methods are prone to loss of information during the optimization process, resulting in artifacts and color deviations in the optimized image and poor quality.

Method used

The image is optimized by connecting multiple deep learning models in series, including optimization model, local enhancement model and compensation model, and the image is optimized, local enhancement and highlighted area information compensation is respectively, and the lost detail information and highlighted area content are reconstructed.

Benefits of technology

Improve the quality of the optimized image, avoid artifacts and color deviations, and enhance the details of the image and the characteristics of the highlighted area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113781320B_ABST
    Figure CN113781320B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, apparatus, terminal device and storage medium, which are applied to the field of image processing. The image processing method provided by the present application includes: using a trained optimization model to perform target type optimization processing on a to-be-processed image to obtain an initial optimized image; performing local enhancement processing on the initial optimized image through a trained local enhancement model to obtain an enhanced image; inputting the enhanced image and the overexposure mask image of the to-be-processed image into a trained compensation model for processing to perform information compensation on the highlighted area of the enhanced image to obtain a compensated image, and the overexposure mask image indicates the highlighted area. The image processing method, apparatus, terminal device and storage medium provided by the present application can improve the quality of the optimized image in the image optimization processing task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, terminal device, and storage medium. Background Art

[0002] Image optimization processing tasks generally include optimization tasks for images such as image retouching and color adjustment, image beautification, image denoising, image super-resolution, and image enhancement. Some optimization tasks are performed on each video frame in a video to be processed, which can also be regarded as image optimization tasks, such as SDR video to HDR video conversion, video denoising, video super-resolution, etc. The optimized image after image processing can better reflect the visual information in the real scene compared with the original image. In the prior art, when performing the above tasks using traditional image processing methods, generally only a deep learning model related to the processing task is used to process the original image. For example, only an image denoising model is used to denoise the original image, and an HDR conversion model is used to convert an SDR video frame into an HDR video frame. When implementing image optimization processing tasks through such methods, a lot of information may be lost, resulting in more artifacts and color deviations in the optimized image, and the quality of the optimized image is poor. Summary of the Invention

[0003] Embodiments of this application provide an image processing method, apparatus, terminal device, and storage medium, which can improve the quality of the optimized image in the image optimization processing task.

[0004] In a first aspect, an embodiment of this application provides a method, which includes: performing target type optimization processing on an image to be processed by using a trained optimization model to obtain an initial optimized image; performing local enhancement processing on the initial optimized image by using a trained local enhancement model to obtain an enhanced image; and inputting the enhanced image and an overexposure mask image of the image to be processed into a trained compensation model for processing to perform information compensation on the highlighted area of the enhanced image to obtain a compensated image, where the overexposure mask image indicates the highlighted area.

[0005] The image processing method provided by this application can first perform target type optimization processing on an image to be processed by using an optimization model to obtain an initial optimized image. For the lost local information in the initial optimized image, a local enhancement model can be used for enhancement processing to reconstruct the lost texture detail information to obtain an enhanced image. Then, based on the highlighted area indicated in the overexposed image corresponding to the image to be processed, the enhanced image is processed by a compensation model to compensate for the lost content information in the overexposed area. This application uses multiple cascaded deep learning models to perform information compensation on the initial optimized image obtained in the image optimization processing task, which can avoid artifacts and color deviations in the optimized image and improve the quality of the optimized image.

[0006] Optionally, the local enhancement model includes: a downsampling module, an upsampling module, and a plurality of residual networks disposed between the downsampling module and the upsampling module.

[0007] Optionally, the method for determining the pixel value of a pixel point in the overexposure mask image includes: according to the formula to determine the pixel value of the pixel point in the overexposure mask image, where I mask (x, y) represents the pixel value of the pixel point at (x, y) in the overexposure mask image, and I s (x, y) represents the pixel value of the pixel point at (x, y) in the image to be processed, and λ represents a preset overexposure threshold.

[0008] Optionally, the compensation model includes a generator; inputting the enhanced image and the overexposure mask image of the image to be processed into the trained compensation model to perform information compensation on the highlighted area of the enhanced image, including: inputting the enhanced image into the trained generator for processing to obtain global exposure information; determining the overexposure information of the highlighted area according to the overexposure mask image of the image to be processed and the global exposure information; using the overexposure information to compensate the highlighted area to obtain a compensated image.

[0009] Optionally, the optimized initial model, the local enhancement initial model, and the compensation initial model are respectively trained to obtain the corresponding optimized model, local enhancement model, and compensation model.

[0010] Optionally, the training method of the generator includes: constructing a generative adversarial network, which includes an initial model of the generator and a discriminator; using a preset loss function and a training set to perform adversarial training on the generative adversarial network to obtain the generator, where the training set includes enhanced image samples, overexposure mask image samples, and compensated image samples corresponding to a plurality of images to be processed; the loss function is used to describe the comprehensive loss value of the absolute error loss value between the compensated image sample and the predicted image, the perceptual loss value between the compensated image sample and the predicted image, and the discriminator loss value of the predicted image; the predicted image refers to the image obtained by multiplying the enhanced image sample after being processed by the initial model of the generator, multiplying by the overexposure mask image sample, and then superimposing with the enhanced image sample.

[0011] Optionally, the target type optimization process refers to HDR conversion processing, the image to be processed is a video frame extracted from an SDR video, and the compensated image output after each video frame in the SDR video is sequentially processed by the optimized model, the local enhancement model, and the compensation model is framed to obtain an HDR video corresponding to the SDR video.

[0012] In a second aspect, an embodiment of the present application provides an image processing apparatus, including: an optimization unit configured to perform target type optimization processing on an image to be processed by using a trained optimization model to obtain an initial optimized image; an enhancement unit configured to perform local enhancement processing on the initial optimized image by using a trained local enhancement model to obtain an enhanced image; a compensation unit configured to input the enhanced image and an overexposure mask image of the image to be processed into a trained compensation model for processing, and perform information compensation on the highlighted area of the enhanced image to obtain a compensated image, where the overexposure mask image indicates the highlighted area.

[0013] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to any one of the above first aspects is implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the method according to any one of the above first aspects is implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on a terminal device, causes the terminal device to execute the method according to any one of the above first aspects.

[0016] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the related descriptions of the beneficial effects brought by the above first aspect and each possible implementation manner of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present application;

[0019] Figure 2 is a schematic structural diagram of an optimization model provided by an embodiment of the present application;

[0020] Figure 3 is a schematic structural diagram of a local enhancement model provided by an embodiment of the present application;

[0021] Figure 4 is a schematic structural diagram of a compensation model provided by an embodiment of the present application;

[0022] Figure 5 It is a schematic diagram of the representation ranges of HDR and SDR color gamuts provided by an embodiment of the present application;

[0023] Figure 6 It is a schematic structural diagram of a compensation initial model provided by an embodiment of the present application;

[0024] Figure 7 It is a flowchart of converting an SDR video to an HDR video provided by an embodiment of the present application;

[0025] Figure 8 It is a schematic diagram for comparing the image processing results of multiple models provided by an embodiment of the present application;

[0026] Figure 9 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application;

[0027] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0028] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0029] In image optimization, generally, a deep learning model related to the processing task is used to process the original image, which results in a lot of information being lost in the optimized image. Exemplarily, in the process of image denoising of a normal original image, the edge information of the original image is generally smoothed to achieve the purpose of denoising. When the contrast of the original image is low, it is difficult for traditional image denoising methods to retain more detailed information while ensuring the denoising effect, resulting in poor quality of the optimized image. Or when the exposure of the original image is high, the information in some highlighted areas is not easily extracted. If the overexposed original image is processed according to the optimization method for normally exposed images, some content information in the highlighted areas will be lost in the optimized image, resulting in color deviation in the optimized image. Or in the task of converting an HDR video, after converting an SDR video frame into an HDR video frame, generally only a neural network is used to obtain the color mapping relationship between the SDR video frames to achieve the HDR conversion of the SDR video frames, resulting in a lot of detailed information and highlighted area information being lost in the HDR video frames, and the obtained HDR video has poor quality.

[0030] In order to improve the quality of image optimization, currently, a large network is usually constructed for training based on the optimization objective. For example, in the HDR video conversion task, in order to balance the low-frequency conversion (color mapping) and high-frequency conversion (detail enhancement) between SDR video frames and HDR video frames, a large network model that can balance color mapping and detail enhancement is usually constructed, and the large network model is trained as a whole so that the large network model can balance the functions of color mapping and detail enhancement. However, the quality of the optimized image obtained in this way has not been significantly improved, especially in the color transition area, where the optimization effect is significantly poor.

[0031] In view of the problems existing in the image optimization processing task, the present application provides an image processing method, which optimizes the image to be processed through a plurality of cascaded deep learning models. Specifically, first, the optimization model is used to perform optimization processing of the target type on the image to be processed to obtain an initial optimized image, then the local enhancement model is used to perform local enhancement processing on the initial optimized image, and then the compensation model is used to compensate the information in the highlighted area of the enhanced image to reconstruct the lost detail information and the content information in the highlighted area of the initial optimized image. By decoupling the image optimization task, for different lost information in the optimized image, a plurality of deep learning models are selected to perform corresponding tasks, and the plurality of deep learning models are cascaded to optimize the image to be processed, thereby compensating for the lost information in the optimized image and improving the quality of the optimized image.

[0032] The technical solution of the present application will be described in detail below with reference to the accompanying drawings. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0033] In one possible implementation, in combination with Figure 1 An exemplary introduction to an image processing method provided by the present application is given. The image processing method can be applied to an image processing device, and the image processing device can be a mobile terminal such as a smart phone, a tablet computer, a camera, etc., or can also be a device capable of processing image data such as a desktop computer, a robot, a server, etc. As Figure 1 shown, an image processing method provided by the present application includes:

[0034] S100, using a trained optimization model to perform optimization processing of the target type on the image to be processed to obtain an initial optimized image.

[0035] In one embodiment, the image to be processed can be any image that needs to be subjected to image retouching and color adjustment, image beautification, image denoising, or image super-resolution and other image processing; it can also be any video frame extracted from a video to be processed that needs to be retouched and color adjusted, beautified, denoised, or video converted. In addition, devices with a camera function such as smartphones, tablets, cameras, desktop computers, and robots can be used to obtain the image to be processed or the video to be processed.

[0036] In a possible implementation, for different types of image processing tasks, the image to be processed can be optimized for the target type based on a deep learning method. In one embodiment, a fully convolutional neural network can be used to convert the image to be processed into an initial optimized image, for example, a fully convolutional neural network with 3 convolutional layers and a convolutional kernel size of 1×1. Other network structures can also be added on the basis of the fully convolutional neural network to form a new network model. Exemplarily, the present application provides an optimization model to optimize the image to be processed for the target type to obtain an initial optimized image.

[0037] The optimization model provided by the present application is as Figure 2 shown. The optimization model includes a main network and a color condition network. Among them, the color condition network includes at least one color condition module (Color Condition Block, CCB) and a feature conversion module connected in sequence. At least one color condition module is used to extract global color feature information from the low-resolution image of the image to be processed. The feature conversion module is used to convert the global color feature information into N sets of adjustment parameters. The N sets of adjustment parameters are respectively used to adjust N intermediate features extracted by the main network during the process of converting the image to be processed into an optimized image, where N is an integer greater than or equal to 1.

[0038] Exemplarily, the image to be processed can be downsampled by a certain multiple (for example, downsampled by 4 times) to obtain a corresponding low-resolution image. Assume that the image to be processed is downsampled by 4 times to obtain a low-resolution image. The low-resolution image has the same size as the image to be processed, but the number of pixels per unit area of the image to be processed is 4 times the number of pixels per unit area of the low-resolution image.

[0039] As Figure 2 shown, the color condition module includes a convolutional layer, a pooling layer, a first activation function, and an IN (Instance Normalization) layer connected in sequence. This color condition module can perform global feature extraction on the input low-resolution image. Compared with the method based on local feature extraction of the image, it can effectively represent the global feature information of the image to be processed, and thus can avoid introducing artificial artifacts in the optimized image.

[0040] The feature transformation module includes a Dropout layer, a convolutional layer, a pooling layer, and N fully connected layers. Among them, the Dropout layer, the convolutional layer, and the pooling layer are connected in sequence and are used to process the global color feature information extracted by at least one color condition module to obtain a conditional vector. The N fully connected layers are respectively used to perform feature transformation on the conditional vector to obtain N sets of adjustment parameters. It should be noted that each fully connected layer processes the conditional vector to obtain a set of adjustment parameters, and finally the number of fully connected layers can be the same as the number of sets of adjustment parameters.

[0041] Exemplarily, such as Figure 2 The optimization model shown includes 4 color condition modules connected in sequence. In the color condition module and the feature transformation module, the size of the convolution kernel in the convolutional layer is 1×1, and the pooling layer uses average pooling. The first activation function is the non-linear activation function LeakyReLU.

[0042] In the embodiment of the present application, the main network includes N Global Feature Modulation (GFM) layers, and the N sets of adjustment parameters are input into the N GFM layers. The GFM layer can adjust the intermediate features input to the GFM layer according to the adjustment parameters.

[0043] In one example, the main network further includes N convolutional layers and N - 1 second activation functions, and the N GFM layers are respectively connected to the output ends of the N convolutional layers. The main network is used to convert the image to be processed into an optimized image, and during the conversion process, the N convolutional layers can be used to extract N intermediate features. The size of the convolution kernel in each convolutional layer is 1×1. The second activation function can be the non-linear activation function ReLU.

[0044] It should be noted that the number of fully connected layers in the color condition network and the number of sets of corresponding generated adjustment parameters should be designed based on the number of convolutional layers in the main network. For example, if the main network includes N convolutional layers, it means that N intermediate features generated by the N convolutional layers need to be adjusted. Therefore, the color condition network needs to output N sets of adjustment parameters corresponding to the N intermediate features, and the main network needs to have N GFM layers to adjust the N intermediate features according to the N sets of adjustment parameters.

[0045] Exemplarily, such as Figure 2As shown, assuming N = 3, the main network includes 3 convolutional (Conv) layers, 3 GFM layers, and 2 second activation function (ReLU) layers. Specifically, the main network, from input to output, successively includes a convolutional layer, a GFM layer, a ReLU layer, a convolutional layer, a GFM layer, a ReLU layer, a convolutional layer, and a GFM layer. Correspondingly, in the color conditional network, the color conditional module includes 4 sequentially connected CCB layers; the feature transformation module may include a Dropout layer, a convolutional (Conv) layer, an average pooling (Avgpool) layer, and 3 fully connected (FC) layers respectively connected to the conditional vector output by the average pooling layer. Each fully connected layer can convert the conditional vector into a corresponding set of adjustment parameters (λ, β), and the color conditional network outputs 3 sets of adjustment parameters (i.e., adjustment parameter 1, adjustment parameter 2, and adjustment parameter 3). Each GFM layer in the main network adjusts the intermediate feature input to this GFM layer according to the corresponding adjustment parameter, which can be expressed as formula (1):

[0046] GFM(x i ) = γ * x i + β (1)

[0047] In formula (1), x i represents the i-th intermediate feature input to the GFM layer; GFM(x i ) represents the adjustment result of the GFM layer on the input intermediate feature x i according to the adjustment parameters (λ, β).

[0048] It can be understood that there are different color mapping relationships between the images to be processed containing different scenarios and the initial optimized image. The optimization model provided in this application extracts the color feature information of the images to be processed as prior information through the color conditional network, which is used to adjust the intermediate features in the main network, so that the optimization model can adaptively output the initial optimized image corresponding to the image to be processed based on the color prior feature information of different images to be processed, and artificial artifacts in the initial optimized image can be avoided.

[0049] In another possible implementation, the images to be processed can also be subjected to target type optimization processing through a color lookup table or traditional digital image processing methods to obtain the initial optimized image.

[0050] The initial optimized image obtained by the method described in the above embodiments has higher color richness, contrast, or clarity compared to the image to be processed. However, during the process of optimizing the image to be processed, the edge texture information and the information of the highlighted part of the initial optimized image may be lost. Therefore, in order to ensure the quality of the initial optimized image, further processing needs to be performed on the initial optimized image to compensate for the detailed information of the initial optimized image and the missing content information in the highlighted area.

[0051] S300, perform local enhancement processing on the initial optimized image through the trained local enhancement model to obtain an enhanced image.

[0052] In the image optimization processing task, when there are problems such as blurred details, uneven brightness distribution, poor contrast, or severe noise pollution in the original image, image enhancement technology can be used to adjust the original image so that the image can better reflect the visual information in the real environment or facilitate subsequent image analysis and processing.

[0053] In a possible implementation manner, the initial optimized image can be enhanced based on a neural network method, or the initial optimized image can be locally enhanced through conventional digital image processing methods. In this embodiment of the application, taking the neural network method as an example, the local enhancement processing process of the initial optimized image is introduced exemplarily.

[0054] In one embodiment, the local enhancement model provided in this application is as Figure 3 shown. The local enhancement model includes a downsampling module, an upsampling module, and a plurality of residual networks arranged between the downsampling module and the upsampling module. Among them, the downsampling module includes at least one set of alternately arranged first convolutional layer (Conv1) and first activation layer (ReLU1), and the upsampling module includes an upsampling layer and at least one set of alternately arranged second activation layer (Conv2) and second convolutional layer (ReLU2). Exemplarily, as Figure 3 shown, the upsampling layer can be a pixel shuffle layer.

[0055] In one embodiment, as Figure 3 shown, the residual network includes a third convolutional layer (Conv3), an activation layer, a fourth convolutional layer (Conv4), and a skip connection. Specifically, for any one residual network, the first image feature input to the residual network is sequentially processed by the first convolutional layer, the activation layer, and the second convolutional layer to obtain a second image feature, and the first image feature and the second image are fused through the skip connection, and at the same time, the fusion result is used as the input of the next layer. Exemplarily, the activation layer can be a non-linear activation function ReLU.

[0056] In the image color mapping task, when converting the image to be processed into the initial optimized image, generally only the color mapping is considered, while the extraction of detailed feature information is ignored, resulting in the loss of some detailed information in the initial optimized image. Therefore, it is necessary to perform local enhancement processing on the initial optimized image. In this application, the initial optimized image is input into the local enhancement model to enhance the edge texture detail information of the initial optimized image and obtain the enhanced image.

[0057] S300, input the enhanced image and the over-exposed mask image of the image to be processed into the trained compensation model for processing, compensate the information of the highlighted area of the enhanced image, and obtain the compensated image. The over-exposed mask image indicates the highlighted area.

[0058] In a possible implementation, the method of deep learning can be used to compensate the information of the over-exposed area by using the trained neural network model. Exemplarily, the embodiment of this application provides a compensation model. The enhanced image and the over-exposed mask image (Over-exposed mask) of the image to be processed are input into the trained compensation model to compensate the information of the highlighted area of the enhanced image.

[0059] Reference Figure 4 , an exemplary description of the compensation model provided by this application is given. The compensation model includes a generator. Specifically, the enhanced image is input into the trained generator for processing to obtain the global exposure information. The over-exposure information of the highlighted area is determined according to the over-exposed mask image of the image to be processed and the global exposure information, and the highlighted area is compensated by using the over-exposure information to obtain the compensated image.

[0060] In one embodiment, the over-exposed mask image of the image to be processed can be obtained by formula (2), that is:

[0061]

[0062] In formula (2), I mask (x,y) represents the pixel value of the pixel point of the over-exposed mask image at (x,y); I S (x,y) represents the pixel value of the pixel point of the image to be processed at (x,y); λ is a preset over-exposure threshold for controlling the over-exposure degree of the image to be processed, and corresponding values can be set according to actual needs. The highlighted area in the image to be processed can be determined according to the pixel value of the pixel point in the over-exposed mask image.

[0063] It should be noted that the generator can be a neural network structure including any convolutional layer for obtaining the global exposure information from the enhanced image. Exemplarily, the structure of the generator (Generator) provided by this application is as Figure 4As shown in the figure, the generator includes: a plurality of downsampling modules connected in sequence and a plurality of upsampling modules corresponding to the plurality of downsampling modules one by one. Among them, the downsampling module includes a convolutional layer and a downsampling layer (DownSample), and the upsampling module includes an upsampling layer (UpSample) and a convolutional layer. In this application, by inputting the enhanced image into the trained generator, global exposure information can be obtained from the enhanced image.

[0064] In one embodiment, the overexposure information of the highlighted area is determined according to the overexposure mask image of the image to be processed and the global exposure information, and the highlighted area is compensated by using the overexposure information to obtain a compensated image, which specifically includes: multiplying the global exposure information and the overexposure mask image pixel by pixel to obtain the overexposure information of the highlighted area; adding the overexposure information and the enhanced image to obtain the compensated image. This process can be expressed by formula (3):

[0065] I H =I mask ×G(I LE )+I LE (3)

[0066] In formula (3), I H represents the compensated image; I mask represents the overexposure mask image; I LE represents the enhanced image; G(I LE ) represents the overexposure information of the highlighted area obtained after the generator processes the enhanced image.

[0067] The image processing method provided by this application uses an optimization model to perform target type optimization processing on the image to be processed to obtain an initial optimized image. For the missing detail information in the initial optimized image, a local enhancement model can be used for enhancement processing to reconstruct the lost texture detail information to obtain an enhanced image. At the same time, global exposure information is extracted from the enhanced image, and the overexposure information of the highlighted area of the enhanced image is determined through the overexposure mask image of the image to be processed. After fusing the overexposure information and the enhanced image, the missing content information in the highlighted part of the enhanced image can be compensated. By optimizing the image to be processed through a series of neural network models connected in series to compensate for the lost information, the highlighted area of the finally obtained optimized image (i.e., the compensated image) has more feature information than the highlighted area of the initial optimized image, and at the same time, the edge texture information is also richer, which can avoid artifacts and color deviations in the optimized image and improve the quality of the optimized image in the image optimization processing task.

[0068] The optimization model, local enhancement model, and compensation model provided by this application are all versatile. On the one hand, these three types of models can be used separately to perform corresponding tasks. Specifically, the optimization model can be applied to any task that requires color optimization or color conversion of the image or video frame to be processed. The local enhancement model can be applied to any task that requires enhancing the texture detail information of the image or video frame. The compensation model can be applied to any task that requires compensating the content information of the highlight area of the image or video frame. On the other hand, the optimization model can be connected in series with any one of the local enhancement model and the compensation model respectively to further enhance the details or the information of the highlight area of the initial optimized image obtained in the image optimization task (such as color optimization or color conversion); the optimization model, local enhancement model, and compensation model can also be connected in series, so that the local enhancement model and the compensation model sequentially compensate the initial optimized image obtained after the target type optimization by the optimization model for the detail information and the highlight area information. The image processing tasks include image editing, image retouching and color adjustment, image coloring, SDR (Standard Dynamic Range) video to HDR (High Dynamic Range) video conversion, image denoising, image super-resolution processing, etc.

[0069] Taking the conversion of SDR video to HDR video as an example, due to the limitations of shooting equipment, there are few existing HDR video resources, and a large number of existing SDR videos need to be converted into HDR videos to meet the needs of users. Figure 5 It is a schematic diagram of the representation ranges of HDR and SDR color gamuts. Among them, BT.709 and BT.2020 are both television parameter standards issued by the ITU (International Telecommunication Union), and DCI-P3 is a color gamut standard formulated by the American film industry for digital cinemas. From Figure 5 it can be seen that among DCI-P3, BT.709, and BT.2020, the color gamut with the largest range is BT.2020, followed by DCI-P3, and the color gamut represented by BT.709 is the smallest. Currently, SDR videos use the BT.709 color gamut, while HDR videos use the BT.2020 color gamut or DCI-P3 color gamut with a wider range. For the same video, whether the HDR video uses the BT.2020 color gamut or the DCI-P3 color gamut, the HDR video can show higher contrast and richer colors than the SDR video.

[0070] Common video conversion methods convert SDR data into HDR data through image encoding technology so that the HDR data can be played on HDR terminal devices. In addition, through super-resolution conversion methods, low-resolution SDR video content needs to be converted into high-resolution HDR video content that meets the HDR video standard. The existing video conversion methods have a high computational cost, and some detail information will be lost in the converted HDR video. If the exposure of the SDR video content is too high, the information in some highlighted areas is not easily extracted. If the overexposed SDR video content is processed according to the optimization method for normally exposed images, some content information in the highlighted areas will be lost in the HDR video, thus affecting the video quality. Compared with the existing methods, the image processing method provided in this application directly processes the SDR video frames by using an optimization model, a local enhancement model, and a compensation model in sequence, converts the SDR video frames into HDR video frames, and further enhances the detail information and the content information of the brightness area of the HDR video frames, avoiding artifacts and color deviations in the HDR video.

[0071] It can be understood that for different tasks, the initial optimization model, the initial local enhancement model, and the initial compensation model can be trained by designing corresponding training sets and loss functions respectively, so as to obtain the optimization model, the local enhancement model, and the compensation model applicable to different tasks.

[0072] Taking the task of converting SDR video to HDR video as an example, the training process and application of the initial color mapping model, the initial local enhancement model, and the initial compensation model provided in this application will be respectively described exemplarily.

[0073] The network structure of the initial optimization model is the same as that of the optimization model shown in Figure 2 In one embodiment, the training process of the initial optimization model is as follows:

[0074] Step 1, obtain a training set.

[0075] For the task of converting SDR video to HDR video based on the optimization model, the training set may include multiple SDR video frame samples and HDR video frame samples corresponding to the multiple SDR video frame samples one by one.

[0076] Specifically, first obtain the SDR video samples and their corresponding HDR video samples. Exemplarily, the SDR video samples and their corresponding HDR video samples can be obtained from public video websites. It is also possible to perform SDR and HDR processing on videos in the same RAW data format respectively to obtain the SDR video samples and their corresponding HDR video samples. It is also possible to use an SDR camera and an HDR camera respectively to capture the corresponding SDR video samples and HDR video samples in the same scene. After obtaining the SDR video samples and their corresponding HDR video samples, perform frame extraction on the SDR video samples and their corresponding HDR video samples respectively to obtain multiple SDR video frame samples, and HDR video frame samples that correspond one-to-one with the multiple SDR video frame samples in terms of time sequence and space.

[0077] Step 2: Use the training set and a preset loss function to train the optimized initial model to obtain an optimized model.

[0078] After building the optimized initial model, input the SDR video frame samples into the main network of the optimized initial model. Perform downsampling processing on the multiple SDR video frame samples respectively to obtain multiple low-resolution images, and input the low-resolution images into the color conditional network of the optimized initial model to obtain adjustment parameters for adjusting the HDR video frames predicted by the optimized initial model.

[0079] The preset loss function f1 is used to describe the L2 loss between the HDR video frames predicted by the optimized initial model and the HDR video frame samples H. It can be expressed as formula (4):

[0080]

[0081] Based on the training set and the above preset loss function, the optimized initial model can be iteratively trained by the gradient descent method until the model converges, and then the trained optimized model can be obtained.

[0082] The structure of the local enhancement initial model is the same as Figure 3 the structure of the local enhancement model shown. In another embodiment, the training process of the local enhancement initial model is as follows:

[0083] Step 1: Obtain the training set.

[0084] First, obtain the SDR video samples and their corresponding HDR video samples. The specific obtaining method can refer to the description in the training process of the optimized initial model above. Then perform frame extraction on the SDR video samples and their corresponding HDR video samples to obtain multiple SDR video frame samples and HDR video frame samples that correspond one-to-one with the multiple SDR video frame samples in terms of time sequence and space.

[0085] For each SDR video frame sample, the SDR video frame sample can be input into the trained optimization model provided in this application or other trained neural network models for HDR conversion processing, or the SDR video frame sample can be subjected to HDR conversion processing through a color lookup table to convert the SDR video frame sample into an initial optimized image sample, which is actually an image sample of HDR data. Therefore, when training the local enhancement initial model, the training set includes multiple training samples, and each training sample includes an initial optimized image sample corresponding to the SDR video frame sample and an HDR video frame sample.

[0086] Step 2: Use the training set and a preset loss function to train the local enhancement initial model to obtain a local enhancement model.

[0087] For each training sample in the training set, input the initial optimized image sample in the training sample into the local enhancement initial model as Figure 3 shown for training. Specifically, sequentially enhance the detailed information of the initial optimized image sample through a downsampling module, multiple residual networks, and an upsampling module to obtain a predicted enhanced image. Iteratively train the loss function based on the predicted enhanced image and the HDR video frame sample corresponding to the initial optimized image sample until the model converges to obtain a local enhancement model.

[0088] Exemplarily, when training the local enhancement initial model, the L2 loss function can be used, and the gradient descent method can be used to iteratively train the loss function.

[0089] In one embodiment, the generator in the compensation initial model can be trained by constructing a generative adversarial network. Use a preset loss function and a training set to perform adversarial training on the generative adversarial network to obtain a generator. Among them, the training set includes enhanced image samples, overexposure mask image samples, and compensation image samples corresponding to multiple images to be processed. The compensation initial model provided in this application is as Figure 6 shown, and this model includes an initial model of the generator and a discriminator, and the initial model of the generator and the discriminator constitute a generative adversarial network. The process of training the compensation initial model is as follows:

[0090] Step 1: Obtain a training set.

[0091] First, obtain SDR video samples and their corresponding HDR video samples. The specific obtaining method can refer to the description in the above training process of the optimization initial model. Then, perform frame extraction on the SDR video samples and their corresponding HDR video samples to obtain multiple SDR video frame samples and HDR video frame samples that are in one-to-one correspondence with the multiple SDR video frame samples in terms of time sequence and space.

[0092] In this embodiment, for each SDR video frame sample, the SDR video frame sample can be subjected to HDR conversion through a trained optimization model or a color lookup table or other trained neural network models to obtain an initial optimized image sample, and the detail information of the initial optimized image sample can be enhanced by using a trained local enhancement model or other trained neural network models to obtain a corresponding enhanced image sample. In addition, the overexposure mask image sample corresponding to the SDR video frame sample can be obtained by using the above formula (2). Therefore, when training the compensation initial model, the training set includes a plurality of training samples, and each training sample includes an enhanced image sample, an overexposure mask image sample, and an HDR video frame sample corresponding to the SDR video frame sample.

[0093] Step 2: Input the enhanced image sample and the overexposure mask image sample in the training set into the compensation initial model for processing to obtain a predicted image.

[0094] Specifically, for each training sample, the enhanced image sample in the training sample is input into the initial model of the generator for processing to obtain global exposure information. After multiplying the global exposure information and the corresponding overexposure mask image sample pixel by pixel, the overexposure information of the highlight area is obtained. The overexposure information is fused with the enhanced image sample to obtain a predicted image.

[0095] Step 3: Input the predicted image and the corresponding HDR video frame sample into the discriminator for iterative training to obtain a compensation model.

[0096] In one embodiment, for each training sample, the predicted image and the corresponding HDR video frame sample in the training sample are input into the discriminator for processing to obtain the discrimination result of the training sample. The initial compensation model is iteratively trained according to the discrimination result of each training sample and a preset loss function to obtain a trained compensation model.

[0097] In this embodiment, the preset loss function is used to describe the absolute error loss value between the HDR video frame sample (i.e., the compensation image sample) and the predicted image The perceptual loss value between the HDR video frame sample (i.e., the compensation image sample) and the predicted image And the discriminator loss value L of the predicted image GAN =-logD(I H ) The comprehensive loss value. The preset loss function L provided in the embodiment of the present application can be expressed as formula (5):

[0098]

[0099] Among them, L1 represents the absolute error loss; L p Represents the perceptual loss; L GANDenotes the generative adversarial loss; I GT Denotes the HDR video frame sample (i.e., the compensated image sample); I H Denotes the predicted image; D(·) denotes the output of the discriminator; α, β, and γ are all hyperparameters.

[0100] Exemplarily, gradient descent can be used for training. When the preset loss function meets certain requirements, it indicates that the model has converged, that is, the initial compensation model has completed training to obtain the compensation model.

[0101] The trained optimization model, local enhancement model, and compensation model can convert the SDR video into an HDR video. Figure 7 This is the flowchart for converting an SDR video into an HDR video provided by the embodiments of this application. Specifically, frame extraction is performed on the SDR video to be processed to obtain SDR video frames. For each SDR video frame, the SDR video frame is input into the trained optimization model for HDR conversion processing, that is, the SDR video frame is converted into HDR data to obtain the initial optimized image. The texture detail information of the initial optimized image is enhanced by the trained local enhancement model to obtain the enhanced image. According to the overexposure mask image of the SDR video frame, and at the same time, the content information of the highlighted area of the enhanced image is compensated by the trained compensation model to obtain the final compensated image. The compensated images corresponding to each SDR video frame are combined to obtain the HDR video corresponding to the SDR video to be processed.

[0102] Next, taking Figure 2 the optimization model shown, Figure 3 the local enhancement model shown, and Figure 4 the compensation model shown as an example of the cascaded model obtained by cascading, combined with Table 1, the performance of the cascaded model provided by this application is described:

[0103] In Table 1, the Residual Network (ResNet), the Cycle Generative Adversarial Network (CycleGAN), and the Pixel-to-Pixel Generative Network (Pixel 2Pixel) are algorithmic models for image-to-image translation. The High Dynamic Range Network (HDRNet), the Conditional Sequential Retouching Network (CSRNet), and the Adaptive 3D Lookup Table (Ada-3DLUT) Network are algorithmic models for photo retouching. The Deep Super-Resolution Inverse Tone-Mapping (Deep SR-ITM) and the GAN-Based Joint Super-Resolution and Inverse Tone-Mapping (JSI-GAN) are algorithmic models for SDR video to HDR video conversion.

[0104] Table 1

[0105] Model Params PSNR SSIM SR-SIM <![CDATA[ΔE ITP > HDR-VDP3 ResNet 1.37M 37.32 0.9720 0.9950 9.02 8.391 Pixel2Pixel 11.38M 25.80 0.8777 0.9871 44.25 7.136 CycleGAN 11.38M 21.33 0.8496 0.9595 77.74 6.941 HDRNet 482K 35.73 0.9664 0.9957 11.52 8.462 CSRNet 36K 35.04 0.9625 0.9955 14.28 8.400 Ada-3DLUT 594K 36.22 0.9658 0.9967 10.89 8.423 Deep SR-ITM 2.87M 37.10 0.9686 0.9950 9.24 8.233 JSI-GAN 1.06M 37.01 0.9694 0.9928 9.36 8.169 Series Model 37.2M 37.21 0.9699 0.9968 9.11 8.569

[0106] As can be seen from Table 1, in terms of performance metrics such as Peak Signal to Noise Ratio (PSNR), structural similarity index measure (SSIM), spectral residual based similarity index measure (SR-SIM), color fidelity ΔE ITP , High Dynamic Range Visible Difference Predictor (HDR-VDP3), etc., the cascaded model provided by this application has very good experimental results.

[0107] Figure 8 This is an example of an optimized image obtained after processing the same picture using each of the models listed in Table 1. From Figure 8 the two examples listed, it can be seen that by using the method provided by this application and optimizing the image based on multiple cascaded models, the optimization effect in the color transition area is significantly better.

[0108] In summary, the method for optimizing a to-be-processed image by multiple cascaded deep learning models provided in this application can reduce information loss during the process of optimizing the image. Compared with the prior art, it can significantly improve the quality of the optimized image.

[0109] Based on the same inventive concept, an embodiment of this application also provides an image processing apparatus. As Figure 9 shown, the apparatus 400 includes: an optimization unit 401, configured to perform target type optimization processing on the to-be-processed image by using a trained optimization model to obtain an initial optimized image. An enhancement unit 402, configured to perform local enhancement processing on the initial optimized image by using a trained local enhancement model to obtain an enhanced image. A compensation unit 403, configured to input the enhanced image and the overexposure mask image of the to-be-processed image into a trained compensation model for processing, and perform information compensation on the highlighted area of the enhanced image to obtain a compensated image, where the overexposure mask image indicates the highlighted area.

[0110] Optionally, the local enhancement model includes: a downsampling module, an upsampling module, and a plurality of residual networks disposed between the downsampling module and the upsampling module.

[0111] Optionally, the method for determining the pixel value of a pixel point in the overexposure mask image includes: according to the formula determine the pixel value of the pixel point in the overexposure mask image, where I mask (x, y) represents the pixel value of the pixel point at (x, y) in the overexposure mask image, and I s (x, y) represents the pixel value of the pixel point at (x, y) in the to-be-processed image, and λ represents a preset overexposure threshold.

[0112] Optionally, the compensation model includes a generator; inputting the enhanced image and the overexposure mask image of the to-be-processed image into the trained compensation model to perform information compensation on the highlighted area of the enhanced image includes: inputting the enhanced image into the trained generator for processing to obtain global exposure information; determining the overexposure information of the highlighted area according to the overexposure mask image of the to-be-processed image and the global exposure information; and using the overexposure information to compensate the highlighted area to obtain a compensated image.

[0113] Optionally, the optimization initial model, the local enhancement initial model, and the compensation initial model are respectively trained to obtain the corresponding optimization model, local enhancement model, and compensation model.

[0114] Optionally, the training method of the generator includes: constructing a generative adversarial network, which includes an initial model of the generator and a discriminator; performing adversarial training on the generative adversarial network using a preset loss function and a training set to obtain the generator, where the training set includes enhanced image samples, overexposure mask image samples, and compensation image samples corresponding to multiple images to be processed; the loss function is used to describe the comprehensive loss value of the absolute error loss value between the compensation image sample and the predicted image, the perceptual loss value between the compensation image sample and the predicted image, and the discriminator loss value of the predicted image; the predicted image refers to the image obtained by multiplying the enhanced image sample processed by the initial model of the generator with the overexposure mask image sample and then superimposing it with the enhanced image sample.

[0115] Optionally, the target type optimization process refers to HDR conversion processing. The image to be processed is a video frame extracted from an SDR video. Each video frame in the SDR video is sequentially processed by an optimization model, a local enhancement model, and a compensation model, and the output compensation images are combined into frames to obtain an HDR video corresponding to the SDR video.

[0116] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.

[0117] Based on the same inventive concept, an embodiment of the present application also provides a terminal device. As Figure 10 shown, the terminal device 500 of this embodiment includes: a processor 501, a memory 502, and a computer program 504 stored in the memory 502 and executable on the processor 501. The computer program 504 can be run by the processor 501 to generate instructions 503, and the processor 501 can implement the steps in the above-mentioned embodiments of various image color optimization methods according to the instructions 503. Alternatively, when the processor 501 executes the computer program 504, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 9 the functions of the unit 401 and the unit 402 shown.

[0118] Exemplarily, the computer program 504 can be divided into one or more modules / units, and one or more modules / units are stored in the memory 502 and executed by the processor 501 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 504 in the terminal device 500.

[0119] Those skilled in the art can understand that Figure 10 merely examples of the terminal device 500, which do not constitute a limitation on the terminal device 500, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device 600 may further include input / output devices, network access devices, buses, etc.

[0120] The processor 501 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0121] The memory 502 can be an internal storage unit of the terminal device 500, such as the hard disk or memory of the terminal device 500. The memory 502 can also be an external storage device of the terminal device 500, such as a plug-in hard disk equipped on the terminal device 500, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 502 can also include both the internal storage unit and the external storage device of the terminal device 500. The memory 502 is used to store computer programs and other programs and data required by the terminal device 500. The memory 502 can also be used to temporarily store data that has been output or will be output.

[0122] The terminal device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0123] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods described in the above method embodiments are implemented.

[0124] The embodiments of the present application further provide a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the methods described in the above method embodiments when executed.

[0125] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0126] The reference to "one embodiment" or "some embodiments" etc. described in the present application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0127] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0128] In addition, in this application, unless otherwise clearly defined and limited, terms such as "connected" and "linked" shall be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. It can be the communication inside two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0129] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them. Although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An image processing method, characterized in that, Including: Performing target type optimization processing on the image to be processed by using the trained optimization model to obtain an initial optimized image; The optimization model includes a main network and a color conditional network. The main network is used to extract intermediate features from the image to be processed, and the color conditional network is used to extract global color feature information from the image to be processed to adjust the intermediate features extracted by the main network; Performing local enhancement processing on the initial optimized image by using the trained local enhancement model to obtain an enhanced image; The local enhancement model includes: a downsampling module, an upsampling module, and a plurality of residual networks arranged between the downsampling module and the upsampling module; Inputting the enhanced image and the overexposure mask image of the image to be processed into the trained compensation model to process the highlighted area of the enhanced image to obtain a compensated image; wherein, the compensation model includes a generator; the overexposure mask image indicates the highlighted area; specifically including: inputting the enhanced image into the trained generator for processing to obtain global exposure information; determining the overexposure information of the highlighted area according to the overexposure mask image of the image to be processed and the global exposure information; compensating the highlighted area by using the overexposure information to obtain a compensated image.

2. The image processing method according to claim 1, characterized in that, The method for determining the pixel value of the pixel point in the overexposure mask image includes: According to the formula to determine the pixel value of the pixel point in the overexposed mask image, where I mask (x, y) represents the pixel value of the pixel point at (x, y) in the overexposed mask image, and I s (x, y) represents the pixel value of the pixel point at (x, y) in the image to be processed, and λ represents a preset overexposure threshold.

3. The image processing method according to claim 1, wherein Training the optimization initial model, the local enhancement initial model, and the compensation initial model respectively to obtain the corresponding optimization model, the local enhancement model, and the compensation model.

4. The image processing method according to claim 1, wherein The training method of the generator includes: Constructing a generative adversarial network, which includes the initial model of the generator and a discriminator; Performing adversarial training on the generative adversarial network by using a preset loss function and a training set to obtain the generator, wherein the training set includes enhanced image samples, overexposure mask image samples, and compensated image samples corresponding to a plurality of images to be processed samples; The loss function is used to describe the comprehensive loss value of the absolute error loss value between the compensated image sample and the predicted image, the perceptual loss value between the compensated image sample and the predicted image, and the discriminator loss value of the predicted image; the predicted image refers to the image obtained by multiplying the enhanced image sample processed by the initial model of the generator, multiplying by the overexposure mask image sample, and then superimposing with the enhanced image sample.

5. The image processing method according to any one of claims 1 to 4, characterized in that, The target type optimization processing refers to HDR conversion processing. The image to be processed is a video frame extracted from an SDR video. The compensated image output after each video frame in the SDR video is sequentially processed by the optimization model, the local enhancement model, and the compensation model is combined into frames to obtain an HDR video corresponding to the SDR video.

6. An image processing apparatus, characterized in that, Including: An optimization unit, configured to perform target type optimization processing on the image to be processed by using the trained optimization model to obtain an initial optimized image; The optimization model includes a main network and a color conditional network. The main network is used to extract intermediate features from the image to be processed, and the color conditional network is used to extract global color feature information from the image to be processed to adjust the intermediate features extracted by the main network; An enhancement unit is configured to perform local enhancement processing on the initial optimized image through a trained local enhancement model to obtain an enhanced image; The local enhancement model includes: a downsampling module, an upsampling module, and a plurality of residual networks disposed between the downsampling module and the upsampling module; A compensation unit is configured to input the enhanced image and the overexposure mask image of the image to be processed into a trained compensation model to process the highlight region of the enhanced image to obtain a compensated image; wherein, the compensation model includes a generator; the overexposure mask image indicates the highlight region; specifically including: inputting the enhanced image into the trained generator for processing to obtain global exposure information; determining the overexposure information of the highlight region according to the overexposure mask image of the image to be processed and the global exposure information; compensating the highlight region by using the overexposure information to obtain a compensated image.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the image processing method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Image processing method and electronic equipment

    CN105100637A

  • Image optimizing method, image optimizing device and terminal

    CN106791471A