Image Processing Method, Image Processing Device, and Readable Storage Medium
By alternately training the first neural network and the discriminant network, the parameters of the first neural network are optimized, and the real-time processing problem caused by the large amount of parameters of the deep learning algorithm model is solved, and image clarity is improved and processing speed is accelerated.
Patent Information
- Application Number
- CN202080002356.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-11-27
AI Technical Summary
The deep learning algorithm model has a large number of parameters, which makes the real-time processing requirements of videos at the display terminal unable to meet, affecting the video viewing experience.
The trained first neural network and discriminative network are trained using alternating training methods, optimize the first neural network by adjusting parameters, reduce its parameter amount, use the trained second neural network and discriminative network to generate images with higher clarity, and adjust the parameters of the first neural network through the total loss function to improve image clarity.
While reducing the amount of parameters, it improves image processing speed and clarity, and is suitable for lightweight terminal devices.
Smart Images

Figure CN114641792B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and particularly relates to an image processing method, an image processing device, and a readable storage medium. Background Art
[0002] Before video transmission, due to bandwidth limitations, compression encoding is required. However, the compressed video will generate various compression noises, which affect people's viewing experience of the video on the display terminal.
[0003] The rise of deep learning technology has brought a technological breakthrough in the direction of video compression and restoration. Through training and learning on a large amount of video data, the restoration effect can be improved well. However, generally, the larger the number of parameters of the deep learning algorithm model and the deeper the network structure, the better the processing effect, which will lead to an excessive amount of computation and cannot meet the real-time processing requirements of the video on the display terminal. Summary of the Invention
[0004] Multiple aspects of the present disclosure provide an image processing method, an image processing device, and a readable storage medium.
[0005] An embodiment of the present disclosure provides an image processing method, including:
[0006] Processing an input image by using a trained first neural network to obtain a target output image; the clarity of the target output image is greater than that of the input image;
[0007] Wherein, the trained first neural network is obtained by training a first neural network to be trained by a first training method, and the first training method includes:
[0008] Alternately training a second neural network to be trained and a discriminant network to be trained to obtain a trained second neural network and a trained discriminant network; wherein, the number of parameters of the trained second neural network is more than that of the first neural network to be trained; the trained second neural network is configured to transform an image with a first clarity received into an image with a second clarity, and the second clarity is greater than the first clarity; the first neural network to be trained includes: a plurality of first feature extraction sub-networks and a first output sub-network located after the plurality of first feature extraction sub-networks, the trained second neural network includes: a plurality of second feature extraction sub-networks and a second output sub-network located after the plurality of second feature extraction sub-networks, and the first feature extraction sub-network corresponds to the second feature extraction sub-network one by one;
[0009] Provide the first sample image to the trained second neural network and the first neural network to be trained respectively, so that the first neural network to be trained outputs a first output image, and the trained second neural network outputs a second output image;
[0010] Provide the first output image to the trained discriminant network, so that the trained discriminant network generates a first discrimination result based on the first output image;
[0011] Adjust the parameters of the first neural network according to the total loss to obtain the updated first neural network; wherein, the total loss includes a first loss, a second loss and a third loss, the first loss is obtained based on the difference between the first output image and the second output image; the second loss is obtained based on the difference between the first discrimination result and the first target result; the third loss is obtained based on the difference between the output images of at least one of the first feature extraction sub-networks and the output images of the corresponding second feature extraction sub-networks.
[0012] In some embodiments, the number of channels of the output image of the first feature extraction sub-network is less than the number of channels of the output image of the corresponding second feature extraction sub-network;
[0013] The first training method further includes: providing the output images of the multiple second feature extraction sub-networks to multiple dimensionality reduction layers one by one, so that each dimensionality reduction layer generates an intermediate image; the number of channels of the intermediate image is the same as the number of channels of the output image of the first feature extraction sub-network;
[0014] Adjusting the parameters of the first neural network according to the total loss function includes: adjusting the parameters of both the first neural network and the dimensionality reduction layers; wherein, the third loss is obtained based on the sum of the differences between each intermediate image and the output image of the corresponding first feature extraction sub-network.
[0015] In some embodiments, the total loss further includes: a fourth loss, the fourth loss is the perceptual loss based on the first output image and the second output image.
[0016] In some embodiments, the perceptual loss between the first output image and the second output image is calculated according to the following formula:
[0017]
[0018] wherein, y1 is the first output image, and y2 is the second output image, For a preset network layer in the trained discriminative network, j is the layer number of the preset network layer in the discriminative network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer.
[0019] In some embodiments, the first loss includes the L1 loss between the first output image and the second output image.
[0020] In some embodiments, the second loss includes the cross-entropy loss between the first discrimination result and the first target result.
[0021] In some embodiments, the third loss term includes the sum of the L2 losses between the output images of each first feature extraction sub-network and the corresponding intermediate images.
[0022] In some embodiments, providing the first output image to the trained discriminative network to enable the trained discriminative network to generate a first discrimination result based on the first output image includes:
[0023] Setting the first output image with a ground truth label and providing the first output image with the ground truth label to the trained discriminative network so that the discriminative network outputs a first discrimination result.
[0024] In some embodiments, in the step of alternately training the second neural network to be trained and the discriminative network to be trained, training the discriminative network to be trained includes:
[0025] Providing a second sample image to the current second neural network to enable the current second neural network to generate a first sharpness-enhanced image;
[0026] Providing the first sharpness-enhanced image and the original sample image corresponding to the second sample image to the current discriminative network, and adjusting the parameters of the current discriminative network according to the loss function of the current discriminative network so that the discriminative network after parameter adjustment outputs a discrimination result that can characterize whether the input of the discriminative network is the output image of the second neural network or the original sample image.
[0027] In some embodiments, in the step of alternately training the second neural network to be trained and the discriminative network to be trained, training the second neural network to be trained includes:
[0028] Providing a third sample image to the current second neural network to enable the current second neural network to generate a second sharpness-enhanced image;
[0029] Inputting the second definition-enhanced image into the parameter-adjusted discriminant network, so that the parameter-adjusted discriminant network generates a second discrimination result based on the second definition-enhanced image;
[0030] Based on the current loss function of the second neural network, the parameters of the second neural network are adjusted to obtain an updated second neural network; the first item in the current loss function of the second neural network is based on the difference between the second clarity-enhanced image and its corresponding original sample image, and the second item in the current loss function of the second neural network is based on the difference between the second discrimination result and the second target result.
[0031] In some embodiments, the first term in the current loss function of the second neural network is λ1LossG1, where λ1 is a preset weight value, and LossG1 is the L1 loss between the second definition-enhanced image and its corresponding original sample image;
[0032] The second term in the current loss function of the second neural network is λ2L D , λ2 is the preset weight, L D is the cross entropy between the second discrimination result and the second target result;
[0033] The third term in the current loss function of the second neural network is λ3 is a preset weight, y is the original sample image corresponding to the second definition-enhanced image, enhancing the image for the second definition;
[0034] is the preset network layer in the preset optimization network, j is the number of layers of the preset network layer in the preset optimization network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer; the preset optimization network adopts the VGG-19 network.
[0035] In some embodiments, the first neural network to be trained includes: a plurality of first upsampling layers, a plurality of first downsampling layers, and a plurality of single-layer convolutional layers, each of the first upsampling layers and each of the first downsampling layers is located between two of the single-layer convolutional layers; the input data of the i-th last single-layer convolutional layer includes the superposition of the output data of the i-th last first upsampling layer and the output data of the i-th positive single-layer convolutional layer; wherein the number of the single-layer convolutional layers is an even number, i is greater than 0 and less than half the number of the single-layer convolutional layers;
[0036] The trained second neural network includes: a plurality of second upsampling layers, a plurality of second downsampling layers, and a plurality of residual blocks. The plurality of second upsampling layers correspond one-to-one to the plurality of first upsampling layers, the plurality of second downsampling layers correspond one-to-one to the plurality of first downsampling layers, and the plurality of residual blocks correspond one-to-one to the plurality of single-layer convolutional layers. The input data of the i-th residual block from the bottom is the superposition of the output data of the i-th second upsampling layer from the bottom and the output data of the i-th residual block from the top.
[0037] The first feature extraction sub-network includes: the first upsampling layer, or the first downsampling layer, or the single-layer convolutional layer; the first output sub-network includes the single-layer convolutional layer; the second feature extraction sub-network includes: the second upsampling layer, or the second downsampling layer, or the residual block; the second output sub-network includes the residual block.
[0038] An embodiment of the present disclosure further provides an image processing device, including a memory and a processor, where a computer program is stored on the memory. When the computer program is executed by the processor, the above-mentioned image processing method is implemented.
[0039] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned image processing method is implemented. Description of the Drawings
[0040] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:
[0041] Figure 1 It is a schematic diagram of a convolutional neural network.
[0042] Figure 2 It is a schematic diagram of the image processing method provided in the embodiment of the present disclosure.
[0043] Figure 3 It is a schematic diagram of the first training method provided in the embodiment of the present disclosure.
[0044] Figure 4 It is a schematic diagram of the network architecture including the first neural network and the second neural network provided in the embodiment of the present disclosure.
[0045] Figure 5 It is an example diagram of a residual block.
[0046] Figure 6 It is a schematic diagram of the structure of the trained discriminant network provided in the embodiment of the present disclosure.
[0047] Figure 7 It is a flowchart of an alternative implementation of step S21 provided in an embodiment of the present disclosure.
[0048] Figure 8 They are effect diagrams before and after image processing using the image processing method of the embodiment of the present disclosure. Detailed implementation manners
[0049] The following will describe in detail the specific implementation manners of the present disclosure with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and understanding the present disclosure, and are not used to limit the present disclosure.
[0050] Before video transmission, due to bandwidth limitations, compression encoding is required. However, the compressed video will generate various compression noises, which affect people's viewing experience of the video on the display terminal. The video compression and repair technology based on deep learning can improve the repair effect of video compression noises. However, the parameter quantity of the algorithm model of deep learning is large, resulting in an excessive calculation amount on the display terminal.
[0051] The main component of the deep learning system is the convolutional neural network. Figure 1 It is a schematic diagram of a convolutional neural network. This convolutional neural network can be used for image processing, which uses images as inputs and outputs, and replaces scalar weights with filters (i.e., convolutions). Figure 1 Only a convolutional neural network with a 3-layer structure is shown in [Figure reference], and the embodiments of the present disclosure are not limited thereto. As Figure 1 shown, the convolutional neural network includes an input layer 101, a hidden layer 102, and an output layer 103. Four input images are input into the input layer 101. There are 3 units in the middle hidden layer 102 to output 3 output images, and there are 2 units in the output layer 103 to output 2 output images.
[0052] As Figure 1 shown, the convolutional layer has weights w ij k and biases b i k , where the weights w ij k represent the convolution kernels, and the biases are scalars added to the output of the convolutional layer. Here, k is the label representing the number of the input layer 101, and i and j are the labels of the units of the input layer 101 and the hidden layer 102 respectively. For example, the first convolutional layer 201 includes a first set of convolution kernels ( Figure 1 the w in ij 1 ) and a first set of biases ( Figure 1 the b in i 1 ). The second convolutional layer 202 includes a second set of convolution kernels ( Figure 1 the w inij 2 ) and the second set of offsets ( Figure 1 b in i 2 ). Generally, each convolutional layer includes dozens or hundreds of convolutional kernels. If the convolutional neural network is a deep convolutional neural network, it may include at least five convolutional layers.
[0053] Such as Figure 1 As shown, the convolutional neural network further includes a first activation layer 203 and a second activation layer 204. The first activation layer 203 is located after the first convolutional layer 201, and the second activation layer 204 is located after the second convolutional layer 202. The activation layer includes an activation function, and the activation function is used to introduce non-linearity into the convolutional neural network so that the convolutional neural network can better solve more complex problems. The activation function may include a rectified linear unit (ReLU) function, a sigmoid function (Sigmoid function), or a hyperbolic tangent function (tanh function), etc. The activation layer can be a separate layer of the convolutional neural network, or the activation layer can also be included in the convolutional layer.
[0054] Figure 1 The convolutional neural network can be used to improve the clarity of an image. The trained convolutional neural network enhances the clarity of the input low-clarity image to obtain a high-clarity image. The training process of the convolutional neural network is a process of optimizing the parameters of the convolutional neural network. Among them, the loss of the convolutional neural network helps to optimize the parameters (weights) of the convolutional neural network. The goal of the training process is to minimize the loss of the neural network by optimizing the parameters of the neural network. Among them, the loss of the neural network is used to measure the quality of the prediction of the network model, that is, to represent the degree of difference between the prediction result and the actual data.
[0055] Embodiments of the present disclosure provide an image processing method. Figure 2 It is a schematic diagram of the image processing method provided in the embodiments of the present disclosure. As Figure 2 shown, the image processing method includes: S10. Process the input image using the trained first neural network to obtain a target output image. The clarity of the target output image is higher than that of the input image.
[0056] It should be noted that in the embodiments of the present disclosure, "clarity" refers to, for example, the clarity of each detailed texture and its boundary in the image. The higher the clarity, the better the visual effect on the human eye. The clarity of the target output image is higher than that of the input image, which means that, for example, the image processing method provided in the embodiments of the present disclosure is used to process the input image, such as performing denoising and / or deblurring processing, so that the obtained target output image is clearer than the input image.
[0057] In an embodiment of the present disclosure, the trained first neural network is obtained by training a first neural network to be trained using a first training method. Figure 3 It is a schematic diagram of the first training method provided in the embodiment of the present disclosure. As Figure 3 shown, the first training method includes:
[0058] S21. Alternately train the second neural network to be trained and the discriminant network to be trained to obtain a trained second neural network and a trained discriminant network.
[0059] In an embodiment of the present disclosure, the parameters of the trained second neural network are more than those of the first neural network to be trained. Figure 4 It is a schematic diagram of a network architecture including the first neural network to be trained and the trained second neural network provided in the embodiment of the present disclosure. The first neural network 10 to be trained includes: a plurality of first feature extraction sub-networks ML1 and a first output sub-network OL1 located after the plurality of first feature extraction sub-networks ML1. The trained second neural network 20 includes: a plurality of second feature extraction sub-networks ML2 and a second output sub-network OL2 located after the plurality of second feature extraction sub-networks ML2. The first feature extraction sub-networks ML1 and the second feature extraction sub-networks ML2 are in one-to-one correspondence.
[0060] In some embodiments, the first neural network 10 to be trained includes: a plurality of first upsampling layers 13, a plurality of first downsampling layers 12, and a plurality of single-layer convolutional layers 11. Each first upsampling layer 13 and each first downsampling layer 12 are located between two single-layer convolutional layers 11. The input data of the i-th single-layer convolutional layer 11 from the bottom includes the output data of the i-th first upsampling layer 13 from the bottom and the superposition of the output data of the i-th single-layer convolutional layer 11 from the top. Among them, the number of single-layer convolutional layers 11 is even, and i is greater than 0 and less than half of the number of single-layer convolutional layers. The second neural network 20 includes: a plurality of second upsampling layers 23, a plurality of second downsampling layers 22, and a plurality of residual blocks 21. The plurality of second upsampling layers 23 and the plurality of first upsampling layers 13 are in one-to-one correspondence. The plurality of second downsampling layers 22 and the plurality of first downsampling layers 12 are in one-to-one correspondence. The plurality of residual blocks 21 and the plurality of single-layer convolutional layers 11 are in one-to-one correspondence. The input data of the i-th residual block 21 from the bottom includes the output data of the i-th second upsampling layer 23 from the bottom and the superposition of the output data of the i-th residual block 21 from the top.
[0061] The single convolutional layer 11, the first upsampling layer 13, and the first downsampling layer 12 all use a 3*3 convolutional kernel, and the number of convolutional kernels is 128 for all. The sampling ratio of the second upsampling layer 23 is the same as that of the first upsampling layer 13, and the sampling ratio of the second downsampling layer 22 is the same as that of the first downsampling layer 12. Exemplarily, both the first upsampling layer 13 and the first downsampling layer 12 are 2-fold sampling. The first downsampling layer 12 and the second downsampling layer 22 may include an inverse Muxout layer, strided convolution, a maxpool layer, or a standard per-channel downsampler (such as bicubic interpolation). The first upsampling layer 13 and the second upsampling layer 23 may include a Muxout layer, strided transposed convolution, or a standard per-channel upsampler (such as bicubic interpolation).
[0062] Figure 5 An example diagram of a residual block is shown as Figure 5 shown. Each residual block 21 includes three sub-residual blocks 21a connected in sequence. Each sub-residual block 21a uses two convolutional layers with a 3*3 convolutional kernel, and an activation layer is connected between the two convolutional layers. In each sub-residual block 21a, its input is superimposed on the output result of the last convolutional layer, so as to serve as the output of the sub-residual block 21a. The activation layer includes an activation function, and the activation function may include a rectified linear unit (ReLU) function, a sigmoid function, or a hyperbolic tangent function (tanh function), etc. The activation layer can be a separate layer of the convolutional neural network, or the activation layer can also be included in the convolutional layer. In the first convolutional network 10, the single convolutional layer 11 is used to replace the residual block 21 in the second neural network 20, thereby reducing the number of parameters of the first convolutional network 10.
[0063] The first feature extraction sub-network ML1 includes: a first upsampling layer 13, or a first downsampling layer 12, or the single-layer convolutional layer 11; the first output sub-network OL1 includes a single-layer convolutional layer 21; the second feature extraction sub-network ML2 includes: a second upsampling layer 23, or a second downsampling layer 22, or a residual block 21; the second output sub-network OL2 includes a residual block 21. Additionally, the number of channels of the output image of the first feature extraction sub-network ML1 is greater than the number of channels of the output image of the second feature extraction sub-network ML2. Exemplarily, the number of channels of the output image of the first feature extraction sub-network ML1 is 128, and the number of channels of the output image of the second feature extraction sub-network ML2 is 32. It should be noted that in a neural network, the images input to each network layer are represented as matrices. The image received by the first layer in the neural network can be an image matrix with three channels of R, G, and B. That is, the image matrix of each channel represents the data of the red component, green component, or blue component of the image. And each network layer is used to extract features from the image. After feature extraction, the output data of the network layer includes multiple matrices, and each matrix represents one channel of the image.
[0064] In the embodiments of the present disclosure, the second neural network to be trained and the discriminative network to be trained are alternately trained so as to compete with each other to obtain the best model. Specifically, the trained second neural network is configured to transform the received image with a first sharpness into an image with a second sharpness, where the second sharpness is greater than the first sharpness. The trained discriminative network is configured to determine the matching degree between the output result of the second neural network and a preset standard image, and this matching degree is between 0 and 1. Among them, when training the second neural network to be trained, by adjusting the parameters of the current second neural network, so that after the output result of the second neural network with adjusted parameters is input into the current discriminative network, the discriminative network outputs a matching degree as close to 1 as possible; when training the discriminative network to be trained, by adjusting the parameters of the current discriminative network, so that after the preset standard image is input into the current discriminative network, the output result of the current discriminative network is as close to 1 as possible (that is, the discriminative network determines its input as a "true" sample), and when the output result of the current second neural network enters the discriminative network, the output result of the discriminative network is as close to 0 as possible (that is, the discriminative network determines its input as a "false" sample). Through the alternate training of the second neural network and the discriminative network, the discriminative network is continuously optimized to try to distinguish the output result of the second neural network from the preset standard image as much as possible, while the second neural network is continuously optimized to make the output result as close to the preset standard image as possible. This method enables the two neural networks to compete and continuously improve based on the better and better results of the other network in each training to obtain an increasingly excellent network model.
[0065] Figure 6The following is a schematic structural diagram of the trained discriminative network provided in the embodiments of the present disclosure. As Figure 6 shown, the trained discriminative network 30 includes multiple convolutional layers 31-34 and a fully connected layer 35. Exemplarily, each of the convolutional layers 31-34 is a 2x downsampling convolutional layer, and each convolutional layer 31-34 is followed by an activation layer. The activation layer includes an activation function, and the activation function may include a rectified linear unit (ReLU) function, a sigmoid function, or a hyperbolic tangent function (tanh function), etc. Each of the convolutional layers 31-34 uses a 3*3 convolutional kernel. The number of channels of the image output by the convolutional layer 31 is 32, the number of channels of the image output by the convolutional layer 32 is 64, the number of channels of the image output by the convolutional layer 33 is 128, and the number of channels of the image output by the convolutional layer 34 is 192. The fully connected layer 35 outputs a 1024*1 vector, and after passing through an activation layer (for example, this activation layer uses sigmoid as the activation function), a value between 0 and 1 is output.
[0066] It should be understood that the structures of the trained discriminative network and the discriminative network to be trained (i.e., the number of convolutional layers and the number of convolutional kernels in the convolutional layers) are the same, and the difference lies in the weights in the convolutional layers.
[0067] It should be noted that Figure 4 、 Figure 5 The number of network layers in the trained first neural network 10 and the trained discriminative network 30 in
[0068] S22. Provide the first sample image to the trained second neural network and the discriminative network to be trained respectively, so that the discriminative network to be trained outputs a first output image, and the trained second neural network outputs a second output image.
[0069] In some examples, the original video can be compressed at a low bit rate (for example, the compression bit rate is 1 Mbps) to obtain a compressed video. Each frame image in the compressed video can be used as a first sample image with noise, and the noise can be Gaussian noise.
[0070] S23. Provide the first output image to the trained discriminative network, so that the trained discriminative network generates a first discrimination result based on the first output image.
[0071] S24. Adjust the parameters of the first neural network according to the total loss to obtain an updated first neural network. The total loss includes a first loss, a second loss, and a third loss. The first loss is obtained based on the difference between the first output image and the second output image; the second loss is obtained based on the difference between the first discrimination result and the first target result; the third loss is obtained based on the difference between the output images of at least one first feature extraction subnet and the output images of the corresponding second feature extraction subnets.
[0072] As described above, the output of the trained discrimination network 30 is a matching degree between 0 and 1. In this case, the first target result is a matching degree close to 1 or equal to 1.
[0073] It should be noted that "adjust the parameters of the first neural network according to the total loss" means adjusting the parameters of the first neural network so that the value of the total loss generally shows a decreasing trend when the first training method is performed multiple times. The number of executions of the first training method can be preset, or when the total loss is less than a preset value, the first training method is no longer performed. It should also be noted that the first sample images used in different executions of the first training method can be different.
[0074] In the embodiments of the present disclosure, the difference between two images is the difference in the low-frequency information of the two images, which can be characterized by the L1 loss value, the mean square error (MSE), the similarity (SSIM), etc.
[0075] In some embodiments, the first loss includes the L1 loss between the first output image and the second output image, specifically x1Loss1, where x1 is a preset weight value, and Loss1 is the L1 loss between the first output image and the second output image, that is, Loss1 = ||y1 - y2||1, y1 is the first output image, and y2 is the second output image.
[0076] In some embodiments, the second loss includes the cross-entropy loss between the first discrimination result and the first target result, specifically x2Loss2, x2 is a preset weight value, and Loss2 is the cross-entropy loss between the first discrimination result of the discrimination network and the first target result.
[0077] Specifically, Loss2 = -[PlogP’+(1 - P)log(1 - P’)], where P is the first target result and P’ is the first discrimination result. In some embodiments, step S23 specifically includes: setting the first output image with a true value label and providing the first output image with the true value label to the trained discrimination network so that the discrimination network outputs the first discrimination result. The true value label is used to indicate that the image is a "true" sample, and the first target result is the probability corresponding to the true value label. For example, the first target result is 1.
[0078] In some embodiments, the third loss is specifically obtained based on the difference between the transformed image of the output image of at least one first feature extraction sub-network and the output image of the corresponding second feature extraction sub-network. For example, the network architecture including the first neural network and the second neural network further includes a plurality of dimensionality reduction layers 40. The dimensionality reduction layers 40 are in one-to-one correspondence with the first feature extraction sub-networks ML1, and the dimensionality reduction layers 40 are configured to perform channel dimensionality reduction on the output images of the corresponding first feature extraction sub-networks to generate intermediate images; the number of channels of the intermediate images is the same as the number of channels of the output images of the second feature extraction sub-networks.
[0079] The first training method further includes: providing the output images of the plurality of second feature extraction sub-networks to the plurality of dimensionality reduction layers one-to-one, so that each dimensionality reduction layer generates an intermediate image; the number of channels of the intermediate image is the same as the number of channels of the output image of the first feature extraction sub-network. In this case, in step S24, the parameters of both the first neural network and the dimensionality reduction layers are adjusted. Among them, the third loss is obtained based on the sum of the differences between each intermediate image and the output image of the corresponding first feature extraction sub-network.
[0080] In some embodiments, the difference between the intermediate image and the output image of the second feature extraction sub-network is represented by the L2 loss between the two. The third loss is x3Loss3, where Loss3 is the sum of the L2 losses between the output image of each first feature extraction sub-network and the corresponding intermediate image. Specifically, Loss3 is calculated according to the following formula:
[0081]
[0082] where x3 is a preset weight value; T is the number of first feature extraction sub-networks, S n (z) is the output image of the nth layer second feature extraction sub-network in the first neural network, G n (z) is the output image of the nth layer first feature extraction sub-network in the second neural network, f(G n (z)) is the intermediate image output by the dimensionality reduction layer corresponding to the nth layer first feature extraction sub-network in the second neural network.
[0083] In the embodiments of the present disclosure, compared with the trained second neural network, the trained first neural network is simplified, and the trained first neural network has fewer parameters and a simpler network structure, so that the trained first neural network occupies less resources (such as computing resources, storage resources, etc.) during its operation, and thus can be applied to lightweight terminals. Moreover, in the total loss used for training the first neural network to be trained, the first loss is obtained based on the difference between the output result of the first neural network and the output result of the second neural network, the second loss is obtained based on the difference between the discrimination result of the trained discrimination network and the first target result, and the third loss is obtained based on the difference between the output images of at least one first feature extraction sub-network and the output images of the corresponding second feature extraction sub-network, so that the performance of the trained first neural network is as close as possible to that of the second neural network. Therefore, the embodiments of the present disclosure can reduce the parameters of the image processing model while ensuring the image processing effect, thereby improving the image processing speed.
[0084] In some embodiments, the total loss further includes: a fourth loss, which is obtained based on the perceptual loss between the first output image and the second output image. Among them, the perceptual loss is used to characterize the difference in the high-frequency information of two images (for example, detailed features such as textures and hairs on the images).
[0085] Optionally, the fourth loss is: x4 is a preset weight value. is the perceptual loss between the first output image and the second output image, which is calculated according to the following formula:
[0086]
[0087] where y1 is the first output image, y2 is the second output image, is a preset network layer in the trained discrimination network, j is the layer number of the preset network layer in the discrimination network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer. It can be understood that, is the output image of the preset network layer after the first output image is input into the trained discrimination network; is the output image of the preset network layer after the second output image is input into the preset optimization network. Optionally, the preset network layer can be a convolutional layer with 128 output image channels.
[0088] Figure 7 This is an optional implementation flowchart of step S21 provided in the embodiments of the present disclosure, as shown in Figure 7As shown, step S21 specifically includes: alternately performing step S21a and step S21b until a preset training condition is reached. The preset training condition is, for example, that the number of alternating steps S21a and S21b reaches a preset number.
[0089] S21a. Provide the second sample image to the current second neural network, so that the second neural network generates a first definition-enhanced image. Provide the first definition-enhanced image and the original sample image corresponding to the second sample image to the current discriminant network, and adjust the parameters of the discriminant network based on the loss function of the current discriminant network so that the output of the discriminant network after the adjustment can represent the discrimination result of whether the input of the discriminant network is the output image of the second neural network or the original sample image.
[0090] S21b. Provide the third sample image to the current second neural network, so that the second neural network generates a second definition-enhanced image. Input the second definition-enhanced image into the discriminant network after parameter adjustment, so that the discriminant network after parameter adjustment generates a second discrimination result based on the second definition-enhanced image. Adjust the parameters of the second neural network based on the loss function of the second neural network to obtain an updated second neural network.
[0091] It should be noted that, if the nth step S21a and the nth step S21b are considered a training round, then in the first training round, the current second neural network is the second neural network to be trained; in each training round after the first, the current second neural network is the second neural network updated in step S21b of the previous training round. In the first training round, the current discriminant network is the discriminant network to be trained; in each training round after the first, the current discriminant network is the discriminant network after parameter adjustment in step S21a of the previous training round.
[0092] The first term in the loss function of the second neural network is based on the difference between the second definition-enhanced image and its corresponding original sample image, and the second term in the loss function of the second neural network is based on the difference between the second discrimination result and the second target result.
[0093] In some embodiments, the first term in the loss function LossG of the second neural network is λ1LossG1, where λ1 is a preset weight value and LossG1 is the L1 loss between the second definition-enhanced image and its corresponding original sample image. Specifically, y is the original sample image corresponding to the second definition-enhanced image, The image is enhanced for the second definition.
[0094] The second term in the loss function of the second neural network is λ2L D, λ2 is the preset weight, L D is the cross entropy between the second discriminant result and the second target result. The second target result indicates that the input to the discriminant network is the original image corresponding to the second definition-enhanced image, that is, it indicates that the input to the discriminant network is a "true" sample. For example, the second target result is 1.
[0095] The third term in the loss function of the second neural network is Among them, λ3 is the preset weight.
[0096] is the preset network layer in the preset optimization network, j is the number of layers in the preset optimization network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer; the preset optimization network uses the VGG-19 network. It should be noted that in a neural network, the image output by each network layer is not a visually visible image, but is represented by a matrix. The image height can be regarded as the number of rows in the matrix, and the image width can be regarded as the number of columns in the matrix.
[0097] That is to say, after the second neural network is trained multiple times, the L1 loss value of the image output by the updated second neural network is as close as possible to that of the original sample image, and the perceptual loss of the image output by the second neural network is as close as possible to that of the original sample image. At the same time, after the image output by the second neural network is provided to the discriminant network, the result output by the discriminant network is close to 1.
[0098] Optionally, in steps S21a and S21b in the same round of training, the second sample image and the third sample image may be the same, while in different rounds of training, the second sample image and the third sample image may be different.
[0099] It should be noted that in each round of training, the training steps of the discriminant network can be performed first, or the training steps of the generative network can be performed first.
[0100] In some examples, the original video can be losslessly compressed to obtain a lossless compressed video, and the image frames in the lossless compressed video image can be used as the original sample images; the original video can be compressed at a first bit rate to obtain a low-loss compressed video, and the image frames in the low-loss compressed video can be used as the second sample image or the third sample image.
[0101] In some examples, the training process of step S21 may use an Adam optimizer with a learning rate of 1e-4.
[0102] In the embodiments of the present disclosure, the trained first neural network has fewer parameters and a simpler network structure than the second neural network, such that the first neural network occupies fewer resources (such as computing resources, storage resources, etc.) during its operation, and thus can be applied to lightweight terminals. Moreover, the training method of the first neural network to be trained can make the performance of the trained first neural network close to that of the trained second neural network. Therefore, the image processing method of the embodiments of the present disclosure can obtain an image with higher clarity while improving the image processing speed.
[0103] Figure 8 The figures on the left and right are the effect diagrams before and after image processing using the image processing method of the embodiments of the present disclosure. Figure 8 The left figure is the input image before processing, and the right figure is the target output image after processing. As Figure 8 shown, after image processing, the clarity of the image is improved. Moreover, compared with the second convolutional network, the compression multiple of the number of parameters of the first convolutional network is greater than 50 times, and the processing speed is increased by about 15 times.
[0104] The present disclosure also provides an image processing device, including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the training method of the above image processing model is implemented.
[0105] The present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the training method of the above image processing model is implemented.
[0106] The above memory and the computer-readable storage medium include but are not limited to the following readable media: such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), and other non-transitory media. Examples of processors include but are not limited to general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, etc.
[0107] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present disclosure, and the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.
Claims
1. An image processing method, comprising: processing an input image using a trained first neural network to obtain a target output image; the clarity of the target output image being greater than that of the input image; wherein, the trained first neural network is obtained by training a first neural network to be trained using a first training method, and the first training method includes: alternately training a second neural network to be trained and a discriminant network to be trained to obtain a trained second neural network and a trained discriminant network; wherein, the parameters of the trained second neural network are more than those of the first neural network to be trained; the trained second neural network is configured to transform an image with a first clarity received into an image with a second clarity, and the second clarity is greater than the first clarity; the first neural network to be trained includes: a plurality of first feature extraction sub-networks and a first output sub-network located after the plurality of first feature extraction sub-networks, the trained second neural network includes: a plurality of second feature extraction sub-networks and a second output sub-network located after the plurality of second feature extraction sub-networks, and the first feature extraction sub-networks and the second feature extraction sub-networks correspond one by one; providing a first sample image to the trained second neural network and the first neural network to be trained respectively, so that the first neural network to be trained outputs a first output image and the trained second neural network outputs a second output image; setting the first output image with a ground truth label and providing the first output image with the ground truth label to the trained discriminant network, so that the trained discriminant network generates a first discrimination result based on the first output image; adjusting the parameters of the first neural network according to the total loss to obtain the updated first neural network; wherein, the total loss includes a first loss, a second loss and a third loss, the first loss is obtained based on the difference between the first output image and the second output image; the second loss is obtained based on the difference between the first discrimination result and a first target result; the third loss is obtained based on the difference between the output images of at least one of the first feature extraction sub-networks and the output images of the corresponding second feature extraction sub-networks; wherein, in the step of alternately training the second neural network to be trained and the discriminant network to be trained, training the discriminant network to be trained includes: providing a second sample image to the current second neural network, so that the current second neural network generates a first clarity improvement image; providing the first clarity improvement image and the original sample image corresponding to the second sample image to the current discriminant network, and adjusting the parameters of the current discriminant network according to the loss function of the current discriminant network, so that the discriminant network after parameter adjustment outputs a discrimination result capable of characterizing whether the input of the discriminant network is the output image of the second neural network or the original sample image; In the step of alternately training the second neural network to be trained and the discriminant network to be trained, training the second neural network to be trained includes: providing a third sample image to the current second neural network so that the current second neural network generates a second sharpness-enhanced image; inputting the second sharpness-enhanced image into the discriminant network after parameter adjustment so that the discriminant network after parameter adjustment generates a second discriminant result based on the second sharpness-enhanced image; adjusting the parameters of the current second neural network based on the loss function of the current second neural network to obtain an updated second neural network.
2. The image processing method according to claim 1, wherein, The number of channels of the output image of the first feature extraction sub-network is less than the number of channels of the output image of the corresponding second feature extraction sub-network; The first training method further includes: providing the output images of the multiple second feature extraction sub-networks to the multiple dimensionality reduction layers one by one so that each dimensionality reduction layer generates an intermediate image; the number of channels of the intermediate image is the same as the number of channels of the output image of the first feature extraction sub-network; Adjusting the parameters of the first neural network according to the total loss function includes: adjusting the parameters of both the first neural network and the dimensionality reduction layer; wherein, the third loss is obtained based on the sum of the differences between each intermediate image and the output image of the corresponding first feature extraction sub-network.
3. The image processing method according to claim 1, wherein The total loss further includes: a fourth loss, which is the perceptual loss based on the first output image and the second output image.
4. The image processing method according to claim 3, wherein, The perceptual loss between the first output image and the second output image is calculated according to the following formula: where y1 is the first output image and y2 is the second output image. is a preset network layer in the trained discriminative network, j is the layer number of the preset network layer in the discriminative network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer.
5. The image processing method according to any one of claims 1 to 4, wherein, The first loss includes the L1 loss between the first output image and the second output image.
6. The image processing method according to any one of claims 1 to 4, wherein, The second loss includes the cross-entropy loss between the first discriminant result and the first target result.
7. The image processing method according to any one of claims 2 to 4, wherein, The third loss term includes the sum of the L2 losses between the output image of each first feature extraction sub-network and the corresponding intermediate image.
8. The image processing method according to claim 1, wherein, During the process of training the second neural network to be trained, the first term in the loss function of the current second neural network is based on the difference between the second sharpness-enhanced image and its corresponding original sample image, and the second term in the loss function of the current second neural network is based on the difference between the second discriminant result and the second target result.
9. The image processing method according to claim 8, wherein, The first term in the loss function of the current second neural network is λ1LossG1, where λ1 is a preset weight value and LossG1 is the L1 loss between the second sharpness-enhanced image and its corresponding original sample image; The second term in the loss function of the current second neural network is λ2L D , where λ2 is a preset weight value, and L D is the cross-entropy between the second discrimination result and the second target result; The third term in the loss function of the current second neural network is λ3 is a preset weight value, y is the original sample image corresponding to the second sharpness-enhanced image, is the second sharpness-enhanced image; For a preset network layer in a preset optimization network, j is the layer number of the preset network layer in the preset optimization network, C is the number of channels of the output image of the preset network layer, H is the height of the output image of the preset network layer, and W is the width of the output image of the preset network layer; the preset optimization network adopts a VGG-19 network.
10. The image processing method according to any one of claims 1 to 4, wherein, The first neural network to be trained includes: a plurality of first upsampling layers, a plurality of first downsampling layers, and a plurality of single-layer convolutional layers. Each of the first upsampling layers and each of the first downsampling layers are located between two of the single-layer convolutional layers; the input data of the i-th last single-layer convolutional layer includes the output data of the i-th last first upsampling layer and the output data of the i-th first single-layer convolutional layer superimposed; wherein, the number of single-layer convolutional layers is even, and i is greater than 0 and less than half of the number of single-layer convolutional layers; The trained second neural network includes: a plurality of second upsampling layers, a plurality of second downsampling layers, and a plurality of residual blocks. The plurality of second upsampling layers correspond one-to-one to the plurality of first upsampling layers, the plurality of second downsampling layers correspond one-to-one to the plurality of first downsampling layers, and the plurality of residual blocks correspond one-to-one to the plurality of single-layer convolutional layers. The input data of the i-th residual block from the bottom is the superposition of the output data of the i-th second upsampling layer from the bottom and the output data of the i-th residual block from the top. The first feature extraction sub-network includes: the first upsampling layer, or the first downsampling layer, or the single-layer convolutional layer; the first output sub-network includes the single-layer convolutional layer; the second feature extraction sub-network includes: the second upsampling layer, or the second downsampling layer, or the residual block; the second output sub-network includes the residual block.
11. An image processing apparatus includes a memory and a processor, and a computer program is stored on the memory, wherein, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Neural network training method, image processing method and image processing device
CN111767979A
Super resolution using a generative adversarial network
US20180075581A1