An Image Joint Processing Method and System Based on Multi-Task Learning
By adopting multi-task learning method in joint image processing, image super-resolution reconstruction and image halftone are jointly optimized, the problems of high computational complexity and poor generalization ability in the prior art are solved, and the effect of efficiently processing two tasks in a single model is achieved.
Patent Information
- Application Number
- CN202410733392.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-06-07
AI Technical Summary
The existing multitasking model has high computational complexity and poor generalization capabilities in image joint processing, which limits its practical application effect.
Multi-task learning is used to jointly optimize image super-resolution reconstruction and image halftone, and a joint loss function is designed to realize the joint processing of the two tasks through shallow feature extraction module, deep feature extraction module, super-segment reconstruction module and halftone reconstruction module.
Simultaneous processing of two tasks, super-resolution reconstruction and half-tone image generation in a single model improves the performance of both tasks, reduces computational complexity and improves generalization capabilities.
Smart Images

Figure CN118587091B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to an image joint processing method and system based on multi-task learning. Background Art
[0002] Raster Image Processor (RIP) software technology converts digital image data into image dot matrix data that can be accepted by a printer, which affects the effect of the image presented on the printing medium and is usually applied to devices such as digital printing, printers, and plotters. There are three key technologies in RIP: image color separation, image super-resolution, and image halftoning.
[0003] Image halftoning technology is the process of converting an image that has undergone color separation and super-resolution processing into a binary image, that is, converting the image into a dot matrix diagram containing only two colors, black and white. Through delicate dot arrangement and size variation, it visually simulates the effect of a continuous tone image, making the printed image look rich in gray levels and color gradations, rather than simply black and white contrast.
[0004] Image super-resolution technology finely reconstructs and restores low-resolution images (LR) through advanced software algorithms to generate high-resolution (HR) images, which are extremely rich in details. This technology plays an indispensable role in many key fields, especially in applications such as satellite remote sensing, medical imaging, security monitoring, and image compression and transmission, all of which have demonstrated its unique value. Although significant progress has been made in image super-resolution reconstruction technology based on deep learning, continuous research and technological innovation are still needed to fully realize its application value in real life. This includes, but is not limited to, improving algorithm efficiency, reducing resource consumption, and optimizing algorithm structures to adapt to different application scenarios.
[0005] Multi-task learning can jointly model the relevant information of multiple tasks by sharing the model's representation and learning process, thereby improving the performance and generalization ability of each task. Currently, most multi-task based image joint models have problems such as complex model structures, large computational amounts, and poor generalization ability, which limit the practical application effects of these methods. Summary of the Invention
[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problems of high computational complexity and poor generalization ability of multi-task models in the prior art.
[0007] To solve the above technical problems, the present invention provides an image joint processing method and system based on multi-task learning, which jointly optimizes image super-resolution reconstruction and image halftoning in a multi-task learning manner, designs a joint loss function, improves the performance of both tasks simultaneously, and realizes the processing of two tasks, namely super-resolution reconstruction and halftone image generation, in a single model. The method includes the following steps:
[0008] S1: Construct an image joint processing model; wherein, the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module, and a halftone reconstruction module. The output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module, and two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module;
[0009] S2: Design a joint loss function to train the image joint processing model to obtain a trained image joint processing model;
[0010] S3: The trained image joint processing model generates a super-resolution image through the super-resolution reconstruction module and generates a halftone image through the halftone reconstruction module.
[0011] In an embodiment of the present invention, the shallow feature extraction module is composed of a noise compensation module and a convolutional layer with a convolution kernel of n×n;
[0012] Wherein, the noise compensation module uses two collaborative convolutions as the feature map of the output of the previous layer network and the map of Gaussian noise respectively, and then obtains the input of the next layer network through additive variation.
[0013] In an embodiment of the present invention, the deep feature extraction module is constructed by sequentially connecting multiple cascaded residual blocks. The cascaded residual block includes a CUNet module and a feature fusion module connected to the CUNet module;
[0014] Wherein, the image joint processing model based on multi-task learning obtains a deep feature image by using the deep feature extraction module, including: inputting a shallow feature image into the cascaded residual block, splicing the identity mapping of the shallow feature image and its output feature after being processed by the CUNet module, and fusing the spliced features through the feature fusion module to obtain a fused feature image as the input of the next CUNet module. After being processed by multiple cascaded residual blocks, a deep feature image is obtained.
[0015] In one embodiment of the present invention, the CUNet module is based on an encoder-decoder architecture and uses M skip connection blocks as basic blocks, where m skip connection blocks are used as encoders and the remaining skip connection blocks are used as decoders. The features obtained by fusing the output features of the m encoders processed by the multi-scale dilated convolution extraction module and the output features of the encoders are jointly used as the input of the decoder, and the multi-level features are concatenated by the decoder.
[0016] The multi-scale dilated convolution extraction module uses the output features of the previous layer of the network, which are processed by 1×1 pointwise convolution and then output through dilated convolutions with dilation rates of 1, 2, and 3, and then the input of the next layer of the network is obtained through feature fusion.
[0017] In one embodiment of the present invention, the super-resolution reconstruction module consists of a convolution with a convolution kernel of n×n and an upsampling block based on orthogonal position encoding;
[0018] Among them, the super-resolution reconstruction module uses the upsampling block based on orthogonal position encoding to obtain a high-resolution image, including:
[0019] Extracting features from the input image;
[0020] Reordering the extracted features, the number of channels of the features is 3(2n + 1) 2 , each pixel point in the features corresponds to 3(2n + 1) 2 feature values, and the feature values are equally divided into three parts corresponding to the features of the RGB three channels as the latent code of the pixel point;
[0021] Select a pixel point in the input image, obtain the boundary point coordinates of the pixel point, and define a query point within the area of the pixel point to obtain the nearest boundary point of the query point, and use the coordinate difference between the query point and the nearest boundary point as the orthogonal basis of the query point;
[0022] Perform matrix multiplication on the orthogonal basis of the query point and the latent code, calculate the RGB pixel value of the query point, and based on the RGB pixel value of the query point, render each pixel value of the high-resolution image, thereby obtaining the high-resolution image.
[0023] In one embodiment of the present invention, the halftone reconstruction module reconstructs the shallow feature image obtained by the shallow feature extraction module and the deep feature image obtained by the deep feature extraction module to obtain a halftone image.
[0024] In one embodiment of the present invention, the expression of the joint loss function L is:
[0025] L = L RH + LSR
[0026] Among them, L RH is the halftone loss, and L SR is the super-resolution loss. The super-resolution loss L SR is:
[0027]
[0028] Among them, ||·|| represents L1 regularization, and I SR represents the super-resolution image reconstructed from the low-resolution image through the image super-resolution model; N represents the number of high-resolution image and low-resolution image pairs in each batch during the training of the image super-resolution model; and respectively represent the values of the i-th pixel point in the super-resolution image I SR and the high-resolution image I HR ; SSIM(I HR , I SR ) represents the structural similarity between the super-resolution image I SR and the high-resolution image I HR .
[0029] In an embodiment of the present invention, the expression of the halftone loss L RH is:
[0030] L RH = ω1L T + ω2L N + ω3L B + ω4L G
[0031] Among them, L T is the tone consistency loss, L N is the blue noise loss, L B is the binary loss, L G is the perceptual consistency loss, and ω1, ω2, ω3, and ω4 are all hyperparameters;
[0032] The tone consistency loss L T is:
[0033] L T = E Ic ∈I{||G(O h ) - G(I c )||2}
[0034] Among them, G(·) is a Gaussian filter, and E Ic ∈I{·} represents the average operator for all input images I c in the training set, and O hRepresents a halftone image;
[0035] The binary loss L B is:
[0036]
[0037] where is the pseudo - halftone image before the binary gate, and C d is a constant matrix of the same size as , where the value of all elements is d = {0, 1};
[0038] The perceptual consistency loss L G is:
[0039] LG = EI c ∈I{||F(O h ) - I c ||2}
[0040] where F(·) represents the inverse halftone pattern.
[0041] In an embodiment of the present invention, the blue noise loss L N is:
[0042] L N = L D + σL AS
[0043] where L D represents the low - frequency component loss in the spectrum, L AS represents the anisotropic loss in the high - frequency region, and σ is a weight factor;
[0044] The low - frequency component loss L D is:
[0045] L D = E Ic ∈I{||(DCT(O h ) - DCT(I c ))⊙M||1}
[0046] where DCT(·) represents the discrete cosine transform, ⊙ represents the element - wise product, and M is a constant binary mask;
[0047] The anisotropic loss L AS in the high - frequency region is:
[0048]
[0049] where P i (f) and P i (f ρrepresent the average power and sampling point power on the i-th frequency band respectively, and θ represents the number of frequency bands.
[0050] Based on the same inventive concept, the present invention also provides an image joint processing system based on multi-task learning, which is used to implement the image joint processing method based on multi-task learning, and specifically includes the following modules:
[0051] A model construction module, used to construct an image joint processing model; wherein, the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module, and a halftone reconstruction module. The output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module, and two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module;
[0052] A model training module, used to design a joint loss function to train the image joint processing model to obtain a trained image joint processing model;
[0053] A result acquisition module, used for the trained image joint processing model to generate a super-resolution image through the super-resolution reconstruction module and generate a halftone image through the halftone reconstruction module.
[0054] The above technical solutions of the present invention have the following advantages compared with the prior art:
[0055] 1. The present invention can solve the problems of high computational complexity and single-scale super-resolution in existing image super-resolution reconstruction algorithms, realizes super-resolution at any scale by using a parameter-free upsampling module, and at the same time introduces a multi-scale dilated convolution extraction module, reducing the computational complexity of the network and improving the generalization ability.
[0056] 2. The present invention performs joint modeling by sharing the model. In the training stage, a joint loss function is designed to optimize the effects of image super-resolution and image halftone simultaneously, thereby improving the performance of both tasks at the same time. In the inference stage, separate task branches are used to implement image super-resolution or image halftone, and the computational amount is not increased additionally. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention and in combination with the accompanying drawings, where
[0058] Figure 1 is a flowchart for implementing an image joint processing method based on multi-task learning provided in an embodiment of the present invention;
[0059] Figure 2It is a schematic structural diagram of an image joint processing model based on multi-task learning provided in an embodiment of the present invention;
[0060] Figure 3 It is a schematic structural diagram of a CUNet module provided in an embodiment of the present invention;
[0061] Figure 4 It is a schematic structural diagram of a skip connection block (DCB) provided in an embodiment of the present invention;
[0062] Figure 5 It is a schematic structural diagram of a multi-scale dilated convolution extraction module (MDCB) provided in an embodiment of the present invention;
[0063] Figure 6 It is an upsampling flow chart of an upsampling block based on orthogonal position encoding (OPE) provided in an embodiment of the present invention;
[0064] Figure 7 is Figure 6 Flow chart of the method for obtaining the orthogonal basis of the query point in;
[0065] Figure 8 It is a schematic structural diagram of an image joint processing model based on multi-task learning for experimental verification provided in an embodiment of the present invention;
[0066] Figure 9 It is a schematic structural diagram of an image joint processing system based on multi-task learning provided in an embodiment of the present invention;
[0067] Explanation of the reference numerals in the specification drawings: 100, model construction module; 200, model training module; 300, result acquisition module. Detailed implementation manners
[0068] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention.
[0069] Embodiment 1
[0070] Referring to Figures 1 - 2 as shown, the present invention proposes an image joint processing method based on multi-task learning, and the method includes the following steps:
[0071] S1: Construct an image joint processing model; wherein, the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module, and a halftone reconstruction module. The output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module. Two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module;
[0072] S2: Design a joint loss function to train the image joint processing model to obtain a trained image joint processing model;
[0073] S3: The trained image joint processing model generates a super-resolution image through the super-resolution reconstruction module and generates a halftone image through the halftone reconstruction module.
[0074] In S1, the image joint processing model based on multi-task learning includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module, and a halftone reconstruction module.
[0075] Specifically, the shallow feature extraction module is composed of a noise compensation module and a convolutional layer with a convolutional kernel of n×n; wherein, the noise compensation module uses two collaborative convolutions as the feature map of the output of the upper-layer network and the map of Gaussian noise respectively, and then obtains the input of the lower-layer network through an additive transformation.
[0076] Preferably, the shallow feature extraction module in this embodiment superimposes a noise map on the input of the original three-channel low-resolution picture, and then uses a convolution with a convolutional kernel of 3×3 and a filter number set to 16 to extract shallow features:
[0077] PF = F SFE (Concat(I LR , noise))
[0078] wherein, F SFE represents the shallow feature extraction operation, I LR represents the input low-resolution image, noise represents the Gaussian noise map, and PF represents the shallow features of the image. Then, the shallow features PF are used as the input of the deep feature module's cascaded residual block to further extract the deep features of the image.
[0079] The deep feature extraction module is constructed by sequentially connecting multiple cascaded residual blocks. The cascaded residual block includes a CUNet module and a feature fusion module connected to the CUNet module;
[0080] Among them, the image joint processing model based on multi-task learning uses the deep feature extraction module to obtain a deep feature image, including: inputting the shallow feature image into the cascaded residual block, splicing the identity mapping of the shallow feature image and its output feature after being processed by the CUNet module, and fusing the spliced features through the feature fusion module to obtain a fused feature image as the input of the next CUNet module. After being processed by multiple cascaded residual blocks, a deep feature image is obtained.
[0081] Preferably, in this embodiment, the deep feature extraction module is constructed by using three cascaded residual blocks (CRB), including a CUNet module and a feature fusion module. Further extract deep features from the output of the shallow feature module:
[0082] DF = F CRB (PF)
[0083] Among them, DF represents the deep feature of the image, and F CRB represents the deep feature extraction operation, which is completed by three CRB modules. Its processing flow is expressed by the formula as follows:
[0084]
[0085] Among them represents the operation of the i-th CUNet module. The Concat operation is to splice the features instead of directly adding them. F Fusion represents the operation of the feature fusion module, which effectively fuses the output of the previous module and the current CUNet, can fuse features at different levels, can receive more useful information while reducing the model complexity, and this method also helps to avoid the problems of gradient disappearance and explosion related to the deep learning model structure. The feature fusion module consists of a convolution with a kernel size of 1×1 and a PReLU.
[0086] Furthermore, the CUNet module is based on the encoder-decoder architecture, using M skip connection blocks as basic blocks, where m skip connection blocks are used as the encoder, and the remaining (M - m) skip connection blocks are used as the decoder. The features obtained by fusing the output features of m encoders processed by the multi-scale dilated convolution extraction module and the output features of the encoder are jointly used as the input of the decoder, and the multi-level features are spliced by the decoder;
[0087] The multi-scale dilated convolution extraction module outputs the output features of the previous layer of the network after being processed by 1×1 pointwise convolution and then passing through dilated convolutions with dilation rates of 1, 2, and 3 respectively, and then obtains the input of the next layer of the network through feature fusion.
[0088] Preferably, the structure of the CUNet moduleFigure 3 As shown, 9 skip connection blocks (DCBs) are used as the basic blocks in the CUNet module. The first 5 DCBs are used as the encoder (E1, E2, E3, E4, E5) for multi-level feature extraction, and the remaining 4 blocks are used as the decoder (D1, D2, D3, D4) for feature fusion.
[0089] When fusing the decoder features, the skip connections of the corresponding level of the decoder are retained. Secondly, the multi-scale dilated convolution extraction module (MDCB) is used to fuse the features output by the five encoders, enabling the network to retain the original features to the greatest extent. The processing flow of the CUNet module is expressed by the following formula:
[0090]
[0091] MO = F MDCB (Concat(EO i ))i∈{1,2,3,4,5}
[0092]
[0093] Where and represent the operations of the i-th encoder and the i-th decoder respectively. CUI represents the input of the CUNet module. EO i and DO i represent the outputs of the i-th encoder and the i-th decoder respectively. F MDCB represents the operation of the MDCB, and MO is the output of the MDCB. To represent the hierarchical correspondence between the encoder and the decoder, the two are connected in the reverse order, and the decoder closest to the encoder is marked with the number 4.
[0094] The skip connection block (DCB) connects the features along the channel dimension, realizing the reuse of features, that is, the features generated by all previous convolution operations will be used as input in the current convolution operation through feature splicing, improving the network performance with fewer parameters and computational costs. The DCB structure is as Figure 4 shown.
[0095] To achieve the goal of model lightweighting, the DCB module used in this embodiment only contains three convolutional layers, which can also effectively perform feature extraction operations. The convolutional layers in DCB include a 3×3 convolution and a PReLU. During the feature extraction process, the original input of each convolutional layer is connected to the outputs of all previous convolutional layers through channel concatenation. Therefore, early feature information will be continuously retained and accessed by subsequent convolutional layers. The outputs of the three convolutional layers are added pixel by pixel to obtain the final output features, which can make good use of features at all levels. The processing process of the DCB module is expressed by the following formula:
[0096] F1 = FC1(I F )
[0097] F2 = FC2(Concat(I F , F1))
[0098] F3 = FC3(Concat(I F , F1, F2))
[0099] O F = F1 + F2 + F3
[0100] Where I F represents the input feature of DCB, FC i represents the i-th convolutional layer operation, F i represents the output of the i-th convolutional layer, and O F represents the final output of the DCB block. The parameter information of the three convolutional layers in DCB is shown in Table 1. All use the PReLU activation function, the convolution kernel size is 3×3, and the output channels of the convolution are all 16. The difference is that the input channels of the convolution are different, which are 16, 32, and 48 respectively.
[0101] Table 1
[0102]
[0103] In the traditional UNet, feature fusion is performed through long connections between the corresponding encoder and decoder. However, after passing through one encoder and decoder, the details of the image often have a certain loss. Moreover, the output channels of the convolutional layers in DCB are only 16, which will exacerbate this problem. Therefore, a multi-scale dilated convolution extraction module (MDCB) is used to fuse the features output by 5 encoders, and then the fused features are respectively input into each encoder, enabling the encoder to learn information with different receptive fields. The structure of MDCB is as Figure 5 shown.
[0104] The MDCB module first compresses the channels through a convolutional layer with a 1×1 convolutional kernel, reducing the feature dimension of the input multi-scale features. Then, it is fed into three branches respectively. These three branches are composed of 3×3 dilated convolutions with dilation rates of 1, 2, and 3. Such a setting can have a broader receptive field without increasing the network parameters. The processing flow of the MDCB module is expressed by the following mathematical formula:
[0105] I fd =F C (PReLU(I f ))
[0106]
[0107] O MDCB =PReLU(F fusion (Concat(B i ))),i∈{1,2,3}
[0108] Among them, I f is the input of the MDCB module, F C represents the compression operation, I fd is the feature output after compression processing, represents the operation of the i-th dilated convolution branch, B i represents the output of the i-th dilated convolution branch, F fusion represents the fusion operation, which is fused using a 3×3 convolution, and O MFEB represents the final output feature of the MDCB module.
[0109] The super-resolution reconstruction module consists of a convolution with an n×n convolutional kernel and an upsampling block based on orthogonal position encoding (OPE). Among them, OPE uses the two-dimensional coordinates of the query point and the local latent code as inputs, and then directly calculates the pixel value corresponding to the query point with a set of orthogonal bases. This method does not require training parameters but directly performs linear operations to reconstruct a high-resolution image. Therefore, compared with the commonly used implicit neural network representation method, the method based on OPE has higher computational efficiency and consumes less memory.
[0110] Among them, the image joint processing model based on multi-task learning uses the super-resolution reconstruction module to obtain a high-resolution image. The upsampling flow chart of the method based on OPE is as Figure 6 shown, including:
[0111] Extract the features in the input image;
[0112] Reorder the extracted features, and the number of channels of the features is 3(2n + 1) 2 , and each pixel point in the feature corresponds to 3(2n + 1)2 One eigenvalue, divide the eigenvalue into three equal parts and correspond to the features of the RGB three channels respectively as the latent code of the pixel point;
[0113] As Figure 7 shown, select a pixel point (the rectangular area composed of S1, S2, S3, and S4) in the input image, obtain the boundary point coordinates of the pixel point, and define a query point q within the area of the pixel point, obtain the nearest boundary point S1 of the query point q, and use the coordinate difference between the query point and the nearest boundary point as the orthogonal basis P of the query point;
[0114] Perform matrix multiplication on the orthogonal basis and the latent code of the query point, calculate the RGB pixel value of the query point, and based on the RGB pixel value of the query point, render each pixel value of the high-resolution image, thereby obtaining the high-resolution image.
[0115] In this embodiment, the halftone reconstruction module reconstructs the shallow feature image obtained by the shallow feature extraction module and the deep feature image obtained by the deep feature extraction module to obtain a halftone image.
[0116] In S2, design a joint loss function to train the image joint processing model based on multi-task learning. The expression of the joint loss function L is:
[0117] L = L RH + L SR
[0118] where, L RH is the halftone loss, and L SR is the super-resolution loss.
[0119] The super-resolution loss L SR is composed of L1 loss and structural similarity loss, and its expression is:
[0120]
[0121] where, ||·|| represents L1 regularization, I SR represents the super-resolution image reconstructed from the low-resolution image through the image super-resolution model; N represents the number of pairs of high-resolution images and low-resolution images included in each batch during the training of the image super-resolution model; and respectively represent the values of the i-th pixel point in the super-resolution image I SR and the high-resolution image I HR ; SSIM(I HR , I SR ) represents the super-resolution image I SR and the high-resolution image IHR Structural similarity
[0122] Furthermore, the halftone loss L RH has the following expression:
[0123] L RH = ω1L T + ω2L N + ω3L B + ω4L G
[0124] where L T is the tone consistency loss, L N is the blue noise loss, L B is the binary loss, L G is the perceptual consistency loss, and the hyperparameters ω1 = 0.6, ω2 = 0.3, ω3 = 0.1, ω4 = 1 are set according to experience.
[0125] The tone consistency loss L T is as follows:
[0126] L T = E Ic ∈I{||G(O h ) - G(I c )||2}
[0127] where G(·) is a Gaussian filter, and E Ic ∈I{·} represents the average operator over all input images I c in the training set, and O h represents the halftone image;
[0128] The binary loss L B is as follows:
[0129]
[0130] where is the pseudo - halftone image before the binary gate, and C d is a constant matrix of the same size as , and all elements take values d = {0, 1};
[0131] The perceptual consistency loss L G is as follows:
[0132] L G = E Ic ∈I{||F(O h ) - I c ||2}
[0133] where F(·) represents the inverse halftone pattern.
[0134] Specifically, the blue noise loss L N is as follows:
[0135] L N = L D + σL AS
[0136] where L D represents the loss of the low-frequency components in the spectrum, L AS represents the anisotropic loss in the high-frequency region, and σ is the weight factor;
[0137] The loss of the low-frequency components L D is as follows:
[0138] L D = E Ic ∈I{‖‖(DCT(O h ) - DCT(I c )) ⊙ M‖‖1}
[0139] where DCT(·) represents the discrete cosine transform, ⊙ represents the element-wise product, and M is a constant binary mask;
[0140] The anisotropic loss L AS in the high-frequency region is as follows:
[0141]
[0142] where P i (f) and P i (f ρ ) represent the average power and the sampling point power on the i-th frequency band respectively, and θ represents the number of frequency bands.
[0143] In this embodiment, the DIV2K dataset is used to train the model, which includes 800 training sets and 100 validation sets. The bilinear interpolation algorithm is used to generate low-resolution images (LR), and then the low-resolution images are randomly cropped into the size of 48×48 as the input data. To expand the dataset, the method of data augmentation (such as random rotation, translation, and flipping) is used to expand the dataset. To verify the reconstruction performance of the model, the super-resolution test benchmark datasets: Set5, Set14, B100, and Urban100 are used.
[0144] During training, RGB images of size 64×64 in the LR images are used as inputs and the corresponding HR images are used as labels. All images are preprocessed by subtracting the average RGB values of the DIV2K dataset to obtain the image mean. The network is built based on the PyTorch framework, and the model is trained using an NVIDIA GeForce RTX 2080Ti GPU. The ADAM optimizer is used to train the model, with β1 = 0.9, β2 = 0.999, the mini-batch size (Batch Size) set to 16, and the learning rate initialized to 10 -4 , and the learning rate is halved every 100 Epochs. The entire model is trained for 1000 Epochs.
[0145] To verify the impact of different components in the network on the model performance, the present invention conducts ablation experiments on the noise compensation module (NCB), the multi-scale dilated convolution extraction module (MDCB), the dense connection block (DCB), and the activation function. To avoid the loss caused by the reduction of parameters, a 3×3 convolution with 64 filters (denoted as Conv-3-64) is used as the baseline model of CUNet, as Figure 8 shown. On this basis, the NCB is first added to the shallow feature extraction module, then Conv-3-64 is replaced with the DCB, followed by replacing the ReLU activation function with the PReLU, and finally the MDCB is added, which is indicated by the dashed line in the figure. The peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM) are used as the main evaluation metrics for measuring the super-resolution reconstruction performance of the model.
[0146] The results of ×4 times on the Set5 test set are shown in Table 2. It can be seen that the reconstruction effect is poor on the baseline model, and the reconstruction effect of the model has been improved to a certain extent after adding noise, indicating that noise is helpful for improving the performance of the model, and the increase in the number of parameters and the amount of computation is not obvious after introducing noise. Subsequent experiments will further illustrate the impact of noise intensity on the model performance. After introducing the DCB module, the performance of the model has been improved by 0.83 / 0.0127, and this improvement shows the effectiveness of dense connection in making up for the insufficient feature extraction ability of the convolutional network with a small number of parameters. The dense connection block not only improves the performance of the model by promoting the information flow between features, but also helps to reduce the number of parameters and the amount of computation of the model, thereby improving the processing efficiency. In addition, after adding the MDCB module and the PReLU activation function, the performance has been improved by 0.14 / 0.0022 and 0.07 / 0.0002 respectively, indicating that these two components are also important for improving the model performance.
[0147] Table 2
[0148]
[0149] To verify the impact of the cascaded UNet module (CUNet) on the performance of the network model proposed in the present invention, CUNet was replaced with a residual network having the same number of parameters for a comparative experiment. The residual network selected the backbone network in the EDSR model, where the number of channels in the convolutional layer was selected as 60 and the number of residual blocks was 8. The comparison of the number of parameters of the two modules is shown in Table 3.
[0150] Table 3
[0151]
[0152] The two networks were trained separately, and the PSNR / SSIM metrics were calculated on the Set5 test set. The experimental results are shown in Tables 4 - 8. At three magnification scales, the reconstruction effect of CUNet is better than that of the model based on the residual network. This is because the U-shaped network combines multi-layer information for deep information fusion, thus having good feature extraction ability.
[0153] Table 4
[0154]
[0155] To verify the impact of the multi-scale feature extraction module on the model performance, three multi-scale modules, namely MFEF, RFB, and MSRB, were separately selected and added to the CRUN network for a comparative experiment. The PSNR / SSIM metrics were calculated on the Set5 dataset. The experimental results are shown in Table 5. RFB and MDCB adopted the method of dilated convolution to expand the receptive field, and their performances were both improved to a certain extent. The MDCB module proposed in the present invention has relatively small number of parameters and computational complexity compared with the other two modules, and its performance is the best, indicating that the method of multi-branch dilated convolution can better increase the receptive field while retaining fine-grained features, thus enabling the network to have better reconstruction performance.
[0156] Table 5
[0157]
[0158] To verify the impact of the input noise intensity on the model performance, the noise intensities were set to 0.1, 0.2, 0.3, 0.4, 0.5, and 1.0 respectively to train the network, and the impact of the increase in noise intensity on the reconstruction effect was observed. The PSNR / SSIM metrics were calculated on the Set5 dataset. The experimental results are shown in Table 6. It can be seen from the table that when the noise intensity is 0.3, the model performance reaches the best. When the noise intensity increases further, the performance decreases instead. Therefore, an appropriate noise intensity is beneficial for the model to learn the super-resolution mapping well from the influence of a small amount of noise, which helps to improve the generalization ability of the model and can be better applied to the super-resolution tasks in real scenarios.
[0159] Table 6
[0160]
[0161]
[0162] To verify the effects of three upsampling modules, namely transposed convolution, sub-pixel convolution, and orthogonal position encoding, on the model performance, the super-resolution reconstruction module in the model was used to train the network with these three upsampling modules respectively, and the PSNR / SSIM metrics were calculated for 4x super-resolution on the Set5 dataset. The experimental results are shown in Table 7. Orthogonal position encoding performed best in both PSNR and SSIM metrics, indicating that the orthogonal position encoding upsampling technique can more effectively restore the details of high-resolution images. Compared with transposed convolution and sub-pixel convolution, it has significant advantages in maintaining image structure information and details. Orthogonal position encoding can more precisely process the high-frequency information in the image, thereby effectively improving the reconstruction performance of the model without increasing additional computational costs.
[0163] Table 7
[0164]
[0165] Embodiment 2
[0166] Based on the same inventive concept as the method described in Embodiment 1, the present invention also provides an image joint processing system based on multi-task learning, which is used to implement the image joint processing method based on multi-task learning described in Embodiment 1, as Figure 9 shown, and the system specifically includes the following modules:
[0167] A model construction module, used to construct an image joint processing model; wherein, the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module, and a halftone reconstruction module. The output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module, and two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module;
[0168] A model training module, used to design a joint loss function to train the image joint processing model to obtain a trained image joint processing model;
[0169] A result acquisition module, used for the trained image joint processing model to generate a super-resolution image through the super-resolution reconstruction module and generate a halftone image through the halftone reconstruction module.
[0170] An image joint processing system based on multi-task learning proposed in this embodiment is used to implement the aforementioned image joint processing method based on multi-task learning. Therefore, the specific implementation manners in the image joint processing system based on multi-task learning can be seen in the embodiment part of the aforementioned image joint processing method based on multi-task learning. For example, the model construction module 100, the model training module 200, and the result acquisition module 300 are respectively used to correspondingly implement steps S1, S2, and S3 in the image joint processing method based on multi-task learning in Embodiment 1. Therefore, the specific implementation manners can refer to the descriptions of the corresponding individual embodiments. To avoid redundancy, they will not be elaborated here.
[0171] As can be seen from the above implementation manners, the present invention proposes to jointly learn two types of tasks by using the method of multi-task learning, and integrates the two tasks of image super-resolution and image halftoning by using a novel framework. In the training stage, joint modeling is performed by sharing weights, and a joint loss function is designed to optimize the performance of both image super-resolution and image halftoning simultaneously. In the inference stage, although the model is trained as a whole, the two tasks are processed independently. This means that when generating a super-resolution image or a halftone image, only the corresponding reconstruction module needs to be used, rather than the entire model participating in the calculation. Therefore, this method does not increase the additional number of model parameters and the amount of calculation.
[0172] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes Figure 1 a process or processes and / or blocks Figure 1 specified in a block or blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 a process or processes and / or blocks Figure 1 specified in a block or blocks.
[0176] Obviously, the above-described embodiments are merely examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for joint image processing based on multi-task learning, characterized in that: The following steps are involved: S1: constructing an image joint processing model; wherein the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module and a halftone reconstruction module, the output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module, and two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module; S2: Designing a joint loss function to train the image joint processing model to obtain a trained image joint processing model; S3: The trained image joint processing model generates a super-resolution image through the super-resolution reconstruction module, and generates a halftone image through the halftone reconstruction module.
2. The image joint processing method based on multi-task learning according to claim 1, characterized in that: The shallow feature extraction module is composed of a noise compensation module and a convolution layer with a convolution kernel of n×n; The noise compensation module uses two collaborative convolutions as the feature map of the output of the previous network layer and the mapping of Gaussian noise, and then obtains the input of the next network layer through additive changes.
3. The image joint processing method based on multi-task learning according to claim 1, characterized in that: The deep feature extraction module is constructed by sequentially connecting a plurality of cascaded residual blocks, wherein the cascaded residual block includes a CUNet module and a feature fusion module connected to the CUNet module; Among them, the image joint processing model based on multi-task learning uses the deep feature extraction module to obtain a deep feature image, including: inputting a shallow feature image into the cascaded residual block, splicing the identity mapping of the shallow feature image and its output features after processing by the CUNet module, fusing the spliced features through the feature fusion module, obtaining a fused feature image as the input of the next CUNet module, and obtaining a deep feature image after processing by multiple cascaded residual blocks.
4. The image joint processing method based on multi-task learning according to claim 3 is characterized in that: The CUNet module is based on an encoder-decoder architecture and uses M jump connection blocks as basic blocks, of which m jump connection blocks are used as encoders and the remaining jump connection blocks are used as decoders. The output features of the m encoders are processed by the multi-scale dilated convolution extraction module and the output features of the encoder are used as the input of the decoder, and the multi-level features are spliced by the decoder. The multi-scale dilated convolution extraction module uses the output features of the previous layer of network to undergo 1×1 point-by-point convolution processing, and then outputs dilated convolutions with expansion rates of 1, 2, and 3, and then obtains the input of the next layer of network through feature fusion.
5. The image joint processing method based on multi-task learning according to claim 1, characterized in that: The super-resolution reconstruction module is composed of a convolution with a convolution kernel of n×n and an upsampling block based on orthogonal position coding; The super-resolution reconstruction module obtains a high-resolution image using the up-sampling block based on orthogonal position coding, including: Extract features from the input image; The extracted features are reordered, and the number of channels of the features is 3 (2n+1) 2 , each pixel in the feature corresponds to 3(2n+1) 2 The eigenvalues are divided into three equal parts, each corresponding to the features of the three RGB channels as the latent code of the pixel; Select a pixel point in the input image, obtain the boundary point coordinates of the pixel point, define a query point in the area of the pixel point, obtain the nearest boundary point of the query point, and use the coordinate difference between the query point and the nearest boundary point as the orthogonal basis of the query point; The orthogonal basis and the latent code of the point to be queried are subjected to matrix multiplication to calculate the RGB pixel value of the point to be queried, and each pixel value of the high-resolution image is rendered based on the RGB pixel value of the point to be queried, thereby obtaining a high-resolution image.
6. The image joint processing method based on multi-task learning according to claim 1, characterized in that: The halftone reconstruction module reconstructs the shallow feature image obtained by the shallow feature extraction module and the deep feature image obtained by the deep feature extraction module to obtain a halftone image.
7. The image joint processing method based on multi-task learning according to claim 1, characterized in that: The expression of the joint loss function L is: L==L RH +L SR Among them, L RH is the halftone loss, L SR is the excess loss, the excess loss L SR for: Among them, ||·|| represents L1 regularization, I SR represents the super-resolution image reconstructed from the low-resolution image through the image super-resolution model; N represents the number of high-resolution image and low-resolution image pairs contained in each batch during the training of the image super-resolution model; and Represent the super-resolution image I SR and high resolution image I HR The value of the i-th pixel in HR ,I SR ) represents the super-resolution image I SR and high resolution image I HR 's structural similarity.
8. The image joint processing method based on multi-task learning according to claim 7, characterized in that: The halftone loss L RH The expression is: L RH =ω1L T +ω2L N +ω3L B +ω4L G Among them, L T is the hue consistency loss, L N is the blue noise loss, L B is the binary loss, L G For perceptual consistency loss, ω1, ω2, ω3, and ω4 are all hyperparameters; The hue consistency loss L T for: L T =E Ic ∈I{||G(O h )-G(I c )||2} Where G(·) is a Gaussian filter, E Ic ∈I{·} represents all input images I in the training set c The average operator, O h Represents a halftone image; The binary loss L B for: in is the pseudo halftone image before the binary gate, C d is with A constant matrix of the same size, where all elements have the value d = {0, 1}; The perceptual consistency loss L G for: L G =NO c ∈I{||F(O h )-I c ||2} Where F(·) represents the inverse halftone mode.
9. The image joint processing method based on multi-task learning according to claim 8, characterized in that: The blue noise loss L N for: L N =L D +σL AS Among them, L D Indicates the loss of low-frequency components in the spectrum, L AS represents the anisotropic loss in the high-frequency region, and σ is the weight factor; The low frequency component loss L D for: L D =E Ic ∈I{||(DCT(O h )-DCT(I c ))⊙M||1} Where DCT(·) represents discrete cosine transform, ⊙ represents element-wise product, and M is a constant binary mask; The anisotropic loss L in the high frequency region AS for: Among them, P i (f) and P i (f ρ ) represent the average power and sampling point power on the ith frequency band, respectively, and θ represents the number of frequency bands.
10. An image joint processing system based on multi-task learning, characterized in that: The system is used to implement the image joint processing method based on multi-task learning according to any one of claims 1 to 9, and specifically includes the following modules: A model construction module, used for constructing an image joint processing model; wherein the image joint processing model includes a shallow feature extraction module, a deep feature extraction module, a super-resolution reconstruction module and a halftone reconstruction module, the output layer of the shallow feature extraction module is connected to the input layer of the deep feature extraction module, and two branches generated by the output layer of the deep feature extraction module are respectively connected to the super-resolution reconstruction module and the halftone reconstruction module; A model training module, used for designing a joint loss function to train the image joint processing model to obtain a trained image joint processing model; The result acquisition module is used for the trained image joint processing model to generate a super-resolution image through the super-resolution reconstruction module and to generate a halftone image through the halftone reconstruction module.
Citation Information
Patent Citations
Image processing method and device
CN112991165A
Construction method, device and application of image super-resolution reconstruction model
CN116993592A