A Differentiable Halftoning Method and Device Based on Deep Self-Supervised Learning
Through the differentiable halftone method of deep self-supervised learning, the halftone model is optimized using the Gumbel-Softmax module and the blue noise loss function, solving the problems of slow speed and poor effect of the halftone algorithm, generating high-quality halftone images and improving computing efficiency, which is suitable for multi-level halftone processing.
Patent Information
- Application Number
- CN202411533978.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-31
AI Technical Summary
The existing halftone algorithms are slow in processing speed and poor ineffectiveness, especially in deep learning applications, which leads to low training efficiency and insufficient image quality.
The differentiable halftone method based on deep self-supervised learning is adopted to construct a deep differentiable halftone model, and reparameterization is used to combine blue noise loss and region confidence aggregation mechanism to optimize the loss function to generate high-quality halftone images.
It realizes efficient generation of high-quality halftone images rich in detail without label data, solves the problem of gradient non-differentiation, improves computing efficiency, and shows good scalability in multi-stage halftone processing.
Smart Images

Figure CN119540086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of halftone technology, and in particular to a differentiable halftone method and device based on deep self-supervised learning. Background Art
[0002] Halftone, as a technology for converting continuous-tone images into their approximate discrete versions, is widely used in the printing field. Halftone processing creates different gray levels through different arrangements of pixel points and pixel densities. Due to the low-pass filtering characteristics of the human visual system, the human eye will perceive the halftone image as a smoothly transitioning continuous-tone image at a certain viewing distance. Halftone algorithms are classified according to the periodicity of pixel point arrangements and whether they are aggregated or dispersed points. Halftone images generated by non-periodic dispersed-point halftone algorithms usually consist of uniformly distributed and discrete black and white dots, and often have blue noise characteristics, which conform to the human eye perception characteristics. Therefore, a halftone pattern matching the blue noise attribute is the key to generating high-quality halftone images. The implementation of color halftone usually decomposes a color image into multiple color channels, performs halftone processing on the color intensity information of each channel separately, and during printing or display, these halftone-processed channels are superimposed in a specific way, and the interaction between the dots will reproduce the color and brightness of the original image. Therefore, the halftone processing method for grayscale images is the basis for extending to color halftone, and understanding and optimizing grayscale halftone processing is crucial for achieving high-quality color halftone.
[0003] In the research of halftone technology, image quality and processing efficiency are the main research focuses. Common halftone algorithms can be divided into three categories: ordered dithering algorithms, error diffusion algorithms, and search-based algorithms. The ordered dithering algorithm systematically divides a continuous-tone image into uniform small regions and compares them with a pre-designed or generated dither matrix for processing. This method provides an effective way for fast image halftone processing with its high parallel processing ability and low computational complexity; however, this speed advantage often comes at the cost of sacrificing the quality of the halftone image. The error diffusion algorithm has made significant progress in visual quality through fine pixel-level processing and error transfer mechanisms, but the serial nature of the algorithm and potential visual artifact problems have become obstacles to further optimization. The search-based method regards halftone processing as a complex optimization problem. This type of method first defines a halftone quality evaluation index that comprehensively considers the characteristics of the human visual system (HVS), and then uses heuristic algorithms, such as simulated annealing algorithms or direct binary search algorithms, to precisely optimize these evaluation indexes; since this method directly adjusts specially designed indexes (such as the mean square error at the pixel level), it can generate halftone images with better quality, but its high computational cost has become the main obstacle restricting the wide application of this technology.
[0004] Deep learning models, especially convolutional neural networks, have attracted much attention due to their outstanding performance in image recognition, generation, and conversion tasks. Supervised learning, unsupervised learning, and reinforcement learning are three common paradigms in deep learning. In halftone processing, convolutional neural networks learn the mapping relationship between paired image data in order to achieve efficient image dithering through a single forward propagation. However, this learning process faces two major challenges:
[0005] (1) Preparing “real” halftone images as training labels is costly, and searching for the best halftone images requires huge computing resources. Existing search-based methods can either only optimize specific metrics or rely on meta-heuristic search algorithms that require parameter tuning for each instance.
[0006] (2) Halftone processing is essentially a one-to-many mapping problem, that is, the same continuous image may correspond to multiple valid halftone representations. If the pixel-level loss function (such as cross entropy) is directly used for optimization, the model may only learn the average representation of the input image, which cannot meet the discrete requirements of halftone processing.
[0007] Therefore, how to achieve halftoning by relying on self-supervision without a large-scale labeled dataset is an important part of halftoning research.
[0008] Another core problem in the application of deep learning to halftone processing stems from the strong binary characteristics of halftone images, which directly leads to the non-differentiability of discrete choices (such as deterministic selection of the best pixel state) during training, thus hindering the back propagation of gradients. Existing research uses generative adversarial networks to train from pre-prepared halftone label datasets, aiming to minimize the Pearson error between the output and the label. Divergence, but because it requires the construction of a large number of labeled data sets, its practicality is limited to a certain extent. On the other hand, the research mainly solves the problem of gradient backpropagation, and one strategy is the gradient pass-through estimator. This method performs actual discrete operations in the forward propagation, and in the backpropagation process, the gradient is directly passed to the input of the discrete variable, ignoring the discontinuity caused by discretization. This method weakens the characteristics of binarization, so it is necessary to introduce a discretization penalty loss to push the output value to the nearest discrete end. However, this greedy binarization rule may damage the optimization from a global perspective. Although the gradient pass-through estimator is simple and easy to use, the inaccuracy of its gradient estimation may affect the optimization quality of the model. The existing technology also uses policy gradient-based methods such as REINFORCE. Since halftone image processing requires sampling of all pixels, its gradient estimation is often accompanied by a high variance through accumulation, resulting in low training efficiency and difficulty in model convergence.
[0009] In summary, the existing halftone algorithms have problems of slow processing speed and poor halftone effect. Summary of the Invention
[0010] Therefore, the technical problem to be solved by the present invention is to overcome the problems of slow processing speed and poor halftone effect in the existing halftone algorithms.
[0011] To solve the above technical problems, the present invention provides a differentiable halftone method based on deep self-supervised learning, including:
[0012] Constructing a deep differentiable halftone model, including a residual network and a Gumbel-Softmax module;
[0013] When training the deep differentiable halftone model, input a continuous-tone grayscale image and a constant grayscale image into the deep differentiable halftone model. After introducing random Gaussian noise into the continuous-tone grayscale image and the constant grayscale image respectively, use the residual network to extract the feature map, and use the Gumbel-Softmax module to reparameterize the feature map output by the residual network, and output halftone images corresponding to the continuous-tone grayscale image and the constant grayscale image respectively; use the loss function to calculate the distance between the continuous-tone grayscale image, the constant grayscale image and the halftone image, and iteratively optimize the parameters of the residual network;
[0014] After the deep differentiable halftone model is trained, replace the Gumbel-Softmax module in it with the argmax function to obtain the target deep differentiable halftone model;
[0015] Use the target deep differentiable halftone model to perform halftoning on the continuous-tone image.
[0016] Preferably, the reparameterization of the feature map output by the residual network using Gumbel-Softmax includes:
[0017] Given a discrete random variable, assume it has K categories, and generate an independent Gumbel noise for each category;
[0018] Use the Softmax function to reparameterize the output feature map of the residual network for discretization and output the halftone image.
[0019] Preferably, an independent Gumbel noise is generated for each pixel in the output feature map under each category, and the formula is:
[0020] G k =-log(-log(U k ))
[0021] where U kRepresents the probability distribution of the k-th category and follows a uniform distribution U~Uniform(0,1); G k Represents the Gumbel noise of the k-th category, k = 1, 2, …, K;
[0022] The Softmax function is used to reparameterize the output feature map of the residual network for discretization, and the formula is:
[0023]
[0024] where, x k is the output value of the k-th category, τ is the temperature parameter, π k and π l represent the probabilities of the k-th category and the l-th category respectively, and G l represents the Gumbel noise of the l-th category;
[0025] The maximum output value among all categories is used as the halftone output value of the current pixel.
[0026] Preferably, the loss function of the deep differentiable halftone model includes blue noise loss, mean square error loss, and structural similarity index loss.
[0027] Preferably, the formula for the blue noise loss is:
[0028] L N = MSE(V c (f ρ ), V h (f ρ ))
[0029] where, L N represents the blue noise loss, V c (f ρ ) represents the ideal blue noise spectrum variance, V h (f ρ ) represents the blue noise spectrum variance of the halftone image, and MSE represents the mean square error;
[0030] The formula for the blue noise spectrum variance of the halftone image is:
[0031]
[0032] where, V h (f ρ ) represents the blue noise spectrum variance of the halftone image, is the power spectrum of the halftone image, P(f ρ ) represents the radially averaged power spectral density of the halftone image, f represents the frequency, f ρ is the radial frequency with a width of ρ, r(fρ ) is f ρ A circular ring with a surrounding width of Δρ = 1, n(r(f ρ )) is the number of discrete frequency samples in r(f ρ );
[0033] The calculation formula for the radial average power spectral density of the halftone image is:
[0034]
[0035] The calculation formula for the power spectrum of the halftone image is:
[0036]
[0037] Among them, h represents the halftone image output by the Gumbel-Softmax module, N represents the total number of pixels in the image, and DFT represents the Fourier transform.
[0038] Preferably, the formula for the mean square error loss is:
[0039] : C = HVS(HVS(h), HVS(c))
[0040] Among them, HVS represents the human visual system, h represents the halftone image, and c represents the continuous tone grayscale image.
[0041] Preferably, optimizing the mean square error loss includes:
[0042] Introducing a regional confidence aggregation mechanism when calculating the mean square error loss, constructing a confidence feature map for pixel-by-pixel matching based on the mean square error of each pixel in the halftone image and the continuous tone grayscale image, then extracting a sampled sub-region from the confidence feature map, and calculating the local descriptor of each pixel;
[0043] Calculating the mean square error between the halftone image and the continuous tone grayscale image according to the local descriptor of each pixel to obtain the mean square error loss.
[0044] Preferably, the formula for the local descriptor is expressed as:
[0045]
[0046] Among them, D i,j represents the local descriptor of the pixel with coordinates (i, j), N(i, j) represents the sampled sub-region centered on the pixel with coordinates (i, j), (u, v) represents the pixel in the u-th row and v-th column of the sampled sub-region, w u,v represents the convolution filter weight of the pixel in the u-th row and v-th column of the sampled sub-region, F i+u,j+vRepresents the mean square error of the pixel with coordinates (i+u, j+v).
[0047] Preferably, the sampled sub-region of the regional confidence aggregation mechanism is consistent with the filter window of the human visual system.
[0048] The present invention also provides a differentiable halftoning device based on deep self-supervised learning, comprising:
[0049] A model construction module for constructing a deep differentiable halftoning model, including a residual network and a Gumbel-Softmax module;
[0050] A training module for training the deep differentiable halftoning model, inputting a continuous-tone grayscale image and a constant grayscale image into the deep differentiable halftoning model. After introducing random Gaussian noise into the continuous-tone grayscale image and the constant grayscale image respectively, using the residual network to extract the feature map, using the Gumbel-Softmax module to reparameterize the feature map output by the residual network, and outputting halftone images corresponding to the continuous-tone grayscale image and the constant grayscale image respectively; calculating the distance between the continuous-tone grayscale image, the constant grayscale image and the halftone image using a loss function, and iteratively optimizing the parameters of the residual network;
[0051] A halftoning module for replacing the Gumbel-Softmax module with an argmax function after the deep differentiable halftoning model is trained to obtain a target deep differentiable halftoning model; using the target deep differentiable halftoning model to perform halftoning on a continuous-tone image.
[0052] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:
[0053] In the differentiable halftoning method based on deep self-supervised learning of the present invention, a Gumbel-Softmax module is introduced during the training process of the deep differentiable halftoning model. Using the reparameterization strategy of Gumbel-Softmax, the sampling process of the discrete distribution is simulated by a smooth and differentiable approximation method, solving the non-differentiable problem caused by halftone discrete selection, realizing effective unbiased gradient estimation during the training process, getting rid of the dependence on labeled data, directly optimizing the halftone evaluation metric. The trained target deep differentiable halftoning model can not only generate high-quality halftone images rich in details, but also has significant advantages in terms of operation efficiency and has good scalability.
[0054] Furthermore, the present invention optimizes the blue noise loss in the loss function of the deep differentiable halftone model. By quantifying and comparing the mean square error of the blue noise spectrum variance between the halftone image output by the deep differentiable halftone model and the ideal blue noise spectrum variance, the anisotropy phenomenon of constant grayscale images is significantly suppressed during the training phase, thereby promoting the natural generation of blue noise characteristics. Moreover, the present invention introduces a regional confidence aggregation mechanism into the mean square error loss. By comprehensively considering the mutual dependence between pixels, the perceptual characteristics of the human visual system are effectively simulated, thereby maintaining the local structural information and texture details of the image, enhancing the information extraction of the model for local region details and global structures, and improving the quality of the halftone image. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To make the content of the present invention more clearly understood, the following further details the present invention based on the specific embodiments of the present invention in conjunction with the drawings, where:
[0056] Figure 1 is the structural diagram of the deep differentiable halftone model of the present invention;
[0057] Figure 2 is the halftone effect of each method on the cochlear organ image, where Figure 2 in (a) is the continuous-tone grayscale image of the cochlear organ, Figure 2 in (b) is the halftone image output by the VAC method, Figure 2 in (c) is the halftone image output by the OVED method, Figure 2 in (d) is the halftone image output by the DBS method, Figure 2 in (e) is the halftone image output by the RVH method, Figure 2 in (f) is the halftone image output by the TRH method, Figure 2 in (g) is the halftone image output by the method of the present invention without using the regional confidence aggregation mechanism, Figure 2 in (h) is the halftone image output by the method of the present invention;
[0058] Figure 3 is the halftone effect of each method on the real butterfly image, where Figure 3 in (a) is the continuous-tone grayscale image of the real butterfly image, Figure 3 in (b) is the halftone image output by the VAC method, Figure 3 in (c) is the halftone image output by the OVED method, Figure 3 in (d) is the halftone image output by the DBS method, Figure 3 in (e) is the halftone image output by the RVH method, Figure 3 in (f) is the halftone image output by the TRH method, Figure 3Among them, (g) is the halftone image output without using the regional confidence aggregation mechanism in the method of the present invention, Figure 3 Among them, (h) is the halftone image output by the method of the present invention;
[0059] Figure 4 It is a comparison chart of the halftone results of different constant grayscale images, as well as the radial average power spectral density and anisotropy measure, where Figure 4 In column (a) among them, it is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant grayscale level of 20, Figure 4 In column (b) among them, it is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant grayscale level of 80, Figure 4 In column (c) among them, it is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant grayscale level of 100, Figure 4 In column (d) among them, it is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant grayscale level of 127;
[0060] Figure 5 It is the Fourier amplitude spectrum of the image with a constant grayscale level of 127 without using blue noise loss and with using blue noise loss, where Figure 5 In (a) among them, it is the Fourier amplitude spectrum without using blue noise loss, Figure 5 In (b) among them, it is the Fourier amplitude spectrum with using blue noise loss;
[0061] Figure 6 It is the MSE loss curve chart of various deep learning methods;
[0062] Figure 7 It is the optimization result chart under the weight parameter values of different structural similarity index losses;
[0063] Figure 8 It is the result chart of two five-level halftone instances, where Figure 8 In column (a) among them, it is the continuous-tone grayscale image, Figure 8 In column (b) among them, it is the five-level halftone result chart, Figure 8 In column (c) among them, it is the output line chart. Detailed implementation manners
[0064] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited do not limit the present invention.
[0065] Embodiment 1
[0066] Refer to Figure 1As shown in the figure, the present invention proposes a differentiable halftoning method based on deep self-supervised learning, including:
[0067] Construct a deep differentiable halftoning model, including a residual network and a Gumbel-Softmax module;
[0068] When training the deep differentiable halftoning model, input the continuous-tone grayscale image and the constant grayscale image into the deep differentiable halftoning model. After introducing random Gaussian noise into the continuous-tone grayscale image and the constant grayscale image respectively, use the residual network to extract the feature map, and use the Gumbel-Softmax module to reparameterize the feature map output by the residual network, and output the halftone images corresponding to the continuous-tone grayscale image and the constant grayscale image respectively; use the loss function to calculate the distance between the continuous-tone grayscale image, the constant grayscale image and the halftone image, and iteratively optimize the parameters of the residual network;
[0069] After the deep differentiable halftoning model is trained, replace the Gumbel-Softmax module in it with the argmax function to obtain the target deep differentiable halftoning model;
[0070] Use the target deep differentiable halftoning model to perform halftoning on the continuous-tone image.
[0071] The training process of the deep differentiable halftoning model is introduced in detail below.
[0072] Since the convolutional neural network has a convolutional paradigm with spatially shared kernels, the inductive bias it brings is prone to the phenomenon of flatness degradation. Specifically, the convolution of the constant signal s(x)≡b with any kernel function k(x) is still a constant signal where μ(k(x)) represents the mean of the kernel function.
[0073] Therefore, given a flat input X, regardless of the parameters of the convolutional neural network, the operations of the convolutional neural network will degenerate into a scaling operation Y = αX. Therefore, on the premise of not destroying the integrity of the original input information, the present invention uses a Gaussian noise map as a spatial variation and introduces it into the feature space to provide sufficient jitter dependence for the model.
[0074] The input of the deep differentiable halftoning model is defined by the following formula:
[0075] I c = Concatenate(c,z)
[0076]
[0077] where c is the continuous-tone grayscale image, c gis a constant grayscale image, z is a Gaussian noise image; Concatenate is a splicing operation; I c is a noise continuous-tone grayscale image, is a noise constant grayscale image. The random noise introduced by the Gaussian noise image helps the depth differentiable halftone model focus on the statistical distribution characteristics of the overall pattern rather than a single pixel value.
[0078] Input the noise continuous-tone grayscale image and the noise constant grayscale image into the residual network, and use the residual network to extract the feature map. The residual network contains multiple residual blocks, and each residual block processes the input features and adds the processed feature map to the original input. These residual blocks help to effectively transmit information in the deep network, thus avoiding the problem of gradient disappearance and enabling the network to maintain the stability of performance when learning deep features.
[0079] Use the Gumbel-Softmax module to reparameterize the feature map output by the residual network to achieve an unbiased estimate of the gradient, and output the halftone images h corresponding to the continuous-tone grayscale image and the constant grayscale image respectively.
[0080] The Gumbel-Softmax module aims to solve the problem that the reparameterization technique cannot be directly applied to discrete data. This strategy is based on two insights: using the Gumbel distribution can parameterize the discrete distribution; the argmax function itself is not continuous, and by using the Softmax function controlled by the temperature parameter, a continuous and differentiable approximation can be provided for the discrete selection process.
[0081] Specifically, in order to discretize the output feature map of the residual network, a discrete random variable X needs to be given. Suppose it has K categories, and the probabilities of these K categories are π1, π2,..., π K , and the unnormalized log probability of each category is expressed as logπ k .
[0082] Each pixel of the output feature map generates an independent Gumbel noise for each category, and the formula is:
[0083] G k =-log(-log(U k ))
[0084] where, U k represents the probability distribution of the kth category and follows the uniform distribution U~Uniform(0,1); G k represents the Gumbel noise of the kth category, k = 1, 2,..., K.
[0085] The traditional halftone method adds Gumbel noise to the unnormalized logarithmic probabilities of each category, and then represents the discrete random variable X through the argmax function. The formula is as follows:
[0086]
[0087] In the processing of discrete variables by the traditional method, the discrete result based on the argmax function is discontinuous. In order to facilitate gradient calculation during the training process, the present invention adopts the Softmax function based on the temperature parameter to reparameterize the output feature map of the residual network, perform discretization, and output a halftone image. The formula is as follows:
[0088]
[0089] where x k is the output value of the k-th category, τ is the temperature parameter, π k and π l represent the probabilities of the k-th category and the l-th category respectively, and G l represents the Gumbel noise of the l-th category.
[0090] The maximum output value among all categories is used as the halftone output value of the current pixel.
[0091] The present invention transfers the dependence of the probabilities π1, π2,..., π K from the non-differentiable random sampling function to the differentiable function composed of Softmax and log operations, and its internal mechanism cleverly continues the noise benefit introduced in the initial Concatenate operation, that is, by introducing noise based on the Gumbel distribution, not only effectively simulates and extends the randomness influence of the initial noise, but also further enhances the model's ability to process discrete outputs.
[0092] Due to the discreteness of halftone, the image filtered by the human visual system (HVS) can effectively simulate the perception characteristics of the human eye. On this basis, the present invention further conducts end-to-end optimization.
[0093] In order to comprehensively evaluate and optimize the quality of the halftone image generated by the model, the loss function of the deep differentiable halftone model includes blue noise loss, mean square error loss, and structural similarity index loss. The formula is as follows:
[0094] L total = L C + w1·L N + w2·L S
[0095] where L total is the loss function of the deep differentiable halftone model, LC is the mean squared error loss, L N is the blue noise loss, w1 is the weight parameter of the blue noise loss, L S is the structural similarity index loss, w2 is the weight parameter of the structural similarity index loss.
[0096] In this embodiment, w1 = 0.02 and w2 = 0.01, which are empirical setting values.
[0097] The loss function of the deep differentiable halftoning model is introduced in detail below.
[0098] In halftoning, the blue noise texture pattern has a non-periodic characteristic while avoiding low-frequency graininess, which conforms to human eye perception. The halftoning that conforms to blue noise should maintain the following spectral characteristics: (1) containing fewer low-frequency components; (2) the energy distribution in the high-frequency region is relatively flat; (3) the anisotropy is very low at all frequencies. Previous studies have shown that when processing constant grayscale images, simply relying on reducing the mean squared error loss and structural similarity loss of the image filtered by the human visual system is not sufficient to ensure good blue noise quality. In addition, the parallel processing characteristics of convolutional neural networks are prone to causing global consistency checkerboard artifacts, which further lead to a decline in the quality of halftone images. In existing studies, the low-frequency components of halftone images are penalized by discrete cosine transform. Although these components are minimized to a certain extent, the problem of excessive anisotropy is not significantly solved. This excessive anisotropy may lead to an undesirable texture bias visually.
[0099] In order to optimize the blue noise characteristics of the generated halftone image and evaluate its quality, and to avoid the model tending to extreme outputs, the present invention introduces a metric method based on power spectrum variance during the model training process to construct the blue noise loss.
[0100] First, calculate the power spectrum of the halftone image output by the Gumbel-Softmax module through discrete Fourier transform (DFT). The formula is:
[0101]
[0102] where h represents the halftone image output by the Gumbel-Softmax module, N represents the total number of pixels in the image, DFT represents the Fourier transform, is the power spectrum of the halftone image.
[0103] Calculate the radially averaged power spectral density of the halftone image. The formula is:
[0104]
[0105] where P(f ρ) represents the radial average power spectral density of a halftone image, f ρ is the radial frequency with width ρ, r(f ρ ) is f ρ the ring with width Δρ = 1 around it, n(r(f ρ )) is the number of discrete frequency samples in r(f ρ ), and f represents frequency.
[0106] Calculate the blue noise spectrum variance, and the formula is:
[0107]
[0108] where V h (f ρ ) represents the blue noise spectrum variance of the halftone image.
[0109] Since variance is used to measure the degree of data dispersion and can explicitly characterize the distribution characteristics and fluctuation degree of halftone dots, the present invention designs a blue noise loss, and the formula is:
[0110] L N = MSE(V c (f ρ ), V h (f ρ ))
[0111] where V c (f ρ ) represents the ideal blue noise spectrum variance, V h (f ρ ) represents the blue noise spectrum variance of the halftone image, and MSE represents the mean square error.
[0112] The blue noise loss can accurately evaluate the effect of the model in maintaining the blue noise attributes of the image by quantifying the mean square error of the blue noise spectrum variance of the halftone image output by the depth differentiable halftone model compared with the ideal blue noise spectrum variance. Since spectral analysis is only meaningful for halftone images with constant grayscale, the present invention optimizes this loss function on additional small batches of constant grayscale images.
[0113] The mean square error loss is used to calculate the mean square error between the halftone image filtered based on the human visual system and the continuous tone grayscale image, and the formula is
[0114] L C = MSE(HVS(h), HVS(c))
[0115] where HVS represents the human visual system, h represents the halftone image, and c represents the continuous tone grayscale image. The filter window size of the human visual system is 11×11.
[0116] The structural similarity index loss is used to evaluate the similarity between a halftone image and a continuous-tone grayscale image, and the formula is:
[0117] L S = 1 - SSIM(h, c)
[0118] where SSIM represents the structural similarity index, h represents the halftone image, and c represents the continuous-tone grayscale image.
[0119] During the training process, the loss function L total is used as the final optimization target of the model to guide the model to optimize the visual similarity, structural similarity, and blue noise characteristics of the image simultaneously during the learning process. In this way, not only can a halftone image visually similar to the original image be generated, but also its ideal blue noise characteristics can be ensured, thus achieving an optimal balance in visual effects.
[0120] In the evaluation of halftone image quality, existing metrics such as mean squared error and structural similarity index can quantify the differences between images, but they are often limited by the one-sided pursuit of pixel-level precision and ignore the impact of the interaction between pixels on the overall perceived quality. These metrics preprocess the target image and the reference image through a human visual system model to simulate the sensitivity of the human eye to image details, and then generate an error map reflecting the image differences and calculate the average to obtain a scalar metric result. The performance of a pixel depends not only on its own value but also on the influence of other pixels within its local window, and the scope of this effect is determined by the window size defined by the HVS filter. However, by directly optimizing the average precision through Generalized Mean Pooling and simplifying the gradient of each network parameter to the average of the gradients at all pixels, this processing method treats the confidence matching costs of all involved pixel pairs equally and fails to fully consider the key impact of the pixel arrangement pattern in the halftone image on the quality of the generated image.
[0121] To address the above limitations, the present invention introduces a regional confidence aggregation mechanism when calculating the mean squared error loss, which is used to abandon the practice of examining pixels in isolation and instead adopt a more global perspective. By aggregating the confidence of local pixels to a unified clustering center for processing, it aims to more comprehensively optimize the texture and structural details of the image by considering the quality evaluation of the local area where the pixels are located.
[0122] Specifically, the regional confidence aggregation mechanism constructs a confidence feature map for pixel-by-pixel matching based on the mean squared error of each pixel in the halftone image and the continuous-tone grayscale image. Then, it extracts the sampled sub-regions from the confidence feature map and calculates the local descriptors for each pixel. Each descriptor not only contains the information of the current pixel but also integrates the confidence values of all pixels within its sampled sub-region. At the same time, the regional confidence aggregation mechanism introduces learnable convolutional filter weights to adjust the contribution degrees of different pixels during the aggregation process, ensuring that the model can adaptively capture the local features of the image.
[0123] The formula representation of the local descriptor is as follows:
[0124]
[0125] where D i,j represents the local descriptor of the pixel with coordinates (i, j), N(i, j) represents the sampled sub-region centered on the pixel with coordinates (i, j), (u, v) represents the pixel at the u-th row and v-th column in the sampled sub-region, and w u,v represents the convolutional filter weight of the pixel at the u-th row and v-th column in the sampled sub-region, and F i+u,j+v represents the mean squared error of the pixel with coordinates (i + u, j + v).
[0126] Preferably, the sampled sub-region of the regional confidence aggregation mechanism is consistent with the HVS filter window, enabling the model to be within the same perceptual range as the HVS, thus more effectively aggregating the confidence information in the neighborhood. In this embodiment, the size of the sampled sub-region is 11x11.
[0127] Meanwhile, the convolutional filter weights can be regarded as part of the convolutional layer in the model and can be learned together with the network parameters. To achieve the convergence goal faster, the initial convolutional filter weight values are set to a Gaussian distribution with a standard deviation of 2. Under this mechanism, the decision output of the model is not only based on the current pixel point but also better captures the context information of the image by marking the correlation between pixels, thereby enhancing the spatial perception ability of the model.
[0128] The mean squared error loss calculates the mean squared error between the halftone image and the continuous-tone grayscale image based on the local descriptors of each pixel calculated by the regional confidence aggregation mechanism.
[0129] In summary, in the training process of the deep differentiable halftone model, the present invention introduces the Gumbel-Softmax module. By using the reparameterization strategy of Gumbel-Softmax, the sampling process of discrete distribution is simulated through a smooth and differentiable approximation method, solving the non-differentiable problem caused by halftone discrete selection. Effective unbiased gradient estimation is achieved during the training process, getting rid of the dependence on labeled data, and directly optimizing for the halftone evaluation metric. The trained deep differentiable halftone model can not only generate high-quality halftone images rich in details, but also has significant advantages in operation efficiency, can be applied to scenarios such as multi-level halftone, has good scalability, and provides an effective solution for image halftoning.
[0130] Furthermore, the present invention optimizes the blue noise loss in the loss function of the deep differentiable halftone model. By quantifying and comparing the mean square error of the blue noise spectral variance of the halftone image output by the deep differentiable halftone model and the ideal blue noise spectral variance, the anisotropy phenomenon of the constant gray-level image is significantly suppressed during the training stage, thus promoting the natural generation of blue noise characteristics. And, the present invention introduces a regional confidence aggregation mechanism into the mean square error loss. By comprehensively considering the mutual dependence between pixels, the perceptual characteristics of the human visual system are effectively simulated, thus maintaining the local structure information and texture details of the image, enhancing the information extraction of the model for local region details and global structure, and improving the quality of the halftone image.
[0131] Embodiment 2
[0132] To verify the effectiveness of the method proposed by the present invention, in this embodiment, relevant ablation experiments are carried out by comparing the performance differences between the method of the present invention and existing halftone methods, and its scalability in different application scenarios is explored.
[0133] 1. Experimental Setup
[0134] In this embodiment, the VOC2012 dataset is used for training and evaluation. Randomly select 13,758 images from it as the training set, and retain 1,684 and 1,683 images as the validation set and the test set for model performance evaluation. The present invention is based on self-supervised learning. By designing a specific loss function, the model does not need to rely on labeled data during training, so that a large number of unlabeled gray-level images can be used for learning. To ensure the consistency and comparability of the experiments, all images are converted to gray-level images before processing, and the proposed method is comprehensively compared with other existing methods. To further prove the effectiveness and generalization ability of the model, tests are carried out on multiple public datasets, including Set5, Set14, B100, B200, General100, Manga109, Urban100, DIV2K datasets.
[0135] The present invention adopts a residual network as the backbone network for halftoning. During training, the batch size is set to 64, and the training images are randomly cropped to 64x64. During the training process, the network is trained by minimizing the loss function L total with hyperparameters w1 = 0.02, w2 = 0.01, and the initial temperature coefficient τ of reparameterization set to 0.5. The entire model is trained using the ADAM optimizer, and the learning rate α is adjusted from 3e-4 to 1e-5 through a cosine annealing schedule. The entire training process is carried out for 400 epochs on an NVIDIA RTX 2080Ti GPU.
[0136] 2. Halftone Quality Evaluation
[0137] To comprehensively evaluate the performance of the proposed method in the halftone image generation task, this embodiment conducts a systematic quantitative evaluation and visual quality analysis, comparing multiple existing halftone methods, including the Void-and-cluster method (VAC), Ostromoukhov's method (OVED), the optimization-based search method (DBS), and the deep neural network-based methods (RVH and TRH), where the present invention / base representation refers to the method of the present invention without using the region confidence aggregation mechanism.
[0138] Table 1 details the quantitative results of all methods on the test dataset. The tonal consistency is measured by the peak signal-to-noise ratio (PSNR) between the halftone filtered by HVS and the input continuous tone, and the structural and texture information of the halftone image is evaluated by the structural similarity index (SSIM).
[0139] Figure 2 shows the halftone effects of each method on the cochlear organ image, where Figure 2 (a) in is the continuous-tone grayscale image of the cochlear organ, Figure 2 (b) in is the halftone image output by the VAC method, Figure 2 (c) in is the halftone image output by the OVED method, Figure 2 (d) in is the halftone image output by the DBS method, Figure 2 (e) in is the halftone image output by the RVH method, Figure 2 (f) in is the halftone image output by the TRH method, Figure 2 (g) in is the halftone image output by the method of the present invention without using the region confidence aggregation mechanism, Figure 2 (h) in is the halftone image output by the method of the present invention.
[0140] Table 1 Quantitative Comparison of Halftone Methods
[0141]
[0142] The experimental results show that the method proposed in the present invention has achieved competitive scores in terms of PSNR and obtained the best performance in terms of SSIM. It is worth noting that although the DBS method obtains the highest PSNR value by optimizing the search strategy, an extremely high PSNR does not represent a visually better halftone effect, and the extreme pursuit of the PSNR index will sacrifice the image structure details. Figure 2 In (d) of [reference], it shows that the result generated by the DBS method loses local details visually. The method of the present invention pays more attention to the balance between the image structure texture and the tone consistency.
[0143] Table 1 also compares the number of parameters of different methods and the running time on the 512x512 pixel "Lenna" test image. The results show that although the VAC method has the shortest running time, the quality of the halftone image it generates is relatively low. On the other hand, the search-based method (DBS) has a long running time due to computationally intensive operations, which limits its feasibility in practical applications. In contrast, the method based on deep neural network can generate halftone images with richer details and higher quality. At the same time, the method of the present invention is superior to the RVH and TRH methods in reducing the number of parameters and shortening the running time, showing higher computational efficiency and lower time overhead. Generally speaking, the method proposed in the present invention is in the leading position in terms of image quality evaluation and computational cost.
[0144] From Figure 2 it can be seen that VAC, OVED, and DBS fail to effectively retain the vein structure details of the image because they ignore the structural similarity of the image when generating halftone images, resulting in poor performance of these methods in dealing with edge details, making the halftone pictures overly blurred and losing details. Although these methods can suppress artifacts to a certain extent, this usually comes at the cost of sacrificing fine texture and edge information. In contrast, the method of the present invention effectively balances the relationship between artifact suppression and detail retention, and more accurately reflects the structural features of the original image. It not only performs well in suppressing artifacts, but also avoids over-blurring and successfully retains the key features of the image, which makes the method of the present invention significantly superior to other methods in the generation quality of halftone images.
[0145] Figure 3 is the halftone effect of each method on the real butterfly image, where Figure 3 in (a) is the continuous-tone grayscale image of the real butterfly image, Figure 3 in (b) is the halftone image output by the VAC method, Figure 3 in (c) is the halftone image output by the OVED method, Figure 3Among them, (d) is the halftone image output by the DBS method, Figure 3 Among them, (e) is the halftone image output by the RVH method, Figure 3 Among them, (f) is the halftone image output by the TRH method, Figure 3 Among them, (g) is the halftone image output by the method of the present invention without using the regional confidence aggregation mechanism, Figure 3 Among them, (h) is the halftone image output by the method of the present invention. It can be observed that obvious checkerboard grid artifacts appear in the halftone images generated by the OVED, RVH, and TRH methods, and the OVED has relatively serious oblique texture features. Although this pattern presents relatively low-frequency features visually, it does not fully meet the ideal characteristics of blue noise. The RVH and TRH methods have the phenomenon of point aggregation, that is, pixel points tend to aggregate in certain areas, and such point distribution does not belong to the ideal random uniform distribution, which further affects the overall visual effect of the halftone image. While optimizing the low-frequency features, these methods fail to effectively handle the blue noise distribution and point dispersion problems. In contrast, the halftone image generated by the method of the present invention not only effectively avoids the generation of checkerboard grid patterns but also significantly reduces the point aggregation phenomenon, and can generate a halftone image with more natural texture.
[0146] To verify the effectiveness of the proposed method, in this embodiment, comparative experiments on the PSNR and SSIM metrics of different methods are carried out on multiple public test sets, and the quantitative index comparison is shown in Table 2.
[0147] Table 2 Test results of public data sets (SSIM / PSNR)
[0148]
[0149]
[0150] As can be seen from Table 2, the method of the present invention shows the best SSIM metric and better PSNR metric on each test set, verifying that the method of the present invention has obvious advantages in maintaining image details and structural integrity, and also indicating its good generalization.
[0151] 3. Ablation experiment
[0152] To achieve differentiability of the network framework during the binarization process, the RVH and TRH methods use a straight-through estimator to enable gradient backpropagation during the binarization process and propose a binarization loss. Although this method promotes discrete selection to some extent, the existence of the binarization loss limits the optimization of the overall image quality. In contrast, the present invention adopts a strategy based on Gumbel-Softmax reparameterization, which allows the model to explore the continuous decision space during training while retaining the discrete selection characteristics. It not only realizes the differentiability of the discrete selection process but also enhances the training stability and efficiency of the model, thus significantly improving the optimization effect of image quality evaluation.
[0153] During the halftoning process, the blue noise characteristic is a key factor affecting the quality of halftone pictures. Figure 4 is a comparison chart of the halftone results of different constant gray-level images, as well as the radial average power spectral density and anisotropy metric, where Figure 4 the (a) column in is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant gray level of 20, Figure 4 the (b) column in is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant gray level of 80, Figure 4 the (c) column in is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant gray level of 100, Figure 4 the (d) column in is the halftone result chart, radial average power spectral density chart, and anisotropy degree chart of the image with a constant gray level of 127. In the halftone result chart of the image, the left, middle, and right are the halftone effects generated by disabling the blue noise loss, using the blue noise loss, and the DBS method, respectively. It can be clearly seen that the halftone image without the blue noise loss has relatively prominent stripe artifacts, while using this loss can significantly solve the stripe artifact problem. At the same time, the optimized halftone image also shows ideal blue noise characteristics in the radial frequency power spectrum, that is, fewer low-frequency components and higher and smoothly transitioning mid- and high-frequency components. In addition, the anisotropy curve of the halftone spectrum with blue noise characteristics is also relatively smooth, avoiding sharp frequency peaks or abrupt changes. This smooth frequency response not only improves the visual quality of the image but also effectively suppresses possible visual artifacts.
[0154] Figure 5 is the Fourier amplitude spectrum of the image with a constant gray level of 127 without using the blue noise loss and with using the blue noise loss, where Figure 5 the (a) in is the Fourier amplitude spectrum without using the blue noise loss, Figure 5 the (b) in is the Fourier amplitude spectrum with using the blue noise loss. From Figure 5As can be clearly observed in (a) of [reference], when the blue noise loss is not used, the halftone dot distribution has an obvious direction preference. After using the blue noise loss, this preference is significantly reduced, showing more ideal blue noise characteristics.
[0155] By introducing a regional confidence aggregation mechanism, the method of the present invention can more comprehensively capture and utilize the confidence information within the local region. The SSIM and PSNR scores of this mechanism on the test dataset are shown to be improved in Table 1.
[0156] Figure 6 It is the MSE loss curve graph of various deep learning methods, further indicating that by comprehensively considering the confidence information between pixels, the method of the present invention can converge more quickly, thereby improving the training efficiency.
[0157] In addition, to further analyze the impact of model hyperparameters on performance, this embodiment conducts a sensitivity analysis on the hyperparameter w2 in the loss function. Figure 7 It is the optimization result graph under the weight parameter values of different structural similarity index losses. In this embodiment, w2 = 0.01 is selected, and the generated halftone image performs well between structural clarity and consistency. According to different preferences for image quality, a higher PSNR score can be obtained by reducing w2, but this may be at the cost of sacrificing structural and texture details.
[0158] 4. Multi-level halftone
[0159] Although the method of the present invention starts from the basic two-level halftone problem, its flexibility and scalability enable it to adapt to more complex image processing scenarios, especially for the expansion of multi-level halftone. The multi-level halftone technology provides more dot options for each region of the image by increasing the gray levels, thereby visually simulating a smoother and more continuous tone.
[0160] To specifically illustrate this point, this embodiment provides a specific example of multi-level halftone here. By adjusting the number of output channels of the model and using the Gumbel-Softmax technique, the one-hot vector output by it is weighted and summed with the pre-defined multi-tone halftone level weights. Since the proposed blue noise loss is only optimized for the binary discretized output, this blue noise solution is not used, and finally the required discrete halftone result is obtained. Through experiments, the SSIM and PSNR scores of five-level halftone on the VOC test set are 0.3216 / 42.825 respectively. Figure 8 It is the result graph of two five-level halftone instances, where Figure 8 the (a) column in [reference] is the continuous-tone gray image, Figure 8 the (b) column in [reference] is the five-level halftone result graph, Figure 8Column (c) in [description] is the output line graph. It can be observed from the output line graph that the output of the model still focuses on the centers of discrete gray levels.
[0161] Although the extended solution proposed in the present invention shows feasibility in the implementation of multi-level halftone algorithms, in order to further improve image quality, several additional complex dimensions must be noted, including selecting a suitable blue noise model to optimize visual quality and effectively alleviating stripe artifacts. To achieve these goals, these considerations need to be incorporated into the existing framework to achieve deeper optimization and improvement. Given the intertwined complexity of multi-dimensional problems, more refined theoretical analysis and experimental verification are required, and these in-depth research works are regarded as the direction of future research.
[0162] In summary, in view of the limitations of slow processing speed and poor halftone effect of current digital halftone algorithms, the present invention solves the non-differentiable problem caused by halftone discrete selection by introducing the reparameterization strategy of the Gumbel-Softmax module, realizing unbiased estimation of gradients in the network backpropagation process. To further enhance the effect of halftone images, a new blue noise loss function is designed to optimize the distribution of halftone dots. At the same time, a regional confidence aggregation module is proposed, which makes the model pay more attention to the interaction information between pixels during training by combining the spatial correlation of pixels. Based on the above strategies, a self-supervised deep differentiable halftone model without label guidance is constructed by optimizing the expected value of halftone quality assessment. Experimental results show that the method proposed in the present invention can generate high-quality halftone images without image labels, effectively retaining the local structural information and texture details of the image while maintaining a high processing speed and low parameter complexity. Moreover, the model can be flexibly extended to multi-level halftone processing to meet the requirements of multi-level print heads.
[0163] Embodiment III
[0164] Based on the differentiable halftone method based on deep self-supervised learning described in Embodiment I, this embodiment provides a differentiable halftone device based on deep self-supervised learning, including:
[0165] A model construction module for constructing a deep differentiable halftone model, including a residual network and a Gumbel-Softmax module;
[0166] A training module, when training the deep differentiable halftone model, inputs continuous-tone grayscale images and constant grayscale images into the deep differentiable halftone model. After introducing random Gaussian noise into the continuous-tone grayscale images and constant grayscale images respectively, it uses a residual network to extract feature maps, and uses the Gumbel-Softmax module to reparameterize the feature maps output by the residual network, and outputs halftone images corresponding to the continuous-tone grayscale images and constant grayscale images respectively; calculates the distances between the continuous-tone grayscale images, constant grayscale images and halftone images using a loss function, and iteratively optimizes the parameters of the residual network.
[0167] A halftone module, after the deep differentiable halftone model is trained, replaces the Gumbel-Softmax module in it with an argmax function to obtain a target deep differentiable halftone model; uses the target deep differentiable halftone model to perform halftoning on continuous-tone images.
[0168] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0169] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.
[0172] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A differentiable halftoning method based on deep self-supervised learning, characterized in that, Including: Construct a deep differentiable halftone model, including a residual network and a Gumbel-Softmax module; When training the deep differentiable halftone model, input continuous-tone grayscale images and constant grayscale images into the deep differentiable halftone model. After introducing random Gaussian noise into the continuous-tone grayscale images and constant grayscale images respectively, use the residual network to extract feature maps, and use the Gumbel-Softmax module to reparameterize the feature maps output by the residual network, and output halftone images corresponding to the continuous-tone grayscale images and constant grayscale images respectively; use the loss function to calculate the distances between the continuous-tone grayscale images and constant grayscale images and their corresponding halftone images respectively, and iteratively optimize the parameters of the residual network; After the deep differentiable halftone model is trained, replace the Gumbel-Softmax module in it with the argmax function to obtain the target deep differentiable halftone model; Use the target deep differentiable halftone model to perform halftoning on continuous-tone images.
2. The differentiable halftoning method based on deep self-supervised learning according to claim 1, characterized in that The reparameterization of the feature maps output by the residual network using Gumbel-Softmax includes: Given a discrete random variable, assume it has K categories, and generate an independent Gumbel noise for each category; Use the Softmax function to reparameterize the output feature maps of the residual network for discretization and output the halftone image.
3. The differentiable halftoning method based on deep self-supervised learning according to claim 2, wherein For each pixel in the output feature map, an independent Gumbel noise is generated for each category, and the formula is: ; Among them, represents the probability distribution of the k-th category and follows a uniform distribution ; represents the Gumbel noise of the k-th category, where k = 1, 2, …, K; Use the Softmax function to reparameterize the output feature maps of the residual network for discretization, and the formula is: ; where, is the output value of the k-th category, is the temperature parameter, and represent the probabilities of the k-th category and the -th category respectively, represents the Gumbel noise of the -th category; Take the maximum output value among all categories as the halftone output value of the current pixel.
4. A differentiable halftoning method based on deep self-supervised learning according to claim 1, characterized in that, The loss function of the deep differentiable halftone model includes blue noise loss, mean squared error loss, and structural similarity index loss.
5. A differentiable halftoning method based on deep self-supervised learning according to claim 4, wherein The formula for the blue noise loss is: ; Among them, represents the blue noise loss, represents the ideal blue noise spectrum variance, represents the blue noise spectrum variance of the halftone image, represents the mean square error; The formula for the blue noise spectrum variance of the halftone image is: ; Among them, represents the blue noise spectral variance of the halftone image, is the power spectrum of the halftone image, represents the radially averaged power spectral density of the halftone image, represents the frequency, is the radial frequency with a width of ; is a circular ring with a width of around it, is the number of discrete frequency samples in; The calculation formula for the radial average power spectral density of the halftone image is: ; The calculation formula for the power spectrum of the halftone image is: ; Among them, represents the halftone image output by the Gumbel-Softmax module, represents the total number of pixels of the image, represents the Fourier transform.
6. A differentiable halftoning method based on deep self-supervised learning according to claim 4, characterized in that, The formula for the mean squared error loss is: ; Among them, represents the human visual system, represents a halftone image, represents a continuous-tone grayscale image.
7. A differentiable halftoning method based on deep self-supervised learning according to claim 6, characterized in that, Optimizing the mean squared error loss includes: When calculating the mean squared error loss, introduce a region confidence aggregation mechanism. Construct a confidence feature map for pixel-by-pixel matching based on the mean squared error of each pixel between the halftone image and the continuous-tone grayscale image, and then extract a sampled sub-region from the confidence feature map to calculate the local descriptor of each pixel; Calculate the mean squared error between the halftone image and the continuous-tone grayscale image according to the local descriptor of each pixel to obtain the mean squared error loss.
8. A differentiable halftoning method based on deep self-supervised learning according to claim 7, characterized in that, The formula representation of the local descriptor is: ; Among them, represents the local descriptor of the pixel with coordinates . represents the sampling sub-region centered on the pixel with coordinates . represents the pixel at the u-th row and v-th column in the sampling sub-region, represents the convolution filter weight of the pixel at the u-th row and v-th column in the sampling sub-region, represents the mean square error of the pixel with coordinates .
9. A differentiable halftoning method based on deep self-supervised learning according to claim 7, characterized in that The sampled sub-region of the region confidence aggregation mechanism is consistent with the filter window of the human visual system.
10. A differentiable halftoning device based on deep self-supervised learning, characterized in that, Including: A model construction module for constructing a deep differentiable halftone model, including a residual network and a Gumbel-Softmax module; A training module, when training the deep differentiable halftone model, inputs a continuous-tone grayscale image and a constant grayscale image into the deep differentiable halftone model. After introducing random Gaussian noise into the continuous-tone grayscale image and the constant grayscale image respectively, uses a residual network to extract feature maps, and uses a Gumbel-Softmax module to reparameterize the feature maps output by the residual network, and outputs halftone images corresponding to the continuous-tone grayscale image and the constant grayscale image respectively; uses a loss function to calculate the distances between the continuous-tone grayscale image and the constant grayscale image and their corresponding halftone images respectively, and iteratively optimizes the parameters of the residual network; A halftone module, after the deep differentiable halftone model is trained, replaces the Gumbel-Softmax module in it with an argmax function to obtain a target deep differentiable halftone model; uses the target deep differentiable halftone model to perform halftoning on continuous-tone images.
Citation Information
Patent Citations
Image halftone method and system based on improved residual network, and medium
CN116934618A
Reversible halftone method and system based on residual neural network
CN118429457A