A method for learning an image demoireing model from unpaired real data
By learning the image demodeling model from unpaired real data, the problem of complex data preparation and large gap in image characteristics in the prior art is solved, and the efficient demodeling effect is achieved, and the performance and flexibility of the model are improved.
Patent Information
- Application Number
- CN202310669860.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-07
AI Technical Summary
The prior art is difficult to effectively deal with violently changing molar patterns, and collecting pairs of real molar and molarless images requires professional equipment and heavy manpower, resulting in complex data preparation for training models and large gap in image features.
By dividing the real molar image into image blocks and grouping, a molar pattern generation frame is introduced to synthesize the pseudo-molar pattern image, paired with the real molar-free image to train the demolar pattern network, and an adaptive denoising method is used to remove low-quality pseudo-molar pattern images.
The manpower for data preparation is reduced, the diversity of the data set is improved, and the demolarization performance on real molar images is significantly improved, showing superior molarization effect.
Smart Images

Figure CN116452463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the training of an artificial neural network for moiré removal, and in particular, to a method for learning an image moiré removal model from unpaired real data. Background Art
[0002] Modern society is filled with electronic screens for presenting images, texts, videos, etc. With the wide use of portable imaging devices such as smart phones, people have become accustomed to using them to quickly record information. A common problem arises from the inherent interference between the color filter array (CFA) of a camera and the LCD sub-pixel layout of a screen, resulting in the captured images being contaminated by some rainbow-shaped stripes, which is also known as moiré (such as Jingyu Yang, et al. Demoiréing for screen-shot images with multi-channel layer decomposition. In IEEE Visual Communications and Image Processing (VCIP), pages 1–4, 2017). These moirés involve different thicknesses, frequencies, layouts, and colors, reducing the perceived quality of the captured images; therefore, the academic and industrial communities have shown great interest in developing moiré removal algorithms to correct this problem.
[0003] Most of the original research on moiré removal is based on image priors (such as Taeg Sang Cho, et al. Image restoration by matching gradient distributions. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 34: 683–694, 2011.) or traditional machine learning methods (such as Jingyu Yang, et al. Textured image demoiréing via signal decomposition and guided filtering. IEEE Transactions on Image Processing (TIP), 26: 3528–3541, 2017.). These methods have been proven insufficient to handle severely varying moiré patterns (Bolun Zheng, et al. Learning frequency domain priors for image demoireing. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 44: 7705–7717, 2021.). Sophisticated convolutional neural networks (CNNs) have become the de facto infrastructure for success in various computer vision tasks, including image demoireing (such as Shanxin Yuan, et al. Aim 2019 challenge on image demoireing: Methods and results. In Proceedings of IEEE / CVF International Conference on Computer Vision Workshop (ICCVW), pages 3534–3545, 2019.). These CNN-based methods are typically trained in a supervised manner on a wide range of moiré-free and moiré images pairs to simulate moiré mapping. However, considering that natural moiré patterns have different thicknesses, frequencies, layouts, and colors, collecting paired images is challenging. Although moiré images and moiré-free images can be easily obtained, most of them are unpaired. Although many studies have attempted to capture image pairs from digital screens (such as Bin He, et al. Fhde 2net: Full high definition demoireing network. (In Proceedings of the European Conference on Computer Vision (ECCV), pages 713–729, 2020.) However, their quality is blocked by three limitations. First, obtaining high-quality image pairs requires professional camera position adjustment and even special hardware (Xin Yu, et al. Towards efficient and scale-robust ultrahigh-definition image demoiréing. In Proceedings of the European Conference on Computer Vision (ECCV), pages 646–662, 2022.). Second, it requires heavy human effort to select well-aligned moiré-free and moiré pairs. Third, in a highly controlled laboratory environment, the captured moiré content is very single. However, image pairs with more diverse moiré patterns are more promising for improving the demoireing model.
[0004] Therefore, learning to synthesize moiré images has recently attracted increasing attention. The shooting simulation method (such as Dantong Niu, et al. Morié attack (ma): A new potential risk of screen photos. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pages 26117–26129, 2021.) simulates the aliasing between the CFA and the LCD sub-pixels of the screen to generate corresponding paired moiré images. However, the synthesized images are insufficient to capture the characteristics of real moiré, resulting in a large domain gap. Recently, Park et al. (Hyunkook Park, et al. Unpaired screen-shot image demoiréing with cyclic moiré learning. IEEE Access, 10:16254–16268, 2022.) introduced a cyclic moiré learning method, which can achieve better performance than shooting simulation. However, the generated pseudo-moiré still fails to accurately simulate real moiré patterns, resulting in limited performance. Summary of the Invention
[0005] The objective of the present invention is to provide a method for learning an image de - moiré model from unpaired real data, synthesizing moiré images, which have the same moiré characteristics as real moiré images and the same detail characteristics as real non - moiré images. The synthesized pseudo - moiré images are paired with real non - moiré images to form an image pair for training a de - moiré network.
[0006] The present invention includes the following steps:
[0007] 1) Image pre - processing: Divide real moiré images into several image patches and group these image patches according to the complexity of the moiré.
[0008] 2) Moiré synthesis network: Introduce a new moiré generation framework to synthesize moiré images with different moiré characteristics similar to real moiré and details similar to real non - moiré images.
[0009] 3) Adaptive denoising: Introduce an adaptive denoising method to remove low - quality pseudo - moiré images that have an adverse effect on the learning of the de - moiré model.
[0010] In step 1), the division of image patches means dividing the moiré image dataset into several non - overlapping image patches, thereby obtaining a set of moiré image patches where N is the number of image patches in the entire set of moiré image patches. Similarly, the non - moiré image dataset can be divided to obtain a set of non - moiré image patches , where M is the number of image patches in the entire set of non - moiré image patches.
[0011] The grouping of image patches means dividing the set of moiré image patches into K subsets, that is, there is where each subset contains moiré image patches with similar complexity, and any two subsets are disjoint.
[0012] The complexity of the moiré is composed of the frequency and color information of the moiré image. Given a moiré image patch the frequency of this moiré image patch can be measured by a Laplacian edge detection operator with a kernel size of 3 , and the color information of this moiré image patch is a linear combination of the mean and standard deviation of the pixel cloud in the color plane of the RGB color space, expressed as:
[0013]
[0014] Among them, μ(·) and σ(·) return the mean and standard deviation of the input. and respectively represent the red, green, and blue color channels of pm.
[0015] In step 2), the moiré synthesis network includes four parts: a moiré feature encoder E m , a generator G m , a discriminator D m and a content encoder E C ; the goal of the present invention is to generate a pseudo-moiré picture which has moiré patterns of p m while retaining the image details of the moiré-free image patches so that a moiré and moiré-free image pair is formed to guide the learning of the existing moiré removal network;
[0016] To achieve this goal, the moiré feature encoder E m extracts the moiré features of the real moiré image patch p m , denoted as F m :
[0017] F m = E m (p m )
[0018] Then, the generator G m takes F m and p f as inputs and synthesizes a pseudo-moiré image patch
[0019]
[0020] where Con(·,·) represents the concatenation operation.
[0021] The discriminator D m cooperates with the generator G m to obtain a better pseudo-moiré image patch in an adversarial training manner. The generator G m is trained to deceive the discriminator D m :
[0022]
[0023] To obtain better training stability, the least squares loss function is used. At the same time, the discriminator D m is trained to distinguish the pseudo-moiré picture from the real moiré picture p m :
[0024]
[0025] In addition, the moiré pattern feature required to be synthesized is consistent with the moiré pattern feature of the real p f :
[0026]
[0027]
[0028] where ||·||1 represents the l1 loss. To better pair and p f , is also expected to have the content details of p f . An additional content encoder E C is introduced to align the content features between and p f :
[0029]
[0030] The total loss function is:
[0031]
[0032] In step 3), the adaptive denoising refers to removing the low-quality pseudo-moiré images that have an adverse effect on the learning of the moiré removal model. It is found that some pseudo-moiré image blocks occasionally have low-quality problems, where the content and details of p f are damaged in . Such noisy data hinders the learning of the moiré removal model. The damaged structure is mainly attributed to the edge information. Therefore, the edge map of each image block is calculated by the Laplace edge detection operator, and the structural difference is calculated by summing the absolute values of the edge differences between each pair of pseudo-images. Low-quality pseudo-moiré will result in a large fraction of the structural difference. As long as the fraction exceeds a threshold, these pseudo-image pairs can be excluded. This threshold is the γ-th percentile of the structural difference among a total of N pseudo-image pairs.
[0033] Compared with the prior art, the present invention has the following prominent advantages:
[0034] 1) By learning a moiré removal model from unpaired real data, the cumbersome work of collecting real-world paired moiré and non-moiré images is avoided, the manpower in the data preparation process is reduced, and the diversity of the data set is improved.
[0035] 2) A large number of experiments show that on the real moiré image dataset, the present invention greatly improves the comparison benchmark, demonstrating the superior performance of the present invention and also inspiring a new moiré generation method for the moiré removal field. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a diagram of the image preprocessing process of the present invention;
[0037] Figure 2 It is a framework diagram of the moiré synthesis network of the present invention.
[0038] Figure 3 It is the visualization result of moiré removal by the MBCNN network on the UHDM dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following embodiments will further illustrate the present invention in conjunction with the accompanying drawings.
[0040] The method framework diagram of the embodiments of the present invention is as shown in Figure 1 and 2 shown.
[0041] 1. Training Instructions
[0042] The embodiments of the present invention include the following steps:
[0043] 1) Image preprocessing: Divide the real moiré image into several image blocks and group these image blocks according to the complexity of the moiré.
[0044] 2) Moiré synthesis network: Introduce a new moiré generation framework to synthesize moiré images with different moiré characteristics similar to real moiré, as well as details similar to real moiré-free images.
[0045] 3) Adaptive denoising: Introduce an adaptive denoising method to remove low-quality pseudo-moiré images that have an adverse effect on the learning of the moiré removal model.
[0046] In step 1), the division of the image blocks means dividing the moiré image dataset into several non-overlapping image blocks, thereby obtaining a moiré image block set where N is the number of image blocks in the entire moiré image block set. Similarly, the moiré-free image dataset can be divided to obtain a moiré-free image block set , and M is the number of image blocks in the entire moiré-free image block set.
[0047] The grouping of the image blocks means the moiré image block set Divided into K subsets, that is, there are where each subset contains moiré image patches with similar complexity, and any two subsets are disjoint.
[0048] The complexity of the moiré is composed of the frequency and color information of the moiré image (Yuxin Zhang, et al. Real-time image demoireing on mobile devices. In Proceedings of the International Conference on Learning Representations (ICLR), 2023.). Given a moiré image patch The frequency of the moiré image patch can be measured by a Laplacian edge detection operator with a kernel size of 3 (David Marr and Ellen Hildreth. Theory of edge detection. Proceedings of the Royal Society of London. Series B. Biological Sciences, 207∶187 - 217, 1980.), and the color information of the moiré image patch is a linear combination of the mean and standard deviation of the pixel cloud in the color plane of the RGB color space (David Hasler and Sabine E Suesstrunk. Measuring colorfulness in natural images. In Human vision and electronic imaging VIII, volume 5007, pages 87 - 95, 2003.), expressed as:
[0049]
[0050] where μ(·) and σ(·) return the mean and standard deviation of the input, and respectively represent the red, green, and blue color channels of p m of.
[0051] In the present invention, K = 4 is set to obtain four subsets of moiré image patches of the same size, and each subset has unique moiré characteristics. The first group contains the first N / 4 small image patches, so its moiré pattern frequency is low and the colors are few. Then, a new metric is used to sort the remaining image patches from smallest to largest. Then, is composed of the first N / 4 image patches, which have low frequency but rich colors. The middle N / 4 image patches form featured by high frequency and rich colors. The N / 4 image patches with the smallest scores have high frequency but fewer colors and form
[0052] In step 2), the moiré synthesis network comprises four parts: a moiré feature encoder E m , a generator G m , a discriminator D m and a content encoder E C . The objective of the present invention is to generate a pseudo-moiré picture that has p m moiré stripes while retaining the image details of the moiré-free image patches , so that a moiré and moiré-free image pair is formed to guide the learning of the existing moiré removal network.
[0053] To achieve this goal, the moiré feature encoder E m extracts the moiré features of the real moiré image patch p m , denoted as F m :
[0054] F m = E m (p m )
[0055] Then, the generator G m takes F m and p f as inputs and synthesizes a pseudo-moiré image patch
[0056]
[0057] where Con(·, ·) represents the concatenation operation.
[0058] The discriminator D m is paired with the generator G mCollaborate to obtain better moiré image patches in the way of adversarial training (Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pages 2672 - 2680, 2014.). The generator G m is trained to deceive the discriminator D m :
[0059]
[0060] To obtain better training stability, the least squares loss function is used (Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley. Least squares generative adversarial networks. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), pages 2794 - 2802, 2017.). Meanwhile, the discriminator D m is trained to distinguish between moiré images and real moiré images p m :
[0061]
[0062] In addition, it is required that the synthesized moiré features are consistent with those of the real p f :
[0063]
[0064]
[0065] where ||·||1 represents the l1 loss. To better pair and p f , is also expected to have pf Content details. An additional content encoder E C is introduced to align and p f in terms of content features:
[0066]
[0067] The total loss function is:
[0068]
[0069] In step 3), the adaptive denoising refers to removing low-quality pseudo moiré images that have an adverse effect on the learning of the moiré removal model. It is found that some pseudo moiré image patches occasionally have low-quality problems, among which, the content and details of p f are damaged in . Such noisy data hinders the learning of the moiré removal model. Fortunately, the damaged structures mainly boil down to edge information. Therefore, the edge map of each image patch is calculated through the Laplace edge detection operator, and the structural difference is calculated by summing the absolute values of the edge differences between each pair of pseudo-images. Low-quality pseudo moiré will result in a large fraction of the structural difference. As long as the fraction exceeds a threshold, these pairs of pseudo-images can be excluded. This threshold is the γ-th percentile of the structural differences among a total of N pairs of pseudo-images. The above process is carried out for each synthesis network and the corresponding γ i is set to remove low-quality pseudo moiré.
[0070] 2. Implementation details
[0071] The present invention uses publicly available datasets including the FHDMi dataset and the UHDM dataset. The present invention uses the training set to train the proposed moiré synthesis network. For image preprocessing, the training images of FHDMi are cropped into 8 image patches. For UHDM involving higher-resolution images, the training images are cropped into 6 image patches. During the training process, the moiré image patch p m and the non-moiré image patch p f are selected from different original images (before image preprocessing) to ensure that they are unpaired.
[0072] The present invention is implemented using the Pytorch framework. E m and E C contain one convolutional layer and two residual blocks. G m contains three convolutional layers, nine residual blocks and two deconvolutional layers, and ends with one convolutional layer to produce the final output. The residual block consists of two convolutional layers, followed by instance normalization and the ReLU function. The convolutional layer for E mand E C has 16 channels for G m has 128 channels. D m Drawing on PatchGAN (Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1125-1134, 2017.), it consists of three convolutional layers with a stride of 2 and two convolutional layers with a stride of 1, followed by an average pooling layer. For the moiré removal model, MBCNN and ESDNet-L (a large version of ESDNet) are used.
[0073] The moiré synthesis network is trained using the Adam optimizer, where the first momentum and the second momentum are set to 0.9 and 0.999 respectively. It is trained for 100 epochs with a batch size of 4 and an initial learning rate of 2×10 -4 , which linearly decays to 0 in the last 50 epochs. Additionally, different random croppings are performed on the image patches after image preprocessing to verify the flexibility of the present invention in synthesizing pseudo-moiré images. The cropping sizes for FHDMi are set to 192×192 and 384×384, and the cropping sizes for UHDM are set to 192×192, 384×384, and 768×768. For the moiré removal model, the same training configuration as the original paper is retained, and for fair comparison, all models are trained for 150 epochs. All networks are initialized using a Gaussian distribution with a mean of 0 and a standard deviation of 0.02. γ1, γ2, γ3, and γ4 for adaptive denoising are empirically set to 50, 40, 30, and 201 respectively. All experiments are run on an NVIDIA A100 GPU.
[0074] Widely used metrics such as peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and LPIPS are adopted to quantitatively evaluate the performance of the moiré removal model.
[0075] 3. Application Areas
[0076] The present invention can be applied to the training of moiré removal neural networks to learn an image moiré removal model from unpaired real data.
[0077] Table 1 shows the results of the MBCNN and ESDNet-L demoireing models on the FHDMi dataset (Bin He, Ce Wang, Boxin Shi, and Ling-Yu Duan. Fhde 2 net: Full high definition demoireing network. In Proceedings of the European Conference on Computer Vision (ECCV), pages 713–729, 2020.).
[0078] As can be seen from Table 1, the performance of the demoireing models trained on the simulated data for shooting is very poor. For example, when MBCNN is trained with a cropping size of 192×192, it only obtains a PSNR of 10.66 dB, which indicates a large domain gap between the pseudo-data and the real data. Both the cyclic learning method and UnDeM of the present invention show better results. In addition, compared with the cyclic learning method, UnDeM of the present invention successfully simulates the moire pattern and thus presents the highest performance. For example, when the crop size is 192 and 384, MBCNN obtains PSNRs of 19.45 dB and 19.89 dB respectively. For ESDNet-L, the PSNR results are 19.38 dB and 19.66 dB respectively. Correspondingly, the SSIM and LPIPS of UnDeM of the present invention also show much better performance than shooting simulation and cyclic learning.
[0079] Table 1
[0080]
[0081] Table 2 shows the results of the MBCNN and ESDNet-L demoireing models on the UHDM dataset (Xin Yu, et al., Towards efficient and scale-robust ultrahigh-definition image demoireing. In Proceedings of the European Conference on Computer Vision (ECCV), pages 646–662, 2022.). As can be seen from Table 2, the demoireing models trained on simulated captures still cannot handle real data, while cyclic learning provides better results. More importantly, the UnDeM of the present invention outperforms these two methods on different networks and training scales. Specifically, when training MBCNN with crop sizes of 192, 384, and 768, the PSNR increases by 0.54 dB, 0.10 dB, and 0.15 dB respectively. For ESDNet-L, the PSNR gains are 0.28 dB, 0.43 dB, and 0.40 dB respectively. Summarizing from Tables 1 and 2, it can be concluded that the transferability of the moiré images synthesized by the present invention in the downstream demoireing task and the efficacy of UnDeM relative to existing methods have been fully demonstrated.
[0082] Figure 3 Figure shows the visualization results of demoireing of the MBCNN network on the UHDM dataset, where Figure (a) shows the moiré image, Figure (e) shows the corresponding moiré-free image, and Figures (b - d) show the demoireing effects of different methods. As Figure 3 shown in Figure (b) therein, the demoireing result of simulated capture shows unnatural high brightness, resulting in the loss of image details. This degradation in visual quality can be attributed to the generally low brightness of simulated capture, which makes the demoireing model learn incorrect brightness relationships between moiré and moiré-free images. As Figure 3 shown in Figure (c) therein, since cyclic learning cannot model moiré patterns, the demoireing model fails to remove moiré. Figure 3 The result of Figure (d) therein shows the efficacy of UnDeM in removing moiré, reflecting the ability of UnDeM to successfully model moiré.
[0083] Table 2
[0084]
[0085] The above embodiments are only preferred embodiments of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. A method for learning an image de - moiré model from unpaired real data, characterized in that It includes the following steps: 1) Image preprocessing: Divide the real moiré image into several image blocks and group these image blocks according to the complexity of the moiré pattern; Grouping the image blocks refers to dividing the moiré image block set into K subsets, that is where each subset contains moiré image blocks with similar complexity, and any two subsets are disjoint; the complexity of the moiré is composed of the frequency and color information of the moiré image; given a moiré image block the frequency of the moiré image block is measured by a Laplacian edge detection operator with a kernel size of 3 and the color information of the moiré image block is a linear combination of the mean and standard deviation of the pixel cloud in the color plane of the RGB color space, expressed as: where μ(·) and σ(·) return the mean and standard deviation of the input, and represent the red, green, and blue color channels of p m respectively; 2) Moiré synthesis network: Introduce a new moiré generation framework to synthesize moiré images with different moiré characteristics similar to real moiré patterns and details similar to real moiré-free images; The moiré synthesis network includes four parts: a moiré feature encoder E m , a generator G m , a discriminator D m and a content encoder E C ; generate a pseudo-moiré image The pseudo-moiré image has moiré patterns with p m , while retaining the image details of the moiré-free image patches , forming a moiré and moiré-free image pair to guide the learning of existing moiré removal networks; Moiré Feature Encoder E m Extract the real moiré image patch p m The moiré feature of, denoted as F m : F m = E m (p m ) Generator G m takes F m and p f as inputs and synthesizes a pseudo - moiré image patch Among them, Con(·,·) represents the concatenation operation; Discriminator D m collaborates with the generator G m to obtain better moiré image patches in an adversarial training manner; the generator G m is trained to deceive the discriminator D m : To obtain better training stability, the least squares loss function is used; meanwhile, the discriminator D m is trained to distinguish between pseudo moiré images and real moiré images p m : In addition, it is required that the moiré pattern feature of the synthesized is consistent with the moiré pattern feature of the real p f : where ‖·‖1 represents the I1 loss; for better pairing and p f , is also expected to have the content details of p f ; an additional content encoder E C is introduced to align and p f the content features between: The total loss function is: 3) Adaptive denoising: Introduce an adaptive denoising method to remove low-quality pseudo-moiré images that have an adverse effect on the learning of the moiré removal model; The adaptive denoising refers to removing low-quality pseudo-Moiré images that have an adverse effect on the learning of the Moiré removal model; some pseudo-Moiré image blocks occasionally have low-quality problems, where f the content and details of are damaged; this damaged structure is mainly attributed to edge information; the edge map of each image block is calculated by the Laplacian edge detection operator, and the structural difference is calculated by summing the absolute values of the edge differences between each pair of pseudo-images; low-quality pseudo-Moiré will result in a large fraction of the structural difference, and as long as the fraction exceeds a threshold, these pairs of pseudo-images can be excluded; this threshold is the γ-th percentile of the structural difference among a total of N pairs of pseudo-images.
2. The method for learning an image de - moiré model from unpaired real data according to claim 1, wherein In step 1), the division of the real moiré image into a number of image blocks means dividing the moiré image dataset into a number of non-overlapping image blocks, obtaining a moiré image block set where N is the number of image blocks in the entire moiré image block set; at the same time, dividing the moiré-free image dataset to obtain a moiré-free image block set where M is the number of image blocks in the entire moiré-free image block set.