Method for removing moire patterns by using ultra-wide-angle image to assist wide-angle image

Through the method of assisting wide-angle images with ultra-wide-angle images, the position and color alignment, key point matching and kernel prediction alignment modules are used, and the adaptive fusion module is combined with the problem of heavy and large molar patterns processing in the existing technology is solved, achieving the effect of efficiently removing molar patterns and retaining image details.

CN120219200APending Publication Date: 2025-06-27HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510291290.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing image demolarization method is high in calculation cost and the inference process takes too long when processing heavy and large molar patterns, and it is difficult to achieve ideal results when processing ultra-high-definition images on the mobile phone.

Method used

The method of using ultra-wide-angle image to assist wide-angle image to remove molar patterns is constructed by acquiring UW images, molar patterns W images and source target images for position and color alignment, and the training data set is constructed, and the alignment module based on key point matching and kernel prediction is used to align image levels and feature levels, combined with the adaptive fusion module for image restoration, and the processing is performed using a lightweight encoder and decoder.

Benefits of technology

Effectively remove molar patterns in wide-angle images, improve image quality, reduce calculation costs, realize rapid processing, retain image texture and color, and is suitable for imaging devices such as mobile phones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219200A_ABST
    Figure CN120219200A_ABST
Patent Text Reader

Abstract

The invention discloses a method for removing moire patterns by using an ultra-wide-angle image to assist a wide-angle image, and belongs to the technical field of image restoration. The method aims at solving the problem that an existing method for removing moire in an image is achieved by means of a huge network, and the reasoning process is complex. Comprising the following steps: carrying out image level alignment on a UW image in a training set to a corresponding moire W image by adopting an alignment module based on key point matching to obtain an aligned UW image; respectively encoding the moire W image and the moire UW image to obtain image features, performing feature level alignment on the W image features and the UW image features by adopting an alignment module based on kernel prediction, and performing fusion by adopting a self-adaptive fusion module to obtain fusion features; decoding to obtain a restored image; and calculating a loss function based on the aligned target image and the restored image, and adjusting network parameters to obtain a restored network for moire removal of the actual moire W image. The method is used for removing the moire patterns of the wide-angle image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for removing moiré patterns from a wide-angle image by using an ultra-wide-angle image as an aid, and belongs to the technical field of image restoration. Background Art

[0002] It is common to use smartphones to record and transmit data from electronic screens, which can provide convenience for users in many scenarios. However, moiré patterns often appear in images when shooting, which appear as wavy, concentric and other complex geometric patterns. This phenomenon is caused by frequency aliasing between the pixel array of the camera sensor and the pixel grid of the display. Moiré patterns not only affect visual quality, but may also cause loss of information about image texture and details. Therefore, the technology to remove moiré patterns has received widespread attention.

[0003] Existing image de-moiré methods have achieved certain results on HD and UHD images, such as improving the multi-scale processing architecture. This method performs well when dealing with light and small moiré, but cannot achieve ideal results when heavy and large moiré covers the image details and texture. The current solution is to design a complex and large network to enhance the model fitting ability, but this will greatly increase the computing cost, especially when processing UHD images on mobile phones, which will cause the inference process to take too long or even be impossible.

[0004] In order to meet the photography needs of users, current mainstream mobile phones are usually equipped with ultra-wide-angle (UW) lenses and wide-angle lenses (W). Among them, the UW lens has a wider field of view and can capture more scenery; the W lens has high resolution and excellent low-light performance. Through observation, it is found that moiré is very sensitive to the focal length of the lens, the distance and angle between the camera and the display. The W lens is a commonly used camera for capturing screen display content. When there is severe frequency aliasing between the W lens sensor and the screen sub-pixels, heavy and large moiré will appear in the W image. Due to the different focal lengths and positions of the UW and W lenses, the UW image usually does not have heavy and large moiré. Therefore, it is necessary to process the moiré that appears in the W image. Summary of the invention

[0005] In view of the problem that the existing methods for removing moiré in images rely on huge networks and have complex reasoning processes, the present invention provides a method for removing moiré by using ultra-wide-angle images to assist wide-angle images.

[0006] A method for removing moiré from a wide-angle image using an ultra-wide-angle image is provided by the present invention, comprising:

[0007] Obtain the UW image, the moiré W image, and the source target image. After performing position alignment and color alignment on the source target image with respect to the moiré W image, obtain the aligned target image. Construct a training dataset from the UW image, the moiré W image, and the aligned target image.

[0008] Use an alignment module based on key point matching to perform image-level alignment of the UW image in the training set with respect to the corresponding moiré W image, obtaining the aligned UW image.

[0009] Use a first encoder to encode the moiré W image to obtain W image features; use a second encoder to encode the aligned UW image to obtain UW image features; use an alignment module based on kernel prediction to perform feature-level alignment of the W image features and the UW image features, and then use an adaptive fusion module for fusion to obtain fusion features; then use a decoder to decode the fusion features to obtain a restored image.

[0010] Calculate a loss function based on the aligned target image and the restored image, and adjust the network parameters of the first encoder, the second encoder, the alignment module based on kernel prediction, and the decoder until the value of the loss function is within a preset threshold, obtaining a restoration network composed of the alignment module based on key point matching and the trained first encoder, second encoder, alignment module based on kernel prediction, and decoder.

[0011] Use the restoration network to process the actual UW image and the actual moiré W image to obtain a restored image of the actual moiré W image after removing moiré.

[0012] According to the method for removing moiré from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the first encoder uses the encoder of the ESDNet network; the second encoder is a lightweight encoder; the decoder uses the decoder of the ESDNet network.

[0013] According to the method for removing moiré from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the method for performing position alignment of the source target image with respect to the moiré W image is as follows:

[0014] First, use the Gluestick algorithm to perform image matching based on key points and key lines to roughly align the position of the source target image with respect to the moiré W image; then use the FlowFormer algorithm to perform pixel-by-pixel position alignment of the source target image after rough position alignment with respect to the moiré W image, obtaining the source target image after position alignment.

[0015] According to the method for removing moiré from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the source target image after position alignment The method for then performing color alignment with respect to the moiré W image is as follows:

[0016] After Gaussian blurring the source target image and the moiré W image, use the least squares method to estimate the 3×3 color correction matrix of the source target image relative to the moiré W image, and align the source target image after position alignment based on the 3×3 color correction matrix Then perform color alignment on the moiré W image to obtain the aligned target image GTI gt .

[0017] According to the method for removing moiré from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the method for performing image-level alignment of the UW image in the training set to the corresponding moiré W image using the alignment module based on key point matching is as follows:

[0018] Use the SuperPoint method to detect 256 key points from the UW image and the moiré W image downsampled by 4 times respectively; use the LightGlue method to match the key points, and perform image-level alignment of the UW image to the corresponding moiré W image to obtain the aligned UW image:

[0019]

[0020] where is the aligned UW image, KMA represents the alignment module based on key point matching, and I m is the moiré W image, and I uw is the UW image.

[0021] According to the method for removing moiré from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the method for performing feature-level alignment of the W image feature and the UW image feature using the alignment module based on kernel prediction is as follows:

[0022] Represent the W image feature as F m , and represent the UW image feature as F uw ;

[0023] Use adaptive average pooling to aggregate the context information of F m and F uw , and then connect them along the channel dimension to obtain W':

[0024] W' = Concat(AdaptivePool(F m ), AdaptivePool(F uw ))

[0025] where AdaptivePool represents adaptive average pooling and Concat represents connection;

[0026] Then apply two convolutional layers to process W' to obtain the initial convolutional kernel weight W:

[0027] W = Softmax(Reshape(Conv(ReLU(Conv(W'))))),

[0028] where Softmax is the normalized exponential function, Reshape is the shape adjustment operation, Conv is the convolution operation, and ReLU is the activation function;

[0029] Perform element-wise multiplication between the initial convolution kernel weights W i and the learnable parameter P i in each weight group dimension of the kernel prediction-based alignment module, and then sum along the group dimension to obtain the content-adaptive convolution kernel θ:

[0030]

[0031] G represents the number of weight group dimensions;

[0032] Process the UW image feature F uw using the content-adaptive convolution kernel θ to obtain the aligned UW image feature

[0033] According to the method for removing moiré patterns from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the method of performing fusion using the adaptive fusion module is as follows:

[0034] Fuse the aligned UW image feature and the W image feature F m to obtain the fused feature F fuse :

[0035]

[0036] where α is a learnable coefficient.

[0037] According to the method for removing moiré patterns from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the first encoder includes a moiré pattern image encoding unit

[0038] for performing bilinear downsampling operation, first convolution operation, and first activation operation on the moiré pattern W image;

[0039] for successively performing a first residual block, a first semantic alignment perception block, a second convolution operation, and a second activation operation on the output;

[0040] for The output is successively subjected to a second residual block, a second semantic alignment perception block, a third convolution operation, and a third activation operation;

[0041] For The output is successively subjected to a third residual block and a third semantic alignment perception block operation.

[0042] According to the method for removing moiré patterns from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the second encoder includes an ultra-wide-angle image encoding unit

[0043] For performing a bilinear downsampling operation, a first convolution operation, and a first activation operation on the aligned UW image;

[0044] For The output is successively subjected to a second convolution operation, a second activation operation, a third convolution operation, a third activation operation, a fourth convolution operation, a fifth convolution operation, and a fourth activation operation;

[0045] For The output is successively subjected to a sixth convolution operation, a fifth activation operation, a seventh convolution operation, a sixth activation operation, an eighth convolution operation, a ninth convolution operation, and a seventh activation operation;

[0046] For The output is successively subjected to a tenth convolution operation, an eighth activation operation, an eleventh convolution operation, a ninth activation operation, and a twelfth convolution operation.

[0047] According to the method for removing moiré patterns from a wide-angle image assisted by an ultra-wide-angle image of the present invention, the alignment module based on kernel prediction includes alignment units KPA1 to KPA3:

[0048] KPA1 is used to respectively perform max pooling on the and outputs, connect the pooled features together, successively perform a first convolution operation, a first activation operation, a second convolution operation, a second activation operation, and a first group sum operation, use the obtained result as a feature kernel, and then perform a convolution operation on the output;

[0049] KPA2 is used to respectively perform max pooling on the and outputs, connect the pooled features together, successively perform a third convolution operation, a third activation operation, a fourth convolution operation, a fourth activation operation, and a second group sum operation, use the obtained result as a feature kernel, and then perform a convolution operation on the Perform a convolution operation on the output;

[0050] KPA3 is used to and Perform max pooling on the outputs of respectively, and concatenate the features after max pooling, and successively perform the fifth convolution operation, the fifth activation operation, the sixth convolution operation, the sixth activation operation and the third group summation operation, and use the obtained result as the feature kernel, and then perform a convolution operation on the output of Perform a convolution operation on the output;

[0051] The adaptive fusion module includes fusion units AF1 to AF3:

[0052] AF1 is used to Multiply the outputs of and KPA1 by the first coefficient respectively and then add them;

[0053] AF2 is used to Multiply the outputs of and KPA2 by the second coefficient respectively and then add them;

[0054] AF3 is used to Multiply the outputs of and KPA3 by the third coefficient respectively and then add them;

[0055] The three output results of the fusion units AF1 to AF3 are decoded to obtain the restored image.

[0056] Advantages of the present invention: The present invention uses dual-camera fusion for image demoireing, that is, uses an ultra-wide-angle UW image to help remove the moire of the wide-angle W image. The UW image can provide information with relatively normal texture and color but lower resolution, improving the moire removal performance of the W image.

[0057] For the two lenses equipped with a mobile phone, when there is moire in the W image, the UW image can usually provide normal color and texture, mainly due to their different focal lengths. The present invention integrates a lightweight UW image encoder into the demoireing network and adopts a fast two-stage image alignment method.

[0058] The method of the present invention utilizes the complementarity of ultra-wide-angle UW and wide-angle W images to enhance the effect of moiré removal from images, and at the same time constructs an efficient moiré removal model for images by means of the differences between the two. In the application stage of the present invention, a W image containing moiré and the corresponding UW image can be input for processing. Compared with the traditional monocular camera moiré removal method, the present invention makes full use of the resources of dual cameras and does not require complex additional processing to deal with severe moiré. The alignment framework proposed by the present invention includes two stages, namely, the key-point matching-based alignment module (KMA) and the kernel prediction-based alignment module (KPA). In the actual scene images, the processing results of the present invention have more thorough moiré removal and better retention of image texture and color, effectively solving problems such as moiré residue and image distortion in other methods. The present invention can be directly applied to imaging devices such as mobile phones, so as to obtain higher-quality images in scenarios where moiré is likely to appear, such as when photographing electronic screens.

[0059] Experiments on the dataset show that, compared with the current state-of-the-art methods, the method of the present invention removes moiré more cleanly and restores details more clearly. Description of the Drawings

[0060] Figure 1 It is a comparison diagram of the method for removing moiré from wide-angle images assisted by ultra-wide-angle images according to the present invention and the existing classic moiré removal methods for images; in the figure, (a) represents a classic moiré removal model for images using a single input, and (b) represents a moiré removal model for images using asymmetric camera fusion based on the method of the present invention;

[0061] Figure 2 It is a schematic diagram of the processing of the training dataset;

[0062] Figure 3 It is an overall architecture diagram of the method for removing moiré from wide-angle images assisted by ultra-wide-angle images according to the present invention;

[0063] Figure 4 It is a network structure diagram of the kernel prediction-based alignment module KPA and the adaptive fusion module AF;

[0064] Figure 5 It is a visual quality comparison of the moiré processing results of the restoration network of the present invention and the existing models under real data Figure 1 ; in the figure, the first column represents the input moiré image, the second column represents the enlarged part of the moiré image, the third column represents the result of the DMCNN method, the fourth column represents the result of the WDNet method, the fifth column represents the result of the MBCNN method, the sixth column represents the result of the ESDNet method, the seventh column represents the result of the method of the present invention, and the eighth column represents the moiré-free image;

[0065] Figure 6Visual quality comparison of the moiré processing results of the restoration network of the present invention and existing models under real data Figure 2 ; In the figure, the first column represents the input moiré image, the second column represents the enlarged part of the moiré image, the third column represents the result of the DMCNN method, the fourth column represents the result of the WDNet method, the fifth column represents the result of the MBCNN method, the sixth column represents the result of the ESDNet method, the seventh column represents the result of the method described in the present invention, and the eighth column represents the moiré-free image. Specific embodiments

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.

[0067] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0068] The present invention will be further described below in conjunction with the accompanying drawings, but it is not a limitation of the present invention.

[0069] Combined with Figures 1 to 4 As shown, the present invention provides a method for removing moiré from a wide-angle image with the assistance of an ultra-wide-angle image, including

[0070] Obtain the UW image, the moiré W image, and the source target image, and perform position alignment and color alignment on the source target image to the moiré W image to obtain the aligned target image; construct a training dataset from the UW image, the moiré W image, and the aligned target image;

[0071] Use an alignment module based on key point matching to perform image-level alignment on the UW image in the training set to the corresponding moiré W image to obtain the aligned UW image;

[0072] Use the first encoder to encode the moiré W image to obtain the W image feature; use the second encoder to encode the aligned UW image to obtain the UW image feature; use an alignment module based on kernel prediction to perform feature-level alignment on the W image feature and the UW image feature, and then use an adaptive fusion module for fusion to obtain the fusion feature; then use a decoder to decode the fusion feature to obtain the restored image;

[0073] Calculate the loss function based on the aligned target image and the restored image, and adjust the network parameters of the first encoder, the second encoder, the kernel prediction-based alignment module, and the decoder until the value of the loss function is within a preset threshold, obtaining a restoration network composed of the key-point matching-based alignment module and the trained first encoder, second encoder, kernel prediction-based alignment module, and decoder;

[0074] Process the actual UW image and the actual moiré W image using the restoration network to obtain the restored image of the actual moiré W image after removing the moiré pattern.

[0075] As an example, the first encoder uses the encoder of the ESDNet network; the second encoder is a lightweight encoder; the decoder uses the decoder of the ESDNet network.

[0076] The ESDNet network is an existing single-image moiré removal network. ESDNet is a multi-scale encoder-decoder architecture, and each scale has a specially designed module. Each ESDNet block contains an extended residual dense module and a plug-and-play semantic alignment scale-aware module (SAM). SAM is mainly composed of two modules, namely the pyramid feature extraction module and the cross-scale dynamic fusion module, which effectively capture multi-scale features and dynamically fuse image features by learning fusion weights. These innovations enable ESDNet to achieve performance beyond previous methods.

[0077] The lightweight encoder can adopt a simple architecture composed of three convolutional blocks, and each convolutional block contains a downsampling layer and three layers of convolution. Compared with ESDNet, this design brings a small amount of additional computational cost.

[0078] Furthermore, as shown in Figure 2 The method for aligning the source target image to the moiré W image is as follows:

[0079] First, use the Gluestick algorithm to perform image matching based on key points and key lines to roughly align the position of the source target image to the moiré W image; then use the FlowFormer algorithm to perform pixel-by-pixel alignment of the position of the source target image after rough alignment to the moiré W image to obtain the source target image after position alignment

[0080] Regarding the production of the dataset:

[0081] The data includes various common screen display scenarios, including natural scene images, web content, documents, research papers, and other materials. Moiré patterns with cluttered colors and complex geometric shapes can seriously affect visual perception and are difficult to remove using previous methods. Therefore, this embodiment is more concerned with collecting images under these moiré patterns. In addition, to collect more different moiré patterns, images can be captured at various angles, distances, and lighting conditions, resulting in moiré images with various shapes, scales, and color characteristics.

[0082] In the training dataset of this embodiment, UW images and moiré W images are captured using a mobile phone camera. Three different brands of mobile phones can be used for image capture to form 9 (3×3) combinations to better implement the training of the system. Approximately 1000 samples are collected for each combination. Finally, there are 8959 samples in total. The training set and the test set contain 7176 and 1783 samples respectively. The images in this dataset exhibit more severe moiré patterns. This significantly increases the difficulty of restoring moiré-free content and provides a challenging benchmark for future work.

[0083] During the specific training process, the main screen content can be cropped from the W image to obtain the moiré input I m . Then, the UW image is cropped at the same cropping ratio to obtain the auxiliary input I uw . To better align I m and the source image in space, this embodiment adopts a two-step image alignment.

[0084] In existing image moiré datasets, the brightness and color of the target image are usually the same as those of the source image. However, this may not be appropriate because different display settings (e.g., brightness, contrast, and color adjustment parameters) can bring various appearances of the displayed images. To ensure that the moiré removal method only eliminates moiré and does not change the color representation of the captured images, this embodiment performs additional color correction.

[0085] After position alignment of the source and target images the method for color alignment with the moiré W image is as follows:

[0086] After Gaussian blurring the source and target images and the moiré W image, the least squares method is used to estimate the 3×3 color correction matrix of the source and target images relative to the moiré W image, representing the linear color mapping between I m and ; based on the 3×3 color correction matrix, the source and target images after position alignment are then color-aligned with the moiré W image to obtain the aligned target image GTI gtAfter performing Gaussian blur on the source target image and the moire W image, calculate the color correction matrix, which can reduce the influence of moire in I m and help achieve more accurate color mapping. After alignment, the target image GTI gt can more faithfully reflect the color state of the scene displayed on the screen.

[0087] Facing complex and irregular moire, single-image moire removal methods (such as ESDNet) are difficult to remove them. This embodiment uses the UW image to assist in removing the moire of the W image. The UW image can be easily captured with a smartphone and can provide additional information with relatively normal texture but lower resolution, thereby improving the performance of moire removal for the W image. The spatial alignment between the UW image and the W image is crucial. Some alignment methods, such as optical flow and deformable convolution, perform poorly or are very slow in this task. To solve this problem, this embodiment proposes a fast and effective two-stage alignment method. First, perform alignment at the image level through key point matching, and then perform alignment at the feature level through convolution kernel prediction. Finally, adaptively fuse the aligned features of the UW and W encoders and input them into the decoder. Finally, moire removal of the image is achieved.

[0088] The alignment framework proposed in this embodiment includes two stages, namely the key point matching-based alignment module KMA and the kernel prediction-based alignment module KPA. The first stage performs rough geometric alignment at the image level, and the second stage performs fine-grained alignment at the feature level. The two-stage alignment method can balance efficiency and performance.

[0089] There are significant differences in the fields of view between the UW and W images. In this case, common optical flow-based alignment methods are not cost-effective, that is, achieving good alignment performance may require a high cost. This embodiment performs sparse key point matching between the two images and then aligns them using the matching cues.

[0090] Combined Figure 3 As shown, the method of performing image-level alignment of the UW image in the training set to the corresponding moire W image using the key point matching-based alignment module is as follows:

[0091] Perform sparse key point matching between the two images and then align them using the matching cues:

[0092] Use the SuperPoint method to detect 256 key points from the UW image and the moire W image downsampled by 4 times respectively; use the LightGlue method to match the key points and perform image-level alignment of the UW image to the corresponding moire W image to obtain the aligned UW image:

[0093]

[0094] In the formula is the aligned UW image, KMA represents the alignment module based on key point matching, and I m is the moiré W image, and I uw is the UW image.

[0095] In this embodiment, the number of key points and the downsampling factor are determined through a large number of experiments, and these settings can achieve a better balance between computational efficiency and alignment performance.

[0096] In the second stage, the KPA module further aligns spatially and I m . It is recommended to perform this operation at the multi-scale feature level because multi-scale features can integrate more fine-grained information and reduce the adverse effects of moiré on alignment. Extract the W feature and the UW feature to estimate the transformation convolution weight of F uw .

[0097] As shown in the combination Figure 4 , the method for feature-level alignment of the W image feature and the UW image feature using the alignment module based on kernel prediction is as follows:

[0098] Represent the W image feature as F m , and represent the UW image feature as F uw ;

[0099] Use adaptive average pooling to aggregate the context information of F m and F uw , and then concatenate along the channel dimension to obtain W':

[0100] W' = Concat(AdaptivePool(F m ), AdaptivePool(F uw ))

[0101] In the formula, AdaptivePool represents adaptive average pooling, and COncat represents concatenation;

[0102] Then apply two convolutional layers to process W' to obtain the initial convolutional kernel weight W:

[0103] W = Softmax(Reshape(Conv(ReLU(Conv(W')00))

[0104] In the formula, Softmax is the normalized exponential function, Reshape is the shape adjustment operation, Conv is the convolutional operation, and ReLU is the activation function; K is the convolutional kernel size;

[0105] Perform an element-wise multiplication between the initial convolutional kernel weights W of each weight group dimension in the kernel prediction-based alignment module i and the learnable parameter P i and then sum along the group dimension to obtain the content-adaptive convolutional kernel θ:

[0106]

[0107] G represents the number of weight group dimensions;

[0108] Process the UW image feature F uw using the content-adaptive convolutional kernel θ to obtain the aligned UW image feature

[0109] which can be regarded as a finer alignment result from to I m .

[0110] Furthermore, the method of using the adaptive fusion module for fusion is as follows:

[0111] Fuse the aligned UW image feature and the W image feature F m to obtain the fused feature F fuse :

[0112]

[0113] where α is a learnable coefficient.

[0114] The learnable coefficient is used to integrate information from UW and W features, ensuring that the structure and context details in the W feature are preserved. The learnable coefficient α is adaptively learned during training. This embodiment effectively integrates the information of W features and UW features, ensuring that the basic structure and context details in the W feature are preserved.

[0115] Combined with Figure 3 as shown, the moiré removal network structure includes a two-stage alignment and fusion module (keypoint matching-based alignment module (KMA), kernel prediction-based alignment module (KPA), and adaptive fusion module (AF)), moiré W image encoder E m , UW image encoder E uw and decoder D.

[0116] The keypoint matching-based alignment module (KMA) uses a pre-trained SuperPoint network to detect 256 keypoints from the 4x downsampled image, and then uses the LightGlue network to match the keypoints to obtain the aligned ultra-wide-angle image.

[0117] As an example, the first encoder includes a moiré image encoding unit

[0118] for performing bilinear downsampling operation, first convolution operation and first activation operation on the moiré W image;

[0119] for the output of which is successively subjected to a first residual block, a first semantic alignment perception block, a second convolution operation and a second activation operation;

[0120] for the output of which is successively subjected to a second residual block, a second semantic alignment perception block, a third convolution operation and a third activation operation;

[0121] for the output of which is successively subjected to a third residual block and a third semantic alignment perception block operation.

[0122] The residual block includes a first convolution operation, adding the result after the first activation operation and the input feature to obtain a first feature, a first dilated convolution operation, adding the result after the second activation operation and the first feature to obtain a second feature, a second convolution operation, adding the result after the third activation operation and the second feature, and adding the result after the third convolution operation and the input feature as the output;

[0123] The semantic alignment perception block includes a first dilated convolution module, a second dilated convolution module, a third dilated convolution module, respectively performing average pooling operations on the results of the three dilated convolution modules and then connecting them together, successively performing a first convolution operation, a first activation operation, a second convolution operation, a second activation operation, a third convolution operation, and then splitting along the channels into three parameters, and performing weighted summation on the results of the three dilated convolution modules.

[0124] The dilated convolution module in the semantic alignment perception block is a first convolution operation, a first activation operation, a first dilated convolution operation, a second activation operation, a second dilated convolution operation, a third activation operation, a third dilated convolution operation, a fourth activation operation, a second convolution operation, a fifth activation operation and a third convolution operation.

[0125] The activation operation in the above modules is the ReLU function.

[0126] The first convolution of the first residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0127] The first dilated convolution of the first residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0128] The second convolution of the first residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0129] The third convolution of the first residual block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0130] The first convolution of the second residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0131] The first expansion convolution of the second residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0132] The second convolution of the second residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0133] The third convolution of the second residual block is a convolution with 96 1×1 convolutional kernels and a stride of 1.

[0134] The first convolution of the third residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0135] The first expansion convolution of the third residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0136] The second convolution of the third residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0137] The third convolution of the third residual block is a convolution with 192 1×1 convolutional kernels and a stride of 1.

[0138] The first convolution of the expansion convolution module of the first semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0139] The first expansion convolution of the expansion convolution module of the first semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0140] The second expansion convolution of the expansion convolution module of the first semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and an expansion rate of 3.

[0141] The third expansion convolution of the expansion convolution module of the first semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0142] The second convolution of the expansion convolution module of the first semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0143] The third convolution of the extended convolution module of the first semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0144] The first convolution of the first semantic alignment perception block is a convolution with 36 1×1 convolutional kernels and a stride of 1.

[0145] The second convolution of the first semantic alignment perception block is a convolution with 36 1×1 convolutional kernels and a stride of 1.

[0146] The third convolution of the first semantic alignment perception block is a convolution with 144 1×1 convolutional kernels and a stride of 1.

[0147] The first convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0148] The first extended convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0149] The second extended convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and an expansion rate of 3.

[0150] The third extended convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0151] The second convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0152] The third convolution of the extended convolution module of the second semantic alignment perception block is a convolution with 96 1×1 convolutional kernels and a stride of 1.

[0153] The first convolution of the second semantic alignment perception block is a convolution with 72 1×1 convolutional kernels and a stride of 1.

[0154] The second convolution of the second semantic alignment perception block is a convolution with 72 1×1 convolutional kernels and a stride of 1.

[0155] The third convolution of the second semantic alignment perception block is a convolution with 288 1×1 convolutional kernels and a stride of 1.

[0156] The first convolution of the extended convolution module of the third semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0157] The first extended convolution of the extended convolution module of the third semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0158] The second dilated convolution of the dilated convolution module of the third semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and a dilation rate of 3.

[0159] The third dilated convolution of the dilated convolution module of the third semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0160] The second convolution of the dilated convolution module of the third semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0161] The third convolution of the dilated convolution module of the third semantic alignment perception block is a convolution with 192 1×1 convolutional kernels, a stride of 1.

[0162] The first convolution of the third semantic alignment perception block is a convolution with 144 1×1 convolutional kernels, a stride of 1.

[0163] The second convolution of the third semantic alignment perception block is a convolution with 144 1×1 convolutional kernels, a stride of 1.

[0164] The third convolution of the third semantic alignment perception block is a convolution with 576 1×1 convolutional kernels, a stride of 1.

[0165] The first convolution in the moiré image encoder is a convolution with 48 5×5 convolutional kernels, a stride of 1, and a padding of 2.

[0166] The second convolution in the moiré image encoder is a convolution with 96 3×3 convolutional kernels, a stride of 2, and a padding of 1.

[0167] The third convolution in the moiré image encoder is a convolution with 192 5×5 convolutional kernels, a stride of 1, and a padding of 2.

[0168] As an example, the second encoder includes an ultra-wide-angle image encoding unit

[0169] for performing bilinear downsampling operation, first convolution operation, and first activation operation on the aligned UW image;

[0170] for successively performing a second convolution operation, a second activation operation, a third convolution operation, a third activation operation, a fourth convolution operation, a fifth convolution operation, and a fourth activation operation on the output of

[0171] for The output is successively subjected to the sixth convolution operation, the fifth activation operation, the seventh convolution operation, the sixth activation operation, the eighth convolution operation, the ninth convolution operation, and the seventh activation operation;

[0172] For The output is successively subjected to the tenth convolution operation, the eighth activation operation, the eleventh convolution operation, the ninth activation operation, and the twelfth convolution operation.

[0173] The activation operations in the above modules are all ReLU functions.

[0174] The first convolution in the ultra-wide-angle image encoder is a convolution with 48 5×5 convolution kernels, a stride of 1, and a padding of 2.

[0175] The second convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0176] The third convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0177] The fourth convolution in the ultra-wide-angle image encoder is a convolution with 48 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0178] The fifth convolution in the ultra-wide-angle image encoder is a convolution with 96 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0179] The sixth convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0180] The seventh convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0181] The eighth convolution in the ultra-wide-angle image encoder is a convolution with 96 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0182] The ninth convolution in the ultra-wide-angle image encoder is a convolution with 192 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0183] The tenth convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0184] The eleventh convolution in the ultra-wide-angle image encoder is a convolution with 32 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0185] The twelfth convolution in the ultra-wide-angle image encoder is a convolution with 192 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0186] As an example, the alignment module based on kernel prediction includes alignment units KPA1 to KPA3:

[0187] KPA1 is used to and perform max pooling on the outputs of respectively, and concatenate the features after max pooling, and then perform a first convolution operation, a first activation operation, a second convolution operation, a second activation operation, and a first group sum operation in sequence. The result obtained is used as the feature kernel, and then perform a convolution operation on the output of;

[0188] KPA2 is used to and perform max pooling on the outputs of respectively, and concatenate the features after max pooling, and then perform a third convolution operation, a third activation operation, a fourth convolution operation, a fourth activation operation, and a second group sum operation in sequence. The result obtained is used as the feature kernel, and then perform a convolution operation on the output of;

[0189] KPA3 is used to and perform max pooling on the outputs of respectively, and concatenate the features after max pooling, and then perform a fifth convolution operation, a fifth activation operation, a sixth convolution operation, a sixth activation operation, and a third group sum operation in sequence. The result obtained is used as the feature kernel, and then perform a convolution operation on the output of;

[0190] The first, third, and fifth activation operations in the above module are all GELU functions.

[0191] The second, fourth, and sixth activation operations in the above module are all Softmax functions.

[0192] The first convolution in the alignment module based on kernel prediction is a convolution with 12 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0193] The second convolution in the alignment module based on kernel prediction is a convolution with 96 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0194] The third convolution in the alignment module based on kernel prediction is a convolution with 24 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0195] The second convolution in the alignment module based on kernel prediction is a convolution with 192 3×3 convolution kernels, a stride of 1, and a padding of 1.

[0196] The first convolution in the alignment module based on kernel prediction is a convolution with 48 1×1 convolution kernels, a stride of 1, and a padding of 1.

[0197] The second convolution in the kernel prediction-based alignment module is a convolution with 384 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0198] The adaptive fusion module includes fusion units AF1 to AF3:

[0199] AF1 is used to add the outputs of

[0200] and KPA1 after multiplying them by the first coefficient respectively; AF2 is used to add the outputs of

[0201] and KPA2 after multiplying them by the second coefficient respectively; AF3 is used to add the outputs of

[0202] and KPA3 after multiplying them by the third coefficient respectively;

[0203] The three output results of the fusion units AF1 to AF3 are decoded by the decoder to obtain the restored image.

[0204] The decoder includes decoding units D1 to D3, and the operations are specifically represented as follows:

[0205] D3 is used to perform a first convolution operation, a first activation operation, a fourth residual block, a fourth semantic alignment perception block, a second convolution operation, and a bilinear upsampling operation on the output of AF3 in sequence;

[0206] D2 is used to perform a third convolution operation, a second activation operation, a fifth residual block, a fifth semantic alignment perception block, a fourth convolution operation, and a bilinear upsampling operation on the concatenated output of AF2 and D3 in sequence;

[0207] D1 is used to perform a fifth convolution operation, a third activation operation, a sixth residual block, a sixth semantic alignment perception block, a sixth convolution operation, and a bilinear upsampling operation on the concatenated output of AF1 and D2 in sequence;

[0208] The activation operations in the above modules are all ReLU functions.

[0209] The first convolution of the fourth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0210] The first expansion convolution of the fourth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and an expansion rate of 2.

[0211] The second convolution of the fourth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0212] The third convolution of the fourth residual block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0213] The first convolution of the fifth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0214] The first dilated convolution of the fifth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0215] The second convolution of the fifth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0216] The third convolution of the fifth residual block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0217] The first convolution of the sixth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0218] The first dilated convolution of the sixth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0219] The second convolution of the sixth residual block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0220] The third convolution of the sixth residual block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0221] The first convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0222] The first dilated convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0223] The second dilated convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and a dilation rate of 3.

[0224] The third dilated convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0225] The second convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0226] The third convolution of the extended convolution module of the fourth semantic alignment perception block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0227] The first convolution of the fourth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0228] The second convolution of the fourth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0229] The third convolution of the fourth semantic alignment perception block is a convolution with 192 1×1 convolutional kernels and a stride of 1.

[0230] The first convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0231] The first extended convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0232] The second extended convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and a dilation rate of 3.

[0233] The third extended convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0234] The second convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0235] The third convolution of the extended convolution module of the fifth semantic alignment perception block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0236] The first convolution of the fifth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0237] The second convolution of the fifth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0238] The third convolution of the fifth semantic alignment perception block is a convolution with 192 1×1 convolutional kernels and a stride of 1.

[0239] The first convolution of the extended convolution module of the sixth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0240] The first extended convolution of the extended convolution module of the sixth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0241] The second dilated convolution of the dilated convolution module in the sixth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 3, and a dilation rate of 3.

[0242] The third dilated convolution of the dilated convolution module in the sixth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, a padding of 2, and a dilation rate of 2.

[0243] The second convolution of the dilated convolution module in the sixth semantic alignment perception block is a convolution with 32 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0244] The third convolution of the dilated convolution module in the sixth semantic alignment perception block is a convolution with 64 1×1 convolutional kernels and a stride of 1.

[0245] The first convolution of the sixth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0246] The second convolution of the sixth semantic alignment perception block is a convolution with 48 1×1 convolutional kernels and a stride of 1.

[0247] The third convolution of the sixth semantic alignment perception block is a convolution with 192 1×1 convolutional kernels and a stride of 1.

[0248] The first convolution in the decoder is a convolution with 64 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0249] The second convolution in the decoder is a convolution with 12 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0250] The third convolution in the decoder is a convolution with 64 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0251] The fourth convolution in the decoder is a convolution with 12 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0252] The fifth convolution in the decoder is a convolution with 64 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0253] The sixth convolution in the decoder is a convolution with 12 3×3 convolutional kernels, a stride of 1, and a padding of 1.

[0254] This embodiment adopts a multi-scale loss function to enhance the network's ability to eliminate multi-scale moiré patterns. It is applied by a multi-scale supervision mechanism to constrain the network to learn consistent features at different spatial scales.

[0255]

[0256] where is the multi-scale loss function, Denote the i-th restored image, Denote the 2-downsampled version of the i-th aligned target image, i-1 where λ p is the weight of the perceptual loss, and φ(·) represents the features extracted by the VGG network.

[0257] In the training stage, the Adam optimizer is used to update the network parameters.

[0258] The restoration network described in this embodiment can be deployed on different imaging devices, which helps to obtain high-resolution images that are cleaner, have a more realistic tone, and are richer in details. Compared with the current mainstream moiré removal methods, this implementation achieves good restoration performance and has a significant improvement in visual effects. This method uses an ultra-wide-angle lens, showing that ultra-wide-angle images can provide valuable information and help effectively remove moiré patterns in wide-angle photos.

[0259] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.

Claims

1. A method for removing moiré patterns from a wide-angle image using an ultra-wide-angle image, characterized in that include, Obtain a UW image, a moiré W image, and a source target image, and perform position alignment and color alignment on the source target image with the moiré W image to obtain an aligned target image; The training dataset is constructed by using UW images, moiré W images and aligned target images; The alignment module based on key point matching is used to align the UW image in the training set with the corresponding moiré W image at the image level to obtain the aligned UW image; The moiré W image is encoded by the first encoder to obtain the W image feature; the aligned UW image is encoded by the second encoder to obtain the UW image feature; The kernel prediction-based alignment module is used to align the W image features and the UW image features at the feature level, and then the adaptive fusion module is used to fuse them to obtain the fused features; Then, a decoder is used to decode the fused features to obtain a restored image; A loss function is calculated based on the aligned target image and the restored image, and network parameters of the first encoder, the second encoder, the alignment module based on kernel prediction, and the decoder are adjusted until the loss function value is within a preset threshold, thereby obtaining a restoration network consisting of the alignment module based on key point matching and the trained first encoder, the second encoder, the alignment module based on kernel prediction, and the decoder; The restoration network is used to process the actual UW image and the actual moiré W image to obtain a restored image of the actual moiré W image after removing the moiré.

2. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 1, characterized in that: The first encoder adopts the encoder of the ESDNet network; the second encoder is a lightweight encoder; and the decoder adopts the decoder of the ESDNet network.

3. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 1, characterized in that: The method for aligning the source target image to the moiré W image is: First, the Gluestick algorithm is used to match the images according to the key points and key lines, and the source target image is roughly aligned with the moiré W image; then the FlowFormer algorithm is used to align the roughly aligned source target image with the moiré W image pixel by pixel to obtain the aligned source target image.

4. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 3, characterized in that: Source and target images after position alignment The method for color alignment to the moiré W image is: After Gaussian blurring the source target image and the moiré W image, the 3×3 color correction matrix of the source target image relative to the moiré W image is estimated using the least squares method. The source target image is aligned based on the 3×3 color correction matrix. Then perform color alignment on the moiré image W to obtain the aligned target image GTI gt .

5. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 4, characterized in that: The method of using the key point matching-based alignment module to align the UW image in the training set to the corresponding moiré W image at the image level is as follows: The SuperPoint method is used to detect 256 key points from the 4-fold downsampled UW image and the moiré W image respectively; the LightGlue method is used to match the key points, and the UW image is aligned to the corresponding moiré W image at the image level to obtain the aligned UW image: In the formula is the aligned UW image, KMA represents the key point matching based alignment module, I m is the moiré W image, I uw Image for UW.

6. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 5, characterized in that: The method of using the kernel prediction-based alignment module to perform feature-level alignment on W image features and UW image features is as follows: Denote the W image feature as F m , the UW image feature is represented as F uw ; Adaptive average pooling aggregation F m and F uw The context information is then connected along the channel dimension to obtain W′: W′=Concat(AdaptivePool(F m ),AdaptivePool(F uw )), In the formula, AdaptivePool means adaptive average pooling, and Concat means concatenation; Apply two more convolutional layers to process W′ and obtain the initial convolution kernel weight W: W=Softmax(Reshape(Conv(ReLU(Conv(W′))))), Where Softmax is the normalized exponential function, Reshape is the shape adjustment operation, Conv is the convolution operation, and ReLU is the activation function; The initial convolution kernel weight W of each weight group dimension in the kernel prediction based alignment module i and learnable parameters P i , and then sum along the group dimension to obtain the content-adaptive convolution kernel θ: G represents the number of weight group dimensions; Based on the content adaptive convolution kernel θ, the UW image feature F uw Processing is performed to obtain the aligned UW image features 7. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 6, characterized in that: The method of using the adaptive fusion module for fusion is: UW image features after alignment and W image features F m Fusion is performed to obtain the fusion feature F fuse : Where α is the learnable coefficient.

8. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 7, characterized in that: The first encoder includes a moiré image encoding unit Used to perform a bilinear downsampling operation, a first convolution operation, and a first activation operation on the moiré W image; Used for The output of is sequentially subjected to a first residual block, a first semantic alignment perception block, a second convolution operation, and a second activation operation; Used for The output of is sequentially subjected to a second residual block, a second semantic alignment perception block, a third convolution operation, and a third activation operation; Used for The output of is sequentially subjected to the third residual block and the third semantic alignment perception block operations.

9. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 8, characterized in that: The second encoder includes an ultra-wide-angle image encoding unit Used to perform a bilinear downsampling operation, a first convolution operation, and a first activation operation on the aligned UW image; Used for The output of is sequentially subjected to a second convolution operation, a second activation operation, a third convolution operation, a third activation operation, a fourth convolution operation, a fifth convolution operation, and a fourth activation operation; Used for The output of is sequentially subjected to a sixth convolution operation, a fifth activation operation, a seventh convolution operation, a sixth activation operation, an eighth convolution operation, a ninth convolution operation, and a seventh activation operation; Used for The output of is subjected to the tenth convolution operation, the eighth activation operation, the eleventh convolution operation, the ninth activation operation, and the twelfth convolution operation in sequence.

10. The method for removing moiré from a wide-angle image using an ultra-wide-angle image as claimed in claim 9, characterized in that: The alignment module based on kernel prediction includes alignment units KPA1 to KPA3: KPA1 is used for and The outputs of are respectively subjected to maximum pooling, and the features after maximum pooling are connected together, and the first convolution operation, the first activation operation, the second convolution operation, the second activation operation and the first group summation operation are performed in sequence. The results are used as the feature kernel, and then The output of is convolved; KPA2 is used for and The outputs of are respectively subjected to maximum pooling, and the features after maximum pooling are connected together, and the third convolution operation, the third activation operation, the fourth convolution operation, the fourth activation operation and the second group of summation operations are performed in sequence. The results are used as feature kernels, and then The output of is convolved; KPA3 is used for and The outputs of are respectively max-pooled, and the features after the max-pooling are connected together, and the fifth convolution operation, the fifth activation operation, the sixth convolution operation, the sixth activation operation and the third group of summation operations are performed in sequence. The results are used as the feature kernel, and then The output of is convolved; The adaptive fusion module includes fusion units AF1 to AF3: AF1 is used for The outputs of and KPA1 are multiplied by the first coefficient and then added; AF2 is used for The outputs of KPA1 and KPA2 are respectively multiplied by the second coefficient and then added; AF3 is used for The outputs of KPA3 and KPA4 are respectively multiplied by the third coefficient and then added; The three output results of the fusion units AF1~AF3 are decoded to obtain the restored image.