Multispectral night vision device intelligent image enhancement system and method

Through frequency domain analysis combining multi-level discrete wavelet transform and two-dimensional fast Fourier transform, combined with a lightweight guiding sub-network and a correction network of U-Net architecture, the problem of artifacts in the generative adversarial network is solved, and efficient and non-invasive image enhancement effects are achieved.

CN120689229AActive Publication Date: 2025-09-23SHANGHAI YUFENG ELECTRONIC INFORMATION TECH DEV CO LTD

Patent Information

Application Number
CN202511204325.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-23
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing generative adversarial network models are prone to produce structured artifacts when generating high-resolution images. Traditional methods find it difficult to effectively remove these artifacts without affecting the image's realism and detail fidelity.

Method used

A frequency domain analysis method combining multi-level discrete wavelet transform and two-dimensional fast Fourier transform is used to generate an artifact signature map. Artifact correction is performed through a lightweight guidance subnetwork and a correction network with a U-Net architecture, and the final corrected image is generated using the residual learning paradigm.

Benefits of technology

Without modifying the existing model, artifacts are effectively removed, the high-frequency details and global structure of the image are preserved, the reliability and quality of the generated images are improved, and the deployment risk and cost are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689229A_ABST
    Figure CN120689229A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent image enhancement system and method for a multispectral night vision device, and relates to the technical field of image enhancement, and the method comprises the steps: firstly decomposing an input generative adversarial network fusion image containing artifacts into sub-band images with different frequencies through the multi-stage discrete wavelet transform; a two-dimensional fast Fourier transform and peak detection algorithm is applied to generate an artifact signature graph capable of accurately representing the spatial position of the artifact; then, the artifact signature graph is refined into a smooth space attention graph by a lightweight guide sub-network; the space attention map guides a lightweight correction network based on a U-Net architecture, so that the correction capability is focused on an artifact area; the correction network follows a residual learning normal form, only one artifact residual image is output, finally, the artifact residual image and the generative adversarial network fusion image containing the artifacts are added pixel by pixel, and a final correction image is obtained. And the fidelity and the credibility of the output result of the existing complex GAN model are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image enhancement technology, and more particularly to an intelligent image enhancement system and method for a multi-spectral night vision device. Background Art

[0002] Generative Adversarial Networks (GANs) have become a fundamental and powerful technology in modern computer vision and image processing, achieving breakthroughs in tasks such as high-quality image synthesis, cross-domain image translation, and multimodal data fusion. The core mechanism of GANs stems from their unique adversarial training framework, in which a generator network and a discriminator network compete and co-evolve in a zero-sum game, ultimately aiming to enable the generator to learn and reproduce the complex distribution of real data. However, ensuring the stability of this adversarial process is extremely challenging, especially when building deep and complex network architectures for generating high-resolution images. A widely recognized and intensively studied systemic challenge is the frequent presence of various "generative artifacts" in the output images of GANs. These artifacts are not simple random noise but rather structured byproducts of specific structural components of the network, often manifesting as unnatural periodic textures, checkerboard patterns, or subtle color distortions. For existing large and complex multispectral fusion GAN models that have already invested a lot of computing resources to train, it is often not feasible to eradicate these artifacts by directly modifying their core architecture. Therefore, developing independent, efficient and non-invasive post-processing techniques to improve and ensure the final fidelity and reliability of generated images has become a crucial research direction.

[0003] In view of this, the present invention proposes a multi-spectral night vision device intelligent image enhancement system and method to solve the above problems. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solution: a multispectral night vision device intelligent image enhancement method, comprising:

[0005] Step F1, based on a preset discrete wavelet transform wavelet type and decomposition level, performing a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image, and combining it with a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate the periodic frequency components caused by the generative artifacts, and generating a single-channel artifact signature map having the same spatial size as the artifact-containing generative adversarial network fusion image;

[0006] Step F2: Input the artifact signature map and the preset guiding sub-network weights, guiding sub-network convolution kernel size, and guiding sub-network output layer convolution kernel size into the guiding sub-network, perform smoothing and nonlinear transformation, and generate a spatial attention map with a value range between [0, 1].

[0007] Step F3, using the spatial attention map and based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module, the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights is guided weighted, and the artifact-containing generative adversarial network fusion image is input to the lightweight correction network to generate an artifact residual map through the residual learning paradigm;

[0008] In step F4, the artifact residual image is added pixel by pixel to the artifact-containing generative adversarial network fusion image to obtain the final corrected image.

[0009] Furthermore, the implementation method of step F1 includes:

[0010] Step F1F1, performing a multi-level two-dimensional discrete wavelet transform based on a wavelet type with a number of layers specified by the decomposition level on the generative adversarial network fusion image containing artifacts to obtain a final low-frequency approximate subband and a set of multi-level high-frequency detail subbands;

[0011] Steps F1 and F2: traverse each high-frequency sub-band image in the multi-level high-frequency detail sub-band set, apply a two-dimensional fast Fourier transform to perform spectrum analysis, and identify abnormal energy peaks using a peak detection algorithm based on a fast Fourier transform peak detection threshold, generating a corresponding sub-band artifact score map for each high-frequency sub-band image;

[0012] In steps F1 and F3, all sub-band artifact score maps are upsampled to the original size of the artifact-containing generative adversarial network fusion image and then accumulated, and the accumulated results are normalized to generate a single-channel artifact signature map.

[0013] Furthermore, the implementation method of step F3 includes:

[0014] Step F3F1, inputting the artifact-containing GAN fused image into the encoder path of a lightweight correction network based on a U-Net architecture and loaded with lightweight correction network weights, wherein the encoder path generates a multi-scale encoder feature map and a bottleneck layer feature map by performing convolution and max pooling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module;

[0015] Steps F3F2: Input the bottleneck layer feature map and the multi-scale encoder feature map into the decoder path loaded with the lightweight correction network weights. Under the guidance of the spatial attention map, the decoder path generates the network raw output by performing convolution and upsampling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module.

[0016] In step F3F3, the original output of the network is defined as the artifact residual map.

[0017] Furthermore,

[0018] The implementation method of step F1F1 includes: step F1F1F1, initializing the generative adversarial network fusion image containing artifacts as the current approximate subband; step F1F1F2, initializing an empty multi-level high-frequency detail subband set; step F1F1F3, when the number of loops is less than the decomposition level, performing a single-level two-dimensional discrete wavelet transform on the current approximate subband based on the wavelet type to obtain a new next-level approximate subband and a current-level high-frequency detail subband including three directions of horizontal, vertical, and diagonal, and updating the next-level approximate subband to the current approximate subband, and adding the current-level high-frequency detail subband to the multi-level high-frequency detail subband set; step F1F1F4, naming the current approximate subband at the end of the loop as the final low-frequency approximate subband; the implementation method of step F1F1F3 includes: if the number of loops is less than the decomposition level, looping through steps F1F1F3F1 to F1F1F3F6: In step F1F1F3F1, using the wavelet and scaling filter corresponding to the wavelet type, first perform a one-dimensional discrete wavelet transform along the rows of the current approximate subband to obtain the result, and then perform a one-dimensional discrete wavelet transform along the columns of the result to obtain the low-frequency approximate subband LL, the horizontal high-frequency detail subband LH, the vertical high-frequency detail subband HL and the diagonal high-frequency detail subband HH; in step F1F1F3F2, assign the LL subband to the next-level approximate subband; in step F1F1F3F3, combine the three subbands LH, HL and HH into the current-level high-frequency detail subband; in step F1F1F3F4, update the next-level approximate subband to the current approximate subband; in step F1F1F3F5, add the current-level high-frequency detail subband to the multi-level high-frequency detail subband set; in step F1F1F3F6, increase the number of loops by one.

[0019] Furthermore, the implementation method of steps F1 and F2 includes:

[0020] Step F1F2F1, initialize the set of all empty sub-band artifact score maps;

[0021] Steps F1, F2, and F2: for each high-frequency subband image in the multi-level high-frequency detail subband set, perform a two-dimensional fast Fourier transform on it to obtain a complex spectrum; calculate the amplitude spectrum of the complex spectrum to obtain an original amplitude spectrum; perform a quadrant shift on the original amplitude spectrum to move the zero-frequency component of the original amplitude spectrum to the center of the spectrum to obtain a centralized amplitude spectrum;

[0022] Steps F1, F2, and F3: Apply a two-dimensional local maximum detection algorithm to the centralized amplitude spectrum and perform screening in combination with a fast Fourier transform peak detection threshold to obtain a list of artifact energy peak coordinates;

[0023] Steps F1, F2, and F4: generating a sub-band artifact score map for the high-frequency sub-band image based on the amplitude and position of each peak in the artifact energy peak coordinate list, and adding the sub-band artifact score map to the set of all sub-band artifact score maps.

[0024] Furthermore, the implementation method of steps F1 and F3 includes:

[0025] Step F1F3F1, initialize the cumulative artifact score map with the same size as the artifact-containing generative adversarial network fusion image and all elements are zero;

[0026] Steps F1, F3, and F2, traverse each sub-band artifact score map in the set of all sub-band artifact score maps;

[0027] Steps F1, F3, and F3: upsampling the subband artifact score map to the original size of the artifact-containing generative adversarial network fusion image through bilinear interpolation according to the decomposition level corresponding to the current subband artifact score map, thereby obtaining an upsampled score map;

[0028] Steps F1, F3, and F4: add the upsampled score map to the cumulative artifact score map pixel by pixel, and update the cumulative artifact score map;

[0029] In steps F1, F3, and F5, after the traversal is completed, the cumulative artifact score map is normalized using the minimum-maximum method so that all its pixel values ​​fall within the interval [0, 1] to obtain the artifact signature map.

[0030] Furthermore, the implementation method of steps F3 and F1 includes:

[0031] Step F3F1F1, constructing a U-Net encoder path and loading the corresponding weight parameters in the lightweight correction network weights for the U-Net encoder path; the overall architecture of the U-Net encoder path and each encoder module contained therein are defined by the lightweight correction network weights;

[0032] Steps F3, F1, and F2 take the artifact-containing generative adversarial network fusion image as the initial input and pass it through each module of the encoder path in sequence. In each module, a convolution operation using the convolution kernel size of the correction network is performed times the number of convolution layers in the correction network module and a maximum pooling operation using the correction network pooling and upsampling kernel size is performed once. The output of each module after convolution is saved as an element of the multi-scale encoder feature map.

[0033] In steps F3F1F3, the output of the deepest layer of the encoder path is defined as the bottleneck layer feature map.

[0034] Furthermore, the implementation method of steps F3 and F2 includes:

[0035] Steps F3, F2, and F1 construct a U-Net decoder path and load the corresponding weight parameters in the lightweight correction network weights into the U-Net decoder path; the overall architecture of the U-Net decoder path and each decoder module it contains are defined by the lightweight correction network weights;

[0036] Steps F3F2F2 take the bottleneck layer feature map as the initial input to the decoder path and pass it through each decoder module in turn. In each decoder module, the input feature map is upsampled, concatenated, and convolved with the guidance of the corresponding level feature map and spatial attention map from the multi-scale encoder feature map to generate a weighted feature map as the input to the next decoder module.

[0037] In steps F3, F2, and F3, the weighted feature map output by the last decoder module is processed by a 1×1 convolution layer using a corrected network convolution kernel size to obtain the original network output.

[0038] Furthermore, the specific implementation method of steps F3F2F2 includes:

[0039] In step F3F2F2F1, a transposed convolution upsampling operation is performed on the output of the previous decoder module, and it is concatenated with the corresponding level feature map from the multi-scale encoder feature map in the channel dimension to obtain the concatenated feature map;

[0040] Step F3F2F2F2, performing convolution operations on the spliced ​​feature map equal to the number of convolution layers in the correction network module, and applying a ReLU activation function after each convolution operation to obtain the output feature map of the current decoder module;

[0041] Step F3F2F2F3, downsample the spatial attention map to the same size as the output feature map of the current decoder module through a bilinear interpolation algorithm to obtain a same-scale attention map, and multiply the same-scale attention map with the output feature map of the current decoder module element by element to obtain the weighted feature map.

[0042] Furthermore,

[0043] The implementation method of step F1F2F2 includes: step F1F2F2F1, performing a two-dimensional fast Fourier transform on the high-frequency sub-band image to obtain a complex spectrum, wherein the calculation method of the transformation includes: for each coordinate point in the frequency domain, traversing all pixel points of the high-frequency sub-band image, multiplying the pixel value of each pixel point with a complex exponential basis function and then accumulating and summing them, wherein the complex exponential basis function is generated by respectively calculating the product of the x-axis coordinate of the coordinate point in the frequency domain and the x-axis coordinate of the pixel point of the high-frequency sub-band image, and the frequency domain The product of the y-axis coordinate of the coordinate point in the image and the y-axis coordinate of the pixel point of the high-frequency sub-band image; normalizing the two products by the corresponding dimensions of the high-frequency sub-band image and then summing them to obtain a sum value; taking the exponent of the result of multiplying the sum value by a preset complex constant; step F1F2F2F2, calculating the absolute value of each element in the complex spectrum to obtain an original amplitude spectrum; step F1F2F2F3, performing a quadrant shift operation on the original amplitude spectrum, by diagonally exchanging the four quadrants of the original amplitude spectrum so that the zero-frequency component is located at the center of the image, to obtain a centralized amplitude spectrum;

[0044] The implementation method of step F1F2F3 includes: step F1F2F3F1, calling the standard two-dimensional peak search algorithm, analyzing the centralized amplitude spectrum to identify all local maxima, and calculating the peak prominence of each maximum value relative to its surrounding background, to obtain a candidate peak list containing the coordinates of each peak and its corresponding prominence value; step F1F2F3F2, screening the candidate peak list according to the condition that the peak prominence is greater than the fast Fourier transform peak detection threshold, to obtain a significant energy peak list; step F1F2F3F3, extracting only the coordinates of all peaks from the significant energy peak list to form an artifact energy peak coordinate list.

[0045] A multispectral night vision device intelligent image enhancement system implements the multispectral night vision device intelligent image enhancement method. The system includes:

[0046] an artifact signature module that performs a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image based on a preset discrete wavelet transform wavelet type and decomposition level, and combines a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate periodic frequency components caused by generative artifacts, thereby generating a single-channel artifact signature map having the same spatial size as the artifact-containing generative adversarial network fusion image;

[0047] The attention module inputs the artifact signature map and the preset guide subnetwork weights, guide subnetwork convolution kernel size, and guide subnetwork output layer convolution kernel size into the guide subnetwork, performs smoothing and nonlinear transformation, and generates a spatial attention map with a value range between [0, 1].

[0048] The artifact residual module uses a spatial attention map and performs guided weighting on the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolutional layers in the correction network module. The artifact residual map is generated by inputting the artifact-containing generative adversarial network fusion image into the lightweight correction network through the residual learning paradigm;

[0049] The correction module adds the artifact residual map and the artifact-containing generative adversarial network fusion image pixel by pixel to obtain the final corrected image.

[0050] Beneficial effects of the multi-spectral night vision device intelligent image enhancement system and method of the present invention:

[0051] The present invention generates an artifact signature map by adopting a hybrid domain analysis method that combines a multi-level discrete wavelet transform (DWT) with a two-dimensional fast Fourier transform (FFT). This feature has the effect of effectively separating the high-frequency details containing artifacts from the main image structure by utilizing the multi-resolution characteristics of the DWT, and then accurately identifying and locating abnormal energy spikes caused by generative artifacts in the frequency domain by utilizing the high sensitivity of the FFT to periodic signals. Therefore, this method can generate a "signature map" that accurately characterizes the spatial position and intensity of the artifact, providing a high-precision guidance basis for subsequent targeted correction, and solving the problem that traditional methods have difficulty in accurately distinguishing artifacts from real image details.

[0052] This paper designs a lightweight guidance subnetwork that smooths and nonlinearly transforms the artifact signature map generated in the previous step to generate a spatial attention map. This feature optimizes the sparse and discontinuous artifact detection results in the original signature map into a spatially smooth and semantically coherent weight mask. This effectively avoids the introduction of new edge mutations or visual artifacts caused by the uneven guidance signal during the correction process, ensuring a natural and smooth transition in the final correction effect.

[0053] This paper adopts a residual learning paradigm and utilizes the aforementioned spatial attention map to guide a lightweight correction network based on the U-Net architecture. This feature results in a lightweight correction network based on the U-Net architecture. By weighting the U-Net decoder feature map with the spatial attention map, the network's correction capability is surgically focused on artifact regions, avoiding unnecessary modifications to artifact-free areas. This design successfully eliminates artifacts while preserving all legitimate, high-frequency details and global structure in the original image with maximum fidelity.

[0054] The overall approach of this invention is a non-invasive post-processing module that obtains the final result by pixel-by-pixel adding the artifact residual map output by the network to the original image containing the artifacts. This feature has the effect of improving the quality of the output image without modifying or retraining any existing complex GAN models. This makes it a modular, computationally efficient, and easy-to-integrate quality assurance layer, significantly enhancing the reliability and trustworthiness of generative AI outputs and reducing the cost and risk of deploying "high-potential" but flawed models into "production-grade" applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Schematic diagram of the method flow of the intelligent image enhancement method for multi-spectral night vision device of the present invention;

[0056] Figure 2 Schematic diagram of the system modules of the multi-spectral night vision device intelligent image enhancement system of the present invention;

[0057] Figure 3 Schematic diagram of the application scenario of the multispectral night vision device intelligent image enhancement method of the present invention. DETAILED DESCRIPTION

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0059] Example 1

[0060] See Figure 1 The multi-spectral night vision device intelligent image enhancement method described in this embodiment includes: steps F1 to F4.

[0061] This section provides a detailed, step-by-step technical implementation of the entire artifact correction process. The entire process is logically divided into two core stages. The core task of the first stage is to generate an artifact signature map that accurately characterizes the spatial distribution of artifacts by performing a hybrid domain processing of the input image that combines multiresolution analysis with frequency domain analysis. The core task of the second stage is to use this signature map to guide a lightweight correction network based on residual learning to generate and apply the artifact residual map in a targeted manner, ultimately producing an artifact-free final corrected image.

[0062] Step F1, based on the preset wavelet type and decomposition level of the discrete wavelet transform, performs a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image, and combines it with a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate the periodic frequency components caused by the generative artifacts, and generates a single-channel artifact signature map with the same spatial size as the artifact-containing generative adversarial network fusion image.

[0063] The artifact-containing generative adversarial network fusion image is a three-channel RGB uint8 array with pixel values ​​ranging from 0 to 255. This data comes from the direct output of a trained generative adversarial network (GAN) that exists as an independent third-party module and is specifically designed for multispectral image fusion tasks.

[0064] The wavelet type is a string whose value is hard-coded to 'db4'. This value was determined based on offline comparative experiments comparing the performance of various standard mother wavelets on image artifact analysis tasks. The db4 wavelet was chosen because it strikes the best balance between good time-frequency localization and computational efficiency. Its smoothness more effectively captures the periodic high-frequency characteristics characteristic of GAN artifacts than the non-overlapping Haar wavelet, while avoiding unnecessary blocking artifacts introduced by the inherent non-smoothness of the analysis tool.

[0065] The decomposition level is an integer hard-coded to 3. This value was determined based on an engineering trade-off analysis aimed at balancing frequency separation capability with spatial localization accuracy. The choice of 3 levels of decomposition is based on the fact that this decomposition depth is sufficient to effectively separate the image signal into frequency components of different scales, thereby separating artifact characteristics from the image's macroscopic structure, while ensuring that the highest-frequency detail subbands still retain sufficient spatial resolution to accurately localize the spatial location of artifacts in the original image.

[0066] The Fast Fourier Transform (FFT) peak detection threshold is a floating-point number determined by offline statistical analysis of a representative sample set of images containing typical GAN-generated artifacts. This analysis involves performing a two-dimensional Fast Fourier Transform (FFT) on the high-frequency subband of each image in the sample set to calculate its magnitude spectrum. Next, for each local energy peak in the spectrum, its prominence relative to the local background noise is calculated. Finally, by analyzing the statistical distribution of the prominence of the two types of peaks, artifact energy spikes and background spectral noise, a value is set that can distinguish them with high confidence, serving as the final FFT peak detection threshold.

[0067] The design of step F1 is based on a core physical insight: the generative artifacts introduced by the upsampling operation in the GAN model often manifest as unnatural, periodic, fine-grained textures with specific frequencies and directions. According to Fourier theory, periodic signals in the spatial domain correspond to discrete, concentrated energy spikes in the frequency domain. Therefore, this step adopts a dual analysis strategy: First, the multiresolution analysis capabilities of the discrete wavelet transform (DWT) are leveraged to effectively separate the image's macroscopic structure from the microscopic details containing the artifacts. Subsequently, the high sensitivity of the two-dimensional fast Fourier transform (FFT) to global periodic patterns is exploited to detect the artifact-induced anomalous energy spikes in the spectrum of each high-frequency subband, thereby achieving precise identification and spatial localization of the artifacts. This hybrid analysis method, combining the localization capabilities of the wavelet transform and the Fourier transform in the time-frequency domain, achieves complementary advantages.

[0068] The implementation method of step F1 includes the following steps: Step F1F1, performing a multi-level two-dimensional discrete wavelet transform (DWT) based on wavelet type, with a number of layers specified by the decomposition level, on the artifact-containing GAN fusion image to obtain a final low-frequency approximation subband and a set of multi-level high-frequency detail subbands. Step F1F2, traversing each high-frequency subband image in the multi-level high-frequency detail subband set, applying a two-dimensional fast Fourier transform to each image for spectral analysis, and identifying abnormal energy spikes using a peak detection algorithm based on a fast Fourier transform peak detection threshold. A corresponding subband artifact score map is generated for each high-frequency subband image. Step F1F3, upsampling all subband artifact score maps to the original size of the artifact-containing GAN fusion image, accumulating them, and normalizing the accumulated results to generate a single-channel artifact signature map.

[0069] The artifact signature map generated in step F1 is a single-channel, two-dimensional floating-point tensor whose spatial dimensions are identical to the input artifact-containing GAN fusion image, and all its pixel values ​​are normalized to the range of 0 to 1. Logically, the artifact signature map serves as a spatialized evidence map, where the intensity value of each pixel directly corresponds to the confidence level of the periodic high-frequency artifact detected at that location through the DWT-FFT hybrid domain analysis. The artifact signature map is used as the core input to the subsequent guidance subnetwork, step F2, which smooths and refines the raw evidence map to generate an attention mask more suitable for guiding correction.

[0070] The implementation method of step F1F1 includes: step F1F1F1, initializing the generative adversarial network fusion image containing artifacts as the current approximate subband. Step F1F1F2, initializing an empty multi-level high-frequency detail subband set. Step F1F1F3, when the number of loops is less than the decomposition level, performing a single-level two-dimensional discrete wavelet transform on the current approximate subband based on the wavelet type to obtain a new next-level approximate subband and the current-level high-frequency detail subband in three directions: horizontal, vertical, and diagonal, and updating the next-level approximate subband to the current approximate subband, while adding the current-level high-frequency detail subband to the multi-level high-frequency detail subband set. Step F1F1F4, naming the current approximate subband at the end of the loop as the final low-frequency approximate subband.

[0071] The implementation method of step F1F1F3 includes: if the number of loops is less than the decomposition level, then iteratively executing steps F1F1F3F1 to F1F1F3F6: In step F1F1F3F1, using the wavelet and scaling filter corresponding to the wavelet type, first perform a one-dimensional discrete wavelet transform along the rows of the current approximation subband to obtain a result. Then, perform a one-dimensional discrete wavelet transform along the columns of the result to obtain a low-frequency approximation subband LL, a horizontal high-frequency detail subband LH, a vertical high-frequency detail subband HL, and a diagonal high-frequency detail subband HH. In step F1F1F3F2, the LL subband is assigned to the next-level approximation subband. In step F1F1F3F3, the LH, HL, and HH subbands are combined to form the current-level high-frequency detail subband. In step F1F1F3F4, the next-level approximation subband is updated to the current approximation subband. In step F1F1F3F5, the current-level high-frequency detail subband is added to the multi-level high-frequency detail subband set. Steps F1F1F3F6, the number of loops increases by one.

[0072] The implementation method of step F1F2 includes: step F1F2F1, initializing an empty set of all subband artifact score maps. Step F1F2F2, performing a two-dimensional fast Fourier transform on each high-frequency subband image in the multi-level high-frequency detail subband set to obtain a complex spectrum; calculating the amplitude spectrum of the complex spectrum to obtain an original amplitude spectrum; performing a quadrant shift on the original amplitude spectrum to move the zero-frequency component of the original amplitude spectrum to the center of the spectrum to obtain a centralized amplitude spectrum. Step F1F2F3, applying a two-dimensional local maximum detection algorithm to the centralized amplitude spectrum, and combining it with a fast Fourier transform peak detection threshold for screening to obtain a list of artifact energy peak coordinates. Step F1F2F4, generating a subband artifact score map for the high-frequency subband image based on the amplitude and position of each peak in the artifact energy peak coordinate list, and adding the subband artifact score map to the set of all subband artifact score maps.

[0073] The implementation method of step F1F2F2 includes: step F1F2F2F1, performing a two-dimensional fast Fourier transform on the high-frequency sub-band image to obtain a complex spectrum, wherein the calculation method of the transformation includes: for each coordinate point in the frequency domain, traversing all pixel points of the high-frequency sub-band image, multiplying the pixel value of each pixel point with a complex exponential basis function and then accumulating and summing them, and the complex exponential basis function is generated by the following method: respectively calculating the product of the x-axis coordinate of the coordinate point in the frequency domain and the x-axis coordinate of the pixel point of the high-frequency sub-band image, and the product of the y-axis coordinate of the coordinate point in the frequency domain and the y-axis coordinate of the pixel point of the high-frequency sub-band image; normalizing the two products with the corresponding dimensions of the high-frequency sub-band image and then summing them to obtain a sum value; taking the exponent of the result of multiplying the sum value by a preset complex constant; its calculation formula is , where Refers to the frequency coordinate The two-dimensional discrete Fourier transform result at , Refers to the spatial coordinates where M and N are the height and width of the high-frequency subband image, respectively, and j is an imaginary unit. Step F1F2F2F2 calculates the absolute value of each element in the complex spectrum to obtain an original amplitude spectrum. Step F1F2F2F3 performs a quadrant shift operation on the original amplitude spectrum, diagonally swapping the four quadrants of the original amplitude spectrum so that the zero-frequency component is located at the center of the image, thereby obtaining a centered amplitude spectrum.

[0074] The implementation method of steps F1F2F3 includes: Step F1F2F3F1, calling a standard two-dimensional peak search algorithm to analyze the centralized amplitude spectrum to identify all local maxima and calculate the peak prominence of each maximum relative to its surrounding background, thereby obtaining a candidate peak list containing the coordinates of each peak and its corresponding prominence value. Step F1F2F3F2, filtering the candidate peak list based on the condition that the peak prominence is greater than the fast Fourier transform peak detection threshold, to obtain a significant energy peak list. Step F1F2F3F3, extracting only the coordinates of all peaks from the significant energy peak list to form a list of artifact energy peak coordinates.

[0075] The implementation method of step F1F3 includes: step F1F3F1, initializing a cumulative artifact score map with the same size as the artifact-containing generative adversarial network fusion image and all elements being zero. Step F1F3F2, traversing each sub-band artifact score map in the set of all sub-band artifact score maps. Step F1F3F3, upsampling the sub-band artifact score map to the original size of the artifact-containing generative adversarial network fusion image through bilinear interpolation according to the decomposition level corresponding to the current sub-band artifact score map, to obtain an upsampled score map. Step F1F3F4, adding the upsampled score map to the cumulative artifact score map pixel by pixel, and updating the cumulative artifact score map. Step F1F3F5, after the traversal is completed, performing minimum-maximum normalization on the cumulative artifact score map so that all its pixel values ​​fall within to interval, and obtain the artifact signature map.

[0076] The theoretical rationale for the hybrid analysis method employed in step F1 is based on a deep understanding of the artifact generation mechanisms of generative adversarial networks (GANs) and the precise application of the properties of classical signal processing transforms. Previous research has clearly demonstrated that the upsampling modules widely used in GANs systematically introduce periodic patterns in the generated image, known as "checkerboard artifacts." These artifacts manifest as isolated, high-energy spikes in the image's frequency domain. This approach first employs the discrete wavelet transform (DWT), a standard multiresolution analysis tool, to decompose the image signal into approximate subbands containing macroscopic structure and detail subbands containing microscopic detail. This decomposition step is crucial because it effectively separates high-frequency components, which are prone to harboring artifacts, from the primary image content. Applying a two-dimensional fast Fourier transform to these separated high-frequency detail subbands significantly improves the detection sensitivity of characteristic artifact energy spikes and reduces the interference of background noise, thereby achieving a robust, physics-based artifact localization method.

[0077] Before proceeding to step F2, the logical necessity of performing this step must be clarified. Although the artifact signature map generated in step F1 can accurately locate artifacts in the frequency domain, it may appear as a sparse point or line structure when directly converted back to the spatial domain and may contain a certain amount of detection noise. If such a raw, potentially spatially discontinuous map is directly used as the multiplicative attention mask for the subsequent correction network, it may introduce harsh edges or new visual discontinuities in the corrected image. Therefore, it is necessary to introduce a lightweight convolutional subnetwork. This network functions as a filter with nonlinear smoothing capabilities obtained through data-driven learning. Its task is to transform the original "artifact evidence strength map" into a spatially smoother and more semantically coherent spatial attention map, thereby ensuring that the subsequent guided correction process can proceed smoothly and naturally.

[0078] Step F2: Input the artifact signature map and the preset guide sub-network weights, guide sub-network convolution kernel size, and guide sub-network output layer convolution kernel size into the guide sub-network for smoothing and nonlinear transformation to generate a value range of to The spatial attention map between .

[0079] The guided subnetwork weights are a set of parameters stored in HDF5 file format. These parameters are obtained by optimizing a lightweight convolutional subnetwork in a separate training task. The goal of this training task is to learn a robust nonlinear mapping from potentially sparse and noisy artifact signature maps to spatially smooth and semantically coherent spatial attention maps.

[0080] The guide subnetwork convolution kernel size is a two-tuple hard-coded to (3, 3). This value is based on established standard practice in deep learning for image processing tasks. A 3×3 convolution kernel is the minimum size that effectively captures local information about a central pixel and its eight neighbors. It provides a well-established balance between maintaining an effective receptive field and controlling computational complexity.

[0081] The kernel size of the output layer of the guidance subnetwork is a two-tuple hard-coded to (1, 1). This value is based on its standard usage in modern convolutional neural network architectures. 1×1 convolution kernels are widely used to linearly combine and reduce features along the channel dimension without changing the spatial dimensionality of the feature map. In this context, the kernel size of the output layer of the guidance subnetwork is used to efficiently aggregate the multi-channel intermediate feature maps into a single-channel output map.

[0082] The artifact signature map generated directly by frequency domain peak detection may have sparse and noisy problems. If it is used directly as an attention map, the correction effect may be too stiff. A lightweight convolutional subnetwork is introduced, which acts as a learnable nonlinear smoothing filter. The network can learn the mapping from the "artifact evidence strength map" to the "optimal correction attention map", making the final spatial attention map smoother and more continuous in space. The final Sigmoid activation function is necessary because it ensures that the output value is strictly located to This is a standard requirement for multiplicative spatial attention masks, ensuring that subsequent operations effectively “re-weight” the features.

[0083] The implementation method of step F2 includes: step F2F1, constructing a guidance sub-network consisting of two convolutional layers using the convolution kernel size of the guidance sub-network and a final convolutional layer using the convolution kernel size of the guidance sub-network output layer, and loading the weights of the guidance sub-network. Step F2F2, passing the artifact signature map as input to the weighted guidance sub-network to obtain the network's original attention output. Step F2F3, applying the Sigmoid activation function to the network's original attention output to normalize the value of each pixel to to interval, and obtain the spatial attention map.

[0084] The spatial attention map generated in step F2 is a single-channel, two-dimensional floating-point tensor whose spatial dimensions are identical to the input artifact-containing generative adversarial network fusion image. Due to the final application of the Sigmoid activation function, all its pixel values ​​are strictly constrained to the range of 0 to 1. Logically, the spatial attention map represents a refined, spatially coherent weight mask, where regions with values ​​close to 1 represent artifact regions that require the strongest correction, while regions with values ​​close to 0 represent artifact-free regions that need to be completely preserved. The purpose of the spatial attention map is to serve as a multiplicative gating signal applied to the feature maps of the decoder path in the subsequent main correction network, that is, step F3, to achieve precise spatial guidance of the network's correction capabilities.

[0085] Before proceeding to step F3, the rationale behind the core architecture selection must be clarified. This solution chooses U-Net as the core architecture for the correction network because it has been widely demonstrated to perform exceptionally well in image-to-image translation tasks. Its signature symmetric encoder-decoder structure, particularly the "skip connection" mechanism, allows the network to simultaneously leverage high-level semantic information from deep encoder layers and features from shallow layers that preserve high-frequency spatial details during decoding. This fusion of multi-scale information is crucial for accurately removing artifacts while losslessly restoring the original image structure and texture. Furthermore, this solution employs a residual learning paradigm, where the network does not directly learn to generate the final artifact-free image, but instead learns to predict the difference between the input and the ideal output, namely the artifact residual map. This design greatly simplifies the network's learning task, as it only needs to fit a correction variable, which is typically sparser and has a smaller range of values, rather than a complex distribution across the entire image. This not only helps accelerate network convergence but also better protects the majority of the original image content from unnecessary modifications.

[0086] Step F3, using the spatial attention map and based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module, the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights is guided weighted, and the artifact-containing generative adversarial network fusion image is input to the lightweight correction network to generate an artifact residual map through the residual learning paradigm.

[0087] The lightweight rectification network weights are a set of parameters stored in the HDF5 file format. These parameters are obtained by training a U-Net architecture through end-to-end optimization in an offline supervised learning task on a dataset containing a large number of paired artifact-containing images and their corresponding artifact-free real images. The training process uses a composite loss function that combines the L1 pixel reconstruction loss, the adversarial loss, the LPIPS perceptual loss, and the feature matching loss. The purpose is to ensure that the artifact residual map generated by the network, when combined with the original image, can produce a final rectified image that is pixel-accurate, perceptually realistic, and free of newly introduced blur.

[0088] The correction network convolution kernel size is a two-tuple, preferably hard-coded to (3, 3). This value is determined by the same method as the guiding sub-network convolution kernel size, that is, following the standard design practice of convolutional neural networks in image processing applications, using 3×3 as the basic unit for capturing local spatial information.

[0089] The pooling and upsampling kernel size for the correction network is a two-tuple, preferably hard-coded to (2, 2). This value is determined based on standard design principles of the U-Net architecture and its variants. In the U-Net encoder path, using a 2×2 max pooling layer with a stride of 2 is a standard way to halve the spatial dimensionality of the feature map; in the decoder path, using a 2×2 upsampling operation with a stride of 2 is a standard way to double the spatial dimensionality of the feature map.

[0090] The number of convolutional layers within a correction network module is an integer, preferably hard-coded to 2. This value is based on the classic design pattern established in the original U-Net paper and its many subsequent successful variants. In each encoder or decoder module of the U-Net, stacking two consecutive convolutional layers at each scale level has been shown to be a robust and efficient balance between effectively increasing the network's nonlinear modeling capabilities while maintaining overall computational efficiency.

[0091] Step F3 uses U-Net as the core network architecture. U-Net is particularly suitable for image-to-image conversion tasks such as image restoration and correction due to its symmetrical encoder-decoder structure and iconic "skip connection". The encoder path captures the contextual information and high-level semantic features of the image through layer-by-layer convolution and pooling; the decoder path gradually restores the spatial resolution of the image through upsampling. The key is that the skip connection directly splices the shallow feature maps in the encoder that have not been deeply compressed and retain high-frequency spatial details with the feature maps of the corresponding level of the decoder. This mechanism enables the network to simultaneously utilize deep semantic information and shallow detail information when performing refined reconstruction, which is crucial for accurately restoring the original image structure while eliminating artifacts.

[0092] The design of the correction network follows the residual learning paradigm. Rather than directly learning to generate a complete "artifact-free image," the network learns the difference between the input and target—the artifact-containing GAN fused image and the true artifact-free image—or the artifact residual map. The final corrected image is obtained by adding the residual map output by the network to the original input image. This design greatly simplifies the learning task because the network only needs to focus on and fit a typically sparse, small-scale correction quantity, rather than the complex distribution of the entire image. This helps accelerate network convergence and better preserve the overall structure and content of the image.

[0093] In order to meet the actual deployment requirements of portable devices, the lightweight guidance sub-network and lightweight correction network can further adopt existing artificial intelligence algorithm optimization technology in engineering implementation. For example, the model volume can be compressed and the computational complexity can be reduced through methods such as model quantization and pruning. The optimized model can be deployed on an embedded chip to achieve low-power real-time processing while ensuring the correction effect, effectively avoiding image processing delays or freezes caused by insufficient computing power of the terminal device. In addition, a system including the aforementioned network can also support offline working mode. In an environment without network connection, the final corrected image and analysis results as key evidence will be temporarily stored locally and automatically synchronized with the command center after the network is restored.

[0094] The lightweight correction network weights are obtained through a multi-objective composite loss function The composite loss function is obtained by joint optimization training. It is constructed by weighted summing of four core loss terms, aiming to balance the four interrelated optimization objectives of pixel-level accuracy, perceptual realism, adversarial stability, and feature distribution matching. Its specific form is:

[0095]

[0096] Where, 、 、 ,and It is the weight hyperparameter of each loss and is set during the model training phase.

[0097] The following will introduce each loss term that constitutes the composite loss function:

[0098] L1 pixel reconstruction loss The construction method includes: calculating the L1 norm distance between the final corrected image and the true artifact-free image to ensure basic pixel-level fidelity. Its formula is , where Refers to the L1 pixel reconstruction loss value, Refers to a true artifact-free image, refers to the final corrected image, It means to find the expectation of all samples in the data set. Denotes the calculation of the L1 norm. This formula quantifies the average absolute difference between the final rectified image and the true artifact-free image at the pixel level, serving as a basis for optimizing network parameters to improve pixel-level fidelity. The calculation process for this formula is as follows: for each image pair in a batch, the absolute value of the difference between the corresponding pixel values ​​of the final rectified image and the true artifact-free image is calculated, and then the absolute differences of all pixels are summed to obtain the sum of the absolute pixel differences of the image pair; finally, the sum of the absolute pixel differences of all image pairs in the batch is averaged to obtain the final L1 loss.

[0099] Adversarial Loss The construction method includes: This loss is based on an independent discriminator network, whose goal is to distinguish the final rectified image from the true artifact-free image. In this scheme, the adversarial loss of the least squares GAN is used to improve training stability. The adversarial loss of the correction network is designed to enable its generated images to "fool" the discriminator. Its specific form is:

[0100]

[0101] Where, is the output of the discriminator for the final rectified image. By minimizing this loss, the rectification network is driven to generate images that are visually more difficult for the discriminator to distinguish between true and false, thereby improving the perceptual quality and realism of the final result.

[0102] The construction method of LPIPS perceptual loss includes: This loss function uses a deep network pre-trained on a large-scale image classification task as a feature extractor. The final corrected image and the true artifact-free image are input into the network respectively, the feature activation maps of multiple intermediate layers are extracted, and then the weighted L2 distance between the feature maps of the corresponding layers is calculated. Compared with the traditional L1 / L2 loss, LPIPS (Learned Perceptual Image Patch Similarity) can better simulate the perceptual judgment of the human visual system, effectively prevent image blur and retain rich texture details. Its calculation formula is , where Refers to the LPIPS perception loss value, is the layer index in the pre-trained network, Indicates that from The feature map extracted by the layer, and are the true artifact-free image and the final corrected image, and It is The height and width of the layer feature map, It is used to calibrate The weights of the layer importance, Indicates the calculation of the square of the L2 norm, Indicates element-by-element multiplication, h and w refer to the The height index and width index of the pixel on the layer feature map. This formula is used to quantify the perceptual similarity of two images in the deep feature space, aiming to guide the network optimization direction so that the generated image is closer to the real image in human visual perception. The calculation process of this formula is as follows: first, the final corrected image and the real artifact-free image are input into a fixed pre-trained network; then, for each selected layer in the network , extract the feature maps corresponding to the two images, calculate their difference, and compare the difference with the calibration weight of the layer Perform element-by-element multiplication; then, calculate the square of the L2 norm of this weighted difference map and divide it by the size of the feature map for normalization; finally, sum the normalized results of all selected layers to obtain the final LPIPS loss.

[0103] Feature matching loss The construction method includes: this loss further utilizes the discriminator network; requires that the feature activations generated by the final corrected image in multiple intermediate layers of the discriminator must match the feature activations generated by the real artifact-free image in the same layer. By forcing the generator to match the statistical distribution of the real data in the feature space, this loss can greatly stabilize the adversarial training process and effectively prevent training problems such as mode collapse. Its calculation formula is , where Refers to the feature matching loss value, is the index of the middle layer of the discriminator network, is the total number of selected intermediate layers, Represents the discriminator The feature map output by the layer, and are the true artifact-free image and the final corrected image, It is The total number of elements in the layer feature map, This formula is used to stabilize the adversarial training process by matching the intermediate feature representations of the generated image and the real image in the discriminator network, and to force the generator to learn the underlying feature distribution of the real data. The calculation process of this formula is as follows: First, the final corrected image and the real artifact-free image are input into the discriminator network; then, for each selected intermediate layer , extract the feature maps corresponding to the two images, calculate the L1 norm between them, and divide it by the number of elements in the feature map of that layer for normalization; finally, sum the normalized results of all selected layers to obtain the final feature matching loss.

[0104] The implementation method of step F3 includes the following steps: Step F3F1: Inputting the artifact-containing generative adversarial network fusion image into the encoder path of a lightweight correction network based on a U-Net architecture and loaded with lightweight correction network weights. The encoder path generates a multi-scale encoder feature map and a bottleneck layer feature map by performing convolution and max pooling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers within the correction network module. Step F3F2: Inputting the bottleneck layer feature map and the multi-scale encoder feature map into the decoder path loaded with lightweight correction network weights. Guided by the spatial attention map, the decoder path generates the original network output by performing convolution and upsampling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers within the correction network module. Step F3F3: Defining the original network output as an artifact residual map.

[0105] The artifact residual map generated in step F3 is a three-channel, two-dimensional floating-point tensor whose spatial dimensions are identical to the input artifact-containing GAN fusion image. Logically, the artifact residual map is the correction network's direct prediction and quantized representation of the generative artifacts present in the input image. It numerically represents the correction required to transform the artifact-containing GAN fusion image to the ideal artifact-free state. The purpose of the artifact residual map is to offset and remove artifacts in the final synthesis step, step F4, by performing pixel-by-pixel arithmetic addition with the original input image to generate the final corrected image.

[0106] The implementation method of step F3F1 includes: step F3F1F1, constructing a U-Net encoder path and loading the corresponding weight parameters in the lightweight correction network weights for the U-Net encoder path; the overall architecture of the U-Net encoder path and the various encoder modules contained therein are defined by the lightweight correction network weights. Step F3F1F2, taking the artifact-containing generative adversarial network fusion image as the initial input, passing it through each module of the encoder path in sequence, performing convolution with the correction network convolution kernel size times the number of convolution layers in the correction network module and the maximum pooling operation with the correction network pooling and upsampling kernel size once in each module, and saving the output of each module after convolution as an element of the multi-scale encoder feature map. Step F3F1F3, defining the output of the deepest layer of the encoder path as the bottleneck layer feature map.

[0107] The implementation method of step F3F1F2 includes: step F3F1F2F1, using the artifact-containing generative adversarial network fusion image as the current feature map. Step F3F1F2F2, initializing an empty multi-scale encoder feature map list. Step F3F1F2F3, for each module in the encoder path, looping through steps F3F1F2F3F1 to F3F1F2F3F3: Step F3F1F2F3F1, looping through the number of convolutional layers in the correction network module on the current feature map, each operation comprising applying a convolutional layer using the correction network convolution kernel size, followed by applying a ReLU activation function, and naming the final result the module output feature map. Step F3F1F2F3F2, adding the module output feature map to the multi-scale encoder feature map list. Step F3F1F2F3F3, applying a max pooling layer using the correction network pooling and upsampling kernel size to the module output feature map, and updating the result as the current feature map for input to the next module.

[0108] The implementation method of step F3F2 includes: step F3F2F1, constructing a U-Net decoder path and loading the corresponding weight parameters from the lightweight correction network weights into the U-Net decoder path; the overall architecture of the U-Net decoder path and the individual decoder modules it contains are defined by the lightweight correction network weights. Step F3F2F2, using the bottleneck layer feature map as the initial input of the decoder path, and passing it through each decoder module in sequence. In each decoder module, the input feature map is upsampled, concatenated, and convolved, guided by the corresponding layer feature map and spatial attention map from the multi-scale encoder feature map, to generate a weighted feature map as the input of the next decoder module. Step F3F2F3, processing the weighted feature map output by the last decoder module through a 1×1 convolution layer using the correction network convolution kernel size to obtain the original network output.

[0109] The specific implementation method of step F3F2F2 includes: step F3F2F2F1, performing a transposed convolution upsampling operation on the output of the previous decoder module, and splicing it with the corresponding level feature map from the multi-scale encoder feature map in the channel dimension to obtain a spliced ​​feature map. Step F3F2F2F2, performing a convolution operation on the spliced ​​feature map equal to the number of convolution layers in the correction network module, and applying the ReLU activation function after each convolution operation to obtain the current decoder module output feature map. Step F3F2F2F3, downsampling the spatial attention map to the same size as the current decoder module output feature map through a bilinear interpolation algorithm to obtain a same-scale attention map, and performing element-by-element multiplication of the same-scale attention map and the current decoder module output feature map to obtain the weighted feature map.

[0110] In step F3, the spatial attention map is applied as the core mechanism for regulating the U-Net decoder feature map. Its theoretical basis is the principle of "conditional computation," which has been successfully verified and applied in the feature-level affine transformation framework FiLM. The FiLM framework proves that using information from one modality as a condition—in this case, the spatial attention map—can effectively perform dynamic, feature-by-feature linear modulation on the features of another modality, in this case, the decoder feature map. The element-by-element multiplication operation used in this scheme is a specific implementation of feature modulation. It plays the role of a learnable spatial gating, precisely guiding the network's repair capabilities. This design successfully applies the theoretical ideas of the FiLM framework to the task of artifact correction, demonstrating the solid theoretical feasibility and rationality of its architectural choice.

[0111] In step F4, the artifact residual image is added pixel by pixel to the artifact-containing generative adversarial network fusion image to obtain the final corrected image.

[0112] The implementation process detailed in this embodiment starts with a generative adversarial network fusion image containing artifacts at the data flow level. Through a hybrid domain analysis process combining discrete wavelet transform and fast Fourier transform, the image is first converted into an artifact signature map that quantifies artifact evidence; then, the artifact signature map is smoothed and refined by a lightweight convolutional network to generate a spatial attention map as a guiding signal; then, the spatial attention map is used to guide a U-Net-based residual network, which processes the original input image and outputs an artifact residual map; finally, by adding the artifact residual map to the original input, the final corrected image is produced as the final deliverable.

[0113] From a technical perspective, this implementation process can efficiently and accurately remove various generative artifacts contained in artifact-containing generative adversarial network fusion images, including unnatural periodic textures and checkerboard patterns. While targeting the repair of artifacts, this process can preserve the content details and global structure of artifact-free areas in the original image with maximum fidelity, avoiding the introduction of new blur or information loss during the correction process. Ultimately, as a non-invasive post-processing module, this solution produces a high-fidelity final corrected image that meets the stringent application requirements in terms of both visual integrity and information accuracy without any modification to the original GAN ​​model.

[0114] Example 2

[0115] See Figure 2 As shown, this embodiment provides a multi-spectral night vision device intelligent image enhancement system, the system comprising:

[0116] an artifact signature module that performs a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image based on a preset discrete wavelet transform wavelet type and decomposition level, and combines a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate periodic frequency components caused by generative artifacts, thereby generating a single-channel artifact signature map having the same spatial size as the artifact-containing generative adversarial network fusion image;

[0117] The attention module inputs the artifact signature map and the preset guide subnetwork weights, guide subnetwork convolution kernel size, and guide subnetwork output layer convolution kernel size into the guide subnetwork, performs smoothing and nonlinear transformation, and generates a spatial attention map with a value range between [0, 1].

[0118] The artifact residual module uses a spatial attention map and performs guided weighting on the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolutional layers in the correction network module. The artifact residual map is generated by inputting the artifact-containing generative adversarial network fusion image into the lightweight correction network through the residual learning paradigm;

[0119] The correction module adds the artifact residual map and the artifact-containing generative adversarial network fusion image pixel by pixel to obtain the final corrected image.

[0120] Example 3

[0121] See Figure 3 This embodiment provides an application scenario of the intelligent image enhancement method for multi-spectral night vision devices of the present invention. In this scenario, a drone equipped with a multi-spectral night vision device first collects multi-spectral raw image data when performing a mission. These raw data are transmitted to an onboard or edge GAN fusion module for processing. The module aims to fuse multimodal data into a more informative image, but this process may introduce generative artifacts, thereby producing a "generative adversarial network fusion image containing artifacts." This artifact-containing generative adversarial network fusion image is then used as input and sent to the intelligent image enhancement method process described in the present invention for processing. The method process of the present invention performs artifact correction on the image and outputs a clear, artifact-free "final corrected image." Finally, the high-quality corrected image is transmitted to the ground control and monitoring center for operators to use in applications with strict requirements on image fidelity, such as real-time analysis, target recognition, or post-event analysis, significantly improving the practicality and reliability of the system.

[0122] Example 4

[0123] This embodiment provides an application scenario of the intelligent image enhancement method for a multi-spectral night vision device of the present invention. When conducting search and rescue at night or in low-visibility environments, a drone can be used as a platform to execute the intelligent image enhancement method described in the present invention. The GAN fusion module on the drone outputs a generative adversarial network fusion image containing artifacts. Subsequently, the aforementioned intelligent image enhancement method is applied in real time to correct the artifacts of the fusion image to obtain a final corrected image. The final corrected image can clearly highlight vital signs, and even if the target is stationary, it can be captured through the weak thermal signal changes of body temperature and breathing, thereby greatly improving the detection sensitivity and positioning accuracy of survivors.

[0124] Example 5

[0125] This embodiment provides an application scenario of the intelligent image enhancement method for multi-spectral night vision devices of the present invention. In military reconnaissance or security monitoring, the multi-spectral images collected by the front-end equipment are fused through the GAN algorithm to enhance the camouflage recognition capability. The fusion process may produce artifacts. The artifact-containing generative adversarial network fusion image is processed by the intelligent image enhancement method described in the present invention, and the final corrected image is output. The final corrected image has high fidelity, so it can effectively resist interference such as infrared decoys or strong light blinding, and can more accurately identify people wearing camouflage clothing or camouflaged equipment, thereby improving battlefield perception and security warning capabilities.

[0126] Example 6

[0127] This embodiment provides an application scenario for the intelligent image enhancement method for multispectral night vision devices of the present invention. During nighttime inspections of substations, chemical plants, and other structures, the thermal infrared image and visible light image of the equipment are fused and then processed using the intelligent image enhancement method described herein to eliminate artifacts, resulting in a final corrected image. Based on this final corrected image, operators or automated systems can accurately identify hotspots in electrical equipment or, by combining visible light signatures, observe damage to the equipment's exterior, providing fault warnings and improving inspection efficiency and safety.

[0128] Example 7

[0129] This embodiment provides an extended application system based on the intelligent image enhancement system of the multispectral night vision device of the present invention. The extended application system uses the intelligent image enhancement system of the present invention as a core preprocessing unit, and further integrates a living target contour reconstruction module. The living target contour reconstruction module receives the final corrected image output by the core preprocessing unit, and inputs the final corrected image into a contour reconstruction network. The contour reconstruction network is a deep neural network, and its structure may include an encoder path for capturing image context and a decoder path for generating completion information, and uses an attention mechanism to transfer effective features between the two. The contour reconstruction network is trained to learn the structural prior knowledge of the target, and can generate a highly realistic completed contour that conforms to the original structure of the target based on the contextual information and multispectral features in the image when receiving the final corrected image containing a partially occluded target.

[0130] Example 8

[0131] This embodiment provides an extended application system based on the intelligent image enhancement system of the multi-spectral night vision device of the present invention. The extended application system also uses the intelligent image enhancement system of the present invention as the core preprocessing unit, and further integrates a long-distance small target enhancement and recognition module. The long-distance small target enhancement and recognition module receives the final corrected image obtained by the core preprocessing unit. In order to enhance the details of the long-distance target, the module can input the regional slices containing potential targets in the final corrected image into a super-resolution reconstruction network. The super-resolution reconstruction network is a deep convolutional neural network, whose network structure includes multiple residual modules for extracting deep features, and includes an upsampling module at the end of the network for reconstructing the feature map into a high-resolution image. The network achieves enhancement of image details by learning the end-to-end mapping from low-resolution images to high-resolution images, thereby supporting the effective recognition of pixel-level targets.

[0132] Anything not described in this application can be achieved by adopting or drawing on existing technologies.

[0133] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0134] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A multispectral night vision device intelligent image enhancement method, characterized in that: include: Step F1, based on a preset discrete wavelet transform wavelet type and decomposition level, performing a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image, and combining it with a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate the periodic frequency components caused by the generative artifacts, and generating a single-channel artifact signature map having the same spatial size as the artifact-containing generative adversarial network fusion image; Step F2: Input the artifact signature map and the preset guiding sub-network weights, guiding sub-network convolution kernel size, and guiding sub-network output layer convolution kernel size into the guiding sub-network, perform smoothing and nonlinear transformation, and generate a spatial attention map with a value range between [0, 1]. Step F3, using the spatial attention map and based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module, the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights is guided weighted, and the artifact-containing generative adversarial network fusion image is input to the lightweight correction network to generate an artifact residual map through the residual learning paradigm; In step F4, the artifact residual image is added pixel by pixel to the artifact-containing generative adversarial network fusion image to obtain the final corrected image.

2. The multispectral night vision device intelligent image enhancement method according to claim 1, characterized in that: The implementation method of step F1 includes: Step F1F1, performing a multi-level two-dimensional discrete wavelet transform based on the wavelet type and the number of layers specified by the decomposition level on the generative adversarial network fusion image containing artifacts, to obtain a final low-frequency approximate subband and a set of multi-level high-frequency detail subbands; Steps F1 and F2: traverse each high-frequency sub-band image in the multi-level high-frequency detail sub-band set, apply a two-dimensional fast Fourier transform to perform spectrum analysis, and identify abnormal energy peaks using a peak detection algorithm based on a fast Fourier transform peak detection threshold, generating a corresponding sub-band artifact score map for each high-frequency sub-band image; In steps F1 and F3, all sub-band artifact score maps are upsampled to the original size of the artifact-containing generative adversarial network fusion image and then accumulated, and the accumulated results are normalized to generate a single-channel artifact signature map.

3. The multispectral night vision device intelligent image enhancement method according to claim 1, characterized in that: The implementation method of step F3 includes: Step F3F1, inputting the artifact-containing GAN fused image into the encoder path of a lightweight correction network based on a U-Net architecture and loaded with lightweight correction network weights, wherein the encoder path generates a multi-scale encoder feature map and a bottleneck layer feature map by performing convolution and max pooling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module; Steps F3F2: Input the bottleneck layer feature map and the multi-scale encoder feature map into the decoder path loaded with the lightweight correction network weights. Under the guidance of the spatial attention map, the decoder path generates the network raw output by performing convolution and upsampling operations based on the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolution layers in the correction network module. In step F3F3, the original output of the network is defined as the artifact residual map.

4. The multispectral night vision device intelligent image enhancement method according to claim 2, characterized in that: The implementation method of step F1F1 includes: step F1F1F1, initializing the generative adversarial network fusion image containing artifacts as the current approximate subband; step F1F1F2, initializing an empty multi-level high-frequency detail subband set; step F1F1F3, when the number of loops is less than the decomposition level, performing a single-level two-dimensional discrete wavelet transform on the current approximate subband based on the wavelet type to obtain a new next-level approximate subband and a current-level high-frequency detail subband including three directions of horizontal, vertical, and diagonal, and updating the next-level approximate subband to the current approximate subband, and adding the current-level high-frequency detail subband to the multi-level high-frequency detail subband set; step F1F1F4, naming the current approximate subband at the end of the loop as the final low-frequency approximate subband; The implementation method of step F1F1F3 includes: if the number of loops is less than the decomposition level, then looping through steps F1F1F3F1 to F1F1F3F6: in step F1F1F3F1, using the wavelet and scaling filter corresponding to the wavelet type, first performing a one-dimensional discrete wavelet transform along the rows of the current approximate subband to obtain a result, and then performing a one-dimensional discrete wavelet transform along the columns of the result to obtain a low-frequency approximate subband LL, a horizontal high-frequency detail subband LH, a vertical high-frequency detail subband HL and Diagonal high-frequency detail subband HH; step F1F1F3F2, assign the LL subband to the next-level approximate subband; step F1F1F3F3, combine the three subbands LH, HL, and HH into the current-level high-frequency detail subband; step F1F1F3F4, update the next-level approximate subband to the current approximate subband; step F1F1F3F5, add the current-level high-frequency detail subband to the multi-level high-frequency detail subband set; step F1F1F3F6, increase the number of loops by one.

5. The multi-spectral night vision device intelligent image enhancement method according to claim 2, characterized in that: The implementation method of steps F1 and F2 includes: Step F1F2F1, initialize the set of all empty sub-band artifact score maps; Steps F1, F2, and F2: for each high-frequency subband image in the multi-level high-frequency detail subband set, perform a two-dimensional fast Fourier transform on it to obtain a complex spectrum; calculate the amplitude spectrum of the complex spectrum to obtain an original amplitude spectrum; perform a quadrant shift on the original amplitude spectrum to move the zero-frequency component of the original amplitude spectrum to the center of the spectrum to obtain a centralized amplitude spectrum; Steps F1, F2, and F3: Apply a two-dimensional local maximum detection algorithm to the centralized amplitude spectrum and perform screening in combination with a fast Fourier transform peak detection threshold to obtain a list of artifact energy peak coordinates; Steps F1, F2, and F4: generating a sub-band artifact score map for the high-frequency sub-band image based on the amplitude and position of each peak in the artifact energy peak coordinate list, and adding the sub-band artifact score map to the set of all sub-band artifact score maps.

6. The multi-spectral night vision device intelligent image enhancement method according to claim 2, characterized in that: The implementation method of steps F1 and F3 includes: Step F1F3F1, initialize the cumulative artifact score map with the same size as the artifact-containing generative adversarial network fusion image and all elements are zero; Steps F1, F3, and F2, traverse each sub-band artifact score map in the set of all sub-band artifact score maps; Steps F1, F3, and F3: upsampling the subband artifact score map to the original size of the artifact-containing generative adversarial network fusion image through bilinear interpolation according to the decomposition level corresponding to the current subband artifact score map, thereby obtaining an upsampled score map; Steps F1, F3, and F4: add the upsampled score map to the cumulative artifact score map pixel by pixel, and update the cumulative artifact score map; In steps F1, F3, and F5, after the traversal is completed, the cumulative artifact score map is normalized using the minimum-maximum method so that all its pixel values ​​fall within the interval [0, 1] to obtain the artifact signature map.

7. The multi-spectral night vision device intelligent image enhancement method according to claim 3, characterized in that: The implementation method of steps F3F1 includes: Step F3F1F1, constructing a U-Net encoder path and loading the corresponding weight parameters in the lightweight correction network weights for the U-Net encoder path; the overall architecture of the U-Net encoder path and each encoder module contained therein are defined by the lightweight correction network weights; Steps F3, F1, and F2 take the artifact-containing generative adversarial network fusion image as the initial input and pass it through each module of the encoder path in sequence. In each module, a convolution operation using the convolution kernel size of the correction network is performed times the number of convolution layers in the correction network module and a maximum pooling operation using the correction network pooling and upsampling kernel size is performed once. The output of each module after convolution is saved as an element of the multi-scale encoder feature map. In steps F3F1F3, the output of the deepest layer of the encoder path is defined as the bottleneck layer feature map.

8. The multi-spectral night vision device intelligent image enhancement method according to claim 3, characterized in that: The implementation method of steps F3 and F2 includes: Steps F3, F2, and F1 construct a U-Net decoder path and load the corresponding weight parameters in the lightweight correction network weights into the U-Net decoder path; the overall architecture of the U-Net decoder path and each decoder module it contains are defined by the lightweight correction network weights; Steps F3F2F2 take the bottleneck layer feature map as the initial input to the decoder path and pass it through each decoder module in turn. In each decoder module, the input feature map is upsampled, concatenated, and convolved with the guidance of the corresponding level feature map and spatial attention map from the multi-scale encoder feature map to generate a weighted feature map as the input to the next decoder module. In steps F3, F2, and F3, the weighted feature map output by the last decoder module is processed by a 1×1 convolution layer using a corrected network convolution kernel size to obtain the original network output.

9. The multi-spectral night vision device intelligent image enhancement method according to claim 8, characterized in that: The specific implementation method of steps F3F2F2 includes: In step F3F2F2F1, a transposed convolution upsampling operation is performed on the output of the previous decoder module, and it is concatenated with the corresponding level feature map from the multi-scale encoder feature map in the channel dimension to obtain the concatenated feature map; Step F3F2F2F2, performing convolution operations on the spliced ​​feature map equal to the number of convolution layers in the correction network module, and applying a ReLU activation function after each convolution operation to obtain the output feature map of the current decoder module; Step F3F2F2F3, downsample the spatial attention map to the same size as the output feature map of the current decoder module through a bilinear interpolation algorithm to obtain a same-scale attention map, and multiply the same-scale attention map with the output feature map of the current decoder module element by element to obtain the weighted feature map.

10. The intelligent image enhancement method for multispectral night vision device according to claim 5, characterized in that: The implementation method of step F1F2F2 includes: step F1F2F2F1, performing a two-dimensional fast Fourier transform on the high-frequency sub-band image to obtain a complex spectrum, wherein the calculation method of the transformation includes: for each coordinate point in the frequency domain, traversing all pixel points of the high-frequency sub-band image, multiplying the pixel value of each pixel point with a complex exponential basis function and then accumulating and summing them, wherein the complex exponential basis function is generated by respectively calculating the product of the x-axis coordinate of the coordinate point in the frequency domain and the x-axis coordinate of the pixel point of the high-frequency sub-band image, and the frequency domain The product of the y-axis coordinate of the coordinate point in the image and the y-axis coordinate of the pixel point of the high-frequency sub-band image; normalizing the two products by the corresponding dimensions of the high-frequency sub-band image and then summing them to obtain a sum value; taking the exponent of the result of multiplying the sum value by a preset complex constant; step F1F2F2F2, calculating the absolute value of each element in the complex spectrum to obtain an original amplitude spectrum; step F1F2F2F3, performing a quadrant shift operation on the original amplitude spectrum, by diagonally exchanging the four quadrants of the original amplitude spectrum so that the zero-frequency component is located at the center of the image, to obtain a centralized amplitude spectrum; The implementation method of step F1F2F3 includes: step F1F2F3F1, calling the standard two-dimensional peak search algorithm, analyzing the centralized amplitude spectrum to identify all local maxima, and calculating the peak prominence of each maximum value relative to its surrounding background, to obtain a candidate peak list containing the coordinates of each peak and its corresponding prominence value; step F1F2F3F2, screening the candidate peak list according to the condition that the peak prominence is greater than the fast Fourier transform peak detection threshold, to obtain a significant energy peak list; step F1F2F3F3, extracting only the coordinates of all peaks from the significant energy peak list to form an artifact energy peak coordinate list.

11. Multi-spectral night vision device intelligent image enhancement system, characterized by: For implementing the multispectral night vision device intelligent image enhancement method according to any one of claims 1 to 10, the system comprises: an artifact signature module that performs a frequency domain joint analysis based on a multi-level discrete wavelet transform and a two-dimensional fast Fourier transform on the artifact-containing generative adversarial network fusion image based on a preset discrete wavelet transform wavelet type and decomposition level, and combines a peak detection algorithm based on a fast Fourier transform peak detection threshold to locate periodic frequency components caused by generative artifacts, thereby generating a single-channel artifact signature map having the same spatial size as the artifact-containing generative adversarial network fusion image; The attention module inputs the artifact signature map and the preset guide subnetwork weights, guide subnetwork convolution kernel size, and guide subnetwork output layer convolution kernel size into the guide subnetwork, performs smoothing and nonlinear transformation, and generates a spatial attention map with a value range between [0, 1]. The artifact residual module uses a spatial attention map and performs guided weighting on the decoder feature map of the lightweight correction network loaded with the lightweight correction network weights based on the architecture defined by the correction network convolution kernel size, the correction network pooling and upsampling kernel size, and the number of convolutional layers in the correction network module. The artifact residual map is generated by inputting the artifact-containing generative adversarial network fusion image into the lightweight correction network through the residual learning paradigm; The correction module adds the artifact residual map and the artifact-containing generative adversarial network fusion image pixel by pixel to obtain the final corrected image.

Citation Information

Patent Citations

  • Image conversion system and method

    CN107633540A

  • High-speed acquisition MRI reconstruction method based on residual self-attention image enhancement

    CN111696168A

  • Underwater image enhancement method based on multi-scale attention mechanism fusion

    CN115034982A

  • Remote sensing image fusion method and system based on fusion correction

    CN117197008A

Cited By

  • Frequency domain enhanced cross-modal pedestrian re-identification method, system and device and medium

    CN121482832A