Image processing method, system, device and storage medium based on frequency enhancement

Through the image packaging processing and the improved LW-ISP method, frequency enhancement, combined with discrete wavelet transformation, the problem of frequency information loss and insufficient recognition performance in intelligent ISP is solved, and better visual effects and recognition performance are achieved.

CN116977229BActive Publication Date: 2025-08-08INST FOR INTERDISCIPLINARY INFORMATION CORE TECH XIAN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311006969.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2025-08-08
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

The existing smart ISPs will lose some positive frequency information in terms of visual effects, and even if the image visual effect is good, it will not produce a good recognition effect.

Method used

The image processing method based on frequency enhancement is adopted, and the image to be processed is packaged and frequency enhancement is performed using the improved LW-ISP method, combined with discrete wavelet transformation, and the frequency information lost during the ISP processing is supplemented, and the space and frequency information are captured.

Benefits of technology

The natural feature compliance after image processing is improved, detailed information is significantly retained, the noise removal effect is enhanced, and the recognition performance is further improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977229B_ABST
    Figure CN116977229B_ABST
Patent Text Reader

Abstract

The present invention relates to an image processing method. In response to the technical problems that existing intelligent ISPs lose some positive frequency information in terms of visual effects, and in terms of recognition performance, even if the visual effects of the image are good, they fail to produce good recognition results, an image processing method, system, device, and storage medium based on frequency enhancement are proposed. The images to be processed are first packaged and processed, and then the packaged images are frequency enhanced using an improved LW-ISP method. Based on deep learning, frequency domain information can be retained and enhanced at the ISP level. The frequency information lost during ISP processing can be supplemented by FCB, which, combined with discrete wavelet transform, can capture spatial and frequency information. Practical verification has shown that the image processing method of the present invention, in terms of image processing, produces images that are more consistent with natural features, can better retain detail information, have significant denoising and enhancement effects, and further improve recognition performance through frequency domain enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing method, and in particular to an image processing method, system, device and storage medium based on frequency enhancement. Background Art

[0002] Deep Neural Networks (DNNs) have achieved results that surpass traditional algorithms in many fields, such as image classification and cancer detection. Image signal processing (ISP) methods, which receive and process raw signals from camera imaging sensors, play a crucial role in image quality. The integration of ISP image preprocessing tasks and DNNs is a current research hotspot. In particular, intelligent ISPs, derived from existing deep learning (DL)-based methods within the ISP pipeline, can learn raw image statistics and implement multi-task joint solutions. Therefore, intelligent ISPs can creatively connect the photo imaging process with subsequent recognition applications.

[0003] However, existing smart ISPs still have the following problems:

[0004] (1) Visual effects: Although it can fully capture the inherent properties of the original image input, the visual effects usually lack realistic details and textures due to the spectral bias in the neural network. Figure 1 As shown in the figure, although the current intelligent ISP is more effective than the traditional ISP, there is still room for improvement in terms of high-frequency and low-frequency information. In addition, some positive frequency information of the original image will be lost after the intelligent ISP. Figure 1 In the figure, column A represents the original image, column B represents the traditional ISP, and column C represents the intelligent ISP.

[0005] (2) Recognition performance: Although progress has been made in restoring or enhancing the visual quality of images, these advances cannot always be used in a useful way for image recognition tasks. Visually appealing images do not necessarily produce good recognition results. Summary of the Invention

[0006] In order to solve the technical problems that some positive frequency information of existing intelligent ISP is lost in visual effects, and in recognition performance, even if the visual effect of the image is good, good recognition effect cannot be produced, an image processing method, system, device and storage medium based on frequency enhancement are proposed.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides an image processing method based on frequency enhancement, comprising the following steps:

[0009] Packing the images to be processed;

[0010] The frequency enhancement of the packed image is performed using the improved LW-ISP method;

[0011] The improved LW-ISP method is specifically as follows:

[0012] The CCBs that receive features from the first level in the LW-ISP method are replaced with FCBs, and the last FCB is replaced with the features received from the first level FGAM in the LW-ISP method with the packed image;

[0013] The characteristics of FCB received from the first and second levels of the improved LW-ISP method are defined as follows:

[0014] The output features received from the first-level FGAM are recorded as L1, the output features received from the first-level DB are recorded as L2, and the output features received from the second-level convolutional layer are recorded as L3;

[0015] The specific processing method in the FCB is:

[0016] Convolution is performed on L2 and L3 respectively to obtain the corresponding global and local features;

[0017] Recombining the global and local features, and fine-tuning the channel size through convolution;

[0018] The fine-tuned features of L2 and L3 are upsampled respectively, and then residual learning and feature channel merging are performed to obtain feature A that matches the L1 dimension;

[0019] Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

[0020] In a second aspect, the present invention proposes an image processing system based on frequency enhancement, including an LW-ISP system, wherein the LW-ISP system includes four downsampling groups arranged in sequence at the first stage and four upsampling groups arranged in sequence at the second stage, wherein the downsampling groups include DB and FGAM, and the first upsampling group includes a convolutional layer and an FC layer; and further includes a packing module;

[0021] The packaging module is used to package the images to be processed;

[0022] The second to fourth upsampling groups each include a convolutional layer and FCB arranged in sequence, which are used to perform frequency enhancement on the packed image;

[0023] The FCB is used to receive the output feature L1 of the first-level FGAM, the output feature L2 of the first-level DB, and the output feature L3 of the convolution layer in the same upsampling group; perform convolution processing on L2 and L3 respectively to obtain corresponding global and local features; perform pixel reorganization on the global and local features, and then fine-tune the channel size through convolution processing; upsample the corresponding fine-tuned features of L2 and L3 respectively, and then perform residual learning and feature channel merging to obtain feature A that matches the L1 dimension;

[0024] Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

[0025] In a third aspect, the present invention proposes a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] This paper proposes an image processing method based on frequency enhancement. The image to be processed is first packaged and then frequency-enhanced using an improved LW-ISP method. This improves upon the existing LW-ISP method and, based on deep learning, is capable of preserving and enhancing frequency domain information at the ISP level. Frequency information lost during ISP processing can be supplemented by FCB, which, combined with discrete wavelet transforms, captures both spatial and frequency information. Practical verification has shown that the image processing method of this invention produces images that are more consistent with natural features, better preserve detail information, and exhibit significant denoising and enhancement effects. Furthermore, frequency domain enhancement further improves recognition performance.

[0029] The present invention also proposes an image processing system based on frequency enhancement for implementing the above-mentioned image processing method. At the same time, it also proposes a storage medium and a device for implementing and applying the above-mentioned method with the help of different hardware media. It has all the advantages of the above-mentioned image processing method and is also convenient for the promotion and application of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1 A comparison of the visual effects of the original image, traditional ISP, and intelligent ISP in the background technology;

[0032] Figure 2 This is a flow chart of the first embodiment of the present invention;

[0033] Figure 3 This is the schematic diagram of the existing LW-ISP;

[0034] Figure 4 Schematic diagram of the reconstruction results of the classification dataset CIFAR-100; (a) is a schematic diagram of the convergence trend of different loss combinations, and (b) is a schematic diagram of the convergence trend after deleting the L_c curve in (a);

[0035] Figure 5 This is the improved LW-ISP schematic diagram;

[0036] Figure 6 This is the schematic diagram of FCB;

[0037] Figure 7 This is a comparison chart of the same RAW image processed by the existing LW-ISP method, Huawei P20, and the method of the present invention;

[0038] Figure 8 The following is a qualitative comparison of the recognition effects of different methods; (a) is the RGB image obtained after processing the input RAW, (b) is the top 5 categories obtained through recognition and their confidence values, (c) is the frequency domain mapping of the corresponding RGB image, and (d) is the brightness statistics of the samples reconstructed in the frequency domain and the GT method. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0040] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0041] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0042] In the description of the embodiments of the present invention, it should be noted that if the terms "upper," "lower," "horizontal," "inner," etc. appear, the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the inventive product is typically placed when in use. These terms are merely for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. In addition, the terms "first," "second," etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0043] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0044] In the description of the embodiments of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0045] The present invention is described in further detail below with reference to the accompanying drawings and embodiments:

[0046] Example 1

[0047] See also Figure 2 , an embodiment of the present invention discloses an image processing method based on frequency enhancement, comprising the following steps:

[0048] S101, packaging the images to be processed.

[0049] S102: performing frequency enhancement on the packed image using an improved LW-ISP method.

[0050] like Figure 3 As shown, it is a schematic diagram of the principle of the existing LW-ISP, which mainly includes the first stage J1 and the second stage J2. In the first stage, the feature map is gradually downsampled at different levels to speed up the calculation speed. Each downsampling is achieved through a downsampling block and an FGAM (Fine-grained Attention Module). The second stage further connects the processed global vector with the feature map of the same size in the first half through a symmetrical jump connection, including a convolutional layer, an FC layer, a convolutional layer, a CCB (Contextual Complement Upsampling Block, contextual completion upsampling block), a convolutional layer, a CCB, a convolutional layer, a CCB, a convolutional layer, a CCB and a convolutional layer arranged in sequence. Four CCBs are set in the second stage for adaptive high-frequency decomposition in the feature space and fusing the corresponding size features of the previous stage. Based on the existing LW-ISP, the present invention has made the following improvements:

[0051] The CCBs that receive features from the first level in the LW-ISP method are replaced with FCBs. At the same time, the input of the last FCB is adjusted, and the features received by the last FCB from the first level FGAM in the LW-ISP method are replaced with the packed image.

[0052] First, the characteristics of FCB received from the first and second levels of the improved LW-ISP method are defined as follows:

[0053] The output features received from the first-level FGAM are denoted as L1, the output features received from the first-level DB are denoted as L2, and the output features received from the second-level convolutional layer are denoted as L3.

[0054] In the improved LW-ISP method, the specific processing method in FCB is as follows:

[0055] Convolution is performed on L2 and L3 respectively to obtain the corresponding global and local features;

[0056] Recombining the global and local features, and fine-tuning the channel size through convolution;

[0057] The fine-tuned features of L2 and L3 are upsampled respectively, and then residual learning and feature channel merging are performed to obtain feature A that matches the L1 dimension;

[0058] Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

[0059] After frequency enhancement using the improved LW-ISP method, a frequency-enhanced RGB image is obtained.

[0060] For the trained deep model, in order to verify the influence of high-frequency information of the image on the classification results, image reconstruction training was performed on the image classification dataset to explore whether the performance of the pre-trained classification model can be exceeded by only applying classification loss and frequency loss. The classification dataset CIFAR-100 was reconstructed and obtained. Figure 4 The reconstruction results shown in the figure, (a) is the convergence trend of different loss combinations, (b) is the convergence trend diagram after deleting the L_c curve, Figure 4 In the figure, L_c represents classification loss, L_hf represents high-frequency loss, L_lf represents low-frequency loss, and Original represents the pre-training of ResNet-18. The horizontal axis represents the training period, and the vertical axis represents the confidence level (%). The reconstruction network trained with high-frequency loss and classification loss has the best performance, surpassing the pre-trained ResNet-18. Figure 4 It can be seen that the high-frequency information of the image is very important for the effective prediction of the model, while the low-frequency information of the image is not always positively fed back into the recognition results.

[0061] Example 2

[0062] The second embodiment is a preferred embodiment based on the first embodiment. The principle of the improved LW-ISP method is as follows: Figure 5 For the sake of convenience, we first Figure 5 The serial numbers in the figure are used to illustrate the network structure. 1 represents the downsampling block, 2 represents the FGAM, 3 represents the convolutional layer, 4 represents the FCB, 5 represents the FC layer, and 6 represents the post-processing upsampling block. The overall encoder-decoder structure is improved on the basis of the existing LW-ISP. The first level J3 progressively downsamples the feature map four times according to different levels. The second level J4 includes four upsampling groups. The main framework is the same as the existing LW-ISP method structure, and the CCB is replaced by the FCB (Frequency Complement Block) as the cascade between the first level J3 and the second level J4. The FCB connection is used as the baseline. The input of the entire improved LW-ISP method is a RAW image, and the output is the corresponding processed RGB image. Through the FCB, frequency enhancement is performed in the entire network structure.

[0063] The output features received by FCB from the FGAM of the first level J3 are denoted as L1 (H×W) for feature complementarity, the output features received from the DB of the first level J3 are denoted as L2 (2H×2W) for frequency decomposition, and the output features received from the convolutional layer of the second level J4 are denoted as L3 (H×W). FCB is used for frequency enhancement during the upsampling process to suppress the loss of image information during scaling, and the RAW input large-scale features in the first level J3 are used to provide semantic information.

[0064] like Figure 6 As shown, FCB first performs sub-pixel convolution on the H×W first-level and H×W second-level layers to obtain global and local features. Re-pixel re-convolution and convolution are then performed, with the convolution operations before and after re-pixel re-convolution used for channel resizing and fine-tuning, respectively. Subsequently, the first UB fuses the obtained features through residual learning to derive coarse, high-resolution features. Figure 6 In the example, UB1 is the first UB and UB2 is the second UB. In UB1, the fine-tuned features of L2 and L3 are upsampled (performed in the second UB), and then residual learning and feature channel merging are performed. The upsampling and residual learning in the second UB can be expressed by the following formula:

[0065]

[0066] Among them, R after It represents upsampling the features after fine-tuning of L2 and L3 respectively, and then performing residual learning on the features. P() represents sub-pixel convolution, and + represents residual learning. It represents the features obtained after L2 fine-tuning. It represents the features obtained after L3 fine-tuning.

[0067] Regarding how to extract the frequency band of an image, this paper uses discrete wavelet transform (DWT). Compared with other frequency analysis methods such as Fourier transform (DFT), DWT can capture both spatial information and frequency information in symbols, making it a more effective low-level visual analysis method. Given a function ψ, let X(ψ) be the set of expansions and shifts of ψ:

[0068] χ(ψ)={ψ jk =2 -j / 2 ψ(2 -j xk)j,k∈Z}

[0069] Here, ψ represents the orthogonal wavelet, χ(ψ) represents the set of dilations and shifts of ψ, k represents the translation parameter, j represents the scaling parameter, x represents the discrete sampled signal, that is, the specific pixel values of the image, and Z represents the set of integers. k and j determine the frequency and position of the wavelet function.

[0070] Using DWT, each image will be decomposed into four frequency bands: LL, LH, HL and HH. Among them, LL represents the low frequency band, LH, HL and HH represent different high frequency bands. DWT and IDWT are represented as Ψ(·) and The 2H×2W features of the first stage and the output of the first UB Perform DWT separately to get the frequency domain representation. Then calculate the key representation map as:

[0071]

[0072] Among them, R FC represents the result of the inverse discrete wavelet transform, Ψ represents the discrete wavelet transform, represents the inverse discrete wavelet transform, represents the output feature L1 received from the first-level FGAM, Indicates feature A.

[0073] In addition, after frequency enhancement using the improved LW-ISP method, the frequency-enhanced image is sequentially subjected to convolution, upsampling, and convolution. Upsampling is performed here to match the size of the target image, scaling the low-resolution feature map to the same resolution as the original input. Furthermore, while upsampling can increase the resolution of the image, it cannot capture high-level semantic information. To preserve useful feature information after upsampling, convolution is added here to help extract and integrate feature information and increase the ability of nonlinear transformations to improve model performance.

[0074] In order to demonstrate the technical effect of the improved LW-ISP of the present invention, the following verification experiments were conducted:

[0075] First, the verification experiment setup is described. The experiment evaluates the effectiveness of the proposed method on the ZurichRAW to RGB (Zurich) dataset, which is currently the largest ISP dataset. In addition, the image denoising and enhancement effects are also evaluated on the SIDD, DND, and LoL datasets. The specific verification results are as follows:

[0076] 1. Image processing effects

[0077] Table 1 shows the quantitative performance comparison results of different methods on the real RAW to RGB mapping problem:

[0078] Table 1 Quantitative performance comparison of different methods on the Zurich dataset from real RAW to RGB mapping problem

[0079]

[0080] To ensure fairness in the quantitative performance comparisons in Table 1, no data augmentation or additional supervision was added to the default experimental settings. For example, the existing LW-ISP method does not account for heterogeneous knowledge extraction. In terms of PSNR, our method improves the baseline (21.31 dB) by 0.32 dB and achieves the current best performance (21.63 dB). Figure 7 The visual effects of the present invention, the previous best model (existing LW-ISP), and Huawei P20 are compared when processing different RAW images. Figure 7 In the figure, A1 is the RAW image to be processed, A2 is the image obtained by Huawei P20, A3 is the image obtained by the existing LW-ISP, and A4 is the image obtained by using the method of the present invention. It can be seen that the images taken by Huawei P20 are usually darker and over-render the sky and other backgrounds. The processed images obtained by the method of the present invention are more consistent with natural features and can produce better details.

[0081] 2. Subtask Results

[0082] For subtasks, the potential of reducing computational cost, denoising and enhancing various subtasks of images is further explored.

[0083] (1) Image denoising. The framework of the proposed method was trained on the SIDD training set and directly evaluated on test images from the SIDD and DND datasets. The quantitative comparison of the SIDD dataset is shown in Table 2:

[0084] Table 2 Comparison of image denoising effects of different methods

[0085]

[0086] As can be seen from Table 2, the proposed method has superior performance compared to other methods. The existing LW-ISP result is 39.20dB, while the proposed method achieves 39.40dB and 0.950 in PSNR and SSIM, respectively. In addition, when RIDNet and VIDNet use additional training data, the proposed method provides better results, but the FLOPs of the proposed method (4.22G) are reduced by 136 times and 40.3 times compared to MPRNet (573.50G) and HINet (170.71G), respectively.

[0087] (2) Image enhancement: The proposed method also performs very well on the LoL dataset, with a PSNR of 20.23 dB, which is better than previous methods such as CRM and the existing LW-ISP.

[0088] 3. Recognition results

[0089] In order to achieve simultaneous measurement of visual effects and recognition performance, a dataset of RAW-RGB Recognition labels is required. This dataset is not yet available. An alternative is to use RGB images to label and transfer the labels to RAW. The transfer to RAW can be completed through the inverse model of ISP. In addition, image annotation data is generated, and the labels of the existing RGB images are generated by a pre-trained recognition model. Using the existing RAW-RGB dataset (Zurich) and the pre-trained recognition model (Swin-B) to generate labels, on the one hand, the amount of labeling is reduced, and on the other hand, the recognition accuracy of the RGB images can be effectively verified. The following comments are made from two aspects: quantitative results and qualitative results:

[0090] Quantitative Results: We first used the Swin-B model pre-trained on ImageNet-1K and ImageNet-22K to generate two different versions of labels for the test images of the Zurich dataset. In this experiment, we directly tested the images processed by the improved ISP model without pre-training on the recognition task. Table 3 shows the recognition results of different methods:

[0091] Table 3 Comparison of recognition results of different methods

[0092]

[0093] As can be seen from Table 3, compared with other methods, the performance of the proposed method is significantly improved, with Top 1 confidence of 8.2% and Top 5 confidence of 4.7%, and Top 5 confidence of 3.1% and 2.1%. Although the existing LW-ISP outperforms PyNET in terms of visual effects, its recognition performance actually declines. Similar trends can be observed in tests with different label qualities.

[0094] Qualitative results: Figure 8 Several examples are provided in the paper to compare the qualitative results of different methods in terms of recognition effect. Among them, (a) is the RGB image obtained after the input RAW is processed. (b) is the Top 5 categories obtained by recognition and their confidence values. (c) is the frequency domain mapping of the corresponding RGB image. (d) is the brightness statistics of the samples reconstructed in the frequency domain and the GT method. Figure 8 It can be seen from the above that the method of the present invention improves the recognition performance through frequency domain enhancement.

[0095] The above experiments demonstrated the potential for improvement using frequency domain information. Our processing method, without advanced assistance, bridges low-level vision and high-level recognition from a frequency perspective. We designed the FCB and constructed a new intelligent ISP algorithm, demonstrating the superiority of our method through experiments.

[0096] Example 3

[0097] To implement the above method, the present invention also proposes a frequency enhancement-based image processing system, including a LW-ISP system. The LW-ISP system includes four downsampling groups arranged sequentially in the first stage and four upsampling groups arranged sequentially in the second stage. The downsampling groups include DB and FGAM, and the first upsampling group includes a convolutional layer and an FC layer. Different from the existing LW-ISP system, the image processing system of the present invention also includes a packaging module:

[0098] The packaging module is used to package the images to be processed;

[0099] The second to fourth upsampling groups each include a convolutional layer and FCB arranged in sequence, which are used to perform frequency enhancement on the packed image;

[0100] The FCB is used to receive the output feature L1 of the first-level FGAM, the output feature L2 of the first-level DB, and the output feature L3 of the convolution layer in the same upsampling group; perform convolution processing on L2 and L3 respectively to obtain corresponding global and local features; perform pixel reorganization on the global and local features, and then fine-tune the channel size through convolution processing; upsample the corresponding fine-tuned features of L2 and L3 respectively, and then perform residual learning and feature channel merging to obtain feature A that matches the L1 dimension;

[0101] Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

[0102] As a preferred embodiment of the image processing system based on frequency enhancement of the present invention, it also includes a post-processing first convolution layer, a post-processing upsampling block and a post-processing second convolution layer connected in sequence, which is used to perform convolution processing, upsampling processing and convolution processing on the frequency enhanced image in sequence.

[0103] As a preferred embodiment of the frequency enhancement-based image processing system of the present invention, the FCB includes a first convolution sub-block, a pixel reorganization sub-module, a second convolution sub-block, a first UB, a DWT sub-module and an IDWT sub-module. The first convolution sub-block is used to perform convolution processing on L2 and L3 respectively to obtain corresponding global and local features. The pixel reorganization sub-module is used to perform pixel reorganization on the global and local features. The second convolution sub-block is used to fine-tune the feature channel size after pixel reorganization. The first UB is used to up-sample the fine-tuned features of L2 and L3 respectively, and then perform residual learning and feature channel merging to obtain feature A that matches the L1 dimension. The DWT sub-module is used to perform discrete wavelet transform on L1 and feature A respectively, and perform difference on the transformed results. The IDWT sub-module is used to perform inverse discrete wavelet transform on the features after difference in the DWT sub-module, and then sum them with feature A.

[0104] One embodiment of the present invention provides a computer device. The computer device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of each of the aforementioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in each of the aforementioned apparatus embodiments are implemented.

[0105] The computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to accomplish the present invention.

[0106] The computer device may be a desktop computer, a notebook computer, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory.

[0107] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0108] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.

[0109] If the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0110] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An image processing method based on frequency enhancement, characterized in that: The following steps are involved: Packing the images to be processed; The frequency enhancement of the packed image is performed using the improved LW-ISP method; The improved LW-ISP method performs frequency enhancement through the cascaded first and second stages. The first stage sequentially sets four downsampling groups, each of which includes a DB and a fine-grained attention module. The second stage sequentially sets four upsampling groups. The first upsampling group includes a convolutional layer and an FC layer, and the second to fourth upsampling groups each include a convolutional layer and a frequency compensation block. The features of the packed image received by the frequency compensation block in the fourth upsampling group are recorded as L1, and the output features received by the frequency compensation blocks in the second and third upsampling groups from the first-level fine-grained attention module are recorded as L1; the output features received by each frequency compensation block from the first-level DB are recorded as L2, and the output features received from the convolutional layer of the upsampling group to which it belongs are recorded as L3; The specific processing method in the frequency compensation block is: Convolution is performed on L2 and L3 respectively to obtain the corresponding global and local features; Recombining the global and local features, and fine-tuning the channel size through convolution; The fine-tuned features of L2 and L3 are upsampled respectively, and then residual learning and feature channel merging are performed to obtain feature A that matches the L1 dimension; Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

2. The image processing method based on frequency enhancement according to claim 1, characterized in that: After frequency enhancement is performed using the improved LW-ISP method, the method further includes: The frequency enhanced image is sequentially subjected to convolution processing, upsampling processing and convolution processing.

3. The image processing method based on frequency enhancement according to claim 1 or 2, characterized in that: The features after fine-tuning of L2 and L3 are upsampled respectively, and then residual learning is performed to obtain the features , which is calculated specifically by the following formula: in, It means that the features after L2 and L3 fine-tuning are up-sampled and then the features are obtained after residual learning. represents sub-pixel convolution, represents residual learning, It represents the features obtained after L2 fine-tuning. It represents the features obtained after L3 fine-tuning.

4. The image processing method based on frequency enhancement according to claim 3, characterized in that: The discrete wavelet transform is specifically calculated by the following formula: in, represents orthogonal wavelet, express The set of expansions and shifts, represents the translation parameter, represents the scaling parameter, represents a discrete sampled signal, Represents a set of integers.

5. The image processing method based on frequency enhancement according to claim 4, characterized in that: The inverse discrete wavelet transform is specifically calculated by the following formula: in, represents the result of the inverse discrete wavelet transform, represents discrete wavelet transform, represents the inverse discrete wavelet transform, represents the output feature L1 received from the first-level fine-grained attention module, Indicates feature A.

6. A frequency enhancement-based image processing system, comprising a LW-ISP system, wherein the LW-ISP system comprises four downsampling groups sequentially arranged in a first stage and four upsampling groups sequentially arranged in a second stage, each downsampling group comprising a DB and a fine-grained attention module, and the first upsampling group comprising a convolutional layer and an FC layer; characterized in that: Also includes packaging modules; The packaging module is used to package the images to be processed; The second to fourth upsampling groups each include a convolutional layer and a frequency compensation block arranged in sequence, for performing frequency enhancement on the packed image; The features of the packed image received by the frequency compensation block in the fourth upsampling group are recorded as L1, and the output features L1 received by the frequency compensation blocks in the second and third upsampling groups from the fine-grained attention module of the first level; each frequency compensation block receives the output features L2 from the first-level DB and the output features L3 from the convolutional layer of the upsampling group to which it belongs; The processing performed by the frequency compensation block includes: performing convolution processing on L2 and L3 respectively to obtain corresponding global and local features; performing pixel reorganization on the global and local features, and then fine-tuning the channel size through convolution processing; upsampling the corresponding fine-tuned features of L2 and L3 respectively, and then performing residual learning and feature channel merging to obtain feature A that matches the L1 dimension; Perform discrete wavelet transform on L1 and feature A respectively, make the difference between the transformed results, and then perform inverse discrete wavelet transform and sum them with feature A.

7. The image processing system based on frequency enhancement according to claim 6, characterized in that: It also includes a post-processing first convolution layer, a post-processing upsampling block, and a post-processing second convolution layer that are sequentially connected to process the frequency-enhanced image.

8. The image processing system based on frequency enhancement according to claim 6, characterized in that: The frequency compensation block includes a first convolution sub-block, a pixel reorganization sub-module, a second convolution sub-block, a first UB, a DWT sub-module and an IDWT sub-module; The first convolution sub-block is used to perform convolution processing on L2 and L3 respectively to obtain corresponding global and local features; The pixel reorganization submodule is used to perform pixel reorganization on the global and local features; The second convolution sub-block is used to fine-tune the feature channel size after pixel reorganization; The first UB is used to upsample the fine-tuned features of L2 and L3 respectively, and then perform residual learning and feature channel merging to obtain feature A that matches the L1 dimension; The DWT submodule is used to perform discrete wavelet transform on L1 and feature A respectively, and perform difference on the transformed results; The IDWT submodule is used to perform inverse discrete wavelet transform on the features obtained by subtracting the features from the DWT submodule, and then sum them with feature A.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Retina OCT image effusion segmentation method based on deep learning

    CN116503593A

  • Text-guided image restoration model and method based on multi-granularity image-text semantic learning

    CN116523799A