Supervision method based on visual neural network, related device, and image processing method
By adopting a supervision method based on visual neural network in the intelligent ISP training stage, the problem of insufficient recognition effect in the existing technology is solved, and the performance improvement of visual neural network and the recognition performance enhancement are achieved.
Patent Information
- Application Number
- CN202311006586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-08-10
AI Technical Summary
The existing supervision methods of smart ISPs during the training stage cannot adapt to the increasingly stringent image recognition requirements, resulting in insufficient recognition effect.
A supervision method based on visual neural network is proposed. During the visual neural network training process, the regularization distance in all frequency domains is minimized in the early stage, and the regularization distance in the high frequency domain is only supervised in the middle and late stages to avoid interference with the learning process.
Through this supervision method, the denoising and recognition performance of visual neural networks has been greatly improved, which can better adapt to strict recognition requirements, and further improve image processing performance through multi-dimensional loss supervision.
Smart Images

Figure CN117315384B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to a neural network supervision method, and particularly relates to a supervision method based on a visual neural network, related devices, and an image processing method. Background Art
[0002] Deep Neural Networks (DNN) have achieved results beyond traditional algorithms in many fields, such as image classification and cancer detection. Image Signal Processing (ISP) methods receive and process the original signals of camera imaging sensors and play a decisive role in the quality of images. Among them, the combination of ISP series image preprocessing tasks and DNN is a current research hotspot. In particular, intelligent ISP derived from existing Deep Learning (DL)-based methods by the ISP pipeline can learn the statistical information of the original image and accept joint solutions for multiple tasks. Therefore, intelligent ISP can creatively link the photo imaging process with subsequent recognition applications.
[0003] Before processing images using ISP, as a preprocessing step, it is also necessary to train ISP to make it in the best state. Therefore, in the ISP training stage, it is particularly important to improve the processing effect of ISP. Currently, there are also corresponding supervision methods in the ISP training stage. However, as the requirements for image recognition effects in various fields are increasing day by day, using existing recognition methods to supervise in the ISP training stage results in insufficient recognition effects of the trained ISP. Therefore, the existing supervision methods can no longer meet the increasingly stringent recognition requirements. Summary of the Invention
[0004] Aiming at the technical problem that in the training stage of intelligent ISP, the existing supervision methods can no longer meet the increasingly stringent recognition requirements in terms of recognition effects, a supervision method based on a visual neural network, related devices, and an image processing method are proposed.
[0005] To achieve the above object, the present invention is implemented by the following technical solutions:
[0006] In the first aspect, the present invention proposes a supervision method based on a visual neural network. During the training process of the visual neural network, for the RGB image output by the visual neural network, it includes the following steps:
[0007] In the first 1 / a stage of the visual neural network training, minimize the regularized distance between the output RGB image and the ground truth in the low-frequency domain and high-frequency domain; where a is greater than 2;
[0008] From the 1 / a stage of visual neural network training to the completion of training, minimize the regularized distance between the output RGB image and the ground truth in the high-frequency domain.
[0009] In a second aspect, the present invention proposes a supervision system based on a visual neural network. During the training process of the visual neural network, for the RGB image output by the visual neural network, it includes a first minimization module and a second minimization module;
[0010] The first minimization module is used to minimize the regularized distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain during the first 1 / a stage of visual neural network training; where a is greater than 2;
[0011] The second minimization module is used to minimize the regularized distance between the output RGB image and the ground truth in the high-frequency domain from the 1 / a stage of visual neural network training to the completion of training.
[0012] In a third aspect, the present invention proposes a computer program product. The computer program product contains instructions that, when executed by a processor, implement the above method.
[0013] In a fourth aspect, the present invention proposes an image processing method. The image is processed by a trained visual neural network; during the training process of the visual neural network, for the RGB image output by the visual neural network, the above-mentioned supervision method based on the visual neural network is used for supervision.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] The present invention proposes a supervision method based on a visual neural network. During the training process of the visual neural network, for the RGB image output by the visual neural network, in the early stage of training, minimizing the regularized distance in all frequency domains helps the performance to rise rapidly. In the middle and late stages of training, only the regularized distance in the high-frequency domain is supervised to avoid interfering with the learning process of the visual neural network. And through actual verification, after training is completed using the supervision method of the present invention, the denoising and recognition performances of the corresponding model of the visual neural network have been greatly improved.
[0016] Furthermore, the present invention also introduces reconstruction loss supervision and structure loss supervision, which improves the image processing performance of the corresponding model of the visual neural network from multiple dimensions.
[0017] The present invention also proposes a supervision system based on a visual neural network for implementing the above supervision method. At the same time, it also proposes a computer program product that can implement the application of the above method by means of different hardware media. The present invention also proposes an image processing method, all of which have all the advantages of the above supervision method and are also convenient for the popularization and application of the present invention.
[0018] Furthermore, for the image processing method proposed by the present invention, an improved LW-ISP method is used to perform frequency enhancement on the image. First, the image to be processed is packed, and then the packed image is subjected to frequency enhancement through the improved LW-ISP method. The existing LW-ISP method is improved, and based on deep learning, it can retain and enhance frequency domain information at the ISP level. The frequency information lost during ISP processing can be supplemented by FCB. Combining discrete wavelet transform, both spatial information and frequency information can be captured. After actual verification, the image processing method of the present invention, in terms of image processing, the processed image is more in line with natural features, can better retain detail information, and has obvious denoising and enhancement effects. Through frequency domain enhancement, the recognition performance is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a comparison diagram of the visual effects of the original image, traditional ISP, and intelligent ISP;
[0021] Figure 2 It is a schematic diagram of the reconstruction results of the classification dataset CIFAR-100; among them, (a) is a schematic diagram of the convergence trend of different loss combinations, and (b) is a schematic diagram of the convergence trend after deleting the L_c curve in (a);
[0022] Figure 3 It is a schematic diagram of the process of Embodiment 1;
[0023] Figure 4 It is a schematic diagram of the process of Embodiment 2;
[0024] Figure 5 It is a schematic diagram of the principle of the existing LW-ISP;
[0025] Figure 6 It is a schematic diagram of the principle of the improved LW-ISP;
[0026] Figure 7 It is a schematic diagram of the principle of FCB;
[0027] Figure 8 It is a schematic diagram of the training results and frequency domain distance of the model corresponding to the improved LW-ISP method; among them, (a) is a schematic diagram of the training results, and (b) is a schematic diagram of the frequency domain distance;
[0028] Figure 9 It is a comparison chart of the existing LW-ISP method, Huawei P20, and the method of the present invention for processing the same RAW image;
[0029] Figure 10 It is a qualitative comparison chart of the recognition effects of different methods; among them, (a) is the RGB image obtained after processing the input RAW, (b) is the top 5 categories and their confidence values obtained through recognition, (c) is the frequency-domain mapping of the corresponding RGB image, and (d) is the brightness statistics of the samples reconstructed in the frequency domain and the GT method. Specific embodiments
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0031] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0032] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0033] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0034] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0035] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if the terms "set", "installed", "connected", "connected" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0036] In recent years, it has become increasingly popular to make network models learn frequency-domain features, especially in low-level vision. Specifically, for example: performing phased processing on the high-frequency information and low-frequency information of images for the enhancement of low-light images. Adding high-frequency residual images to blurred images to enrich detail and texture information. Designing a focal frequency loss during image synthesis. Improving the conversion performance between images through wavelet loss, etc. Although there are many related studies, regarding how to improve the reconstruction quality of RAW images based on the frequency domain, there are still few related studies.
[0037] In a limited dataset, there is a correlation between high-frequency information and the semantics represented by pictures. However, for a trained deep model, it is not clear how much different frequency components affect the performance. As Figure 1 shown,
[0038] As Figure 1 shown in the last column, the high-frequency information of the image is crucial for the classification result. In addition, image reconstruction training was performed on the image classification dataset (CIFAR-100
[35] ) to explore whether it is possible to recover or even exceed the performance of the pre-trained classification model only using classification loss and frequency loss.
[0039] As Figure 2 shown, (a) is the convergence trend of different loss combinations, and (b) is the schematic diagram of the convergence trend after deleting the L_c curve. Figure 2 Among them, L_c represents the classification loss, L_hf represents the high-frequency loss, L_lf represents the low-frequency loss, Original represents the pre-training of ResNet-18. The abscissa in the figure represents the training epoch, and the ordinate represents the confidence (%)). The reconstruction network trained with high-frequency loss and classification loss has the best performance, exceeding the pre-trained ResNet-18. It can be seen from Figure 2 that the high-frequency information of the image is very important for the effective prediction of the model, while the low-frequency information of the image is not always positively fed back to the recognition result, especially in the setting without any pixel-wise supervision.
[0040] Based on the understanding of high-frequency information, the present invention proposes a supervision method based on a visual neural network, related devices, and an image processing method. The following further describes the present invention in detail with reference to the accompanying drawings and embodiments:
[0041] Embodiment 1
[0042] Refer to Figure 3 , Embodiment 1 of the present invention is a basic embodiment of a supervision method based on a visual neural network. During the training process of the visual neural network, for the RGB image output by the visual neural network, supervision is performed through the following steps:
[0043] S101. In the first 1 / a stage of the visual neural network training, minimize the regularization distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain; where a is greater than 2;
[0044] S102. From the 1 / a stage of the visual neural network training to the completion of training, minimize the regularization distance between the output RGB image and the ground truth in the high-frequency domain.
[0045] The core inventive point of Embodiment 1 is that in the early stage of training, minimizing the regularization distance in all frequency domains helps the performance to rise rapidly. In the middle and late stages of training, only supervise the regularization distance in the high-frequency domain to avoid interfering with the learning process of the visual neural network.
[0046] It should be noted that the "low-frequency domain" and "high-frequency domain" are relative concepts well-known in the field of image processing, and no specific range is defined in this field. It can be understood according to the conventional understanding in this field.
[0047] Embodiment 2
[0048] As Figure 4 shown, as a preferred embodiment of a supervision method based on a visual neural network of the present invention, supervision is specifically performed through the following steps:
[0049] S201. In the first 1 / a stage of the visual neural network training, minimize the regularization distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain; where a is greater than 2;
[0050] S202. From the 1 / a stage of the visual neural network training to the completion of training, minimize the regularization distance between the output RGB image and the ground truth in the high-frequency domain;
[0051] S203. Reconstruction loss supervision
[0052] Given a training sample I, the RGB image and the ground truth output by the visual neural network are denoted as f(I) and J respectively. The mean absolute error (MAE) is used to measure the difference L between f(I) and J r :
[0053] L r = |f(I) - J|
[0054] Make the training result of the visual neural network meet the preset reconstruction loss;
[0055] If the calculated reconstruction loss L r does not meet the preset reconstruction loss requirement, continue to train the visual neural network until the reconstruction loss L r meets the preset reconstruction loss requirement.
[0056] S204, Structural loss supervision
[0057] Use the multi-scale structural similarity (MS-SSIM) loss to increase the dynamic range of the reconstructed image:
[0058] Calculate the structural loss L between the output RGB image and the ground truth through the following formula s , make the training result of the visual neural network meet the preset structural loss:
[0059]
[0060] where l M represents the luminance similarity; a represents the weight coefficient used to balance the importance of luminance, contrast, and structural similarity; M represents the maximum scale of the image; j represents the specific scale at which the original image is divided; c j represents the contrast similarity at the j-th scale; s j represents the structural similarity at the j-th scale.
[0061] If the calculated structural loss L s does not meet the preset structural loss requirement, continue to train the visual neural network until the structural loss L s meets the preset structural loss requirement.
[0062] It should be noted that the above steps S201 and S202 together constitute the minimization supervision. The order of minimization supervision, reconstruction loss supervision, and structural loss supervision may not be executed in the order in Embodiment 2, may also be performed simultaneously, or executed in any other order.
[0063] In some other embodiments of the present invention, based on local and global correction and perceptual quality of image signal preprocessing, multiple loss functions are used for training, and combined with the reconstruction loss L r , the structural loss L s and the minimization result L FAS are jointly supervised by the following formula:
[0064] L overall =αL r +L s +βL FAS
[0065] wherein, L overall represents the comprehensive reconstruction loss, structural loss and the supervised result of minimization. As an optimal solution, α = 0.3 and β = 0.6, which can improve by about 0.1 dB. α and β are two hyperparameters for balancing the magnitudes of L r and L FAS .
[0066] A batch within the range of 2 to 12 will not have an obvious impact on the final result, but only affect the training speed. Herein, the batch size refers to the number of data input to the model at one time during the model training process. In addition, an appropriate initial learning rate is required. When the learning rate is greater than 3×10 -4 , the model will gradually collapse. When the learning rate is low, the training process will be insufficient.
[0067] Embodiment III
[0068] The present invention also proposes an image processing method, which processes an image through a trained visual neural network. During the training process of the visual neural network, for the RGB image output by the visual neural network, the aforementioned supervision method based on the visual neural network is used for supervision.
[0069] Embodiment IV
[0070] As a preferred embodiment of the image processing method proposed by the present invention, the specific process of the visual neural network for processing the image is as follows:
[0071] S101, perform packaging processing on the image to be processed.
[0072] S102, perform frequency enhancement on the packaged image through the improved LW-ISP method.
[0073] Such as Figure 5As shown in the figure, it is a schematic diagram of the principle of the existing LW-ISP, mainly including the first stage J1 and the second stage J2. In the first stage, the feature map is gradually downsampled at different levels to speed up the calculation. Each downsampling is achieved through a downsampling block and a FGAM (Fine-grained Attention Module). In the second stage, through symmetric skip connections, the processed global vector is further connected to the feature map of the same size in the first half, including a convolutional layer, an FC layer, a convolutional layer, a CCB (Contextual Complement Upsampling Block), a convolutional layer, a CCB, a convolutional layer, a CCB, a convolutional layer, a CCB, and a convolutional layer set in sequence. Four CCBs are set in the second stage to perform adaptive high-frequency decomposition in the feature space and fuse the corresponding size features of the previous stage. Based on the existing LW-ISP, the present invention has made the following improvements:
[0074] All CCBs that receive features from the first level in the LW-ISP method are replaced with FCBs. At the same time, the input of the last FCB is adjusted, and the feature received by the last FCB from the first-level FGAM in the LW-ISP method is replaced with the image after packing processing.
[0075] First, define the features received by the FCB from the first level and the second level in the improved LW-ISP method as follows:
[0076] The output feature received from the first-level FGAM is denoted as L1, the output feature received from the first-level DB is denoted as L2, and the output feature received from the second-level convolutional layer is denoted as L3.
[0077] In the improved LW-ISP method, the specific processing method in the FCB is as follows:
[0078] Perform convolutional processing on L2 and L3 respectively to obtain the corresponding global and local features;
[0079] Perform pixel recombination on the global and local features, and then fine-tune the channel size through convolutional processing;
[0080] Upsample the features after fine-tuning corresponding to L2 and L3 respectively, then perform residual learning and feature channel merging to obtain feature A that matches the dimension of L1;
[0081] Perform discrete wavelet transform on L1 and feature A respectively, take the difference of the transformed results, and then sum with feature A after inverse discrete wavelet transform.
[0082] After frequency enhancement through the improved LW-ISP method, a frequency-enhanced RGB image is obtained.
[0083] The principle of the improved LW-ISP method is specifically as follows Figure 6 As shown, for the sake of convenience in explanation, first, the serial numbers in Figure 6 are explained. 1 represents the downsampling block, 2 represents the FGAM, 3 represents the convolutional layer, 4 represents the FCB, 5 represents the FC layer, and 6 represents the post-processing upsampling block. The overall structure adopts an encoder-decoder architecture, which is improved based on the existing LW-ISP. The first stage J3 performs four times of downsampling feature mapping progressively at different levels. The second stage J4 includes four upsampling groups. The main framework has the same structure as the existing LW-ISP method. The CCB is replaced by the FCB (Frequency Complement Block), and the FCB is connected as the baseline for the cascade between the first stage J3 and the second stage J4. The input of the entire improved LW-ISP method is the RAW image, and the output is the corresponding processed RGB image. Through the FCB, frequency enhancement is performed in the entire network structure.
[0084] The output feature received by the FCB from the FGAM of the first stage J3 is denoted as L1(H×W) for feature complementarity. The output feature received by the FCB from the DB of the first stage J3 is denoted as L2(2H×2W) for frequency decomposition. The output feature received by the FCB from the convolutional layer of the second stage J4 is denoted as L3(H×W). During the upsampling process, the FCB is used for frequency enhancement to suppress the loss of image information during scaling, and the large-scale features of the RAW input in the first stage J3 are used to provide semantic information.
[0085] As Figure 7 shown, the FCB first performs sub-pixel convolution on the H×W of the first stage and the H×W of the second stage to obtain global and local features. Then, pixel recombination and convolution are carried out. The convolution operations before and after pixel recombination are used for channel size adjustment and fine-tuning respectively. Subsequently, the first UB fuses the obtained features through residual learning to derive rough high-resolution features. Figure 7 In
[0086]
[0087] where, R after represents the feature obtained by respectively performing upsampling on the features corresponding to L2 and L3 after fine-tuning, and then performing residual learning. P() represents sub-pixel convolution, + represents residual learning, represents the feature obtained after fine-tuning the corresponding L2, represents the feature obtained after fine-tuning the corresponding L3.
[0088] Regarding how to extract the frequency domain of an image, the present invention uses discrete wavelet transform (DWT). Compared with other frequency analysis methods such as Fourier analysis (DFT), DWT can capture both the spatial information in the symbol and the frequency information in the symbol, and is a more effective low-level visual analysis method. Given a function ψ, let X(ψ) be the set of dilations and shifts of ψ:
[0089] χ(ψ) = {ψ jk = 2 -j / 2 ψ(2 -j x - k)| j,k ∈ Z}
[0090] where ψ represents an orthogonal wavelet, χ(ψ) represents the set of dilations and shifts of ψ, k represents the translation parameter, j represents the scaling parameter, x represents the discrete sampling signal, that is, the specific pixel value of the image, and Z represents the set of integers. k and j determine the frequency and position of the wavelet function.
[0091] Using DWT, each image will be decomposed into four frequency domains: LL, LH, HL, and HH. Among them, LL represents the low-frequency domain, and LH, HL, and HH represent different high-frequency domains respectively. Denote DWT and IDWT as Ψ(·) and Perform DWT on the 2H×2W features in the first stage and the output of the first UB respectively to obtain the frequency domain representation. Then the key characterization map is calculated as:
[0092]
[0093] where R FC represents the result of the inverse discrete wavelet transform, Ψ represents the discrete wavelet transform, represents the inverse discrete wavelet transform, represents the output feature L1 received from the first-level FGAM, R a ′ fter represents feature A.
[0094] In addition, after frequency enhancement by the improved LW-ISP method, convolution processing, upsampling processing, and convolution processing are sequentially performed on the frequency-enhanced graph. The upsampling processing is performed here to match the size of the target image and enlarge the low-resolution feature map to the same resolution as the original input. In addition, although the upsampling operation can increase the resolution of the image, it cannot capture high-level semantic information. In order to retain useful feature information after upsampling, convolution processing is added here after upsampling to help extract and integrate feature information and increase the ability of non-linear transformation to improve the performance of the model.
[0095] Based on Embodiment 4, the present invention designs a new frequency-domain supervision method for the training process of visual neural networks. As Figure 8 shown, it is a schematic diagram of the training results and frequency-domain distance of the improved LW-ISP method corresponding model. As can be seen from (a), the rapid improvement of the model performance in the early stage benefits from the common convergence of all frequency domains, especially when mainly learning the low-frequency components (LL) at this time. As can be seen from (b), in the middle and late stages of training, when the model performance chases the SOTA, it no longer focuses on the low-frequency components (LL), and the high-frequency domain (HH) plays a dominant role. Therefore, the frequency-domain supervision designed by the present invention changes with the training stage. Representing the DWT (Discrete Wavelet Transform) as Ψ(·), the high-frequency domain and low-frequency domain of the image x can be respectively deduced as Ψ H (x) and Ψ L (x). After synchronizing the RGB image output by the visual neural network model with the ground truth J, minimize the difference between the two:
[0096]
[0097] where Ψ represents the wavelet discrete transform, L represents the low-frequency domain, H represents the high-frequency domain, st represents the training stage of the visual neural network, st = 0 represents the first 1 / a stage, st = 1 represents from the 1 / a stage to the end of training, and L FAS represents the minimization result.
[0098] Corresponding to the full-frequency domain convergence trend in the early training stage (st = 0), minimizing the regularization distance of all frequency domains helps the rapid increase of performance. Corresponding to the high-frequency convergence trend in the middle and late stages of training (st = 1), only supervise the regularization distance of the high-frequency domain. Experiments show that adding low-frequency supervision at this time will significantly interfere with the learning process of the model.
[0099] In some other embodiments of the present invention, it has been experimentally verified that when a = 5, the robustness is the best, and the supervision effect of the present invention can reach the best.
[0100] Regarding the influence of the FCB in Embodiment 4 of the present invention and the supervision method of the present invention, as shown in Table 1. The network structure adopts the improved LW-ISP. The following baseline includes the main structure of the improved LW-ISP. Except for the skip connection, it can reach 21.3076dB and 0.8567dB in PSNR and MS-SSIM.
[0101] Table 1 Ablation study results table of network structure
[0102]
[0103] In Table 1, FCB-UB represents the second UB in the FCB, FCB-FC represents the first UB in the FCB, and FAS represents the supervision method of the present invention. The supervision method of the present invention can be extended to different low-level vision tasks.
[0104] In addition, ablation experiments were conducted on the supervision method of the present invention in image denoising, specifically in the SIDD dataset for image denoising experiments. As shown in Table 2, the early training stage is abbreviated as Stage I, and the middle and late training stages are abbreviated as Stage II. Among them, the LH, HL, and HH frequency domains are collectively referred to as the high-frequency domain (HF), while the LL frequency domain is called the low-frequency domain (LF). All frequency domains are abbreviated as A11. It can be seen from Table 2 that higher performance can be obtained by adopting the supervision method of the present invention. In addition, concentrating the model on learning the low-frequency domain components has little effect, and adding it in the middle and late stages of training will even interfere with normal convergence.
[0105] Table 2 Ablation experiment results of the supervision method of the present invention in image denoising
[0106]
[0107] For the improved LW-ISP in the preferred solution of the present invention, in order to prove the technical effect of the improved LW-ISP, the following verification experiments were conducted:
[0108] First, the verification settings of the verification experiment are described. In the experiment, the effectiveness of the method of the present invention was evaluated on the ZurichRAW to RGB (abbreviated as Zurich) dataset, which is currently the largest ISP dataset. In addition, the effects of image denoising and enhancement were also evaluated on the SIDD, DND, and LoL datasets. The specific verification results are as follows:
[0109] 1. Image processing effect
[0110] Table 3 shows the quantitative performance comparison results of different methods on the real RAW to RGB mapping problem:
[0111] Table 3 Quantitative performance comparison table of the Zurich dataset on the real RAW to RGB mapping problem under different methods
[0112]
[0113] To ensure the fairness of the quantitative performance comparison in Table 3, data augmentation and additional supervision are not added in the default experimental settings. For example, the existing LW-ISP method does not count in heterogeneous knowledge extraction. In terms of PSNR, the method of the present invention brings an improvement of 0.32 dB to the baseline (21.31 dB) and achieves the current best result (21.63 dB). Figure 9The visual effects of the present invention, the previous best model (existing LW-ISP), and the Huawei P20 were compared when processing different RAW images. Figure 9 In Figure 9 , A1 is the RAW image to be processed, A2 is the image obtained by the Huawei P20, A3 is the image obtained by the existing LW-ISP, and A4 is the image obtained by using the present invention. It can be seen that the images taken by the Huawei P20 are usually darker and over-rendered on the sky and other backgrounds. The processed images obtained by using the method of the present invention are more in line with natural features and can produce better details.
[0114] 2. Sub-task results
[0115] For sub-tasks, further explore the potential of reducing computational costs, denoising, and enhancing various sub-tasks of images.
[0116] (1) In terms of image denoising. Train the framework of the method of the present invention on the training set of SIDD and directly evaluate it on the test images of the SIDD and DND data sets. The quantitative comparison of the SIDD data set is shown in Table 4:
[0117] Table 4 Comparison of image denoising effects of different methods
[0118]
[0119] As can be seen from Table 4, the present invention has excellent performance compared with other methods. The result of the existing LW-ISP is 39.20dB, and the method of the present invention reaches 39.40dB and 0.950 in PSNR and SSIM respectively. In addition, when RIDNet and VIDNet use additional training data, the method of the present invention provides better results, but the FLOP (4.22G) of the method of the present invention is reduced by 136 times and 40.3 times compared with MPRNet (573.50G) and HINet (170.71G) respectively.
[0120] (2) In terms of image enhancement. The performance of the method of the present invention on the LoL data set is also very beneficial, and the PSNR can reach 20.23dB, which is better than previous methods such as CRM and the existing LW-ISP.
[0121] 3. Recognition results
[0122] To simultaneously measure visual effects and recognition performance, a dataset with RAW-RGB Recognition labels is required, which is not yet available. An alternative is to use RGB images to label the tags and transfer the tags to RAW. The transfer to RAW can be accomplished through the inverse model of the ISP. Additionally, annotation data for the images is generated, and labels for existing RGB images are generated through a pre-trained recognition model. Using the existing RAW-RGB dataset (Zurich) and the pre-trained recognition model (Swin-B) to generate labels reduces the amount of labeling on the one hand and can effectively verify the recognition accuracy of RGB images on the other hand. The following is a review from two aspects: quantitative results and qualitative results.
[0123] Quantitative results: First, the Swin-B model pre-trained on ImageNet-1K and ImageNet-22K respectively was used to generate two different versions of labels for the test images of the Zurich dataset. In this experiment, we directly tested the images processed by the improved ISP model without pre-training for the recognition task. Table 5 shows the recognition results of different methods:
[0124] Table 5 Comparison table of recognition results of different methods
[0125]
[0126] As can be seen from Table 5, compared with other methods, the performance of the method of the present invention has greatly improved the Top 1 confidence and Top 5 confidence. The Top 1 is 8.2% and 4.7%, and the Top 5 is 3.1% and 2.1%. Although the existing LW-ISP outperforms PyNET in terms of visual effects, in fact, its recognition performance has declined, and a similar trend can be observed in tests with different label qualities.
[0127] Qualitative results: Several examples are provided in Figure 10 to compare the qualitative results of different methods in terms of recognition effects. Among them, (a) is the RGB image obtained after processing the input RAW. (b) is the Top 5 categories and their confidence values obtained through recognition. (c) is the frequency-domain mapping of the corresponding RGB image. (d) is the luminance statistics of the samples reconstructed in the frequency domain and the GT method. As can be seen from Figure 10 the method of the present invention improves the recognition performance through frequency-domain enhancement.
[0128] Through the above experiments, the improvement potential of frequency-domain information is verified. The processing method of the present invention connects low-level vision and high-level recognition from the frequency perspective without advanced assistance. The FCB is designed, a new intelligent ISP algorithm is constructed, and the superiority of the method of the present invention is proven through experiments.
[0129] The present invention also provides a computer program product, which includes instructions that, when executed by a processor, implement the steps in the above-described clearing method embodiments. Alternatively, when the processor executes the computer program, it implements the functions of the various modules / units in the above-described system embodiments.
[0130] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in a memory and executed by the processor to complete the present invention.
[0131] The carrier for implementing the above computer program product can be a computer device. The computer program product can also be stored in a computer storage medium.
[0132] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory.
[0133] The processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0134] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by invoking the data stored in the memory.
[0135] If the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0136] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image processing method that processes an image through a trained visual neural network; Characterized in that: During the training process of the visual neural network, for the RGB image output by the visual neural network, the following supervision method based on the visual neural network is used for supervision: In the first 1 / a stage of the visual neural network training, minimize the regularization distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain; where a is greater than 2; From the 1 / a stage of the visual neural network training to the completion of training, minimize the regularization distance between the output RGB image and the ground truth in the high-frequency domain; The method for processing an image through a trained visual neural network includes: Performing a packaging process on the image to be processed; Performing frequency enhancement on the packaged image through an improved LW-ISP method; The specific improved LW-ISP method is: The improved LW-ISP method includes a first stage and a second stage. Among them, the first stage includes multiple groups of downsampling blocks and a fine-grained attention module; the second stage includes a convolutional layer, an FC layer, four groups of convolutional layers, an FCB, and a convolutional layer connected in sequence; the feature received by the last FCB is the packaged image; Define the features received by the FCB from the first stage and the second stage of the improved LW-ISP method as follows: The output feature received from the FGAM in the first stage is denoted as L1, the output feature received from the DB in the first stage is denoted as L2, and the output feature received from the convolutional layer in the second stage is denoted as L3; The specific processing method in the FCB is: Perform convolutional processing on L2 and L3 respectively to obtain corresponding global and local features; Perform pixel recombination on the global and local features, and then fine-tune the channel size through convolutional processing; Perform upsampling on the features of L2 and L3 after fine-tuning respectively, and then perform residual learning and feature channel merging to obtain a feature A that matches the dimension of L1; Perform discrete wavelet transform on L1 and feature A respectively, take the difference of the transformed results, and then sum with feature A after the inverse discrete wavelet transform.
2. The image processing method according to claim 1, Characterized in that: The a is equal to 5.
3. The image processing method according to claim 1 or 2, Characterized in that, It further includes reconstruction loss supervision: Calculate the reconstruction loss between the output RGB image and the ground truth through the following formula : Among them, represents the output RGB image, represents the ground truth, represents the training sample; Make the training result of the visual neural network meet the preset reconstruction loss requirement; otherwise, continue to train the visual neural network until the reconstruction loss meets the preset reconstruction loss requirement.
4. The image processing method according to claim 1 or 2, Characterized in that, It further includes structural loss supervision: Calculate the structural loss between the output RGB image and the ground truth through the following formula : Among them, represents the luminance similarity, represents the weight coefficient, represents the maximum scale of the image, represents the specific scale at which the original image is divided, represents the th contrast similarity at a certain scale, represents the th structural similarity at a certain scale; Make the training result of the visual neural network meet the preset structural loss requirement; otherwise, continue to train the visual neural network until the structural loss meets the preset structural loss requirement.
5. The image processing method according to claim 1 or 2, Characterized in that: It further includes reconstruction loss supervision and structural loss supervision; Calculate the reconstruction loss between the output RGB image and the ground truth by the following formula :[[]]END]] Calculate the structural loss between the output RGB image and the ground truth by the following formula :[[]] The minimization of the regularization distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain, and the minimization of the regularization distance between the output RGB image and the ground truth in the high-frequency domain are specifically minimized by the following formula: Among them, represents wavelet discrete transform, represents the low-frequency domain, represents the high-frequency domain, represents the visual neural network training stage, represents the first 1 / a stage, represents from the 1 / a stage to the completion of training, represents the minimization result; The reconstruction loss , the structure loss and the minimization result are jointly supervised by the following formula: Among them, represents the comprehensive reconstruction loss, the structural loss, and the minimized supervision result, , ; Make the training result of the visual neural network meet the preset co-supervision requirements, otherwise, continue to train the visual neural network until it meets the preset co-supervision requirements.
6. An image processing system that processes an image through a trained visual neural network; Characterized in that: During the training process of the visual neural network, for the RGB images output by the visual neural network, the first minimization module and the second minimization module are used for supervision: The first minimization module is used to minimize the regularized distance between the output RGB image and the ground truth in the low-frequency domain and the high-frequency domain during the first 1 / a stage of the visual neural network training; where a is greater than 2; The second minimization module is used to minimize the regularized distance between the output RGB image and the ground truth in the high-frequency domain from the 1 / a stage to the end of the visual neural network training; A method for processing images by the trained visual neural network includes: Performing a packing process on the image to be processed; Performing frequency enhancement on the packed image by an improved LW-ISP method; The specific improved LW-ISP method is: The improved LW-ISP method includes a first stage and a second stage. The first stage includes multiple groups of downsampling blocks and a fine-grained attention module; the second stage includes a convolutional layer, an FC layer, four groups of convolutional layers, an FCB, and a convolutional layer connected in sequence; the feature received by the last FCB is the packed image; Define the features received by the FCB from the first stage and the second stage of the improved LW-ISP method as follows: The output feature received from the FGAM in the first stage is denoted as L1, the output feature received from the DB in the first stage is denoted as L2, and the output feature received from the convolutional layer in the second stage is denoted as L3; The specific processing method in the FCB is: Performing convolutional processing on L2 and L3 respectively to obtain corresponding global and local features; Performing pixel recombination on the global and local features, and then fine-tuning the channel size through convolutional processing; Performing upsampling on the features of L2 and L3 after corresponding fine-tuning respectively, then performing residual learning and feature channel merging to obtain a feature A that matches the dimension of L1; Performing discrete wavelet transform on L1 and feature A respectively, taking the difference of the transformed results, and then summing with feature A after the inverse discrete wavelet transform.
7. The image processing system according to claim 6, wherein: It further includes a reconstruction loss supervision module and a structure loss supervision module; The reconstruction loss supervision module is used to calculate the reconstruction loss between the output RGB image and the ground truth through the following formula :[[]]END]] Among them, represents the output RGB image, represents the ground truth; The training result of the visual neural network meets the preset reconstruction loss requirement; otherwise, continue to train the visual neural network until the reconstruction loss meets the preset reconstruction loss requirement; The structural loss supervision module is used to calculate the structural loss between the output RGB image and the ground truth through the following formula : Among them, represents the luminance similarity, represents the weight coefficient, represents the maximum scale of the image, represents the specific scale into which the original image is divided, represents the th contrast similarity at a scale, represents the th structural similarity at a scale; The training result of the visual neural network meets the preset structural loss requirement; otherwise, continue to train the visual neural network until the structural loss meets the preset structural loss requirement.
8. A computer program product, wherein: The computer program product contains instructions that, when executed by a processor, implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Image deblurring algorithm based on knowledge distillation and deep neural network
CN114677304A
Single image defogging method based on physical prior and deep learning
CN115719319A