High-noise image recovery method and device
By layering the image, multi-view rendering and denoising of the background layer, and suppressing the light effect layer, solving the problem of noise recovery in low illumination and high dynamic range environments, improving image recovery quality and efficiency.
Patent Information
- Application Number
- CN202311800922.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot effectively eliminate noise and restore images in low illumination and high dynamic range environments, especially under complex motion and high noise, and the light effect affects the noise distribution inconsistently.
The input image is layered and divided into background layer and light effect layer. The background layer is denoised by multi-view rendering, the light effect layer is suppressed, and the processed image is fused.
It improves the quality and aesthetics of image recovery, enhances the image recovery efficiency in low illumination and high dynamic range environments, and reduces the impact of light effects on noise distribution.
Smart Images

Figure CN120259137A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of communications, and in particular, to a method and device for restoring high-noise images. Background Art
[0002] Currently, among the methods for restoring high-noise images, the method based on deep learning is very effective for restoring images. Existing deep learning-based image denoising is usually divided into single-frame image denoising and multi-frame image denoising. Among them, single-frame image denoising only requires one frame of image for denoising, and the method is relatively simple. However, in extremely low-light environments, camera imaging has problems such as serious detail loss, low contrast, and low signal-to-noise ratio, which greatly increases the difficulty of denoising. Even a very complex single-frame denoising network has unsatisfactory results in extremely low light; multi-frame denoising can overcome many inherent problems such as long exposure, and use inter-frame information to better separate the real signal and the noise signal for image denoising, but it requires strict alignment of the images, and frame alignment or motion estimation is quite difficult and resource- and time-consuming. Most multi-frame noise reduction assumes that the object is in simple motion and can achieve pre-alignment, but it cannot be achieved for complex motion and high noise, especially in extremely low light, the effect of using conventional multi-frame denoising is also unsatisfactory.
[0003] In addition, the difficulty of the low-light restoration task lies not only in the low number of photons, low signal-to-noise ratio, and complex noise sources, but also the light effect pollution affects the denoising effect. Light effects usually refer to phenomena such as glare, floodlight, and glow caused by uneven illumination. The existence of light effects results in inconsistent noise distribution in the whole image, affecting the image denoising effect and restoration quality. It is known that existing image enhancement technologies usually separate the original image into the form of multiplying the reflection layer and the brightness layer based on the Retinx idea to suppress the light effect in the brightness layer in order to alleviate the influence of the light effect on the image brightness and contrast adjustment. However, the layering operation is after the image denoising process and does not consider the influence of the light effect on the noise distribution.
[0004] In summary, the related technologies cannot effectively eliminate noise and restore images in environments with low illuminance and high dynamic range. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for restoring high-noise images, so as to at least solve the problem that in the related technologies, noise cannot be effectively eliminated and images cannot be restored in environments with low illuminance and high dynamic range.
[0006] According to an embodiment of the present invention, a method for restoring a high-noise image is provided, including: performing hierarchical processing on an input image to obtain a background layer and a light effect layer; performing noise reduction processing on the background layer to obtain a clean background layer, and performing suppression processing on the light effect layer to obtain a darkened light effect layer; fusing the clean background layer and the darkened light effect layer to obtain a restored final image.
[0007] According to another embodiment of the present invention, a device for restoring a high-noise image is provided, including: a hierarchical module for performing hierarchical processing on an input image to obtain a background layer and a light effect layer; a core network module for performing noise reduction processing on the background layer to obtain a clean background layer, and performing suppression processing on the light effect layer to obtain a darkened light effect layer; a fusion module for fusing the clean background layer and the darkened light effect layer to obtain a restored final image.
[0008] According to still another embodiment of the present invention, a computer-readable storage medium is further provided, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] Through the above embodiments of the present invention, a method for restoring a high-noise image is provided. By performing hierarchical processing on an input image to obtain a background layer and a light effect layer, and then performing noise reduction processing on them, the hierarchical processing is performed on the input image before noise reduction, where the hierarchical operation is not after the noise reduction processing, which can reduce the influence of the light effect on the noise distribution. And the noise reduction processing is performed on the background layer, and the suppression processing is performed on the light effect layer. The images after the noise reduction processing and the suppression processing are fully fused to obtain a restored final image. It solves the problem that in the related art, in an environment with low illuminance and high dynamic range, noise cannot be effectively eliminated and the image cannot be restored, and further achieves the effects of enhancing the quality and beauty of image restoration and improving the image restoration efficiency. Description of the Drawings
[0010] Figure 1 is a hardware structure block diagram of a computer terminal for a method for restoring a high-noise image according to an embodiment of the present invention;
[0011] Figure 2 is a flowchart of a method for restoring a high-noise image according to an embodiment of the present invention;
[0012] Figure 3 is a structure block diagram of a device for restoring a high-noise image according to an embodiment of the present invention;
[0013] Figure 4 is a flow schematic diagram of a method for restoring a high-noise image according to a scenario embodiment of the present invention;
[0014] Figure 5 is a flowchart of image layering according to an embodiment of the scenario of the present invention;
[0015] Figure 6 is a flowchart of light effect suppression according to an embodiment of the scenario of the present invention;
[0016] Figure 7 is a flowchart of transmittance estimation according to an embodiment of the scenario of the present invention;
[0017] Figure 8 is a schematic diagram of the background layer multi-view rendering noise reduction neural network architecture according to an embodiment of the scenario of the present invention;
[0018] Figure 9 is a schematic diagram of the background layer multi-view rendering noise reduction neural network architecture according to an embodiment of the scenario of the present invention. Detailed implementation manners
[0019] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0021] In related technologies, image denoising based on deep learning is often divided into single-frame image denoising and multi-frame image denoising. Most multi-frame denoising assumes that the object is in simple motion and can achieve pre-alignment, but it cannot be achieved for complex motion and high noise, especially the effect of using conventional multi-frame denoising under extremely low light is also not ideal.
[0022] The voxel-based multi-view rendering method has powerful functions in presenting novel scene views. By estimating the radiance and volume density of continuous 5D positions (three-dimensional spatial positions and 2D viewing directions), it restores appearance information from multiple source views. This method not only has an implicit alignment ability, but also enables some consensus relationships to be used among multiple rays to complement each other's information, and has excellent performance in extremely low illumination and extremely noisy image restoration tasks. The Noise Aware Neural Radiance Field (NAN) integrates information from multiple images and has an inherent ability to handle noise. It is a good multi-frame denoiser. Under high noise levels, it can handle large motions and occlusions and reaches the current best (State Of The Art, SOTA) of multi-frame denoising, but it also does not consider the problems brought by light effects.
[0023] In view of this, in the embodiments of the present invention, the input image is hierarchically processed before noise reduction. First, the image is divided into a background layer and a light effect layer. The background layer is mainly denoised for multiple frames by a method based on multi-view rendering; the light effect layer is suppressed, and finally the processed images are fused to obtain the final processing result.
[0024] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 is a hardware structure block diagram of a computer terminal for a high-noise image restoration method according to an embodiment of the present invention. As Figure 1 shown, the computer terminal may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than Figure 1 shown in
[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the high-noise image restoration method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0027] In this embodiment, a high-noise image restoration method running on the above computer terminal is provided. Figure 2 It is a flowchart of the high-noise image restoration method according to an embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:
[0028] Step S202: Perform a layering process on the input image to obtain a background layer and a light effect layer;
[0029] In an exemplary embodiment, in the above step S202, performing a layering process on the input image to obtain a background layer and a light effect layer includes: performing a layering process on the input image according to relative smoothness to obtain the background layer and the light effect layer.
[0030] In an exemplary embodiment, performing a layering process on the input image to obtain a background layer and a light effect layer includes: respectively constructing gradient functions of the background layer and the light effect layer according to the relative smoothness; according to the gradient functions, obtaining an objective optimization function of the weight of the relative smoothness by minimizing the negative logarithmic probability; solving the weight of the relative smoothness to perform a layering process on the input image.
[0031] In the actual implementation process, it is extremely challenging to effectively eliminate noise and restore details in a non-ideal environment (such as low photon number, light pollution effect, extremely high dynamic range, complex noise sources, etc.). For a dark image contaminated by the light effect, the input image is subjected to a layering process, which is divided into a background layer and a light effect layer. Different from the traditional Retinx-based idea of separating the original image into a reflection layer and a brightness layer multiplied together, which does not consider the influence form of the light effect on the noise distribution. The embodiment of the present invention fully considers the problem that the existence of the light effect will cause the noise distribution of the whole image to be inconsistent, first performs a layering process on the input image, and then performs a noise reduction process, so that the image noise reduction effect and restoration quality are better.
[0032] Step S204: Perform a noise reduction process on the background layer to obtain a clean background layer, and perform a suppression process on the light effect layer to obtain a darkened light effect layer;
[0033] In an exemplary embodiment, in the above step S204, noise reduction processing is performed on the background layer to obtain a clean background layer, including: performing noise reduction processing on the background layer by means of multi-image rendering to obtain a clean background layer.
[0034] In an exemplary embodiment, noise reduction processing is performed on the background layer to obtain a clean background layer, including: obtaining light based on the pixels of the background layer, where the light includes multiple sampling points; obtaining the color value and density value of each sampling point; and performing noise reduction processing on the background layer according to the color value and density value to obtain a clean background layer.
[0035] In the actual implementation process, for the background layer image after layering, multi-view rendering is used for denoising. After determining the view to be denoised, the camera parameter pose information of the input view background layer is obtained. First, the camera parameters are normalized so that the average camera pose of the background layer is the same as the world coordinate system to obtain the target view. The distance between the poses of the remaining input views and the target view position is calculated, and the N closest views are selected as the source views. A ray is constructed for each pixel of each source view, and there are many sampling points on the ray. The color value and density value of the sampling points on each ray in the multi-view are solved through a multi-view rendering denoising neural network, and noise reduction processing is performed on the background layer to obtain a clean background layer.
[0036] In an exemplary embodiment, before obtaining light based on the pixels of the background layer, where the light includes multiple sampling points, it further includes: obtaining the camera parameter pose information of the background layer and normalizing the camera parameter pose information so that the average camera pose of the background layer is the same as the world coordinate system.
[0037] In an exemplary embodiment, suppression processing is performed on the light effect layer to obtain a darkened light effect layer, including: estimating the global atmospheric light value and transmittance by means of image dehazing to perform suppression processing on the light effect layer to obtain a darkened light effect layer.
[0038] In the actual implementation process, for the suppression of the light effect, the atmospheric light value and transmittance can be estimated by means of image dehazing.
[0039] Step S206, fusing the clean background layer and the darkened light effect layer to obtain the restored final image.
[0040] In an exemplary embodiment, in the above step S206, fusing the clean background layer and the darkened light effect layer to restore and obtain the final image, including: performing weighted fusion on the clean background layer and the darkened light effect layer according to a preset weight to restore and obtain the final image.
[0041] Through the above steps, a high-noise image restoration method is provided. By performing hierarchical processing on the input image, a background layer and a light effect layer are obtained, and then noise reduction processing is performed on them. The hierarchical processing of the input image is carried out before noise reduction, where the hierarchical operation is not after the noise reduction processing. This can reduce the influence of the light effect on the noise distribution, and perform noise reduction processing on the background layer and suppression processing on the light effect layer. The images after noise reduction processing and suppression processing are fully fused to obtain the final restored image. It solves the problem that the related technology cannot effectively eliminate noise and restore the image in an environment with low illuminance and high dynamic range, and thus achieves the effects of enhancing the quality and beauty of image restoration and improving the image restoration efficiency.
[0042] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0043] In this embodiment, a high-noise image restoration device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0044] Figure 3 is a structural block diagram of a high-noise image restoration device according to an embodiment of the present invention. As Figure 3 shown, the high-noise image restoration device 30 includes: a hierarchical module 310, a core network module 320, and a fusion module 330. Among them, the hierarchical module 310 is used to perform hierarchical processing on the input image to obtain a background layer and a light effect layer. The core network module 320 is used to perform noise reduction processing on the background layer to obtain a clean background layer, and perform suppression processing on the light effect layer to obtain a darkened light effect layer. The fusion module 330 is used to fuse the clean background layer and the darkened light effect layer to obtain the final restored image.
[0045] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form. In the actual implementation process, the module naming and function division in the above high-noise image restoration device can be adjusted according to the actual situation, as long as the steps of the high-noise image restoration method in the above embodiments can be implemented, which will not be elaborated here.
[0046] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0047] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0048] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated here.
[0049] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0050] In order to enable those skilled in the art to better understand the technical solution of the present invention, the following will be elaborated in combination with specific scenario embodiments.
[0051] Scenario Embodiment 1
[0052] In the scenario embodiment of the present invention, a high-noise image restoration method based on multi-view rendering is provided. Figure 4 It is the flow schematic diagram of the high-noise image restoration method according to the scenario embodiment of the present invention, as Figure 4As shown, first, the input image is subjected to layer processing, dividing the image into a background layer and a light effect layer. The background layer is denoised using multi-view rendering; the light effect layer is suppressed and darkened. Finally, the processed images are fused to obtain the final processed result, i.e., the final image.
[0053] Figure 5 is a flowchart of image layering according to an embodiment of the scenario of the present invention, as Figure 5 shown, including:
[0054] Step S501, input the noisy image.
[0055] For each input image, layer processing is performed on the relatively smooth input image I, defined as the background layer and the light effect layer, expressed as:
[0056] I = J + G (1)
[0057] Among them, the J layer represents the high-frequency part of the image, i.e., the background layer; the G layer represents the low-frequency part of the image, i.e., the light effect layer. The G layer is smoother than the J layer, so the image can be layered using relative smoothness.
[0058] In the embodiment of the scenario of the present invention, since the G layer is smoother than the J layer, the gradient probability of the J layer is greater. The gradient probabilities of the J layer and the G layer are defined as P1(x) and P2(x) respectively, expressed as:
[0059]
[0060]
[0061] Among them, x is the gradient value, z is the normalization coefficient, and σ1 and σ2 are small operators that cause the narrow Gaussian to decline rapidly. However, by using the maximum operator in P1 and adding ε, the probability approaching zero is prevented.
[0062] To solve the layer separation problem, according to the gradient probability, the objective optimization function for obtaining the weight of relative smoothness is obtained by minimizing the negative logarithmic probability. Minimizing the negative logarithmic probability is expressed as:
[0063]
[0064]
[0065] Among them, x is the gradient value, σ1 and σ2 are small operators that cause the narrow Gaussian to decline rapidly, C1 and C2 are intermediate constants in the operation process, without actual meaning, and have been omitted in the subsequent calculation process. Assuming that L1 and L2 are two independent layers, then P(L1,L2) = P(L1)·P(L2); assuming that the outputs after passing through the filter are also two independent layers, then P(Lt ) = ∏ i P t (f i *L) i , t ∈ {1, 2}, and L2 = I - L1, where L1 and L2 are intermediate variables introduced to describe the operation principle. In the embodiments of the present invention, L1 is the J layer and L2 is the G layer.
[0066] The objective optimization function is expressed as:
[0067]
[0068] where i is the pixel index, L1 is the target layer (J layer), and I is the input image. Denote (L1 * f j ) i , which means filtering the image using a differential operator (filter). λ is the relative smoothness weight that determines the smoothness of the G layer after image layering. By setting the corresponding weight value, the background layer and the light effect layer are obtained after separation.
[0069] It should be noted that the two directional first - order derivative filters and one second - order Laplacian filter used in the calculation process are expressed as:
[0070]
[0071] Step S502, determine whether the pixel index i is less than i_max. If so, go to step S503; if not, go to step S507. Here, i_max is the maximum value of i preset according to the actual situation.
[0072] Step S503, fix the background layer J, introduce an auxiliary variable g for each pixel, and continuously iterate and update.
[0073] Step S504, fix the auxiliary variable g, calculate and update the background layer J, and alternately solve for the auxiliary variable and the background layer J.
[0074] In the actual implementation process, the semi - quadratic splitting method is used. An auxiliary variable g is introduced for each pixel, and through continuous iteration, the auxiliary variable g and the background layer are alternately solved.
[0075] Step S505, normalize the background layer J.
[0076] In the actual implementation process, normalization is used to limit the optimized background layer J layer within the effective data space.
[0077] Step S506, let i = i + 1, and return to step S502.
[0078] Step S507, output the background layer and the light effect layer.
[0079] After completing the image layering according to the above steps, for the optical effect layer, the idea of image defogging is adopted to suppress it, and the main steps are as follows:
[0080] For each foggy image, the atmospheric scattering model is expressed as:
[0081] I(x) = t(x)J(x) + (1 - t(x))A (8)
[0082] Where, I(x) is the foggy image, J(x) is the scene brightness, A is the global atmospheric light value, and t(x) is the scene transmission rate, that is, the transmittance.
[0083] The transmittance t(x) is expressed as:
[0084] t(x) = e -βd(x) (9)
[0085] Where, β is the medium extinction coefficient, and d(x) is the scene depth.
[0086] Based on formula (8), the scene brightness J(x) is recovered from I(x), which is expressed as:
[0087]
[0088] Where, I(x) is the foggy image, ∈ is a small constant to prevent the denominator from being zero, δ is an intermediate parameter for fine-tuning the defogging effect, t(x) is the scene transmission rate, that is, the transmittance, and A is the global atmospheric light value.
[0089] Substitute the estimated atmospheric light value and transmittance into formula (10) to suppress the optical effect, and the obtained J(x) is the darkened optical effect layer.
[0090] Figure 6 It is the flowchart of optical effect suppression according to the scene embodiment of the present invention, as Figure 6 shown, including the following steps:
[0091] Step 1, after obtaining the darkened optical effect layer by layering, estimate the atmospheric light value, where the atmospheric light value estimation is expressed as:
[0092] A = min(I(x,y) - 1, 255) (11)
[0093] Where, (x,y) represents the image coordinate point with the most severe fog, which can be set customarily.
[0094] Step 2, estimate the transmittance, and through boundary constraint and context regularization constraint, obtain the objective function formula to be optimized, which is expressed as:
[0095]
[0096] Among them, W j is the weight function, t is the transmittance, which is calculated based on the squared difference between the color vectors of two adjacent pixels, expressed as:
[0097]
[0098] By introducing an auxiliary variable u j , using the semi - quadratic splitting method, through continuous iterative calculations, alternately solve for u j and t. Figure 7 is the flowchart of transmittance estimation according to the embodiment of the present invention scenario. As Figure 7 shown, the process of transmittance estimation includes the following steps:
[0099] Step S701, initialize parameters.
[0100] Step S702, determine whether the pixel number i is less than i_max. If so, enter step S703; if not, enter step S706.
[0101] Step S703, fix the transmittance t, use the semi - quadratic splitting method, introduce an auxiliary variable u j for each pixel and continuously iterate and update.
[0102] Step S704, fix the auxiliary variable u j , calculate and update the transmittance t.
[0103] Step S705, pixel number i = i + 1.
[0104] Step S706, output the estimated transmittance t.
[0105] After completing image layering, for denoising the background layer, for the background layer image after layering, adopt the method of multi - view rendering for denoising. After determining the view to be denoised, obtain the pose information of the camera parameters of the background layer of the input view. First, normalize the camera parameters so that the average pose of the camera of the background layer is the same as the world coordinate system to obtain the target view. Calculate the distance between the poses of the remaining input views and the position of the target view, and select the N views closest as the source views. Build a ray for each pixel of each source view, and there are many sampling points on the ray. Solve for the color value c m and density value ρ m of the sampling points on each ray in the multi - view through the multi - view rendering denoising neural network. Among them, the weight of each ray sampling point is obtained through formula (14), and the pixel value after denoising of each ray is obtained through formula (15). The formulas are expressed as:
[0106]
[0107]
[0108] Among them, m is the sampling point order, M is the number of sampling points, and c m is the color value of the sampling point, and ρ m is the density value of the sampling point, and ω m is the weight of the sampling points of each ray, is the pixel value of each ray after denoising.
[0109] During the training calculation process, noise is added to the clean multi-view through a noise model, and then the multi-view rendering denoising neural network is used to solve the color value and density value of the ray sampling points corresponding to the noisy views. The multi-view rendering is completed by formulas (14) and (15), and the pixel value is supervised and learned through the clean image pixel value. Finally, the noisy multi-view is passed through the trained neural network model to obtain the clean background layer image. Among them, the noise model is expressed as:
[0110]
[0111] Among them, x is the image coordinate, and σ r is the signal-independent read-noise parameter, and σ s is the signal-dependent shot-noise parameter, is the Gaussian distribution, is the clean image, and I n (x) is the noisy image.
[0112] Scene Embodiment Two
[0113] In this Scene Embodiment Two, the background layer multi-view rendering denoising in Scene Embodiment One is introduced in detail.
[0114] Figure 8 is the schematic diagram of the background layer multi-view rendering denoising neural network architecture according to the scene embodiment of the present invention. As Figure 8 shown, it includes a front-end backbone network, a data processing module, and a denoising core network. Among them, the core network can correspond to some functions of the core network module in the above embodiments of the present invention. It should be clear that the above modules are only differences in naming and function division, and will not be elaborated here.
[0115] In this Scene Embodiment Two, as Figure 8 shown, the data processing module includes a normalization module, a view selection module, a preprocessing module, and a noise addition module to complete the preprocessing of the data.
[0116] The denoising core network includes a coarse stage (coarse processing stage) and a fine stage (fine processing stage). Each stage includes a ray sampling module, an MLP module, and a rendered image module. The neural network part is mainly composed of MLP and CNN.
[0117] After the data processing module obtains the input view camera parameters and pose information in advance, it first normalizes the camera parameters so that the average pose of the cameras of all input views is the same as the world coordinate system. After obtaining the target view, it calculates the distance between the poses of the remaining input views and the position of the target view, and selects the N closest views as the source views (the first view needs to be replaced with the target view). Finally, through the preprocessing module and the noise addition module, the preprocessing of the data is completed.
[0118] From the extracted features and the data processed by the data processing module, first generate the sampling points (64) of the source view through the camera pose information. Obtain the color value c_m and density value ρ_m of the sampling points through the MLP module in the coarse stage, and obtain the final color value through the rendered image module. Next is the fine stage. Using the sampling point information in the coarse stage, the same process renders the final color value under the increased number of sampling points (128) to obtain the final output.
[0119] Figure 9 It is a schematic diagram of the background layer multi-view rendering denoising neural network architecture according to the scenario embodiment of the present invention, as Figure 9 shown. The front-end backbone network includes a noise preprocessing network and a feature extraction network. Use Pre_Net and Features_Net as the backbone networks for noise preprocessing and feature extraction respectively. Pre_Net uses a 3*3 convolutional kernel, the output channels remain unchanged, there is no activation function, and the weights of each channel are initialized using a Gaussian kernel. The spatial and channel dimensions of the input burst image are retained through Pre_Net, and the input noise is effectively reduced; the output preliminary denoising information enters Features_Net to further extract features. It consists of an asymmetric Unet. First, the resolution is halved through a 7*7 convolution, and the number of channels becomes 64. Then, the resolution is further reduced to 1 / 4 and 1 / 8 respectively using 3*3 convolutions, and the number of channels is 64, 128, and 256 respectively. Then, the resolution is enlarged to 1 / 8 and 1 / 4 using bilinear interpolation and 3*3 convolutions. Instance Norm normalization and ReLU activation functions are used after each convolution. Finally, 1*1 convolution is used to split the coarse / fine features from the channel dimension. Compared with the information after passing through Pre_Net, the resolution of the coarse / fine features here is reduced to 1 / 4, and the number of channels is 32.
[0120] The multi-view rendering denoising neural network first performs supervised training using the dataset LLFF, and then completes inference and testing based on the trained weight parameters to achieve denoising processing of the background layer in multi-view rendering. The training process is divided into four steps:
[0121] In the first step, according to the input multi-view images, the camera parameter information (intrinsic and extrinsic parameters) is calculated through the software COLMAP: it is an N*17 matrix (N is the number of pictures), and the 17 parameters include the rotation matrix (3*3), translation vector (3*1) in the camera pose parameter (camera to world, c2w), the H and W of the image, the focal length f of the camera, and the range of the scene (the nearest and farthest distances from the scene points to the camera center under this camera view).
[0122] In the second step, the information input to the training network is obtained through the feature extraction network of Pre-Net+Features-Net and data preprocessing. After loading the training dataset LLFF, the camera parameters (image height and width, focal length, scene range, etc.) are adjusted according to the scaling factor and normalized to facilitate the construction of the sampling point information of the source view in the world coordinate system. During training, a random selection is made to choose a certain picture of a certain scene as the target view, and then the source view is selected through camera pose calculation in this scene, and the selected view is randomly cropped and flipped. Finally, the selected source view is transformed to the linear domain through inverse white balance and inverse gamma, and noise is added through the noise model of formula (10). Respectively take σ r and σ s , and randomly add noise for training in the log rectangle domain range (std = [-3.0, -0.5, -2.0, -0.5]), and select the noise parameters corresponding to the six gain points of 1, 2, 4, 8, 16, and 20 to perform noise addition tests on the clean images of the source view.
[0123] Step 3: Obtain the ray sampling points corresponding to the multi-view image pixels from the processed camera parameters and the input network features. The internal and external camera parameters of the target view first generate a ray matrix for all pixels (including the direction of the rays, the origin of the rays, which is the position of the camera, and the coordinates of the sampling points). During training, some rays of the target view are randomly selected and then projected into N source views according to the camera parameters. A mask is set to eliminate invalid points (such as those projected behind the camera or outside the image), and at the same time, the noise estimation value of the source view, the features before and after passing through the Pre-Net, the coarse features, and the ray-diff (the unit vector of the difference between the vectors formed by the camera positions of the target view and the source view to the sampling points of the target view and the sum of their inner products concatenated in the channel dimension) are also projected. Similarly, the ray sampling points in the fine stage are generated in the same way as in the coarse stage, but the information in the coarse stage is utilized. The final sampling point weights in the coarse stage generate the sampling points of the target view in the fine stage. Finally, the number of sampling points in the fine stage is the sum of those in the coarse stage and the fine stage.
[0124] Step 4: From the ray sampling points calculated in Step 3, obtain the density values and color values of the ray sampling points through the multi-view rendering denoising neural network MLP coarse / fine and complete the voxel rendering, and finally obtain the pixel values of the denoised image. The coarse features, noise standard deviation estimation, direction features, mask, original noisy RGB values, and Pre-Net feat obtained in Steps 2 and 3 are used as the inputs to the multi-view rendering denoising neural network in the coarse stage. First, feature fusion is performed, and then the density (m is the sampling point order) and color values of the sampling points are calculated. There is an improvement here compared to the original multi-view rendering method when calculating the color values of the sampling points. Using the idea of a 3D kernel, when calculating the RGB blending weights, the spatial neighborhood of each projected pixel is applied with the weight spatial kernel of each projected pixel. The fine stage is the same as the coarse stage. Utilizing the information such as the sampling points in the coarse stage, the final sampling point weights in the coarse stage generate the sampling points of the target view in the fine stage, and the density and color of the sampling points in the fine stage are obtained through a similar process. Finally, the pixel values of the selected pixels after denoising are output by the integral equations in discrete form of Equations (14) and (15).
[0125] In the testing and inference stages, N test images and poses are input and sequentially pass through the front-end backbone network, the data processing module, and the denoising core network. Different from training, the inference is performed H*W / chunk_size (the number of pixel values rendered each time) times, and finally a rendered image is obtained, realizing the denoising processing of the multi-view rendering on the background layer.
[0126] The noisy image is layered to obtain a background layer and a light effect layer. The background layer is denoised, and the light effect layer is suppressed. By fully utilizing the implicit alignment information of multiple views, the processed images are finally weighted and fused to obtain an image for low-illuminance and high-noise restoration based on multi-view rendering.
[0127] In summary, the embodiments of the present invention provide a high-noise image restoration method and apparatus. In the embodiments of the present invention, compared with the existing image enhancement work based on the Retinx idea, the embodiments of the present invention separate the uneven light effect layer by using relative smoothness and implement the noise reduction task on the background layer after layering, avoiding the influence of the light effect layer on the noise distribution. The method based on multi-view rendering restores the appearance information from multiple source views by estimating the radiance and volume density of continuous 5D positions (three-dimensional spatial positions and 2D viewing directions). By calculating the position information of the scene views in advance through the software colmap, the implicit alignment ability can be naturally guaranteed during the noise reduction process, enabling some consensus relationships to be used among multiple rays to complement each other's information and maximizing the restoration of the true signal details in extremely low-illuminance and extremely noisy images. In the embodiments of the present invention, by layering the input image, multi-view rendering denoising is performed on the background layer, the light effect layer is suppressed, and the implicit alignment information of multiple views is fully utilized, improving the quality and aesthetics of image restoration in low-illuminance and high-noise situations.
[0128] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A high-noise image restoration method, characterized in that, including: performing layer processing on the input image to obtain a background layer and a light effect layer; performing noise reduction processing on the background layer to obtain a clean background layer, and performing suppression processing on the light effect layer to obtain a darkened light effect layer; fusing the clean background layer and the darkened light effect layer to obtain a restored final image.
2. The method according to claim 1, wherein The performing layer processing on the input image to obtain a background layer and a light effect layer includes: performing layer processing on the input image according to relative smoothness to obtain the background layer and the light effect layer.
3. The method according to claim 2, wherein The performing layer processing on the input image to obtain a background layer and a light effect layer includes: respectively constructing gradient functions of the background layer and the light effect layer according to the relative smoothness; obtaining an objective optimization function of the weight of the relative smoothness by minimizing the negative log probability according to the gradient function; solving the weight of the relative smoothness to perform layer processing on the input image.
4. The method according to claim 1, wherein The performing noise reduction processing on the background layer to obtain a clean background layer includes: performing noise reduction processing on the background layer by means of multi-image rendering to obtain a clean background layer.
5. The method according to claim 1, wherein The performing noise reduction processing on the background layer to obtain a clean background layer includes: acquiring light rays according to the pixels of the background layer, where the light rays include multiple sampling points; acquiring the color value and density value of each sampling point; performing noise reduction processing on the background layer according to the color value and the density value to obtain a clean background layer.
6. The method according to claim 5, characterized in that Before acquiring light rays according to the pixels of the background layer, where the light rays include multiple sampling points, the method further includes: acquiring the camera parameter pose information of the background layer, and performing normalization processing on the camera parameter pose information so that the average camera pose of the background layer is the same as the world coordinate system.
7. The method according to claim 1, wherein The performing suppression processing on the light effect layer to obtain a darkened light effect layer includes: estimating the global atmospheric light value and transmittance by means of image defogging to perform suppression processing on the light effect layer to obtain a darkened light effect layer.
8. The method according to claim 1, wherein The fusing the clean background layer and the darkened light effect layer to restore and obtain a final image includes: performing weighted fusion on the clean background layer and the darkened light effect layer according to a preset weight to restore and obtain a final image.
9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the method described in any one of claims 1 to 8.
10. A high-noise image restoration device, characterized in that, including: a layering module for performing layer processing on the input image to obtain a background layer and a light effect layer; a core network module for performing noise reduction processing on the background layer to obtain a clean background layer and performing suppression processing on the light effect layer to obtain a darkened light effect layer; a fusion module for fusing the clean background layer and the darkened light effect layer to obtain a restored final image.