A video denoising method and device

CN120751273BActive Publication Date: 2026-08-07HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-08-16
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]有鉴于此,本申请提供了一种视频去噪方法,以解决视频质量差的问题,其公开的技术方案如下:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751273B_ABST
    Figure CN120751273B_ABST
Patent Text Reader

Abstract

The application provides a video denoising method and device. A first signal-to-noise ratio is obtained for a first frame of original image, and then the resolution of a second frame of original image is determined based on the first signal-to-noise ratio, that is, the resolution of the next frame of original image is determined according to the signal-to-noise ratio of the current frame of original image. When the signal-to-noise ratio is high, the denoising process is relatively simple, and a raw image with high resolution is selected to improve the video quality. When the signal-to-noise ratio is low, the denoising process is relatively difficult, and a raw image with low resolution but small noise is selected to reduce the difficulty of denoising. Moreover, the denoising process starts from the second frame of original image. A fusion coefficient is determined according to the shooting information of the last frame of image, and the last frame of denoised image and the current frame of original image with the second resolution are fused according to the fusion coefficient to obtain the denoised image of the current frame. The fusion coefficient is the proportion of the feature of the last frame of denoised image in the feature of the current frame of denoised image. In this way, the difference in clarity between adjacent raw images can be reduced, and the overall quality of the video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a video denoising method and apparatus. Background Technology

[0002] With technological advancements, users have increasingly higher demands for the video quality captured by mobile phones and other devices. However, video signals are often mixed with various noise signals, causing the video to become blurry and its quality to degrade. Summary of the Invention

[0003] In view of this, this application provides a video denoising method to solve the problem of poor video quality, and the disclosed technical solution is as follows:

[0004] In a first aspect, this application provides a video denoising method applied to an electronic device. The method includes: starting video recording in response to a video recording operation; performing denoising processing on a first frame of the original image obtained from the recording to obtain a first frame of denoised image, wherein the first frame of the original image is an image with a first resolution; obtaining a first signal-to-noise ratio (SNR) corresponding to the first frame of the original image based on the first frame of the original image and the first frame of denoised image; determining a second resolution corresponding to a second frame of the original image based on the first SNR, and obtaining a second frame of the original image with the second resolution, wherein the second resolution is different from the first resolution and is positively correlated with the first SNR; and obtaining a second frame of the original image based on the recording information. The first denoised image is fused with the second original image at a second resolution according to the fusion coefficients corresponding to the frame images. The shooting information includes prior knowledge or shooting parameters of the previous frame image. The fusion coefficient is the proportion of the features of the first denoised image in the features of the second denoised image. A second signal-to-noise ratio (SNR) is obtained based on the second denoised image and the second original image. This SNR is used to determine the resolution of the third original image, and the third original image is processed according to the processing procedure for the second frame image until a stop operation is received to stop video recording, resulting in a denoised video. For example, with this scheme, when the SNR is high, the denoising process is simpler, and a high-resolution raw image can be selected, thus improving video quality. When the SNR is low, the denoising process is more difficult, and a low-resolution but less noisy raw image can be selected to reduce the difficulty of denoising, thereby improving the denoising effect and efficiency. Furthermore, this scheme reduces the sharpness difference between adjacent raw images by fusing the previous frame and the current frame, improving the overall video quality.

[0005] In one possible implementation of the first aspect, the captured information is the signal-to-noise ratio (SNR); obtaining the fusion coefficient corresponding to the second frame image based on the captured information includes: obtaining the fusion coefficient of the second frame image based on the first SNR corresponding to the first original frame image, wherein the fusion coefficient is positively correlated with the first SNR. The fusion ratio in this scheme is determined by the SNR value. For cases with a low SNR value, more of the denoised image from the previous frame is fused. For cases with a high SNR value, more of the noisy raw image from the current frame is fused, thus improving video quality.

[0006] In one possible implementation of the first aspect, obtaining the fusion coefficient corresponding to the second frame image based on the shooting information, and fusing the first denoised image with the second frame original image at the second resolution according to the fusion coefficient to obtain the second denoised image, includes: inputting the first denoised image, the second frame original image at the second resolution, and the first signal-to-noise ratio of the first image into a denoising neural network for denoising processing; the denoising neural network determining the fusion coefficient based on the first signal-to-noise ratio of the input first frame image; the denoising neural network extracting a first image feature from the first denoised image and extracting a second image feature from the second frame original image at the second resolution; and the denoising neural network fusing the first image feature and the second image feature according to the fusion coefficient to obtain the second denoised image.

[0007] In one possible implementation of the first aspect, the training process of the denoising neural network includes: creating an initial denoising neural network; inputting a first frame denoised image and its signal-to-noise ratio (SNR) from a sample video, along with a second frame original image, into the initial denoising neural network for denoising processing to obtain a second frame sample denoised image; determining a target fusion coefficient that matches the first SNR based on the mapping relationship between the SNR and the fusion coefficient; fusing the features of the first sample image and the features of the second sample image based on the target fusion coefficient to obtain a second frame target denoised image, wherein the first sample image features are extracted from the first frame denoised image by the initial denoising neural network, and the second sample image features are extracted from the second frame original image by the initial denoising neural network; obtaining the error between the second frame target denoised image and the second frame sample denoised image, and adjusting the parameters of the denoising neural network according to the error; repeating the processing of the second frame image for any frame after the second frame image in the sample video until the error between the sample denoised image and the target denoised image corresponding to the same frame is within a preset range, thus ending the training process of the denoising neural network.

[0008] In one possible implementation of the first aspect, the electronic device includes an image sensor and an image signal processor. The image sensor supports dynamically switching the shooting resolution during video capture. Determining the second resolution corresponding to the second frame of the raw image based on a first signal-to-noise ratio (SNR), and obtaining the second frame of the raw image at the second resolution, includes: the image signal processor querying the mapping relationship between the SNR range and resolution to obtain the second resolution corresponding to the first SNR; when the image signal processor determines that the second resolution is different from the first resolution, it sends a resolution switching command to the image sensor; the image sensor responds to the resolution switching command by switching the resolution of the captured image to the second resolution and sends the second frame of the raw image at the second resolution to the image signal processor. This scheme allows the image signal processor to notify the image sensor to switch its output mode when it determines that the output mode of the image sensor needs to be switched, thereby achieving dynamic switching of the raw image resolution. Moreover, this scheme controls the resolution of the next frame based on the SNR of the previous frame of the raw image. When the SNR of the previous frame is low, the next frame can choose an output mode with lower resolution but lower noise, thereby reducing the difficulty of noise reduction. When the SNR of the previous frame is high, the next frame can choose an output mode with higher resolution, improving video quality.

[0009] In one possible implementation of the first aspect, obtaining the second signal-to-noise ratio corresponding to the second original image based on the second denoised image and the second original image includes: the image signal processor calculates the second signal-to-noise ratio corresponding to the second original image based on the second denoised image and the second original image, and determines the third resolution of the third original image based on the second signal-to-noise ratio.

[0010] In one possible implementation of the first aspect, the image signal processor maintains a mapping relationship between the signal-to-noise ratio (SNR) range and resolution for each different output mode. The image signal processor queries this mapping relationship to obtain the second resolution corresponding to the first SNR. This includes: the image signal processor queries the mapping relationship between the SNR range and resolution corresponding to the output mode of the first frame of the original image to obtain the second resolution corresponding to the first SNR. As can be seen, in this scheme, the image signal processor maintains the mapping relationship between the SNR range and resolution for each different output mode. This allows the correct resolution corresponding to the SNR of the first frame of the original image to be obtained simply by querying the mapping relationship corresponding to the output mode of the first frame, improving the timeliness of determining the resolution of the next frame and also improving the accuracy of the resolution.

[0011] In one possible implementation of the first aspect, the image signal processor maintains a first mapping relationship between the signal-to-noise ratio (SNR) and resolution corresponding to a first output mode, where the resolution of the original image output under the first output mode is a third resolution. The image signal processor queries the mapping relationship between the SNR range and resolution to obtain a second resolution corresponding to the first SNR, including: the image signal processor converts the first SNR corresponding to the first frame of the original image at the first resolution into a second SNR corresponding to the third resolution; the image signal processor queries the mapping relationship between the SNR range corresponding to the third resolution and resolution to obtain a second resolution corresponding to the second SNR. It is evident that in this scheme, the image signal processor only needs to maintain the mapping relationship between SNR and resolution under one output mode. Since the SNRs corresponding to different resolutions are multiples of each other, the SNR of a specific output mode that needs to be queried can be converted to the output mode corresponding to the mapping relationship maintained by the image signal processor, thereby obtaining the accurate resolution corresponding to that SNR. Moreover, this scheme does not require maintaining the mapping relationship between SNR and resolution under all output modes in the image signal processor, thus saving storage space in the image signal processor.

[0012] In one possible implementation of the first aspect, the electronic device includes an image sensor, an image signal processor, and a camera algorithm library. Obtaining the first signal-to-noise ratio (SNR) corresponding to the first frame of the original image based on the first frame of the original image and the first frame of the denoised image includes: the image signal processor performing denoising processing on the first frame of the original image output by the image sensor to obtain the first frame of the denoised image; the image signal processor transmitting the first frame of the denoised image and the first frame of the original image to the camera algorithm library; and the camera algorithm library calculating the first SNR corresponding to the first frame of the original image based on the first frame of the denoised image and the first frame of the original image. As can be seen, in this scheme, the denoising algorithm in the image signal processor is used to denoise the first frame of the original image (i.e., the first frame raw image) to obtain the first frame of the denoised image. The first frame of the denoised image is then calculated by the camera algorithm library based on the first frame of the denoised image and the first frame of the original image provided by the image signal processor. This eliminates the need to modify the processing logic in the image signal processor, reducing the complexity of the implementation.

[0013] In one possible implementation of the first aspect, the image sensor does not support dynamically switching the shooting resolution during video capture. Determining the second resolution corresponding to the second frame of the original image based on the first signal-to-noise ratio (SNR), and obtaining the second frame of the original image at the second resolution, includes: the camera algorithm library queries the mapping relationship between SNR and resolution to obtain the second resolution corresponding to the first SNR, where the second resolution is lower than the first resolution; the image sensor captures the second frame of the original image at the first resolution and passes it to the camera algorithm library; the camera algorithm library downsamples the second frame of the original image at the first resolution to obtain the second frame of the original image at the second resolution. Therefore, this scheme is suitable for scenarios where the image sensor does not support dynamically switching the raw image resolution. In such scenarios, the camera algorithm library can convert the raw image at the first resolution output by the image sensor to obtain the raw image at the second resolution, supporting only downward adjustment of the raw image resolution. This reduces the performance requirements of the image sensor hardware and expands the applicability of the scheme.

[0014] In one possible implementation of the first aspect, obtaining the second signal-to-noise ratio corresponding to the second original image based on the second denoised image and the second original image includes: the camera algorithm library calculates the second signal-to-noise ratio corresponding to the second original image based on the second original image at the second resolution and the second denoised image at the second resolution.

[0015] Secondly, this application also provides an electronic device, which includes: one or more processors, a memory, and a touch screen; the memory is used to store program code; the processor is used to run the program code, so that the electronic device implements the video denoising method as described in any of the first aspects.

[0016] Thirdly, this application also provides a computer-readable storage medium having instructions stored thereon, which, when executed on an electronic device, cause the electronic device to perform a video denoising method as described in any of the first aspects.

[0017] Fourthly, this application also provides a computer program product, characterized in that it stores instructions that, when the computer program product is run on an electronic device, cause the electronic device to implement the video denoising method as described in any of the first aspects.

[0018] Fifthly, this application also provides a chip system, including: at least one processor and an interface, the interface being used to receive code instructions and transmit them to the at least one processor; the at least one processor executes the code instructions to implement the video denoising method of any one of the first aspects. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of a video noise reduction process provided in an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of a denoising neural network provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of a video noise reduction process based on an electronic device software architecture provided in an embodiment of this application;

[0023] Figure 5 The embodiments provided in this application are related to Figure 4 A flowchart illustrating the corresponding video denoising method;

[0024] Figure 6 This is a schematic diagram of another video noise reduction process based on electronic device software architecture provided in an embodiment of this application;

[0025] Figure 7 The embodiments provided in this application are related to Figure 6 Flowchart of the corresponding video denoising method Detailed Implementation

[0026] The terms "first," "second," and "third," etc., used in this application specification, claims, and drawings are used to distinguish different objects, not to limit a specific order.

[0027] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0028] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0029] Figure 1 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0030] Electronic devices can be mobile phones, tablets, wearable electronic devices, in-vehicle electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, smart screens, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), projectors, etc. This application does not impose any restrictions on the specific type of electronic device.

[0031] Electronic devices may include processors, external memory interfaces, internal memory, universal serial bus (USB) interfaces, charging management modules, power management modules, batteries, antenna 1, antenna 2, mobile communication modules, wireless communication modules, audio modules, speakers, receivers, microphones, headphone jacks, sensor modules, buttons, motors, indicators, cameras, displays, and subscriber identification module (SIM) card interfaces, etc. The sensor modules may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, proximity sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0032] A processor may include one or more processing units. For example, a processor may include at least one of the following processing units: application processor (AP), graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). These different processing units may be independent devices or integrated devices.

[0033] The processor may also include memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or that are used repeatedly. If the processor needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces processor waiting time, and thus improves system efficiency.

[0034] In this embodiment, the processor can execute the video denoising method provided in this embodiment. During the video recording process controlled by the camera application, the output mode of the current frame's Raw image (i.e., the resolution of the next frame's Raw image) is selected based on the signal-to-noise ratio of the previous frame image. Furthermore, the denoised Raw image of the previous frame and the noisy Raw image of the current frame are fused to obtain the denoised image of the current frame. The above processing is performed on each captured Raw image to finally obtain the denoised video.

[0035] The wireless communication function of electronic devices can be realized through devices such as antenna 1, antenna 2, mobile communication module, wireless communication module, modem processor and baseband processor.

[0036] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0037] Mobile communication modules can provide solutions for wireless communication applications, including 2G / 3G / 4G / 5G / 6G, in electronic devices.

[0038] Wireless communication modules can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0039] Electronic devices can implement display functions through GPUs, displays, and application processors. A GPU is a microprocessor for image processing, connected to both the display and the application processor. GPUs are used to perform mathematical and geometric calculations and for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information. The display can be used to show images or videos.

[0040] Electronic devices can achieve shooting functions through ISPs, cameras, video codecs, GPUs, displays, and application processors.

[0041] A camera is used to capture still images or videos. Incident light passes through a lens and converges to the focal point of the lens, causing the object being photographed to be imaged on an image sensor. The image sensor converts the light signal into an electrical signal, which is then processed by an image ISP (Image Signal Processor) to convert it into a digital image signal. The digital image signal is further processed and converted into a standard image signal in formats such as RGB and YUV, which can then be transmitted to a display screen for display. The image ISP can perform algorithmic optimization of image noise, brightness, and color. It can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the image ISP can be integrated into the camera. The electronic device may include one or N cameras, where N is a positive integer greater than 1.

[0042] Video codecs are used to compress or decompress digital video. Electronic devices can support one or more video codecs. This allows the electronic device to play or record video in various encoded formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, and MPEG 4.

[0043] A digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals.

[0044] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0045] Internal memory can be used to store executable program code, which includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. For example, in this embodiment, the processor can perform video noise reduction processing by executing instructions stored in the internal memory.

[0046] A touch sensor, also known as a "touch device," is a component of an electronic device. Touch sensors can be mounted on a display screen, and the touch sensor and display screen together form a touchscreen, also called a "touchscreen." The touch sensor detects touch operations applied to or near it. It then transmits the detected touch operation to an application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen. In some embodiments, the touch sensor may also be located on the surface of the electronic device, in a different position than the display screen.

[0047] It should be noted that, Figure 1The structure shown does not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include a... Figure 1 The components shown may include more or fewer components, or the electronic device may include... Figure 1 The components shown may be a combination of certain components, or the electronic device may include... Figure 1 Sub-components of some of the components shown. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0048] The following will combine Figure 2 and Figure 3 This application describes the video noise reduction process provided in its embodiments.

[0049] Figure 2 This is a schematic diagram of a video noise reduction process provided in an embodiment of this application.

[0050] like Figure 2 As shown, after the electronic device receives the user's video capture operation, the image sensor captures the first frame of the raw image. The raw image usually contains noise signals and can be called a noisy image raw0.

[0051] In some embodiments of this application, the resolution of the noisy image raw0 can be a default resolution. For example, in one scenario, the image sensor supports multiple different output modes, such as Hex-pattern, quadr-pattern, and binning-pattern, with the resolution decreasing sequentially. The resolutions of different output modes are multiples of each other; for example, in one example, the resolution of Hex-pattern is four times that of quadr-pattern, and the resolution of quadr-pattern is four times that of binning-pattern.

[0052] (1) The first frame of noisy image raw0 is denoised to obtain the first frame of denoised image raw0. The denoising process can use existing denoising algorithms, and this application does not limit it.

[0053] (2) Based on the noisy image raw0 of the first frame and its corresponding denoised image, the corresponding signal-to-noise ratio (SNR) value is calculated, that is, the SNR of the first frame image, which can be denoted as SNR0.

[0054] In some embodiments, the SNR value of an image can be calculated according to the following formula:

[0055]

[0056] In the above formula, f(x,y) represents the pixel value corresponding to pixel point (x,y) in the noisy raw image. This represents the pixel value corresponding to pixel point (x, y) in the denoised raw image. Furthermore, this application does not limit the method used to calculate the SNR value of the image.

[0057] (3) Decision on the resolution of the next frame raw image based on the SNR0 of the first frame raw image, and input the obtained next frame raw image with the resolution into the denoising neural network. At the same time, the SNR0 of the first frame and the denoised raw0 of the first frame are input into the denoising neural network.

[0058] If the SNR value is high, choose a higher resolution mode. For example, if the current mode is binning-pattern, you can choose Hex-pattern or quadr-pattern; if the current mode is quadr-pattern, you can choose Hex-pattern. If the SNR value is low, choose a lower resolution mode with lower noise. For example, if the current mode is quadr-pattern, you can choose binning-pattern.

[0059] (4) The denoising neural network is based on the SNR0 of the first frame, and fuses the denoised image raw0 of the first frame and the noisy image raw1 of the second frame to obtain the denoised image raw1 of the second frame.

[0060] (5) The SNR of the second frame image is obtained from the noisy second frame image raw1 and the denoised second frame image raw1, and can be denoised as SNR1.

[0061] (6) Taking the nth frame as an example, the resolution of the nth frame is selected based on the SNR(n-1) value of the (n-1)th frame. Then, the denoising neural network fuses the noisy nth frame and the denoised n-1th frame based on the SNR(n-1) value of the (n-1)th frame to obtain the denoised nth frame. The processing of subsequent images is the same as that of the nth frame until the last frame.

[0062] The video denoising method provided in this embodiment determines the sharpness (i.e., resolution) of the current frame's raw image based on the SNR value of the previous frame. For example, the higher the SNR value of the previous frame, the higher the sharpness of the raw image output mode can be selected for the current frame; conversely, the lower the SNR value of the previous frame, the lower the sharpness but lower noise of the raw image output mode can be selected for the current frame. That is, the SNR value is positively correlated with sharpness. This allows the most suitable raw image to be dynamically selected as the input to the denoising neural network based on the different signal-to-noise ratios during the recording of the same video. For example, when the signal-to-noise ratio is high, the denoising process is simpler, so a high-resolution raw image can be selected, thereby improving video quality; when the signal-to-noise ratio is low, the denoising process is more difficult, so a low-resolution but low-noise raw image can be selected, reducing the difficulty of denoising and thus improving the denoising effect and efficiency.

[0063] Furthermore, a fusion coefficient α can be determined based on the SNR value of the previous frame. This coefficient is then used to fuse the features of the noiseless image from the previous frame and the features of the noisy image from the current frame, resulting in the denoised image for the current frame. Using raw images of different resolutions as input to the denoising process may cause slight variations in the sharpness of adjacent raw images. Fusing the previous and current frames reduces the sharpness difference between adjacent raw images, improving the overall video quality. In this application, the fusion ratio is determined by the SNR value. For frames with low SNR values, more of the denoised image from the previous frame is fused; for frames with high SNR values, more of the noisy raw image from the current frame is fused, thus improving video quality.

[0064] In other embodiments of this application, the fusion coefficient α can also be determined based on other parameters such as ISO value and ambient light intensity. For example, the lower the ambient light intensity, the higher the ISO value; furthermore, the lower the SNR value of the image, the higher the fusion coefficient α can be. This application does not impose any special limitations on the parameters used to determine the fusion coefficient.

[0065] The following will combine Figure 3 This section introduces the process of denoising raw images using a denoising neural network. For example... Figure 3 As shown, the denoising neural network includes a feature extraction module, a fusion coefficient decision module, and a feature fusion module. The inputs of the denoising neural network include the denoised image raw(n-1) of the previous frame, the noisy image rawn of the current frame, and the SNR(n-1) of the previous frame image. The output is the features of the denoised image rawn of the current frame.

[0066] The feature extraction module is used to extract features from the input image. This module extracts image features from the previous frame's denoised image raw(n-1), which can be denoted as feature A, and extracts image features from the current frame's noisy image rawn, which can be denoted as feature B. Both feature A and feature B are matrices, and are referred to as feature A and feature B for ease of description.

[0067] The fusion coefficient decision module is used to determine the fusion coefficient α between feature A of the denoised image in the previous frame and feature B of the noisy image in the current frame, based on the SNR value of the previous frame. Since the SNR values ​​of two adjacent frames are not significantly different, the SNR value of the raw image in the previous frame can be used to determine the fusion coefficient α.

[0068] The larger the fusion coefficient α, the greater the proportion of features from the denoised image of the previous frame in the fused image, and the smaller the proportion of features from the noisy image of the current frame. Conversely, the smaller the fusion coefficient α, the smaller the proportion of features from the denoised image of the previous frame in the fused image, and the greater the proportion of features from the noisy image of the current frame. The SNR value of the previous frame is negatively correlated with the fusion coefficient α; that is, the larger the SNR value of the previous frame, the smaller α, and vice versa.

[0069] The feature fusion module is used to fuse the features A of the previous frame's denoised image and the features B of the current frame's noisy image according to the fusion coefficient α to obtain the current frame's denoised image. For example, the features of the current frame's denoised image are A*α+B*(1-α). Finally, the current frame's image signal is obtained based on the features of the current frame's denoised image.

[0070] The training process of the denoising neural network is as follows:

[0071] (1) Create the initial denoising neural network.

[0072] (2) Input the first frame of the denoised image and its corresponding SNR value, as well as the second frame of the noisy image, into the created denoising neural network. After processing by the denoising neural network, the features of the second frame of the denoised image are output.

[0073] In one exemplary embodiment, a first frame of noisy image is denoised using an existing denoising algorithm to obtain a first frame of denoised image. Further, the SNR value of the first frame is calculated based on the first frame of denoised image and the first frame of noisy image. Then, the first frame of denoised image, the second frame of noisy image, and the SNR of the first frame are input into a denoising neural network. The denoising neural network extracts image features from the first frame of denoised image and the second frame of noisy image, respectively, and determines a fusion coefficient α based on the SNR value of the first frame. Finally, the features of the first frame of denoised image and the second frame of noisy image are fused based on α to obtain the features of the second frame of denoised image.

[0074] (3) Based on the mapping relationship between the SNR value and the fusion coefficient α, the target fusion coefficient α corresponding to the SNR value of the first frame is obtained. Furthermore, based on the feature extraction module in the denoising neural network, the features of the denoised image of the first frame and the noisy image of the second frame are extracted, and the features of the target denoised image corresponding to the second frame are obtained from the target fusion coefficient α.

[0075] In addition, the SNR value of the second frame image is calculated based on the denoised target image and the noisy image of the second frame.

[0076] In an exemplary embodiment, the image features of the target denoised image corresponding to the second frame image can be obtained by using A*α+B*(1-α) based on the feature extraction module in the current denoising neural network, the feature A extracted from the first frame denoised image and the feature B extracted from the second frame noisy image, and the target fusion coefficient α.

[0077] (4) Calculate the error between the features of the denoised image corresponding to the second frame output by the denoising neural network and the features of the target image. Adjust the parameters in the fusion coefficient decision module of the denoising neural network according to the error, that is, update the fusion coefficient decision module of the denoising neural network.

[0078] For each subsequent frame in the sample video, steps (3) and (4) above are repeated until the error corresponding to the same frame is within a preset range, thus obtaining the trained denoising neural network. There can be multiple sample videos, for example, multiple videos of different scene types.

[0079] Scenario 1: The image sensor supports dynamically switching the resolution of the raw image during video recording.

[0080] In one scenario, the image sensor supports dynamically switching the sharpness of the raw image during video recording. In this scenario, the ISP triggers the image sensor to switch the raw image sharpness based on the SNR value of the previous frame. The following will combine... Figure 4 and Figure 5 This section introduces the video noise reduction process for this type of scenario. Figure 4 This is a schematic diagram of a video noise reduction process based on a software architecture of an electronic device provided in an embodiment of this application. Figure 5 This is a flowchart of the video noise reduction processing method provided in the embodiments of this application.

[0081] exist Figure 1 The illustrated electronic device has an operating system running on its hardware components, on which applications can be installed. The operating system of the electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of the electronic device.

[0082] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided from top to bottom into the Application layer, the Framework layer, the Hardware Abstraction Layer (HAL), and the Kernel layer.

[0083] The application layer may include a series of application packages. For example, an application package may include applications such as a camera and a gallery, but this application does not limit this.

[0084] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes predefined functions. For example, it can include camera access interfaces such as camera management and camera devices. Camera management can provide an access interface for managing cameras. Camera devices can provide an interface for accessing cameras.

[0085] The HAL layer encapsulates Linux kernel drivers, provides interfaces to higher-level systems, and shields them from the implementation details of the underlying hardware. In this embodiment, the HAL layer may include a camera HAL and other hardware device abstraction layers. The camera HAL can call algorithms from a camera algorithm library. The camera algorithm library may include algorithms for image processing.

[0086] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera device driver, audio driver, sensor driver, GPU driver, and NPU driver.

[0087] The hardware layer can include image sensors, image signal processors (ISPs), GPUs, and NPUs.

[0088] The following will combine Figure 4 Introducing the video noise reduction process:

[0089] ① Processing of the first frame of raw image

[0090] After receiving the user's touch video shooting control, the camera application generates a shooting command and passes it sequentially through the camera access interface, the camera hardware abstraction layer (i.e., the camera HAL), and the camera device driver to the image sensor. The image sensor responds to the shooting command, obtains the first frame raw image (i.e., noisy raw0), and passes it to the ISP.

[0091] The ISP denoises the noisy raw0 image to obtain denoised raw0, and obtains the SNR value (SNR0) of the first frame raw image based on both the denoised raw0 and noisy raw0. The ISP then sequentially transmits the processed first frame image (after denoising and parameter optimization) through the camera device driver, camera HAL, and camera access interface to the camera application so that the camera application can display a video preview. Simultaneously, the ISP determines the resolution of the next raw frame based on the SNR0 value, and sends a resolution switching command to the image sensor if a resolution change is needed. The image sensor responds to the resolution switching command by adjusting the resolution of the obtained raw image, i.e., changing the resolution of the next raw frame.

[0092] In one exemplary embodiment, the ISP maintains a mapping relationship between different SNR values ​​and output modes. As previously mentioned, different output modes output raw images with different resolutions. The ISP obtains the output mode corresponding to the current frame's SNR value by querying this mapping relationship, thus determining the resolution of the next frame's raw image.

[0093] The SNR values ​​calculated from raw images at different resolutions are proportional. For example, if the resolution of the Hex-pattern is four times that of the quadr-pattern, the SNR value of the Hex-pattern is twice that of the quadr-pattern. The SNR value of the quadr-pattern is twice that of the binning-pattern.

[0094] In an exemplary embodiment, the ISP can maintain the mapping relationship between the SNR values ​​of each plotting pattern and the plotting pattern, such as the mapping relationship between the SNR value range of each Hex-pattern and the plotting pattern, the mapping relationship between the SNR value range of each quadr-pattern and the plotting pattern, and the mapping relationship between the SNR value range of each binning-pattern and the plotting pattern.

[0095] In another exemplary embodiment, the ISP can maintain a mapping relationship between the SNR value of any output mode and the output mode. For example, the ISP maintains a mapping relationship between each SNR value under the binning-pattern and the output mode. For example, under the binning-pattern, if the SNR value of the current frame raw image is in the range of (0, 10], the resolution of the next frame raw image can be selected from a lower resolution output mode, such as binning-pattern. If the SNR value of the current frame raw image is in the range of (10, 20], the resolution of the next frame raw image can be selected from a higher resolution output mode, such as quadr-pattern. If the SNR value of the current frame raw image is in the range of (20, 30], the resolution of the next frame raw image can be selected from a higher resolution output mode, such as Hex-pattern.

[0096] In one exemplary embodiment, the ISP includes processing logic for switching the raw image resolution (i.e., the output mode of the raw image) based on the SNR value. After obtaining the SNR value of the first frame of the raw image, the ISP triggers the execution of this processing logic to obtain the resolution corresponding to the SNR value. This resolution is the optimal resolution for the next frame of the raw image. Furthermore, the ISP generates a resolution switching command including this resolution and transmits it to the image sensor. In addition, the ISP may also include other processing algorithm logic, such as optimizing parameters like exposure and color temperature of the shooting scene.

[0097] With the ISP maintaining a mapping relationship between the SNR values ​​of each output mode and the output mode, after the ISP calculates the SNR value corresponding to the first frame raw image, it can query the mapping relationship between the SNR value of the first frame raw image and the output mode to obtain the output mode corresponding to the SNR value of the first frame (i.e., the resolution of the second frame raw image).

[0098] When the ISP maintains a mapping relationship between the SNR value and the output mode of an image output mode, after the ISP calculates the SNR value of the first raw frame, if the output mode of the first raw frame differs from the output mode of the SNR value in the mapping relationship maintained by the ISP, then the SNR value of the first raw frame is converted to the output mode corresponding to the SNR value in the mapping relationship maintained by the ISP. For example, if the first raw frame is a raw image obtained in quadr-pattern mode and has an SNR value of 20, and the ISP maintains a mapping relationship between each SNR value and the output mode in binning-pattern mode, then the SNR of the first frame (20) needs to be converted to binning-pattern mode. If the resolution in quadr-pattern mode is four times that in binning-pattern mode, then the SNR value of the first raw frame in binning-pattern mode is 10.

[0099] In addition, the ISP will pass the denoised Raw0 and SNR0 to the camera algorithm library via the camera device driver and camera HAL, so that the camera algorithm library can perform other processing on the denoised Raw0 to obtain the processed image, and pass this image to the camera application via the camera HAL and camera access interface so that the camera application can display the first frame of the video.

[0100] ② Processing of raw images in frames 2 to n

[0101] In video shooting mode, after the image sensor obtains the first frame of raw image, it will continue to the next frame of raw image (i.e. the second frame with noise raw1) and transmit it to the ISP.

[0102] In this embodiment, the image sensor responds to the resolution switching command transmitted by the ISP by switching the resolution of the second frame raw image, that is, outputting a raw image with a second resolution. Further, the ISP converts or processes the noisy raw1 image and then transmits it sequentially through the camera device driver and the camera HAL to the camera algorithm library.

[0103] In some embodiments, the camera algorithm library runs with Figure 3 The denoising neural network model shown is illustrated. For example, a camera algorithm library can invoke the NPU to execute the denoising neural network algorithm.

[0104] The camera algorithm library takes the received first frame (denoised image raw0), the second frame (noisy image raw1), and the SNR value of the first frame as inputs to the denoising neural network. The denoising neural network outputs the second frame (denoised image raw1), also known as the processed image. The camera algorithm library can also perform other processing on the denoised image raw1. For details on the process of the denoising neural network obtaining its output based on the input, please refer to [link to relevant documentation]. Figure 3The relevant descriptions will not be repeated here.

[0105] Furthermore, the camera algorithm library transmits the processed image of the second frame sequentially through the camera HALL and the camera access interface to the camera application, so that the camera application can display the processed image of the second frame in the preview window.

[0106] In addition, the camera algorithm library will also pass the obtained denoised raw1 of the second frame to the ISP via the camera HAL and camera device driver, so that the ISP can calculate the SNR value of the second frame based on the noisy raw image of the second frame and the raw1 of the second frame, which can be called SNR1. Furthermore, the ISP selects the resolution of the raw image of the third frame based on SNR1.

[0107] The signal-to-noise ratio (SNR) of images with different resolutions is proportional. The SNR calculated for each frame needs to be unified to the same resolution. For example, the resolution corresponding to the mapping relationship between SNR range and resolution stored in the ISP. For example, the ISP stores the SNR and resolution under quadr-pattern. The SNR of the raw image under binning-pattern is converted to quadr-pattern.

[0108] Furthermore, the algorithms in the camera algorithm library can be implemented based on the hardware layer, such as GPUs or NPUs.

[0109] After receiving the stop shooting command, the camera application sends the stop shooting command to the image sensor layer by layer. The image sensor responds to the stop shooting command, and then passes the processed last frame of the image to the camera application layer by layer, and finally sends the resulting video to the gallery.

[0110] The following is combined with Figure 5 This section introduces the video denoising method and workflow for scenario one. For example... Figure 5 As shown, the method may include the following steps:

[0111] S101a, the camera application receives the user's operation on the video shooting control, generates a video shooting command, and transmits it to the camera HAL.

[0112] S101b, the camera HAL transmits video capture commands to the image sensor.

[0113] Combination Figure 4 As can be seen, the video capture command generated by the camera application is transmitted to the camera HAL via the camera access interface. The camera HAL then transmits this video capture command to the image sensor via the camera device driver.

[0114] S102, the image sensor responds to the video capture command and captures the first frame image at a first resolution.

[0115] The first resolution can be any of the three resolutions supported by the image sensor.

[0116] S103, the image sensor sends a noisy image raw0 with a first resolution to the ISP.

[0117] The first frame image with the first resolution obtained by the image sensor is transmitted to the ISP for corresponding processing.

[0118] S104, the ISP performs denoising processing on the noisy image raw0 to obtain the denoised image raw0 corresponding to the first frame.

[0119] The ISP can use existing denoising algorithms to denoise the noisy raw0 to obtain the first frame denoised image raw0. In addition, the ISP can also perform other processing on the noisy raw0, which will not be detailed here.

[0120] S105, the ISP obtains the SNR value of the first frame based on the denoised image and the noisy image of the first frame.

[0121] For example, the SNR value corresponding to the raw image of the first frame can be calculated based on the denoised image and the noisy image corresponding to the first frame and according to the aforementioned SNR calculation formula, and can be denoted as SNR0.

[0122] In S106a, the ISP transmits the denoised image and SNR value of the first frame to the camera algorithm library layer by layer.

[0123] Combination Figure 4 The ISP transmits the denoised raw0 and SNR of the first frame sequentially through the camera device driver and the camera HAL to the camera algorithm library for image post-processing.

[0124] S106b, the camera algorithm library transmits the processed image of the first frame to the camera HAL.

[0125] S106c, the camera HAL transmits the processed image of the first frame to the camera application.

[0126] Combination Figure 4 The camera HAL transmits the processed image of the first frame to the camera application via the camera access interface, so that the camera application can display the processed image of the first frame in the preview window.

[0127] S107, the ISP determines the resolution of the next frame's raw image based on the SNR0 of the first frame.

[0128] S108 generates a resolution switching command and sends it to the image sensor when the resolution of the next frame of raw image is different from the resolution of the first frame.

[0129] S109, the image sensor responds to the resolution switching command and switches to shooting at the second resolution, and transmits the noisy image raw1 of the second resolution to the ISP.

[0130] S110, the ISP passes the noisy second-resolution image raw1 to the camera algorithm library layer by layer.

[0131] After the ISP processes the received second raw image (e.g., optimizes the exposure and color temperature of the scene), it is passed to the camera algorithm library via the camera device driver and the camera HAL.

[0132] S111a, the camera algorithm library uses the SNR of the first frame to fuse the features of the denoised image of the first frame and the noisy image of the second frame to obtain the processed image of the second frame.

[0133] In an exemplary embodiment of this application, the camera algorithm library includes a denoising neural network model. A denoised first frame image, a noisy second frame image, and the SNR value of the first frame are input into the denoising neural network, which outputs a denoised second frame image. For the processing procedure of the denoising neural network, please refer to [link to relevant documentation]. Figure 3 The corresponding descriptions will not be repeated here.

[0134] In addition, the camera algorithm library can also include other post-processing algorithms, and the denoised image can be further processed to obtain a processed image.

[0135] S111b, the camera algorithm library sends the processed image of the second frame to the ISP so that the ISP can obtain the SNR value of the second frame.

[0136] After the camera algorithm library obtains the second frame processed image through denoising, the second frame denoised image is sequentially transmitted to the ISP through the camera HAL and camera device driver, so that the ISP can calculate the SNR value of the second frame image based on the second frame processed image and the second frame noisy image, and further determine the resolution of the third frame raw image based on the SNR value of the second frame.

[0137] S112a, the camera algorithm library transmits the processed image of the second frame to the camera HAL.

[0138] S112b, the camera HAL transmits the processed image of the second frame layer by layer to the camera application for display.

[0139] Combination Figure 4 The camera HAL transmits the processed second frame image to the camera application via the camera access interface, so that the camera application can display the processed second frame image in the preview window.

[0140] The image processing for frames 3 through n is the same as that for frame 2, and will not be repeated here.

[0141] The video denoising method provided in this embodiment selects the resolution of the current frame's raw image based on the SNR value of the previous frame, where SNR is positively correlated with sharpness. This allows the image sensor to dynamically select the most suitable raw image as input to the denoising neural network based on the different signal-to-noise ratios (SNR) during the recording of the same video segment. For example, when the SNR is high, the denoising process is simpler, so a high-resolution raw image can be selected, thereby improving video quality; when the SNR is low, the denoising process is more difficult, so a low-resolution raw image with less noise can be selected, reducing the difficulty of denoising and thus improving the denoising effect and efficiency. Furthermore, a fusion coefficient α is determined based on the SNR value of the previous frame. This fusion coefficient is used to fuse the features of the noiseless image from the previous frame and the features of the noisy image from the current frame to obtain the denoised image corresponding to the current frame. Using raw images of different resolutions as input to the denoising process may cause slight variations in the sharpness of adjacent raw images. By fusing the previous and current frames, the sharpness difference between adjacent raw images is reduced, improving the overall video quality. The fusion ratio in this application is determined by the SNR value. For frames with low SNR values, more of the denoised image from the previous frame is fused, while for frames with high SNR values, more of the noisy raw image from the current frame is fused, thus improving video quality.

[0142] Scenario 2: The image sensor does not support dynamically switching the resolution of the raw image during video recording.

[0143] If the image sensor does not support dynamic switching of raw image resolution during video recording, a low-resolution raw image can be obtained from the high-resolution raw image output by the image sensor by the camera algorithm library.

[0144] The following will combine Figure 6 and Figure 7 This section describes the video noise reduction process in this scenario. Figure 6 This is a schematic diagram of a video noise reduction process based on a software architecture of an electronic device provided in an embodiment of this application; Figure 7 This is a flowchart of the video noise reduction processing method provided in the embodiments of this application.

[0145] ① Processing of the first frame of raw image

[0146] like Figure 6 As shown, the image sensor responds to the shooting instructions passed down layer by layer by the camera application to obtain the first frame raw image (i.e., noisy raw0) and passes it to the ISP for processing (such as denoising) to obtain denoised raw0. The ISP then passes the noisy raw0 and denoised raw0 sequentially through the camera device driver and the camera HAL to the camera algorithm library. Here, denoised raw0 can be the image obtained after denoising by the ISP, or the image obtained after denoising and other post-processing.

[0147] The camera algorithm library calculates the SNR value of the first frame, i.e., SNR0, based on the denoised raw0 and noisy raw0 of the first frame. Furthermore, the camera algorithm library determines the output mode (i.e. the resolution of the raw image) of the second frame raw image based on the SNR0 value.

[0148] In an exemplary embodiment, the SNR value of the first frame, i.e., SNR0, can be calculated based on the noisy raw0 and the denoised raw0 using an algorithm in the camera algorithm library.

[0149] In addition, the camera algorithm library can process the denoised raw0 to obtain the first frame of the processed image, and then pass the first frame of the processed image to the camera application through the camera HAL and the camera access interface in sequence, so that the camera application can display the first frame of the processed image in the preview window.

[0150] ② Processing of the second frame raw image

[0151] In this scenario, the image sensor always captures video at the first resolution (i.e., the resolution corresponding to the first output mode). When the image sensor outputs the second frame raw image (i.e., the first resolution noisy raw1), it is sequentially transmitted to the camera algorithm library via the camera device driver and the camera HAL.

[0152] The camera algorithm library determines the second resolution (i.e., the resolution corresponding to the second output mode) of the second frame raw image based on the SNR0 of the first frame. If the second resolution is lower than the first resolution of the first frame raw image, a corresponding algorithm is used to convert the noisy raw1 at the first resolution output by the image sensor to noisy raw1 at the second resolution. For example, a downsampling algorithm is used to downsample the first resolution raw image to obtain the second resolution raw image.

[0153] Furthermore, the camera algorithm library performs denoising processing on the noisy raw1 image at the second resolution to obtain the denoised image of the second frame, or the processed image of the second frame. Additionally, the camera algorithm library can calculate the SNR value of the second frame based on the processed image of the second frame and the noisy raw image, so that the camera algorithm library can determine the output mode (or resolution) of the third frame based on this SNR value.

[0154] Meanwhile, the camera algorithm library transmits the processed image of the second frame sequentially through the camera HAL and the camera access interface to the camera application, so that the camera application can display the image in the preview window.

[0155] In this scenario, the camera algorithm library can store the mapping relationship between different SNR ranges and image output modes (or resolutions). The camera algorithm library determines the image output mode (or resolution) of the next frame based on the SNR value of the current frame in the same way as the ISP determination method in Scenario 1, and will not be repeated here. The processing of other captured frame images is the same as the processing of the second frame mentioned above, and will not be repeated here. After stopping shooting, the final video is sent to the image library.

[0156] The following will combine Figure 7 This section introduces the video denoising method and workflow for Scenario 2, such as... Figure 7 As shown, the method may include the following steps:

[0157] S201a, the camera application receives the user's operation on the video shooting controls, generates a video shooting command, and transmits it to the camera HAL.

[0158] S201b, the camera's HAL transmits video capture commands to the image sensor.

[0159] S202, the image sensor responds to the video capture command and captures a video at a first resolution.

[0160] S203, the image sensor sends the first frame of noisy image raw0 to the ISP.

[0161] S204a, the ISP performs denoising processing on the noisy image of the first frame to obtain the denoised image of the first frame.

[0162] For the first frame of raw image, existing denoising algorithms can be used to denoise it, resulting in a denoised image.

[0163] S204b, the ISP transmits the first frame of the noisy image and the first frame of the denoised image to the camera HAL.

[0164] In S204c, the camera HAL transmits the first noisy image and the first denoised image to the camera algorithm library.

[0165] S205, the camera algorithm library obtains the SNR value of the first frame based on the denoised image of the first frame and the noisy image of the first frame.

[0166] In addition, other post-processing algorithms in the camera algorithm library can be used to further process the denoised image of the first frame.

[0167] S206a, the camera algorithm library passes the first frame of denoised image to the camera HAL.

[0168] S206b, the camera HAL transmits the first frame of the denoised image to the camera application.

[0169] See Figure 6The camera HAL transmits the first frame of the denoised image to the camera application through the camera access interface, so that the camera application can display the image in the preview window.

[0170] S207, the image sensor sends the second noisy image to the ISP.

[0171] After receiving the video capture command, the image sensor continuously captures images at the first resolution according to the specified frame rate. The second raw frame image here is still an image at the first resolution. After obtaining the noisy second frame image, the image sensor sends it to the ISP for processing, such as converting it into a digital image signal.

[0172] S208a, the ISP transmits the second noisy image to the camera HAL.

[0173] S208b, the camera HAL transmits the second noisy image to the camera algorithm library.

[0174] The ISP transmits the processed noisy second frame image to the camera algorithm library via the camera device driver and the camera HAL.

[0175] S209, when the camera algorithm library determines that the second resolution corresponding to the second frame is lower than the first resolution of the first frame based on the SNR value of the first frame, the second frame with a noisy image of the second frame with the second resolution is obtained based on the noisy image of the second frame with the first resolution.

[0176] The process of determining the resolution of the second frame based on the SNR value of the first frame in this step is similar to... Figure 5 The same applies to S107, so it will not be repeated here. The camera algorithm library can store the mapping relationship between SNR values ​​and resolutions. Based on this mapping relationship, the resolution corresponding to the SNR value of the first frame can be retrieved, i.e., the second resolution.

[0177] When the second resolution is lower than the first resolution, a downsampling algorithm can be used to downsample the noisy second frame of the first resolution image to obtain the noisy second frame of the second resolution image.

[0178] S210, the camera algorithm library fuses the denoised image of the first frame and the noisy image of the second frame with the second resolution based on the SNR value of the first frame to obtain the processed image of the second frame.

[0179] The specific implementation process of this step and Figure 5 The implementation process of S111a is the same, and will not be repeated here. In addition, the camera algorithm library may also include other post-processing algorithms.

[0180] S211a, the camera algorithm library transmits the processed image of the second frame to the camera HAL.

[0181] S211b, the camera HAL transmits the processed image of the second frame to the camera application.

[0182] The camera algorithm library transmits the processed second frame image sequentially through the camera HAL and the camera access interface to the camera application, so that the camera application can display the image in the preview window.

[0183] The processing of the other captured frames is the same as that of the second frame described above, and will not be repeated here. After stopping recording, the final video is sent to the image library.

[0184] The video denoising method provided in this embodiment selects the resolution of the current frame's raw image based on the SNR value of the previous frame, with the SNR value being positively correlated with the resolution. This allows the camera algorithm library to dynamically obtain the most suitable raw image as input to the denoising neural network based on the different signal-to-noise ratios during the recording of the same video segment, thereby improving the denoising effect and efficiency. Furthermore, a fusion coefficient α is determined based on the SNR value of the previous frame, and features from the noiseless image of the previous frame and the noisy image of the current frame are fused according to this coefficient to obtain the denoised image corresponding to the current frame. Using raw images of different resolutions as input to the denoising process for adjacent raw images may cause slight variations in the sharpness of adjacent raw images. By fusing the previous and current frames, the sharpness difference between adjacent raw images is reduced, improving the overall video quality. The fusion ratio in this application is determined by the SNR value. When the SNR value of the previous frame is low, more of the denoised image from the previous frame is fused; when the SNR value of the previous frame is high, more of the noisy raw image from the current frame is fused, improving video quality. Moreover, this embodiment is applicable to scenarios where the image sensor does not support dynamic resolution switching, which reduces the dependence of the method on the hardware functions of the image sensor and improves the applicability of the method.

[0185] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0186] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video denoising method, characterized in that, Applied to electronic devices, the method includes: In response to a video recording operation, video recording begins. The first frame of the original image is denoised to obtain the first frame of the denoised image. The first frame of the original image is an image with a first resolution. A first signal-to-noise ratio (SNR) corresponding to the first frame original image is obtained based on the first frame original image and the first frame denoised image; a second resolution corresponding to the second frame original image is determined based on the first SNR, and a second frame original image with the second resolution is obtained, wherein the second resolution is different from the first resolution and the second resolution is positively correlated with the first SNR; The fusion coefficient corresponding to the second frame original image is obtained based on the first signal-to-noise ratio, and the first frame denoised image and the second frame original image are fused according to the fusion coefficient to obtain the second frame denoised image; the fusion coefficient is the proportion of the features of the first frame denoised image in the features of the second frame denoised image, and the fusion coefficient is negatively correlated with the first signal-to-noise ratio; The second signal-to-noise ratio (SNR) corresponding to the second original image is obtained based on the second denoised image and the second original image. The second SNR is used to determine the resolution of the third original image, and the third original image is processed according to the processing procedure of the second original image until a stop operation is received to stop video recording, thereby obtaining the denoised video.

2. The method according to claim 1, characterized in that, The step of obtaining the fusion coefficient corresponding to the second frame original image based on the first signal-to-noise ratio, and fusing the first frame denoised image with the second frame original image according to the fusion coefficient to obtain the second frame denoised image includes: The first frame of denoised image, the second frame of original image and the first signal-to-noise ratio are input into the denoising neural network for denoising processing; The denoising neural network determines the fusion coefficients based on the first signal-to-noise ratio of the input; The denoising neural network extracts a first image feature from the first frame of denoised image and a second image feature from the second frame of original image at the second resolution; The denoising neural network fuses the first image features and the second image features according to the fusion coefficient to obtain the second frame denoised image.

3. The method according to claim 2, characterized in that, The training process of the denoising neural network includes: Create the initial denoising neural network; The first frame of the denoised image and signal-to-noise ratio in the sample video, as well as the second frame of the original image, are input into the initial denoising neural network for denoising processing to obtain the second frame of the sample denoised image. Based on the mapping relationship between signal-to-noise ratio and fusion coefficient, a target fusion coefficient that matches the first signal-to-noise ratio is determined. The first sample image features and the second sample image features are fused based on the target fusion coefficient to obtain the second frame target denoised image. The first sample image features are extracted from the first frame denoised image by the initial denoising neural network, and the second sample image features are extracted from the second frame original image by the initial denoising neural network. The error between the second frame target denoised image and the second frame sample denoised image is obtained, and the parameters of the denoising neural network are adjusted according to the error; For any frame after the second frame in the sample video, the processing procedure for the second frame is repeated until the error between the sample denoised image and the target denoised image corresponding to the same frame is within a preset range, thus ending the training process of the denoising neural network.

4. The method according to any one of claims 1-3, characterized in that, The electronic device includes an image sensor and an image signal processor, wherein the image sensor supports dynamically switching the shooting resolution during video recording; Determining the second resolution corresponding to the second frame original image based on the first signal-to-noise ratio, and obtaining the second frame original image at the second resolution, includes: The image signal processor queries the mapping relationship between the signal-to-noise ratio range and the resolution to obtain the second resolution corresponding to the first signal-to-noise ratio; When the image signal processor determines that the second resolution is different from the first resolution, it sends a resolution switching command to the image sensor; The image sensor responds to the resolution switching command by switching the resolution of the captured image to the second resolution and sends the second frame of the original image at the second resolution to the image signal processor.

5. The method according to claim 4, characterized in that, The step of obtaining the second signal-to-noise ratio corresponding to the second frame original image based on the second frame denoised image and the second frame original image includes: The image signal processor calculates the second signal-to-noise ratio corresponding to the second frame original image based on the second frame denoised image and the second frame original image, and determines the third resolution of the third frame original image based on the second signal-to-noise ratio.

6. The method according to claim 4, characterized in that, The image signal processor maintains a mapping relationship between signal-to-noise ratio (SNR) ranges and resolutions for different output modes; the image signal processor queries the mapping relationship between SNR ranges and resolutions to obtain a second resolution corresponding to the first SNR, including: The image signal processor queries the mapping relationship between the signal-to-noise ratio range and resolution corresponding to the output mode of the first frame original image to obtain the second resolution corresponding to the first signal-to-noise ratio.

7. The method according to claim 4, characterized in that, The image signal processor maintains a first mapping relationship between the signal-to-noise ratio and resolution corresponding to the first output mode, and the resolution of the original image output in the first output mode is a third resolution; The image signal processor queries the mapping relationship between the signal-to-noise ratio range and the resolution to obtain the second resolution corresponding to the first signal-to-noise ratio, including: The image signal processor converts the first signal-to-noise ratio corresponding to the first frame of the original image at the first resolution into the second signal-to-noise ratio corresponding to the third resolution. The image signal processor queries the mapping relationship between the signal-to-noise ratio range corresponding to the third resolution and the resolution to obtain the second resolution corresponding to the second signal-to-noise ratio.

8. The method according to any one of claims 1-3, characterized in that, The electronic device includes an image sensor, an image signal processor, and a camera algorithm library; Obtaining the first signal-to-noise ratio corresponding to the first frame original image based on the first frame original image and the first frame denoised image includes: The image signal processor performs denoising processing on the first frame of the original image output by the image sensor to obtain the first frame of the denoised image. The image signal processor transmits the first frame denoised image and the first frame original image to the camera algorithm library; The camera algorithm library calculates the first signal-to-noise ratio corresponding to the first frame original image based on the first frame denoised image and the first frame original image.

9. The method according to claim 8, characterized in that, The image sensor does not support dynamically switching the shooting resolution during video recording; The step of determining the second resolution corresponding to the second frame original image based on the first signal-to-noise ratio, and obtaining the second frame original image at the second resolution, includes: The camera algorithm library queries the mapping relationship between signal-to-noise ratio and resolution to obtain the second resolution corresponding to the first signal-to-noise ratio, and the second resolution is lower than the first resolution; The image sensor captures a second frame of the original image at the first resolution and transmits it to the camera algorithm library; The camera algorithm library downsamples the second frame original image at the first resolution to obtain the second frame original image at the second resolution.

10. The method according to claim 9, characterized in that, The second signal-to-noise ratio (SNR) of the second frame original image is obtained based on the second frame denoised image and the second frame original image, including: The camera algorithm library calculates the second signal-to-noise ratio corresponding to the second frame original image based on the second frame original image at the second resolution and the second frame denoised image at the second resolution.

11. An electronic device, characterized in that, The electronic device includes: one or more processors, a memory, and a touch screen; the memory is used to store program code; the processor is used to run the program code, causing the electronic device to implement the video denoising method as described in any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, It stores instructions that, when executed on an electronic device, cause the electronic device to perform the video denoising method as described in any one of claims 1 to 10.

13. A computer program product, characterized in that, It stores instructions that, when the computer program product is run on the electronic device, cause the electronic device to implement the video denoising method as described in any one of claims 1 to 10.

14. A chip system, characterized in that, include: At least one processor and an interface, the interface being used to receive code instructions and transmit them to the at least one processor; The at least one processor executes the code instructions to implement the video denoising method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Video image noise reduction method and device and computer storage medium

    CN115965537A

  • Integrated sensor with frame memory and programmable resolution for light adaptive imaging

    US5909026A