Video denoising method and device
By dynamically adjusting the image resolution and signal-to-noise ratio fusion technology, the quality degradation problem caused by noise in the video signal is solved, and the video quality is improved and the denoising efficiency is increased.
Patent Information
- Application Number
- CN202411135286.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-08-16
AI Technical Summary
The noise signals mixed in the video signal lead to the degradation of video quality, and the existing technology is difficult to effectively improve the video quality.
Video denoising is achieved by responding to video shooting operations, dynamically adjusting image resolution using a denoising neural network and an image signal processor, performing image fusion based on the signal-to-noise ratio, dynamically switching the resolution of the image sensor, and optimizing the denoising process through the fusion coefficient.
It improves video quality, reduces the difficulty of denoising and improves denoising efficiency, reduces the clarity difference between adjacent frames, and improves the overall video quality.
Smart Images

Figure CN120751273A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a video denoising method and device. Background Art
[0002] With the development of technology, users have higher and higher requirements for the quality of videos shot with mobile phones and other terminals. However, video signals are often mixed with various noise signals, resulting in blurry videos and degraded quality. Summary of the Invention
[0003] In view of this, the present application provides a video denoising method to solve the problem of poor video quality. The disclosed technical solution is as follows:
[0004] In a first aspect, the present application provides a video denoising method, which is applied to an electronic device, the method comprising: starting to shoot a video in response to a video shooting operation; performing denoising on a first frame of original image obtained by shooting to obtain a first frame of denoised image, where the first frame of original image is an image of a first resolution; obtaining a first signal-to-noise ratio corresponding to the first frame of original image based on the first frame of original image and the first frame of denoised image; determining a second resolution corresponding to the second frame of original image based on the first signal-to-noise ratio, and obtaining a second frame of original image with the second resolution, where the second resolution is different from the first resolution and the second resolution is positively correlated with the first signal-to-noise ratio; obtaining a second frame of original image with the second resolution according to shooting information. The first denoised image is fused with the second original image at a second resolution according to the fusion coefficient to obtain a second denoised image. The shooting information includes prior knowledge or shooting parameters of the previous image. The fusion coefficient is the ratio of the features of the first denoised image to the features of the second denoised image. A second signal-to-noise ratio (SNR) corresponding to the second original image is obtained based on the denoised image and the second original image. The second SNR is used to determine the resolution of the third original image. The third original image is processed according to the processing of the second image until a stop operation is received, stopping video capture and obtaining a denoised video. For example, with this solution, when the SNR is high, the denoising process is simpler and a high-resolution raw image can be selected, thereby improving video quality. When the SNR is low, the denoising process is more difficult and a low-resolution raw image with less noise can be selected, reducing the difficulty of denoising and improving the effectiveness and efficiency of the denoising process. Furthermore, by fusing the previous and current frames, this solution reduces the difference in clarity between adjacent raw images, thereby improving overall video quality.
[0005] In one possible implementation of the first aspect, the shooting information is a signal-to-noise ratio; obtaining a fusion coefficient corresponding to the second image frame based on the shooting information includes: obtaining the fusion coefficient for the second image frame based on a first signal-to-noise ratio corresponding to the first original image frame, where the fusion coefficient is positively correlated with the first signal-to-noise ratio. In this solution, the fusion ratio is determined by the SNR value. For low SNR values, more of the denoised image of the previous frame is fused. For high SNR values, more of the noisy raw image of the current frame is fused, thereby improving video quality.
[0006] In a possible implementation of the first aspect, a fusion coefficient corresponding to the second frame image is obtained according to the shooting information, and the first frame denoised image is fused with the second frame original image of the second resolution according to the fusion coefficient to obtain the second frame denoised image, including: inputting the first frame denoised image, the second frame original image of the second resolution and the first signal-to-noise ratio of the first frame image into a denoising neural network for denoising processing; the denoising neural network determines the fusion coefficient based on the first signal-to-noise ratio of the input first frame image; the denoising neural network extracts the first image feature from the first frame denoised image, and extracts the second image feature from the second frame original image of the second resolution; the denoising neural network fuses the first image feature and the second image feature according to the fusion coefficient to obtain the second frame denoised image.
[0007] In a possible implementation of the first aspect, the training process of the denoising neural network includes: creating an initial denoising neural network; inputting the first frame denoised image and signal-to-noise ratio of the sample video and the second frame original image into the initial denoising neural network for denoising to obtain a second frame sample denoised image; determining a target fusion coefficient that matches the first signal-to-noise ratio based on a mapping relationship between the signal-to-noise ratio and the fusion coefficient; fusing the first sample image features and the second sample image features based on the target fusion coefficient to obtain a second frame target denoised image, wherein the first sample image features are extracted from the first frame denoised image by the initial denoising neural network, and the second sample image features are extracted from the second frame original image by the initial denoising neural network; obtaining an error between the second frame target denoised image and the second frame sample denoised image, and adjusting the parameters of the denoising neural network according to the error; for any frame after the second frame image in the sample video, repeating the processing process of the second frame image until the error between the sample denoised image and the target denoised image corresponding to the same frame is within a preset range, thereby terminating the training process of the denoising neural network.
[0008] In a possible implementation of the first aspect, an electronic device includes an image sensor and an image signal processor, wherein the image sensor supports dynamic switching of capture resolution during video capture. Determining a second resolution corresponding to a second frame of raw image based on a first signal-to-noise ratio, and obtaining the second frame of raw image at the second resolution includes: the image signal processor querying a mapping relationship between a signal-to-noise ratio range and resolution to obtain the second resolution corresponding to the first signal-to-noise ratio; when the image signal processor determines that the second resolution is different from the first resolution, sending a resolution switching instruction to the image sensor; and the image sensor switching the resolution of the captured image to the second resolution in response to the resolution switching instruction, and sending the second frame of raw image at the second resolution to the image signal processor. In this solution, when the image signal processor determines that the image sensor's output mode needs to be switched, it notifies the image sensor to switch the output mode, thereby achieving dynamic switching of raw image resolution. Furthermore, in this solution, the resolution of a subsequent frame is controlled based on the signal-to-noise ratio of a previous frame of raw image. When the signal-to-noise ratio of the previous frame is low, a lower-resolution but lower-noise output mode can be selected for the next frame, thereby reducing the difficulty of noise reduction. When the signal-to-noise ratio of the previous frame is high, a higher-resolution output mode can be selected for the next frame, thereby improving video quality.
[0009] In a possible implementation of the first aspect, obtaining a second signal-to-noise ratio corresponding to the second frame original image based on the second frame denoised image and the second frame original image includes: an image signal processor calculating the second signal-to-noise ratio corresponding to the second frame original image based on the second frame denoised image and the second frame original image, and determining a third resolution of the third frame original image based on the second signal-to-noise ratio.
[0010] In one possible implementation of the first aspect, the image signal processor maintains a mapping relationship between the signal-to-noise ratio range and the resolution for each different image output mode; the image signal processor queries the mapping relationship between the signal-to-noise ratio range and the resolution to obtain the second resolution corresponding to the first signal-to-noise ratio, including: the image signal processor queries the mapping relationship between the signal-to-noise ratio range and the resolution corresponding to the image output mode of the first frame of the original image to obtain the second resolution corresponding to the first signal-to-noise ratio. It can be seen that in this solution, the image signal processor maintains the mapping relationship between the signal-to-noise ratio range and the resolution for each different image output mode. In this way, it is only necessary to query the mapping relationship corresponding to the image output mode of the first frame of the original image to obtain the resolution correctly corresponding to the signal-to-noise ratio of the first frame of the original image, thereby improving the timeliness of determining the resolution of the next frame and also improving the accuracy of the resolution.
[0011] In a possible implementation of the first aspect, an image signal processor maintains a first mapping relationship between a signal-to-noise ratio (SNR) and a resolution corresponding to a first output mode, where the resolution of the original image outputted in the first output mode is a third resolution; the image signal processor queries a mapping relationship between a SNR range and a resolution to obtain a second resolution corresponding to the first SNR, including: the image signal processor converts the first SNR corresponding to a first frame of the original image at the first resolution into a second SNR corresponding to the third resolution; and the image signal processor queries a mapping relationship between a SNR range and a resolution corresponding to the third resolution to obtain the second resolution corresponding to the second SNR. As can be seen, in this solution, the image signal processor only needs to maintain a mapping relationship between the SNR and the resolution for one output mode, and since the SNRs corresponding to different resolutions are multiples, the SNR for the output mode currently being queried can be converted to the output mode corresponding to the mapping relationship maintained by the image signal processor, thereby obtaining the accurate resolution corresponding to the SNR. Furthermore, this solution eliminates the need to maintain mapping relationships between the SNR and the resolution for all output modes in the image signal processor, thereby saving storage space in the image signal processor.
[0012] In a possible implementation of the first aspect, an electronic device includes an image sensor, an image signal processor, and a camera algorithm library; obtaining a first signal-to-noise ratio corresponding to the first frame of the original image based on the first frame of the original image and the first frame of the denoised image includes: the image signal processor denoising the first frame of the original image output by the image sensor to obtain the first frame of the denoised image; the image signal processor passing the first frame of the denoised image and the first frame of the original image to the camera algorithm library; and the camera algorithm library calculating the first signal-to-noise ratio corresponding to the first frame of the original image based on the first frame of the denoised image and the first frame of the original image. It can be seen that in this solution, the denoising algorithm in the image signal processor is used to denoise the first frame of the original image (i.e., the first frame of the raw image) to obtain the first frame of the denoised image. The camera algorithm library further calculates the first signal-to-noise ratio based on the first frame of the denoised image and the first frame of the original image provided by the image signal processor. In this way, there is no need to modify the processing logic in the image signal processor, reducing the complexity of the solution implementation.
[0013] In a possible implementation of the first aspect, the image sensor does not support dynamic switching of shooting resolution during video capture; determining the second resolution corresponding to the second frame of the original image based on the first signal-to-noise ratio, and obtaining the second frame of the original image at the second resolution include: the camera algorithm library queries the mapping relationship between the signal-to-noise ratio and the resolution to obtain the second resolution corresponding to the first signal-to-noise ratio, and the second resolution is lower than the first resolution; the image sensor captures the second frame of the original image at the first resolution and passes it to the camera algorithm library; the camera algorithm library downsamples the second frame of the original image at the first resolution to obtain the second frame of the original image at the second resolution. It can be seen that this solution is suitable for scenarios where the image sensor does not support dynamic switching of the raw image resolution. In this scenario, the camera algorithm library can convert the raw image of the first resolution output by the image sensor into a raw image of the second resolution, and only supports downward adjustment of the resolution of the raw image. In this way, the performance requirements for the image sensor hardware are reduced and the scope of application of the solution is improved.
[0014] In a possible implementation of the first aspect, obtaining a second signal-to-noise ratio corresponding to the second frame original image based on the second frame denoised image and the second frame original image includes: a camera algorithm library calculates the second signal-to-noise ratio corresponding to the second frame original image based on the second frame original image at the second resolution and the second frame denoised image at the second resolution.
[0015] In a second aspect, the present application also provides an electronic device, which includes: one or more processors, a memory and a touch screen; the memory is used to store program code; the processor is used to run the program code, so that the electronic device implements the video denoising method as any one of the first aspects.
[0016] In a third aspect, the present application further provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the video denoising method as described in any one of the first aspects.
[0017] In a fourth aspect, the present application further provides a computer program product, characterized in that instructions are stored thereon, and when the computer program product is run on an electronic device, the electronic device implements the video denoising method as described in any one of the first aspects.
[0018] In a fifth aspect, the present application also provides a chip system comprising: at least one processor and an interface, the interface being used to receive code instructions and transmit them to at least one processor; at least one processor runs the code instructions to implement any video denoising method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of a video noise reduction process provided by an embodiment of the present application;
[0021] Figure 3 Schematic diagram of the structure of a denoising neural network provided in an embodiment of the present application;
[0022] Figure 4 This is a schematic diagram of a video noise reduction process based on an electronic device software architecture provided by an embodiment of the present application;
[0023] Figure 5 The embodiment of this application provides Figure 4 Flowchart of the corresponding video denoising method;
[0024] Figure 6 Schematic diagram of another video noise reduction process based on the software architecture of an electronic device provided in an embodiment of the present application;
[0025] Figure 7 The embodiment of this application provides Figure 6 Flowchart of the corresponding video denoising method DETAILED DESCRIPTION
[0026] The terms "first", "second" and "third" in the specification, claims and drawings of this application are used to distinguish different objects rather than to limit a specific order.
[0027] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0028] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0029] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0030] The electronic device can be a mobile phone, a tablet computer, a wearable electronic device, an in-vehicle electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a smart screen, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, etc. The embodiments of the present application do not impose any restrictions on the specific type of electronic device.
[0031] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, antenna 1, antenna 2, a mobile communication module, a wireless communication module, an audio module, a speaker, a receiver, a microphone, an earphone interface, a sensor module, a button, a motor, an indicator, a camera, a display, and a subscriber identification module (SIM) card interface, etc. The sensor module may include a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0032] A processor may include one or more processing units. For example, a processor may include at least one of the following processing units: an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and a neural-network processing unit (NPU). Different processing units may be independent devices or integrated devices.
[0033] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.
[0034] In an embodiment of the present application, a processor can execute the video denoising method provided in an embodiment of the present application. When controlling a camera application to record a video, the processor selects the output mode of the current frame's raw image (i.e., the resolution of the next raw image) based on the signal-to-noise ratio of the previous frame. Furthermore, the processor fuses the previous denoised raw image with the current noisy raw image to obtain the current denoised image. The above processing is performed on each raw image captured to ultimately obtain a denoised video.
[0035] The wireless communication function of the electronic device can be realized by components such as antenna 1, antenna 2, mobile communication module, wireless communication module, modem processor and baseband processor.
[0036] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0037] The mobile communication module can provide wireless communication solutions including 2G / 3G / 4G / 5G / 6G applied in electronic devices.
[0038] The wireless communication module can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0039] Electronic devices can achieve display functionality through a GPU, display, and application processor. A GPU is a microprocessor for image processing that connects the display and application processor. The GPU performs mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information. The display can be used to display images or videos.
[0040] Electronic devices can achieve shooting functions through ISP, camera, video codec, GPU, display and application processor.
[0041] The camera is used to capture still images or videos. The incident light can be focused on the focal point of the lens through the lens, so that the photographed object is imaged on the image sensor. The image sensor can convert the optical signal into an electrical signal, which is transmitted to the ISP for processing and converted into a digital image signal. The digital image signal is further processed and converted into an image signal in standard RGB, YUV and other formats, which can be transmitted to the display for display. The ISP can perform algorithmic optimization on the noise, brightness and color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera. The electronic device may include 1 or N cameras, where N is a positive integer greater than 1.
[0042] Video codecs are used to compress or decompress digital video. Electronic devices may support one or more video codecs. This allows them to play or record videos in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0043] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals.
[0044] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic devices, such as image recognition, face recognition, speech recognition, and text comprehension.
[0045] The internal memory can be used to store computer-executable program code, which includes instructions. The processor executes the instructions stored in the internal memory to perform various functional applications and data processing of the electronic device. For example, in this embodiment, the processor can perform video noise reduction processing by executing the instructions stored in the internal memory.
[0046] A touch sensor, also known as a "touch-sensitive device," can be located on a display screen. The touch sensor and the display screen together form a touch screen, also known as a "touch screen." The touch sensor is used to detect touch operations applied to or near the touch sensor. The touch sensor can transmit the detected touch operations to an application processor to determine the type of touch event. Visual output related to the touch operations can be provided through the display screen. In other embodiments, the touch sensor can also be located on the surface of the electronic device, in a location different from that of the display screen.
[0047] It should be noted that Figure 1The structure shown does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include Figure 1 More or fewer components than those shown, or the electronic device may include Figure 1 Combinations of some of the components shown, or an electronic device may include Figure 1 Subassemblies of some of the components shown. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0048] The following will be combined Figure 2 and Figure 3 The video noise reduction process provided in the embodiments of the present application is introduced.
[0049] Figure 2 This is a schematic diagram of a video noise reduction processing process provided by an embodiment of the present application.
[0050] like Figure 2 As shown, after the electronic device receives the user's operation of shooting a video, the image sensor shoots and obtains the first frame of the original image. The original image usually contains a noise signal and can be called a noisy image raw0.
[0051] In some embodiments of the present application, the resolution of the noisy image raw0 may be a default resolution. For example, in a scenario, the image sensor supports multiple different image output modes, such as Hex-pattern, quadr-pattern, and binning-pattern, with resolutions decreasing in sequence. The resolutions of different image output modes are multiples of each other. For example, in one example, the resolution of Hex-pattern is four times that of quadr-pattern, which in turn is four times that of binning-pattern.
[0052] (1) Denoising the first frame of the noisy image raw0 to obtain the first frame of the denoised image, denoised raw0. The denoising process can use an existing denoising algorithm, which is not limited in this application.
[0053] (2) Based on the first frame of noisy image raw0 and its corresponding denoised image, the corresponding signal-to-noise ratio (SNR) value is calculated, that is, the SNR of the first frame image, which can be recorded as SNR0.
[0054] In some embodiments, the SNR value of the image can be calculated according to the following formula:
[0055]
[0056] In the above formula, f(x,y) represents the pixel value corresponding to the pixel point (x,y) in the noisy raw image. Indicates the pixel value corresponding to the pixel point (x, y) of the denoised raw image. In addition, this application does not limit the method for calculating the SNR value of the image.
[0057] (3) The resolution of the next raw image is determined based on the SNR0 of the first raw image, and the next raw image of the obtained resolution is input into the denoising neural network. At the same time, the SNR0 of the first frame and the denoised raw0 of the first frame are input into the denoising neural network.
[0058] If the SNR value is high, select a mode with higher resolution. For example, if the current mode is binning-pattern, select Hex-pattern or quadr-pattern. If the current mode is quadr-pattern, select Hex-pattern. If the SNR value is low, select a mode with lower resolution and lower noise. For example, if the current mode is quadr-pattern, select binning-pattern.
[0059] (4) Based on the SNR0 of the first frame, the denoising neural network fuses the denoised raw0 of the first frame and the noisy raw1 of the second frame to obtain the denoised image raw1 of the second frame.
[0060] (5) The SNR of the second frame image is obtained based on the second frame noisy image raw1 and the second frame denoised image raw1, which can be recorded as SNR1.
[0061] (6) Taking the nth frame as an example, the resolution of the nth frame is selected based on the SNR(n-1) value of the n-1th frame. The denoising neural network then fuses the noisy image of the nth frame with the denoised image of the n-1th frame based on the SNR(n-1) value of the n-1th frame to obtain the denoised image of the nth frame. The processing of subsequent images is the same as that of the nth frame until the last frame.
[0062] The video denoising method provided in this embodiment determines the clarity (i.e., resolution) of the raw image of the current frame based on the SNR value of the previous frame. For example, the larger the SNR value of the previous frame, the higher the clarity of the raw image output mode that can be selected for the current frame. Conversely, the smaller the SNR value of the previous frame, the lower the clarity but low-noise raw image output mode that can be selected for the current frame. That is, the SNR value is positively correlated with the clarity. This allows the most appropriate raw image to be dynamically selected as the input of the denoising neural network based on the different signal-to-noise ratios in the recording of the same video. For example, when the signal-to-noise ratio is high, the denoising process is relatively simple and a high-resolution raw image can be selected, thereby improving video quality. When the signal-to-noise ratio is low, the denoising process is relatively difficult and a low-resolution but low-noise raw image can be selected, reducing the difficulty of denoising and thereby improving the noise reduction effect and efficiency.
[0063] Furthermore, the fusion coefficient α can be determined according to the SNR value of the previous frame, and the features of the noise-free image of the previous frame and the features of the noisy image of the current frame can be fused according to the fusion coefficient to obtain the denoised image corresponding to the current frame. Using raw images of different resolutions as input for the denoising process for adjacent raw images may cause a small jump in the clarity of adjacent raw images. By fusing the previous frame and the current frame, the clarity difference between adjacent raw images is reduced, thereby improving the overall quality of the video. The fusion ratio in this application is determined by the SNR value. For frames with low SNR values, more denoised images of the previous frame are fused, and for frames with high SNR values, more noisy raw images of the current frame are fused, thereby improving the video quality.
[0064] In other embodiments of the present application, the fusion coefficient α may be determined based on other parameters such as ISO value, ambient light brightness, etc. For example, the lower the ambient light brightness, the higher the ISO value, and the lower the SNR value of the image, the larger the fusion coefficient α may be. The present application does not impose any particular limitation on the parameters used to determine the fusion coefficient.
[0065] The following will be combined Figure 3 This paper introduces the process of denoising raw images using denoising neural networks. Figure 3 As shown in Figure 1, the denoising neural network includes a feature extraction module, a fusion coefficient decision module, and a feature fusion module. The input of the denoising neural network includes the denoised image of the previous frame raw(n-1), the noisy image of the current frame rawn, and the SNR(n-1) of the previous frame. The output is the features of the denoised image of the current frame rawn.
[0066] The feature extraction module is used to extract features from the input image. This module extracts image features from the previous denoised image raw(n-1), which can be recorded as feature A, and image features extracted from the current noisy image rawn, which can be recorded as feature B. Both feature A and feature B are matrices and are recorded as feature A and feature B for convenience.
[0067] The fusion coefficient decision module is used to determine the fusion coefficient α between feature A of the previous denoised image and feature B of the current noisy image based on the SNR value of the previous frame. Since the SNR values of two adjacent frames are similar, the SNR value of the previous raw image can be used to determine the fusion coefficient α.
[0068] The larger the fusion coefficient α, the more features from the previous frame's denoised image are included in the fused image, and the fewer features from the current frame's noisy image are included in the fused image. The smaller the fusion coefficient α, the fewer features from the previous frame's denoised image are included in the fused image, and the more features from the current frame's noisy image are included in the fused image. The SNR value of the previous frame is negatively correlated with the fusion coefficient α: that is, the larger the SNR value of the previous frame, the smaller α is, and the smaller the SNR value of the previous frame, the larger α is.
[0069] The feature fusion module is used to fuse feature A from the previous denoised image with feature B from the current noisy image using a fusion coefficient α to produce the denoised image of the current frame. For example, the feature of the denoised image of the current frame is A*α+B*(1-α). Ultimately, the image signal of the current frame is obtained based on the features of the denoised image of the current frame.
[0070] The training process of the denoising neural network is as follows:
[0071] (1) Create an initial denoising neural network.
[0072] (2) The denoised image of the first frame in the sample video and its corresponding SNR value, as well as the noisy image of the second frame are input into the created denoising neural network. After being processed by the denoising neural network, the features of the denoised image of the second frame are output.
[0073] In one exemplary embodiment, the first noisy image is denoised using an existing denoising algorithm to obtain a first denoised image. Furthermore, the SNR value of the first frame is calculated based on the first denoised image and the first noisy image. Then, the first denoised image, the second noisy image, and the SNR of the first frame are input into a denoising neural network. The denoising neural network extracts image features from the first denoised image and the second noisy image, respectively, and determines a fusion coefficient α based on the SNR value of the first frame. The features of the first denoised image and the second noisy image are then fused based on α to obtain features of the second denoised image.
[0074] (3) According to the mapping relationship between the SNR value and the fusion coefficient α, the target fusion coefficient α corresponding to the SNR value of the first frame is obtained. Further, based on the feature extraction module in the denoising neural network, the features of the denoised image of the first frame and the noisy image of the second frame are extracted, and the target fusion coefficient α is used to obtain the features of the target denoised image corresponding to the second frame.
[0075] In addition, the SNR value corresponding to the second frame image is calculated according to the target denoised image corresponding to the second frame and the noisy image of the second frame.
[0076] In an exemplary embodiment, based on the feature A extracted from the denoised image of the first frame and the feature B extracted from the noisy image of the second frame by the feature extraction module in the current denoising neural network, and the target fusion coefficient α, the image features of the target denoised image corresponding to the second frame image can be obtained using A*α+B*(1-α).
[0077] (4) Calculate the error between the features of the denoised image corresponding to the second frame output by the denoising neural network and the features of the target image, and adjust the parameters in the fusion coefficient decision module in the denoising neural network according to the error, that is, update the fusion coefficient decision module in the denoising neural network.
[0078] For each subsequent frame in the sample video, the above steps (3) and (4) are repeated until the error corresponding to the same frame is within the preset range, and a trained denoising neural network is obtained. The sample videos here can be multiple, for example, multiple videos of different scene types.
[0079] Scenario 1: The image sensor supports dynamic switching of the raw image resolution during video capture
[0080] In one scenario, the image sensor supports dynamic switching of the raw image clarity during video capture. In this scenario, the ISP triggers the image sensor to switch the raw image clarity based on the SNR value of the previous frame. Figure 4 and Figure 5 This section describes the video noise reduction process for this scenario. Figure 4 is a schematic diagram of a video noise reduction processing process based on a software architecture of an electronic device provided in an embodiment of the present application, Figure 5 This is a flowchart of a video noise reduction processing method provided in an embodiment of the present application.
[0081] exist Figure 1 The hardware components of the electronic device shown run an operating system, on which applications can be installed. The operating system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application uses the Android system with a layered architecture as an example to illustrate the software structure of the electronic device.
[0082] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided from top to bottom into the application layer (Application), the application framework layer (Framework), the hardware abstraction layer (HAL), and the kernel layer (Kernel).
[0083] The application layer may include a series of application packages. For example, an application package may include applications such as a camera and a gallery, but this application does not limit this.
[0084] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include camera access interfaces such as camera management and camera device. Camera management provides an access interface for managing the camera. Camera device provides an interface for accessing the camera.
[0085] The HAL layer encapsulates the Linux kernel driver, provides an interface to the kernel, and shields the implementation details of the underlying hardware. In the embodiments of the present application, the HAL layer may include a camera HAL and other hardware device abstraction layers. The camera HAL can call algorithms in the camera algorithm library. The camera algorithm library may include algorithms for image processing.
[0086] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera device driver, audio driver, sensor driver, GPU driver, and NPU driver.
[0087] The hardware layer may include image sensors, image signal processors (ISPs), GPUs, and NPUs.
[0088] The following will be combined Figure 4 Introducing the video noise reduction process:
[0089] ① Processing of the first frame of raw image
[0090] After receiving the user's touch operation of the video shooting control, the camera application generates a shooting command and passes it to the image sensor through the camera access interface, the camera hardware abstraction layer (i.e., camera HAL), and the camera device driver. In response to the shooting command, the image sensor obtains the first frame of raw image (i.e., noisy raw0) and passes it to the ISP.
[0091] The ISP denoises the noisy raw0 to obtain denoised raw0, and obtains the SNR value of the first frame of the raw image, namely SNR0, based on the denoised raw0 and the noisy raw0. The ISP passes the processed image of the first frame (such as denoising, parameter optimization, etc.) through the camera device driver, camera HAL and camera access interface in sequence to the camera application so that the camera application can display a preview of the video. At the same time, the ISP also determines the resolution of the next frame of raw based on the SNR0 value, and passes a resolution switching instruction to the image sensor if the resolution of the raw image needs to be switched. In response to the resolution switching instruction, the image sensor switches the clarity of the raw image, that is, changes the resolution of the next frame of raw image.
[0092] In one exemplary embodiment, the ISP maintains a mapping between different SNR values and output modes. As previously mentioned, different output modes produce different raw image resolutions. The ISP queries this mapping to determine the output mode corresponding to the current frame's SNR value, thereby determining the resolution of the next raw image frame.
[0093] The SNR values calculated for raw images of different resolutions are multiples of each other. For example, if the resolution of the hex pattern is four times that of the quadr pattern, the SNR value of the hex pattern is twice that of the quadr pattern. The SNR value of the quadr pattern is also twice that of the binning pattern.
[0094] In an exemplary embodiment, the ISP may maintain a mapping relationship between the SNR values of various output modes and the output modes, for example, a mapping relationship between the SNR value ranges of the Hex-pattern and the output modes, a mapping relationship between the SNR value ranges of the quadr-pattern and the output modes, and a mapping relationship between the SNR value ranges of the binning-pattern and the output modes.
[0095] In another exemplary embodiment, the ISP may maintain a mapping relationship between the SNR value of any output mode and the output mode. For example, the ISP maintains a mapping relationship between each SNR value and the output mode under the binning-pattern. For example, under the binning-pattern, if the SNR value of the current frame raw image is in the range of (0,10], the resolution of the next frame raw image may select a lower resolution output mode, such as binning-pattern. If the SNR value of the current frame raw image is in the range of (10,20], the resolution of the next frame raw image may select a higher resolution output mode, such as quadr-pattern. If the SNR value of the current frame raw image is in the range of (20,30], the resolution of the next frame raw image may select a higher resolution output mode, such as Hex-pattern.
[0096] In one exemplary embodiment, the ISP includes processing logic for switching the raw image resolution (i.e., the output mode of the raw image) based on the SNR value. After the ISP obtains the SNR value of the first raw image frame, it triggers the execution of this processing logic to obtain the resolution corresponding to the SNR value. This resolution is the optimal resolution for the next raw frame. Furthermore, the ISP generates a resolution switching instruction including this resolution and transmits it to the image sensor. In addition, the ISP may also include other processing algorithm logic, such as optimizing parameters such as exposure and color temperature of the shooting scene.
[0097] When the ISP maintains a mapping relationship between the SNR value of each output mode and the output mode, after the ISP calculates the SNR value corresponding to the first frame of the raw image, it can query the mapping relationship between the SNR value corresponding to the output mode of the first frame of the raw image and the output mode to obtain the output mode corresponding to the SNR value of the first frame (that is, the resolution of the second frame of the raw image).
[0098] When the ISP maintains a mapping relationship between the SNR value of an output mode and the output mode, after calculating the SNR value of the first raw image frame, if the output mode of the first raw image frame differs from the output mode of the SNR value in the mapping relationship maintained by the ISP, the SNR value of the first raw image frame is converted to the output mode corresponding to the SNR value in the mapping relationship maintained by the ISP. For example, if the first raw image frame is obtained in quadr-pattern mode and has an SNR value of 20, and the ISP maintains a mapping relationship between SNR values and output modes in binning-pattern mode, the SNR value of the first raw image frame of 20 needs to be converted to binning-pattern mode. If the resolution in quadr-pattern mode is four times that of binning-pattern mode, the SNR value of the first raw image frame in binning-pattern mode is 10.
[0099] In addition, the ISP also passes the denoised Raw0 and SNR0 to the camera algorithm library through the camera device driver and camera HAL, so that the camera algorithm library can perform other processing on the denoised Raw0 to obtain a processed image. The ISP then passes this image to the camera application through the camera HAL and the camera access interface, so that the camera application can display the first frame of the video.
[0100] ② The 2nd to nth frame raw image processing process
[0101] In video capture mode, after obtaining the first raw image frame, the image sensor will continue to generate the next raw image frame (i.e., the second frame with noise, raw1) and transmit it to the ISP.
[0102] In this embodiment of the present application, the image sensor switches the resolution of the second raw image frame in response to the resolution switching instruction transmitted by the ISP, that is, outputs a raw image at the second resolution. Furthermore, the ISP converts or processes the noisy raw1 and transmits it to the camera algorithm library through the camera device driver and the camera HAL.
[0103] In some embodiments, the camera algorithm library runs Figure 3 The denoising neural network model shown. Exemplarily, the camera algorithm library can call the NPU to execute the denoising neural network algorithm.
[0104] The camera algorithm library inputs the received denoised image raw0 of the first frame and the noisy image raw1 of the second frame, as well as the SNR value of the first frame, into the denoising neural network. The denoising neural network outputs the denoised image raw1 of the second frame, or the processed image. The camera algorithm library can also perform other processing on the denoised image raw1. The process of obtaining output based on input of the denoising neural network can be seen in Figure 3The relevant description will not be repeated here.
[0105] Furthermore, the camera algorithm library transmits the second processed image frame to the camera application through the camera HALL and the camera access interface in sequence, so that the camera application displays the second processed image frame in the preview window.
[0106] In addition, the camera algorithm library also passes the obtained denoised raw1 of the second frame to the ISP through the camera HAL and the camera device driver, so that the ISP can calculate the SNR value of the second frame based on the noisy raw image of the second frame and raw1 of the second frame, which can be called SNR1. Furthermore, the ISP selects the resolution of the raw image of the third frame based on SNR1.
[0107] The signal-to-noise ratios of images with different resolutions are exponentially related. The signal-to-noise ratio calculated for each frame needs to be unified to the same resolution, for example, the resolution corresponding to the mapping between the SNR range and resolution stored in the ISP. For example, the ISP stores the SNR and resolution under the quadr-pattern, and the calculated signal-to-noise ratio of the raw image under the binning-pattern is converted to the quadr-pattern.
[0108] In addition, the algorithms in the camera algorithm library can be implemented based on the GPU or NPU at the hardware layer.
[0109] After the camera application receives the stop shooting operation, it sends the stop shooting instruction to the image sensor layer by layer. The image sensor responds to the stop shooting instruction, and then passes the last frame of processed image to the camera application layer by layer, and sends the final video to the gallery.
[0110] The following combination Figure 5 This section introduces the video denoising method process for scenario 1. Figure 5 As shown, the method may include the following steps:
[0111] S101a: The camera application receives a user's operation on a video shooting control, generates a video shooting instruction, and transmits it to the camera HAL.
[0112] S101b, the camera HAL transmits a video shooting instruction to the image sensor.
[0113] Combine Figure 4 It can be seen that the video capture instruction generated by the camera application is passed to the camera HAL via the camera access interface. The camera HAL passes the video capture instruction to the image sensor via the camera device driver.
[0114] S102 : The image sensor captures a first frame of image at a first resolution in response to a video capture instruction.
[0115] The first resolution may be any one of three resolutions supported by the image sensor.
[0116] S103 : The image sensor sends a noisy image raw0 of a first resolution to the ISP.
[0117] The first frame image of the first resolution obtained by the image sensor is transmitted to the ISP for corresponding processing.
[0118] S104 , the ISP performs denoising processing on the noisy image raw0 to obtain a denoised image raw0 corresponding to the first frame.
[0119] The ISP can use an existing denoising algorithm to denoise the noisy raw0 to obtain the first frame of denoised image raw0. In addition, the ISP can also perform other processing on the noisy raw0, which will not be detailed here.
[0120] S105 , the ISP obtains the SNR value of the first frame according to the denoised image and the noisy image of the first frame.
[0121] For example, the SNR value corresponding to the first frame of the raw image can be calculated based on the denoised image and the noisy image corresponding to the first frame according to the aforementioned SNR calculation formula, which can be recorded as SNR0.
[0122] S106a: The ISP transmits the denoised image and SNR value of the first frame to the camera algorithm library layer by layer.
[0123] Combine Figure 4 ,ISP passes the denoised raw0 and SNR of the first frame to the camera algorithm library through the camera device driver and camera HAL for image post-processing.
[0124] S106b: The camera algorithm library transmits the processed image of the first frame to the camera HAL.
[0125] S106c: The camera HAL transmits the first processed image frame to the camera application.
[0126] Combine Figure 4 The camera HAL transmits the first processed image frame to the camera application via the camera access interface, so that the camera application displays the first processed image frame in the preview window.
[0127] S107: The ISP determines the resolution of the next raw image frame based on the SNR0 of the first frame.
[0128] S108 : When the resolution of the next raw image frame is different from the resolution of the first frame, a resolution switching instruction is generated and sent to the image sensor.
[0129] S109 : The image sensor switches to shooting at the second resolution in response to the resolution switching instruction, and transmits the noisy image raw1 of the second resolution to the ISP.
[0130] S110 , the ISP transmits the noisy image raw1 of the second resolution to the camera algorithm library layer by layer.
[0131] After the ISP processes the received second frame of raw image (for example, optimizing the exposure and color temperature of the captured scene), it passes it to the camera algorithm library through the camera device driver and camera HAL.
[0132] S111a: The camera algorithm library fuses the features of the denoised image of the first frame and the noisy image of the second frame according to the SNR of the first frame to obtain the processed image of the second frame.
[0133] In an exemplary embodiment of the present application, the camera algorithm library includes a denoising neural network model. The denoised image of the first frame, the noisy image of the second frame, and the SNR value of the first frame are input to the denoising neural network, which outputs the denoised image of the second frame. The processing process of the denoising neural network can be seen in Figure 3 The corresponding description will not be repeated here.
[0134] In addition, the camera algorithm library may also include other post-processing algorithms, and the denoised image obtained may be further processed to obtain a processed image.
[0135] S111b, the camera algorithm library transmits the processed image of the second frame to the ISP, so that the ISP obtains the SNR value of the second frame.
[0136] After the camera algorithm library obtains the second frame of processed image after denoising, it passes the second frame of denoised image to the ISP through the camera HAL and the camera device driver in sequence, so that the ISP calculates the SNR value of the second frame of image based on the second frame of processed image and the second frame of noisy image, and further determines the resolution of the third frame of raw image based on the SNR value of the second frame.
[0137] S112a: The camera algorithm library transmits the processed image of the second frame to the camera HAL.
[0138] S112b: The camera HAL passes the processed image of the second frame layer by layer to the camera application for display.
[0139] Combine Figure 4 The camera HAL transmits the second processed image frame to the camera application through the camera access interface, so that the camera application displays the second processed image frame in the preview window.
[0140] The image processing process of the 3rd to nth frames is the same as that of the 2nd frame, and will not be repeated here.
[0141] The video denoising method provided in this embodiment selects the resolution of the current frame's raw image based on the SNR value of the previous frame. The SNR value is positively correlated with the clarity. This allows the image sensor to dynamically select the most appropriate raw image as input to the denoising neural network based on the signal-to-noise ratio (SNR) of the same video recording. For example, when the SNR is high, the denoising process is simpler, and a high-resolution raw image can be selected, thereby improving video quality. When the SNR is low, the denoising process is more difficult, and a low-resolution raw image with less noise can be selected, reducing the difficulty and improving the denoising effect and efficiency. Furthermore, a fusion coefficient α is determined based on the SNR value of the previous frame. The features of the noise-free image of the previous frame are fused with the features of the noisy image of the current frame based on this fusion coefficient to obtain the denoised image corresponding to the current frame. Using adjacent raw images of different resolutions as input to the denoising process may result in small jumps in the clarity of adjacent raw images. By fusing the previous and current frames, the difference in clarity between adjacent raw images is reduced, thereby improving the overall video quality. The fusion ratio in this application is determined by the SNR value. For frames with low SNR values, more denoised images of the previous frame are fused, and for frames with high SNR values, more noisy raw images of the current frame are fused, thereby improving video quality.
[0142] Scenario 2: The image sensor does not support dynamic switching of raw image resolution during video capture.
[0143] If the image sensor does not support dynamic switching of the raw image resolution during video capture, the camera algorithm library can obtain a low-resolution raw image based on the high-resolution raw image output by the image sensor.
[0144] The following will be combined Figure 6 and Figure 7 This section describes the video noise reduction process in this scenario. Figure 6 Schematic diagram of a video noise reduction process based on a software architecture of an electronic device provided in an embodiment of the present application; Figure 7 This is a flowchart of a video noise reduction processing method provided in an embodiment of the present application.
[0145] ① Processing of the first frame of raw image
[0146] like Figure 6 As shown in the figure, the image sensor responds to the capture instructions passed down layer by layer by the camera application, obtains the first raw image frame (i.e., noisy raw0), and passes it to the ISP for processing (such as denoising) to obtain denoised raw0. The ISP passes the noisy raw0 and denoised raw0 to the camera algorithm library via the camera device driver and camera HAL, respectively. The denoised raw0 here can be the image obtained after ISP denoising, or the image obtained after denoising and other post-processing.
[0147] The camera algorithm library calculates the SNR value of the first frame, that is, SNR0, based on the denoised raw0 and noisy raw0 of the first frame. Furthermore, the camera algorithm library determines the output mode of the raw image of the second frame (that is, the resolution of the raw image) based on the SNR0 value.
[0148] In an exemplary embodiment, an algorithm for obtaining SNR in a camera algorithm library may be used to calculate the SNR value of the first frame, ie, SNR0, based on the noisy raw0 and the denoised raw0.
[0149] In addition, the camera algorithm library can process the denoised raw0 to obtain the first frame of processed image, and pass the first frame of processed image to the camera application through the camera HAL and the camera access interface in sequence, so that the camera application can display the first frame of processed image in the preview window.
[0150] ② Processing of the second frame of raw image
[0151] In this scenario, the image sensor always captures video at the first resolution (i.e., the resolution corresponding to the first output mode). After the image sensor outputs the second frame of raw image (i.e., the first resolution noisy raw1), it is passed to the camera algorithm library via the camera device driver and the camera HAL.
[0152] The camera algorithm library determines the second resolution of the raw image of the second frame (i.e., the resolution corresponding to the second output mode) based on the SNR0 of the first frame. If the second resolution is lower than the first resolution of the raw image of the first frame, the camera algorithm library uses a corresponding algorithm to convert the noisy raw image of the first resolution output by the image sensor to a noisy raw image of the second resolution. For example, a downsampling algorithm is used to downsample the raw image of the first resolution to obtain the raw image of the second resolution.
[0153] Furthermore, the camera algorithm library performs denoising on the noisy raw1 at the second resolution to obtain a denoised image frame 2, or a processed image frame 2. Furthermore, the camera algorithm library can calculate an SNR value for frame 2 based on the processed image frame 2 and the noisy raw image, so that the camera algorithm library can determine an output mode (or resolution) for frame 3 based on the SNR value.
[0154] At the same time, the camera algorithm library passes the processed image of the second frame to the camera application through the camera HAL and the camera access interface in sequence, so that the camera application displays the image in the preview window.
[0155] In this scenario, the camera algorithm library can store the mapping between different SNR ranges and image output modes (or resolutions). The camera algorithm library determines the output mode (or resolution) for the next frame based on the SNR value of the current frame in the same way as the ISP determination in scenario 1, and will not be repeated here. The processing of the remaining captured frames is the same as the processing of frame 2 described above and will not be repeated here. After stopping shooting, the final video is sent to the image library.
[0156] The following will be combined Figure 7 Introduce the video denoising method process for scenario 2, such as Figure 7 As shown, the method may include the following steps:
[0157] S201a: The camera application receives a user's operation on a video shooting control, generates a video shooting instruction, and transmits it to the camera HAL.
[0158] S201b, the camera HAL transmits a video shooting instruction to the image sensor.
[0159] S202 : The image sensor shoots at a first resolution in response to a video shooting instruction.
[0160] S203 : The image sensor sends the first frame of noisy image raw0 to the ISP.
[0161] S204a: The ISP performs denoising on the first noisy image frame to obtain a first denoised image frame.
[0162] The first frame of the raw image can be denoised using an existing denoising algorithm to obtain a denoised image.
[0163] S204b: The ISP transmits the first noisy image frame and the first denoised image frame to the camera HAL.
[0164] S204c: The camera HAL passes the first noisy image frame and the first denoised image frame to the camera algorithm library.
[0165] S205 , the camera algorithm library obtains the SNR value of the first frame according to the denoised image of the first frame and the noisy image of the first frame.
[0166] In addition, other processing can be performed on the denoised image of the first frame based on other post-processing algorithms in the camera algorithm library.
[0167] S206a: The camera algorithm library transmits the first denoised image frame to the camera HAL.
[0168] S206b: The camera HAL transmits the first denoised image frame to the camera application.
[0169] See also Figure 6The camera HAL transmits the first denoised image frame to the camera application through the camera access interface, so that the camera application displays the image in the preview window.
[0170] S207: The image sensor sends the second noisy image frame to the ISP.
[0171] After receiving the video capture command, the image sensor continues capturing video at the first resolution at the specified frame rate. The second raw image frame is still an image at the first resolution. The image sensor obtains the second noisy image frame and sends it to the ISP for processing, such as converting it into a digital image signal.
[0172] S208a: The ISP transmits the second noisy image frame to the camera HAL.
[0173] S208b: The camera HAL transmits the second noisy image frame to the camera algorithm library.
[0174] The ISP passes the processed second noisy image frame to the camera algorithm library through the camera device driver and camera HAL.
[0175] S209: When the camera algorithm library determines, based on the SNR value of the first frame, that the second resolution corresponding to the second frame is lower than the first resolution of the first frame, a second noisy image of the second frame with a second resolution is obtained based on the noisy image of the second frame with the first resolution.
[0176] The process of determining the resolution of the second frame according to the SNR value of the first frame in this step is similar to Figure 5 The camera algorithm library can store a mapping relationship between SNR values and resolutions, and query the resolution corresponding to the SNR value of the first frame based on the mapping relationship, that is, the second resolution.
[0177] In the case that the second resolution is lower than the first resolution, a downsampling algorithm may be used to downsample the second noisy image frame of the first resolution to obtain the second noisy image frame of the second resolution.
[0178] S210: The camera algorithm library fuses the denoised image of the first frame and the noisy image of the second frame with the second resolution according to the SNR value of the first frame to obtain a processed image of the second frame.
[0179] The specific implementation process of this step is Figure 5 The implementation process of S111a is the same as that of S111a in FIG.
[0180] S211a, the camera algorithm library transmits the processed image of the second frame to the camera HAL.
[0181] S211b: The camera HAL transmits the second processed image frame to the camera application.
[0182] The camera algorithm library passes the processed second frame image to the camera application through the camera HAL and the camera access interface in sequence, so that the camera application displays the image in the preview window.
[0183] The processing of the other frames obtained by shooting is the same as the processing of the second frame mentioned above, which will not be repeated here. After stopping shooting, the final video is sent to the gallery.
[0184] The video denoising method provided in this embodiment selects the resolution of the current frame's raw image based on the SNR value of the previous frame's image, and the SNR value is positively correlated with the resolution. This allows the camera algorithm library to dynamically obtain the most suitable raw image as input to the denoising neural network based on the different signal-to-noise ratios during the recording of the same video segment, thereby improving the noise reduction processing effect and efficiency. In addition, the fusion coefficient α is determined based on the SNR value of the previous frame, and the features of the noise-free image of the previous frame and the noisy image of the current frame are fused according to the fusion coefficient to obtain the denoised image corresponding to the current frame. Using raw images of different resolutions as input to the denoising process for adjacent raw images may cause small jumps in the clarity of adjacent raw images. By fusing the previous frame and the current frame, the clarity difference between adjacent raw images is reduced, thereby improving the overall video quality. The fusion ratio in this application is determined by the SNR value. When the SNR value of the previous frame is low, more denoised images of the previous frame are fused, and when the SNR value of the previous frame is high, more noisy raw images of the current frame are fused, thereby improving video quality. Moreover, this embodiment is applicable to scenarios where the image sensor does not support dynamic resolution switching, which reduces the dependence of the method on the hardware functions of the image sensor and increases the scope of application of the method.
[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment. The aforementioned storage medium includes: various media that can store program code, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0186] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A video denoising method, characterized in that: Applied to electronic equipment, the method includes: Starting to capture a video in response to a video capturing operation; Performing denoising on a first frame of original image obtained by shooting to obtain a first frame of denoised image, where the first frame of original image is an image of a first resolution; Obtaining a first signal-to-noise ratio corresponding to the first frame of original image based on the first frame of original image and the first frame of denoised image; determining a second resolution corresponding to a second frame of the original image based on the first signal-to-noise ratio, and obtaining a second frame of the original image at the second resolution, where the second resolution is different from the first resolution and is positively correlated with the first signal-to-noise ratio; Obtaining a fusion coefficient corresponding to the second frame of image based on shooting information, and fusing the first denoised image frame with the second frame of original image at the second resolution according to the fusion coefficient to obtain the second denoised image frame, wherein the shooting information includes prior knowledge or shooting parameters of the previous frame of image; the fusion coefficient is a ratio of features of the first denoised image frame to features of the second denoised image frame; A second signal-to-noise ratio corresponding to the second frame of original image is obtained based on the second denoised image and the second frame of original image. The second signal-to-noise ratio is used to determine the resolution of the third frame of original image. The third frame of original image is processed according to the processing process of the second frame of image until a stop operation is received to stop video capture, thereby obtaining a video after noise reduction processing.
2. The method according to claim 1, characterized in that The shooting information is a signal-to-noise ratio; and obtaining a fusion coefficient corresponding to the second frame image according to the shooting information includes: A fusion coefficient of the second frame image is obtained according to a first signal-to-noise ratio corresponding to the first frame original image, and the fusion coefficient is positively correlated with the first signal-to-noise ratio.
3. The method according to claim 2, characterized in that The obtaining of a fusion coefficient corresponding to the second frame image according to the shooting information, and fusing the first frame denoised image with the second frame original image of the second resolution according to the fusion coefficient to obtain the second frame denoised image, includes: Inputting the first denoised image, the second original image at the second resolution, and the first signal-to-noise ratio of the first image into a denoising neural network for denoising processing; The denoising neural network determines the fusion coefficient based on a first signal-to-noise ratio of the input first frame image; The denoising neural network extracts a first image feature from the first denoised image frame, and extracts a second image feature from the second original image frame of the second resolution; The denoising neural network fuses the first image feature and the second image feature according to the fusion coefficient to obtain the second frame denoised image.
4. The method according to claim 3, characterized in that The training process of the denoising neural network includes: Create an initial denoising neural network; Inputting the first frame denoised image and the signal-to-noise ratio of the sample video and the second frame original image into the initial denoising neural network for denoising to obtain the second frame sample denoised image; Determining a target fusion coefficient that matches the first signal-to-noise ratio based on a mapping relationship between the signal-to-noise ratio and the fusion coefficient; fusing the first sample image features and the second sample image features based on the target fusion coefficient to obtain a second frame of target denoised image, wherein the first sample image features are extracted from the first frame of denoised image by the initial denoising neural network, and the second sample image features are extracted from the second frame of original image by the initial denoising neural network; Obtaining an error between the second frame target denoised image and the second frame sample denoised image, and adjusting parameters of the denoising neural network according to the error; For any frame after the second frame image in the sample video, the processing process of the second frame image is repeatedly executed until the error between the sample denoised image and the target denoised image corresponding to the same frame is within a preset range, thereby terminating the training process of the denoising neural network.
5. The method according to any one of claims 1 to 4, characterized in that The electronic device includes an image sensor and an image signal processor, wherein the image sensor supports dynamic switching of shooting resolution during video shooting; Determining a second resolution corresponding to a second frame of original image based on the first signal-to-noise ratio, and obtaining a second frame of original image with the second resolution, includes: The image signal processor queries a mapping relationship between a signal-to-noise ratio range and a resolution to obtain a second resolution corresponding to the first signal-to-noise ratio; When the image signal processor determines that the second resolution is different from the first resolution, sending a resolution switching instruction to the image sensor; The image sensor switches the resolution of the captured image to the second resolution in response to the resolution switching instruction, and sends a second frame of original image with the second resolution to the image signal processor.
6. The method according to claim 5, characterized in that The obtaining, according to the second denoised image and the second original image, a second signal-to-noise ratio corresponding to the second original image, includes: The image signal processor calculates a second signal-to-noise ratio corresponding to the second frame original image based on the second denoised image and the second frame original image, and determines a third resolution of the third frame original image based on the second signal-to-noise ratio.
7. The method according to claim 5, characterized in that The image signal processor maintains a mapping relationship between a signal-to-noise ratio range and a resolution in each different image output mode; the image signal processor queries the mapping relationship between the signal-to-noise ratio range and the resolution to obtain a second resolution corresponding to the first signal-to-noise ratio, including: The image signal processor queries a mapping relationship between a signal-to-noise ratio range and a resolution corresponding to an output mode of the first frame of original image, and obtains a second resolution corresponding to the first signal-to-noise ratio.
8. The method according to claim 5, characterized in that The image signal processor maintains a first mapping relationship between a signal-to-noise ratio and a resolution corresponding to a first image output mode, wherein the resolution of the original image output in the first image output mode is a third resolution; The image signal processor queries a mapping relationship between a signal-to-noise ratio range and a resolution to obtain a second resolution corresponding to the first signal-to-noise ratio, including: The image signal processor converts the first signal-to-noise ratio corresponding to the first frame of the original image with the first resolution into a second signal-to-noise ratio corresponding to the third resolution; The image signal processor queries a mapping relationship between a signal-to-noise ratio range corresponding to the third resolution and resolution to obtain a second resolution corresponding to the second signal-to-noise ratio.
9. The method according to any one of claims 1 to 4, characterized in that The electronic device includes an image sensor, an image signal processor and a camera algorithm library; Obtaining a first signal-to-noise ratio corresponding to the first frame of original image based on the first frame of original image and the first frame of denoised image includes: The image signal processor performs denoising on the first frame of the original image output by the image sensor to obtain the first frame of denoised image; The image signal processor transmits the first frame of denoised image and the first frame of original image to the camera algorithm library; The camera algorithm library calculates a first signal-to-noise ratio corresponding to the first frame of original image based on the first frame of denoised image and the first frame of original image.
10. The method according to claim 9, characterized in that The image sensor does not support dynamic switching of shooting resolution during video shooting; The determining a second resolution corresponding to the second frame of the original image based on the first signal-to-noise ratio, and obtaining the second frame of the original image with the second resolution, includes: The camera algorithm library queries the mapping relationship between the signal-to-noise ratio and the resolution to obtain a second resolution corresponding to the first signal-to-noise ratio, where the second resolution is lower than the first resolution; The image sensor captures a second frame of original image with the first resolution and transmits the image to the camera algorithm library; The camera algorithm library downsamples the second frame original image of the first resolution to obtain a second frame original image of the second resolution.
11. The method according to claim 9 or 10, characterized in that Obtaining a second signal-to-noise ratio corresponding to the second frame original image according to the second frame denoised image and the second frame original image includes: The camera algorithm library calculates a second signal-to-noise ratio corresponding to the second frame original image according to the second frame original image of the second resolution and the second frame denoised image of the second resolution.
12. An electronic device, characterized in that: The electronic device includes: one or more processors, a memory and a touch screen; the memory is used to store program code; the processor is used to run the program code, so that the electronic device implements the video denoising method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that Instructions are stored thereon, and when the instructions are executed on an electronic device, the electronic device executes the video denoising method according to any one of claims 1 to 11.
14. A computer program product, characterized in that Instructions are stored thereon, and when the computer program product is run on an electronic device, the electronic device implements the video denoising method according to any one of claims 1 to 11.
15. A chip system, characterized in that: include: at least one processor and an interface, wherein the interface is configured to receive code instructions and transmit the code instructions to the at least one processor; The at least one processor runs the code instructions to implement the video denoising method according to any one of claims 1-1.
Citation Information
Patent Citations
Coding method of fine and classified video of space domain classified noise / signal ratio
CN101018333A
Video noise reduction method and device, and computer readable storage medium
CN111583151A
Video image noise reduction method and device and computer storage medium
CN115965537A
Video denoising method and device based on optical flow motion detection, and computer storage medium
CN115967777A
Video coding and decoding reference image resampling method combining loop filtering and super-resolution
CN118714355A