A video denoising method and apparatus

By fusing the information of the current frame and adjacent frames in the video data, combining time domain correlation and transformation technology, the existing video denoising methods have been solved, and the efficient and lightweight denoising effect is achieved, and the video image quality is improved.

CN115063301BActive Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110221439.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-27
Publication Date
2025-07-18
Estimated Expiration
2041-02-27

AI Technical Summary

Technical Problem

The existing video denoising technology is difficult to effectively improve the denoising quality under the cyclic recursion method, and the conventional methods are computationally expensive and costly, requiring a lot of manual debugging.

Method used

By obtaining the information of the current frame and adjacent frames in the video data, the characteristics are extracted and the fusion weight is determined, the current frame and adjacent frame images are fused, and the noise effect is reduced through color and frequency transformation.

Benefits of technology

It realizes the improvement of denoising effect under lightweight conditions, reduces noise, reduces calculation complexity, enhances video image quality, reduces ghosting and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063301B_ABST
    Figure CN115063301B_ABST
Patent Text Reader

Abstract

The present application provides a video denoising method and apparatus, which are used to improve the denoising effect while achieving lightweight denoising, obtain a clearer image, and further obtain video data with better image quality. The method includes: acquiring a current frame and a first fused image in the video data, where the first fused image includes information of at least one frame adjacent to the current frame in the video data in a preset order; extracting features from the current frame and the first fused image to obtain a first feature and a second feature, and then determining a first fusion weight and a second fusion weight, where the weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image; fusing the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image; and denoising the second fused image to obtain a denoised image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular, to a video denoising method and apparatus. Background Art

[0002] Video denoising in computational photography is crucial for video imaging quality. The imaging of a terminal device is limited by the hardware performance of the optical sensor of the terminal device. Due to the imperfection of the acquisition process, the formation of digital images is always affected by different forms of noise and degradation. Image restoration algorithms such as denoising, demosaicing, super-resolution, etc. must be used to restore the degraded input to a high-quality input.

[0003] For video denoising, the spatio-temporal fusion of video information can usually be processed in a cyclic recursive manner, but the denoising quality is low. Therefore, how to improve the video denoising quality has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a video denoising method, a video processing method and apparatus, which are used to improve the denoising effect while achieving lightweight denoising, obtain a clearer image, and further obtain video data with better image quality.

[0005] In view of this, in a first aspect, this application provides a video denoising method, including: First, obtain a current frame and a first fused image. The current frame is any frame image after the first frame arranged in a preset order in the video data, that is, the current frame is a non-first frame in the video data. The first fused image includes information of at least one frame adjacent to the current frame in the video data according to a preset order; extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature; Subsequently, determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The first fusion weight and the second fusion weight can be set according to foreground and background. The weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image; fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image; denoise the second fused image to obtain a denoised image.

[0006] Therefore, in the embodiments of the present application, by fusing the current frame and the first fused image, it is equivalent to combining the temporally correlated information between adjacent frames in the scene, smoothing the noise in the current frame, reducing the noise in the image, and obtaining a second fused image with less noise. It is equivalent to using the correlation between images in the video in the time domain to smooth the noise in each frame of the image, improving the denoising effect of each frame in the video data, and also avoiding ghosting, thereby obtaining video data with better image quality and improving the user experience.

[0007] In a possible implementation manner, after denoising the second fused image, the method further includes: fusing the second fused image and the denoised image to obtain an updated denoised image.

[0008] When denoising an image, an over-smoothing phenomenon may occur. At this time, compared with the denoised image, the second fused image has more details. Therefore, by fusing the details included in the second fused image and the denoised image, an image with richer details is obtained, improving the image quality.

[0009] In a possible implementation manner, fusing the second fused image and the denoised image includes: extracting features from the second fused image to obtain a third feature; extracting features from the denoised image to obtain a fourth feature; determining a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature, where the third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value are negatively correlated; fusing the third feature and the fourth feature according to the third fusion weight and the fourth fusion weight to obtain an updated denoised image.

[0010] Generally, there is more noise in the high frequency than in the low frequency. In this embodiment, the frequency of the pixel point and the corresponding weight value are negatively correlated, which means that the higher the frequency value of the pixel point, the lower its corresponding weight, thereby reducing the noise carried in the high frequency component and obtaining a denoised image.

[0011] In a possible implementation manner, before extracting features from the current frame and extracting features from the first fused image, the above method may further include: performing color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component, and the first chrominance component and the first luminance component form a new current frame, and performing color transformation on the first fused image to obtain a second chrominance component and a second luminance component, and the second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel.

[0012] The above-mentioned extracting features from the current frame to obtain the first feature and extracting features from the first fused image to obtain the second feature may include: extracting features from a new current frame to obtain the first feature, and extracting features from the new first fused image to obtain the second feature.

[0013] Therefore, in the embodiments of the present application, color removal processing is also performed on the current frame, which is equivalent to a color removal operation, reducing the correlation between color channels, reducing the subsequent denoising complexity, improving the denoising efficiency and effect, and obtaining an image with better image quality.

[0014] In a possible implementation manner, after denoising the second fused image, the above method may further include: performing color transformation on the denoised image through an inverse color transformation matrix to obtain an updated denoised image, and the inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0015] In the embodiments of the present application, if an inverse color transformation is performed after a color removal-related transformation before denoising, the color in the image can be restored to obtain an image with color.

[0016] In a possible implementation manner, before extracting features from the current frame and extracting features from the first fused image, the method further includes: performing wavelet transform on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component, the first low-frequency component and the first high-frequency component form a new current frame, and performing wavelet transform on the first fused image to obtain a second low-frequency component and a second high-frequency component, the second low-frequency component and the second high-frequency component form a new first fused image, and the wavelet coefficients are preset coefficients or obtained by training at least one convolution kernel.

[0017] Wavelet transform can be understood as performing a decorrelation transform on the pixel points of an image in the frequency dimension. Generally, based on the frequency representation, the effective part and noise of the image are separated into different frequency components, making denoising simpler. It is equivalent to discretely distributing the frequencies of the pixel points to facilitate obtaining a better subsequent denoising effect.

[0018] In a possible implementation manner, after denoising the second fused image, the above method may further include: performing inverse wavelet transform on the denoised image through inverse wavelet coefficients to obtain an updated denoised image, and the inverse wavelet coefficients are the inverse matrix of the wavelet coefficients. Thus, the inverse wavelet transform can accurately restore the high-frequency and low-frequency components to obtain a clearer denoised image.

[0019] In a possible implementation, determining the first fusion weight corresponding to the current frame according to the first feature and the second feature may include: calculating the shot noise and the read noise according to the shooting parameters of the device used to shoot the video data; determining the first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature.

[0020] That is, when calculating the first fusion weight, the noise level is also combined, so as to adapt to scenarios with different noise levels, and accurate denoising can be performed under different noise levels, and the generalization ability is strong.

[0021] In a possible implementation, determining the first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature may specifically include: calculating the noise variance of each pixel point in the current frame according to the shot noise and the read noise; extracting the fifth feature from the noise variance; determining the first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature. This noise variance can be used to accurately determine the noise level of the current frame, so as to facilitate accurate denoising in the subsequent process and improve the denoising effect.

[0022] In a possible implementation, denoising the second fusion image may include: denoising the second fusion image by combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame to obtain a denoised image.

[0023] Therefore, in the implementation manner of the present application, the second fusion image can be denoised by combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame, which is equivalent to providing more reference information for denoising, improving the denoising effect, and the calculated data can be reused, improving the effective utilization rate of the data generated in each step.

[0024] In a possible implementation, denoising the second fusion image by combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame may specifically include: combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame to calculate the variance of each pixel point in the second fusion image to obtain a fusion map variance; then using the fusion map variance and the second fusion image as inputs of a denoising model to output a denoised image.

[0025] During the denoising process of the video data, the fusion map variance decreases as the number of denoised frames increases, that is, it means that the noise in the fusion image is getting smaller and smaller, and finally the noise in the obtained denoised image is also getting less and less, so as to obtain an image with less noise and better image quality.

[0026] In a possible implementation, the above-mentioned taking the fused graph variance and the second fused image as the input of the denoising model may include: taking the current frame, the fused graph variance and the second fused image as the input of the denoising model, and outputting a denoised image, where the denoising model is used to remove the noise in the input image.

[0027] In this embodiment, the fused graph variance can be taken as the input of the denoising model, so that the denoising model can determine the noise level through the fused graph variance, achieve a better denoising effect, and obtain a denoised image with better image quality.

[0028] In a possible implementation, the above-mentioned method may further include: performing at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image; performing denoising processing on at least one downsampled frame and at least one downsampled fused image to obtain a multi-scale fused image; fusing the denoised image and the multi-scale fused image to obtain an updated denoised image.

[0029] In the implementation manner of the present application, the current frame and the first fused image can be downsampled once or multiple times, and the downsampled frames and downsampled fused images at different scales can be iteratively processed, so as to perform denoising processing at different scales and make the denoising effect of the finally obtained denoised image better.

[0030] In a possible implementation manner, any denoising process during the denoising process of at least one downsampled frame and at least one downsampled fusion image, that is, the denoising process of one scale, may include: determining the weight corresponding to the first downsampled frame according to the features extracted from the first downsampled frame and the features extracted from the first downsampled fusion image to obtain a fifth fusion weight, and determining the weight corresponding to the first downsampled fusion image to obtain a sixth fusion weight, where the first downsampled frame is any one of the at least one downsampled frame, and the first downsampled fusion image is a frame with the same scale as the first downsampled frame among the at least one downsampled fusion image; determining the weight corresponding to the second downsampled fusion image according to the features extracted from the second downsampled fusion image to obtain a seventh fusion weight, where the second downsampled fusion image incorporates the information of the images with scales smaller than the first downsampled frame among the at least one downsampled frame; fusing the first downsampled frame, the first downsampled fusion image, and the second fusion image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain a third downsampled fusion image, and the upsampled image of the third downsampled image is used for fusing with the images with scales larger than the first downsampled frame among the at least one downsampled frame; denoising the third downsampled fusion image to obtain a first downsampled denoised image; upsampling the first downsampled denoised image to obtain an upsampled denoised image, and the upsampled image is used for combining with the fusion image with the same scale as the upsampled image to perform denoising to obtain an image with a scale larger than the first downsampled frame.

[0031] Therefore, in the implementation manner of the present application, during the denoising process of each scale, the processing results of other scales can be fused for iterative denoising, which can improve the denoising effect and obtain an image with less noise.

[0032] In a possible implementation manner, the above-mentioned at least one downsampling of the current frame and the first fusion image may include: performing at least one wavelet transform on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image.

[0033] Therefore, in the implementation manner of the present application, downsampling can be performed by means of wavelet transform, and at the same time, the frequencies of pixel points can be discretely distributed in the frequency space to reduce the difficulty of denoising and improve the denoising effect.

[0034] In a second aspect, the present application provides a video denoising device, which may include:

[0035] An acquisition module, configured to acquire a current frame and a first fusion image, where the current frame is any frame image after the first frame arranged in a preset order in the video data, and the first fusion image includes the information of at least one frame adjacent to the current frame in the video data according to the preset order;

[0036] A temporal fusion module, configured to extract features from a current frame to obtain a first feature, and extract features from a first fused image to obtain a second feature;

[0037] The temporal fusion module is further configured to determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image;

[0038] The temporal fusion module is further configured to fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image;

[0039] A denoising module, configured to denoise the second fused image to obtain a denoised image.

[0040] In a possible implementation manner, the video denoising device may further include:

[0041] A refinement module, configured to fuse the second fused image and the denoised image after denoising the second fused image to obtain an updated denoised image.

[0042] In a possible implementation manner, the denoising module is specifically configured to: extract features from the second fused image to obtain a third feature; extract features from the denoised image to obtain a fourth feature; determine a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature. The third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value have a negative correlation; fuse the third feature and the fourth feature according to the third fusion weight and the fourth fusion weight to obtain an updated denoised image.

[0043] In a possible implementation manner, the video denoising device may further include: a color - related transformation module, configured to, before extracting features from the current frame and extracting features from the first fused image, perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component, and the first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fused image to obtain a second chrominance component and a second luminance component, and the second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel;

[0044] The temporal fusion module is specifically configured to extract features from the new current frame to obtain a first feature, and extract features from the new first fused image to obtain a second feature.

[0045] In a possible implementation, the video denoising device may further include:

[0046] An inverse color correlation transformation module, configured to, after denoising the second fused image, perform color transformation on the denoised image through an inverse color transformation matrix to obtain an updated denoised image, where the inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0047] In a possible implementation, the video denoising device may further include:

[0048] A de-frequency correlation transformation module, configured to, before the time-domain fusion module extracts features from the current frame and extracts features from the first fused image, perform wavelet transform on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component, where the first low-frequency component and the first high-frequency component form a new current frame, and perform wavelet transform on the first fused image to obtain a second low-frequency component and a second high-frequency component, where the second low-frequency component and the second high-frequency component form a new first fused image, and the wavelet coefficients are preset coefficients or obtained by training at least one convolution kernel.

[0049] In a possible implementation, the video denoising device may further include:

[0050] An inverse frequency correlation transformation module, configured to, after the denoising module denoises the second fused image, perform inverse wavelet transform on the denoised image through inverse wavelet coefficients to obtain an updated denoised image, where the inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

[0051] In a possible implementation, the time-domain fusion module is specifically configured to: calculate shot noise and read noise according to the shooting parameters of the device used to shoot the video data; determine the first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature.

[0052] In a possible implementation, the time-domain fusion module is specifically configured to: calculate the noise variance of each pixel point in the current frame according to the shot noise and the read noise; extract a fifth feature from the noise variance; determine the first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature.

[0053] In a possible implementation, the denoising module is specifically configured to: denoise the second fused image by combining the first fusion weight, the second fusion weight, the first fused image, and the current frame to obtain the denoised image.

[0054] In a possible implementation, the denoising module is specifically configured to: calculate the variance of each pixel in the second fused image by combining the first fusion weight, the second fusion weight, the first fused image, and the current frame, to obtain the fused image variance; use the fused image variance and the second fused image as the inputs of the denoising model, and output the denoised image.

[0055] In a possible implementation, the denoising module is specifically configured to use the current frame, the fused image variance, and the second fused image as the inputs of the denoising model, and output the denoised image, where the denoising model is used to remove the noise in the input image.

[0056] In a possible implementation, the video denoising device may further include: a downsampling module;

[0057] The downsampling module is configured to perform at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image;

[0058] The denoising module is further configured to perform denoising processing on at least one downsampled frame and at least one downsampled fused image to obtain a multi-scale fused image;

[0059] The denoising module is further configured to fuse the denoised image and the multi-scale fused image to obtain an updated denoised image.

[0060] In a possible implementation, any denoising process during the process of the denoising module performing denoising processing on at least one downsampled frame and at least one downsampled fused image may include:

[0061] Determine the weight corresponding to the first downsampled frame based on the features extracted from the first downsampled frame and the features extracted from the first downsampled fusion image, to obtain the fifth fusion weight, and determine the weight corresponding to the first downsampled fusion image, to obtain the sixth fusion weight. The first downsampled frame is any one of at least one downsampled frame, and the first downsampled fusion image is a frame with the same scale as the first downsampled frame among at least one downsampled fusion image; determine the weight corresponding to the second downsampled fusion image based on the features extracted from the second downsampled fusion image, to obtain the seventh fusion weight. The second downsampled fusion image incorporates the information of images with scales smaller than the first downsampled frame among at least one downsampled frame; fuse the first downsampled frame, the first downsampled fusion image, and the second fusion image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight, to obtain the third downsampled fusion image. The upsampled image of the third downsampled image is used to fuse with images with scales larger than the first downsampled frame among at least one downsampled frame; denoise the third downsampled fusion image to obtain the first downsampled denoised image; upsample the first downsampled denoised image to obtain the upsampled denoised image. The upsampled image is used to combine with the fusion image with the same scale as the upsampled image for denoising to obtain an image with a scale larger than the first downsampled frame.

[0062] In a possible implementation manner, the downsampling module is specifically configured to perform at least one wavelet transform on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image.

[0063] In a third aspect, an embodiment of the present application provides a video denoising device, including: a processor and a memory. The processor and the memory are interconnected by a line. The processor calls the program code in the memory to execute the functions related to processing in the video denoising method shown in any item of the first aspect above. Optionally, the video denoising device may be a chip.

[0064] In a fourth aspect, an embodiment of the present application provides a video denoising device, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is configured to execute the functions related to processing in the first aspect or any optional implementation manner of the first aspect above.

[0065] Fifth aspect, the present application provides a video processing method, including: First, obtain a current frame and a first fused image. The current frame is any frame image after the first frame arranged in a preset order in video data, and the first fused image includes information of at least one frame adjacent to the current frame in the video data according to the preset order; Subsequently, extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature; Then determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image; According to the first fusion weight and the second fusion weight, fuse the current frame and the first fused image to obtain a second fused image.

[0066] Therefore, in the embodiment of the present application, when fusing the current frame and the first fused image, the foreground part in the current frame and the background part in the first fused image can be referred to, so that ghosting in the second fused image can be reduced, and noise included in the current frame can also be reduced by fusing images, obtaining a clearer image.

[0067] In a possible implementation manner, before extracting features from the current frame and extracting features from the first fused image, the current frame can also be subjected to color transformation through a color transformation matrix to obtain a first chrominance component and a first luminance component. The first chrominance component and the first luminance component form a new current frame, and the first fused image is subjected to color transformation to obtain a second chrominance component and a second luminance component. The second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel; The above-mentioned extracting features from the current frame to obtain a first feature and extracting features from the first fused image to obtain a second feature may include: extracting features from the new current frame to obtain a first feature, and extracting features from the new first fused image to obtain a second feature.

[0068] Therefore, in the embodiment of the present application, color removal processing is also performed on the current frame, which is equivalent to a color removal operation, reducing the correlation between color channels. If noise removal is required, the subsequent noise removal complexity can be reduced, the noise removal efficiency and effect can be improved, and an image with better image quality can be obtained.

[0069] In a possible implementation manner, the above method may further include: performing color transformation on the second fused image through an inverse color transformation matrix to obtain an updated second fused image. The inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0070] In the embodiments of the present application, after performing a color-related transformation, an inverse color transformation can be further performed to restore the color in the image and obtain an image with color.

[0071] In a possible implementation, before extracting features from the current frame and extracting features from the first fused image, the above method further includes: performing a wavelet transform on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component, the first low-frequency component and the first high-frequency component forming a new current frame, and performing a wavelet transform on the first fused image to obtain a second low-frequency component and a second high-frequency component, the second low-frequency component and the second high-frequency component forming a new first fused image, where the wavelet coefficients are preset coefficients or obtained by training at least one convolutional kernel.

[0072] The wavelet transform can be understood as performing a decorrelation transformation on the pixel points of the image in the frequency dimension. Generally, based on the frequency representation, the effective part and noise of the image are separated into different frequency components. If denoising is required subsequently, it can make the denoising simpler. It is equivalent to discretely distributing the frequencies of the pixel points to facilitate achieving a better denoising effect subsequently.

[0073] In a possible implementation, the above method may further include: performing an inverse wavelet transform on the second fused image using inverse wavelet coefficients to obtain an updated second fused image, where the inverse wavelet coefficients are the inverse matrix of the wavelet coefficients. Thus, the inverse wavelet transform can accurately restore the high-frequency and low-frequency components to obtain a clearer image.

[0074] In a possible implementation, determining the first fusion weight corresponding to the current frame according to the first feature and the second feature includes: calculating the shot noise and the read noise according to the shooting parameters of the device used to shoot the video data; determining the first fusion weight corresponding to the current frame in combination with the shot noise, the read noise, the first feature, and the second feature.

[0075] It is equivalent to also combining the noise level when calculating the first fusion weight, so as to adapt to scenarios with different noise levels, and accurate denoising can be performed under different noise levels, with strong generalization ability.

[0076] In a possible implementation, determining the first fusion weight corresponding to the current frame in combination with the shot noise, the read noise, the first feature, and the second feature may include: calculating the noise variance of each pixel point in the current frame according to the shot noise and the read noise; extracting a fifth feature from the noise variance; determining the first fusion weight corresponding to the current frame in combination with the fifth feature, the first feature, and the second feature. The noise variance can be used to accurately determine the noise level of the current frame. If denoising is required, it can facilitate accurate denoising subsequently and improve the denoising effect.

[0077] In a possible implementation, the above method may further include: performing at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image; fusing each frame in the at least one downsampled frame with the downsampled fused image of the same scale to obtain a multi-scale fused image; fusing the second fused image and the multi-scale fused image to obtain an updated second fused image.

[0078] In the embodiments of the present application, the current frame and the first fused image can be downsampled one or more times, and the downsampled frames and downsampled fused images of different scales can be iteratively processed, so that a clearer image can be fused and the noise in the image can be reduced.

[0079] In a possible implementation, fusing any one frame in the at least one downsampled frame with the downsampled fused image of the same scale may include: determining the weight corresponding to the first downsampled frame according to the features extracted from the first downsampled frame and the features extracted from the first downsampled fused image to obtain a fifth fusion weight, and determining the weight corresponding to the first downsampled fused image to obtain a sixth fusion weight. The first downsampled frame is any one frame in the at least one downsampled frame, and the first downsampled fused image is a frame in the at least one downsampled fused image with the same scale as the first downsampled frame; determining the weight corresponding to the second downsampled fused image according to the features extracted from the second downsampled fused image to obtain a seventh fusion weight. The second downsampled fused image incorporates the information of the images in the at least one downsampled frame with a scale smaller than the first downsampled frame; fusing the first downsampled frame, the first downsampled fused image, and the second downsampled fused image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain a third downsampled fused image, and the upsampled image of the third downsampled image is used to fuse with the images in the at least one downsampled frame with a scale larger than the first downsampled frame.

[0080] Therefore, in the embodiments of the present application, during the fusion process for each scale, the processing results of other scales can be fused for iterative fusion, improving the quality of the fused image and obtaining an image with less noise.

[0081] In a possible implementation, performing at least one downsampling on the current frame and the first fused image includes: performing at least one wavelet transform on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image. Therefore, in the embodiments of the present application, downsampling can be performed by means of wavelet transform, and at the same time, the frequencies of the pixel points can be discretely distributed in the frequency space to achieve the effect of denoising.

[0082] In a sixth aspect, the present application provides a video processing device, including:

[0083] An acquisition module, configured to acquire a current frame and a first fused image, where the current frame is any frame image after the first frame arranged in a preset order in video data, and the first fused image includes information of at least one frame adjacent to the current frame in the video data according to the preset order;

[0084] A temporal fusion module, configured to extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature;

[0085] The temporal fusion module is further configured to determine a first fusion weight and a second fusion weight according to the first feature and the second feature, where the first fusion weight includes the weight corresponding to the current frame, the second fusion weight includes the weight corresponding to the first fused image, the weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image;

[0086] The temporal fusion module is further configured to fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image.

[0087] In a possible implementation manner, the above device may further include:

[0088] A de-color correlation transformation module, configured to perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component before the temporal fusion module extracts features from the current frame and extracts features from the first fused image, where the first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fused image to obtain a second chrominance component and a second luminance component, where the second chrominance component and the second luminance component form a new first fused image, and the color transformation matrix is a preset matrix or obtained by training at least one convolution kernel;

[0089] The temporal fusion module is specifically configured to extract features from the new current frame to obtain a first feature, and extract features from the new first fused image to obtain a second feature.

[0090] In a possible implementation manner, the above device may further include: an inverse color correlation transformation module, configured to perform color transformation on the second fused image through an inverse color transformation matrix to obtain an updated second fused image, where the inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0091] In a possible implementation, the above-mentioned device may further include: a de-frequency correlation transformation module, configured to perform wavelet transformation on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component before the time-domain fusion module extracts features from the current frame and extracts features from the first fusion image. The first low-frequency component and the first high-frequency component form a new current frame, and perform wavelet transformation on the first fusion image to obtain a second low-frequency component and a second high-frequency component. The second low-frequency component and the second high-frequency component form a new first fusion image. The wavelet coefficients are preset coefficients or obtained by training at least one convolutional kernel.

[0092] In a possible implementation, the above-mentioned device may further include: an inverse frequency correlation transformation module, configured to perform inverse wavelet transformation on the second fusion image through inverse wavelet coefficients to obtain an updated second fusion image. The inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

[0093] In a possible implementation, the time-domain fusion module is specifically configured to: calculate shot noise and read noise according to the shooting parameters of the device used to shoot the video data; determine the first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature.

[0094] In a possible implementation, the time-domain fusion module is specifically configured to: calculate the noise variance of each pixel point in the current frame according to the shot noise and the read noise; extract a fifth feature from the noise variance; determine the first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature.

[0095] In a possible implementation, the above-mentioned device may further include: a downsampling module, configured to perform at least one downsampling on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fusion image;

[0096] The time-domain fusion module is further configured to fuse each frame in at least one downsampled frame and the downsampled fusion image of the same scale to obtain a multi-scale fusion image;

[0097] The time-domain fusion module is further configured to fuse the second fusion image and the multi-scale fusion image to obtain an updated second fusion image.

[0098] In a possible implementation, the time-domain fusion module fusing any one of at least one downsampled frame and a downsampled fusion image of the same scale may include: determining a weight corresponding to the first downsampled frame based on the features extracted from the first downsampled frame and the features extracted from the first downsampled fusion image to obtain a fifth fusion weight, and determining a weight corresponding to the first downsampled fusion image to obtain a sixth fusion weight, where the first downsampled frame is any one of at least one downsampled frame, and the first downsampled fusion image is a frame in at least one downsampled fusion image with the same scale as the first downsampled frame; determining a weight corresponding to the second downsampled fusion image based on the features extracted from the second downsampled fusion image to obtain a seventh fusion weight, where the second downsampled fusion image incorporates information of images in at least one downsampled frame with a scale smaller than the first downsampled frame; fusing the first downsampled frame, the first downsampled fusion image, and the second downsampled fusion image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain a third downsampled fusion image, and the upsampled image of the third downsampled image is used for fusion with images in at least one downsampled frame with a scale larger than the first downsampled frame.

[0099] In a possible implementation, the downsampling module is specifically configured to perform at least one wavelet transform on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image.

[0100] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium including instructions that, when running on a computer, cause the computer to execute the method in any one of the optional implementations in the first aspect or the fifth aspect above.

[0101] In an eighth aspect, an embodiment of the present application provides a computer program product containing instructions that, when running on a computer, cause the computer to execute the method in any one of the optional implementations in the first aspect or the fifth aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1 is a schematic diagram of an artificial intelligence main framework applied in the present application;

[0103] Figure 2 is a schematic diagram of the structure of a convolutional neural network provided by the present application;

[0104] Figure 3 is a schematic diagram of the structure of another convolutional neural network provided by the present application;

[0105] Figure 4 is a schematic diagram of an application scenario provided by the present application;

[0106] Figure 5 is a schematic diagram of another application scenario provided by the present application;

[0107] Figure 6 It is a schematic diagram of a system architecture provided by this application;

[0108] Figure 7A It is a schematic flowchart of a video processing method provided by this application;

[0109] Figure 7B It is a schematic flowchart of a video denoising method provided by this application;

[0110] Figure 8 It is a schematic flowchart of another video denoising method provided by this application;

[0111] Figure 9 It is a schematic diagram of a color - related transformation removal method provided by this application;

[0112] Figure 10 It is a schematic diagram of a frequency - related transformation removal method provided by this application;

[0113] Figure 11 It is a schematic diagram of a time - domain fusion method provided by this application;

[0114] Figure 12 It is a schematic diagram of a denoising method provided by this application;

[0115] Figure 13 It is a schematic diagram of a refinement step provided by this application;

[0116] Figure 14 It is a schematic diagram of an inverse color transformation method provided by this application;

[0117] Figure 15 It is a schematic flowchart of an inverse wavelet transform provided by this application;

[0118] Figure 16 It is a schematic diagram of a multi - scale processing method provided by this application;

[0119] Figure 17 It is a schematic diagram of the structure of a video denoising device provided by this application;

[0120] Figure 18 It is a schematic diagram of the structure of a video processing device provided by this application;

[0121] Figure 19 It is a schematic diagram of the structure of another video denoising device provided by this application;

[0122] Figure 20 It is a schematic diagram of the structure of a video processing device provided by this application;

[0123] Figure 21It is a schematic structural diagram of a chip provided by this application. Specific embodiments

[0124] Next, the technical solutions in the embodiments of this application will be described with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the protection scope of this application.

[0125] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 shown is a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence theme framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0126] (1) Infrastructure

[0127] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and is supported through the basic platform. Communicate with the outside through sensors; the computing power is provided by intelligent chips, such as hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), or field programmable gate array (FPGA); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0128] (2) Data

[0129] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0130] (3) Data processing

[0131] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0132] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the data.

[0133] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information to perform machine thinking and solve problems according to the reasoning control strategy. The typical function is search and matching.

[0134] Decision-making refers to the process of making decisions after the intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.

[0135] (4) General capabilities

[0136] After the data undergoes the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0137] (5) Intelligent products and industry applications

[0138] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which is the encapsulation of the overall artificial intelligence solution, productizing the intelligent information decision-making and realizing the landing application. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0139] The embodiments of this application involve some applications related to neural networks. To better understand the solutions of the embodiments of this application, some related terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.

[0140] The embodiments of this application involve related applications in the fields of neural networks and images. To better understand the solutions of the embodiments of this application, some related terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.

[0141] (1) Neural network

[0142] A neural network can be composed of neural units, and a neural unit can refer to an operation unit with x s and intercept 1 as inputs. The output of this operation unit can be as shown in formula (1-1):

[0143]

[0144] where s = 1, 2, …… n, n is a natural number greater than 1, W s is the weight of x s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0145] (2) Deep neural network

[0146] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple intermediate layers. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the intermediate layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all intermediate layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer.

[0147] Although the DNN looks complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: where, is the input vector, is the output vector, is the offset vector, w is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs such a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and offset vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient w as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscript corresponds to the third-layer index 2 of the output and the second-layer index 4 of the input.

[0148] In summary, the coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is defined as

[0149] It should be noted that there is no W parameter in the input layer. In a deep neural network, more intermediate layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can perform more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).

[0150] (3) Convolutional Neural Network

[0151] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. The convolutional layer refers to the layer of neurons in a convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some adjacent-layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn reasonable weights. Additionally, the direct benefit of sharing weights is to reduce the connections between layers of the convolutional neural network while also reducing the risk of overfitting.

[0152] (4) Loss function: It can also be called the cost function, a measure that compares the predicted output of a machine learning model for a sample with the true value (also called the supervised value) of the sample, that is, it is used to measure the difference between the predicted output of a machine learning model for a sample and the true value of the sample. This loss function usually can include loss functions such as mean squared error, cross-entropy, logarithm, and exponent. For example, the mean squared error can be used as the loss function, defined as Specifically, the specific loss function can be selected according to the actual application scenario.

[0153] (5) Gradient: The derivative vector of the loss function with respect to the parameters.

[0154] (6) Stochastic gradient: In machine learning, since the number of samples is very large, the loss function calculated each time is calculated from the data obtained by random sampling, and the corresponding gradient is called the stochastic gradient.

[0155] (7) Backpropagation (BP): An algorithm for calculating the gradient of the model parameters according to the loss function and updating the model parameters.

[0156] (8) Foreground, background

[0157] Generally, the foreground can be understood as the main body included in the image, or the object to be concerned about, etc., or it can also be called an instance. The background is the other area in the image except the foreground. For example, if an image including a traffic light is captured, the foreground (or called an instance) in the image is the area where the traffic light is located, and the background is the area in the image except the instance. Another example is that if an image of a road is captured during the driving of a vehicle, other vehicles, lane lines, traffic lights, roadblocks, pedestrians, etc. in the image are instances, and the part except the instances is the background.

[0158] (9) R (red), G (green), B (blue)

[0159] Among them, R represents red, G represents green, and B represents blue. Each image can be represented by the color values of these three channels. For example, an RGB image represents an image with three color channels, and an RGGB image represents an image with four color channels, where two color channels are G.

[0160] (10) YUV

[0161] YUV is a color encoding method, often used in various video processing components. When using YUV to encode photos or videos, considering the user's perception ability, it allows reducing the bandwidth of chrominance. "Y" represents luminance (Luminance or Luma, that is, the grayscale value), and "U" and "V" represent chrominance (Chrominance or Chroma), which is used to describe the image color and saturation and is used to specify the color of pixels.

[0162] Exemplarily, CNN is a commonly used neural network. For example, in the following embodiments of the present application, CNN can be used for steps such as feature extraction or fusion. For ease of understanding, the structure of the convolutional neural network is introduced below by way of example.

[0163] A CNN is a deep neural network with a convolutional structure. A CNN is a deep learning architecture, which refers to multiple levels of learning at different levels of abstraction through machine learning algorithms. As a deep learning architecture, a CNN is a feed-forward artificial neural network, in which each neuron responds to overlapping regions in the input image. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weights are the convolutional kernels. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of the image is the same as that of other parts. That is to say, the image information learned in a certain part can also be used in another part. Therefore, for all positions on the image, the same learned image information can be used. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels, the richer the image information reflected by the convolution operation.

[0164] The convolutional kernels can be initialized in the form of matrices of random sizes, and during the training process of the convolutional neural network, the convolutional kernels can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network while reducing the risk of overfitting.

[0165] A convolutional neural network can use the error backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial super-resolution model during the training process, making the reconstruction error loss of the super-resolution model smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the error loss information is backpropagated to update the parameters in the initial super-resolution model, thereby making the error loss converge. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.

[0166] The following is combined with Figure 2The structure of the CNN will be introduced in detail by way of example. As described in the above basic concept introduction, the convolutional neural network is a deep neural network with a convolutional structure and is a deep learning architecture. The deep learning architecture refers to performing multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, the CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network can respond to the input image.

[0167] As Figure 2 shown, the convolutional layer / pooling layer 120 may include layers such as examples 121 - 126. In one implementation, layer 121 is a convolutional layer, layer 122 is a pooling layer, layer 123 is a convolutional layer, layer 124 is a pooling layer, 125 is a convolutional layer, and 126 is a pooling layer; in another implementation, 121 and 122 are convolutional layers, 123 is a pooling layer, 124 and 125 are convolutional layers, and 126 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.

[0168] Taking the convolutional layer 121 as an example, the convolutional layer 121 may include a lot of convolutional operators. The convolutional operator is also called a kernel, and its role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, the convolutional operator can be a weight matrix, and this weight matrix is usually predefined. During the convolutional operation on the image, the weight matrix usually processes the input image pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction, so as to complete the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix and the depth dimension of the input image are the same. During the convolutional operation, the weight matrix will extend to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, a single weight matrix is not used, but multiple weight matrices with the same dimension are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract the edge information of the image, another weight matrix is used to extract the specific color of the image, and another weight matrix is used to blur the unwanted noise in the image, etc. These multiple weight matrices have the same dimension, and the feature maps extracted by these multiple weight matrices with the same dimension also have the same dimension. Then, the multiple extracted feature maps with the same dimension are combined to form the output of the convolutional operation.

[0169] Generally, the weight values in the weight matrix need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can extract information from the input image, thereby helping the convolutional neural network 100 to make correct predictions.

[0170] When the convolutional neural network 100 has multiple convolutional layers, the initial convolutional layer (such as 121) often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network 100 increases, the later convolutional layers (such as 126) extract more and more complex features, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.

[0171] Pooling layer:

[0172] Since it is often necessary to reduce the number of training parameters, a pooling layer is often periodically introduced after the convolutional layer, that is, each layer of 121-126 shown in Figure 2 120 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In the process of image processing, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of the pixel values in the image within a specific range. The max pooling operator can take the pixel with the largest value within a specific range as the result of max pooling. In addition, just as the size of the weight matrix in the convolutional layer should be related to the size of the image, the operators in the pooling layer should also be related to the size of the image. The size of the image output after being processed by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel point in the image output by the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.

[0173] Neural network layer 130:

[0174] After being processed by the convolutional layer / pooling layer 120, the convolutional neural network 100 is still not sufficient to output the required output information. Because as mentioned above, the convolutional layer / pooling layer 120 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network 100 needs to use the neural network layer 130 to generate one or a group of outputs with the number of classes required. Therefore, the neural network layer 130 can include multiple hidden layers (such as Figure 2131, 132 to 13n) and output layer 140. In the present application, the convolutional neural network is: the selected starting network is deformed at least once to obtain a serial network, and then obtained according to the trained serial network. The convolutional neural network can be used for image recognition, image classification, image super-resolution reconstruction, etc.

[0175] After the multiple hidden layers in the neural network layer 130, that is, the last layer of the entire convolutional neural network 100 is the output layer 140, which has a loss function similar to the classification cross entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 100 (such as Figure 2 The propagation from 110 to 140 is forward propagation), and the reverse propagation (such as Figure 2 The propagation from 140 to 110 is called back propagation) and then the weight values and biases of the aforementioned layers will begin to be updated to reduce the loss of the convolutional neural network 100 and the error between the result output by the convolutional neural network 100 through the output layer and the ideal result.

[0176] It should be noted that if Figure 2 The convolutional neural network 100 shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models, for example, Figure 3 The multiple convolutional layers / pooling layers shown are operated in parallel, and the features extracted respectively are input to the full neural network layer 130 for processing.

[0177] The details of the CNN mentioned below in this application can be exemplified by referring to the aforementioned Figure 2 or Figure 3 The CNN shown.

[0178] In the implementation mode of the present application, the video can be denoised to improve the display quality of the video. The video denoising method provided in the present application can be executed by a terminal or by a server. For example, the method provided in the present application can be deployed in a user's mobile phone, camera, video surveillance, television, server or image signal processor (ISP), etc., to denoise the collected or received video data.

[0179] For example, the video denoising method provided in this application can be applied to smart city scenarios, such as Figure 4 As shown, low-quality video data collected by various monitoring devices can be collected and stored in a memory. When playing the video data, the video denoising method provided by the present application can be used to denoise the video data to obtain clearer video data and improve the user's viewing experience.

[0180] For another example, the video denoising method provided by the present application can be applied to various video shooting scenarios. For example, a user can use a terminal to shoot a video and save it locally. Before the user plays the video using the terminal, the video denoising method provided by the present application can be used to denoise the stored video data, so as to obtain higher-quality video data and improve the user's viewing experience.

[0181] For still another example, the video denoising method provided by the present application can be applied to a video live broadcast scenario, such as Figure 5 shown, the server can send a video stream to the client used by the user. After the client receives the data stream sent by the server, the data stream can be denoised by using the video denoising method provided by the present application, so as to obtain video data with higher image quality and improve the user's viewing experience.

[0182] In addition, the video denoising method provided by the present application can also be applied to scenarios such as an autonomous driving scenario and image enhancement, which will not be elaborated here one by one.

[0183] The video denoising method provided by the embodiments of the present application can be executed on a server or on a terminal device. The terminal device can be a mobile phone with image processing functions, a tablet personal computer (TPC), a media player, a smart TV, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a camera, a video camera, a smart watch, a wearable device (WD), or an autonomous driving vehicle, etc. The embodiments of the present application do not limit this.

[0184] Exemplarily, the system architecture to which the video denoising method provided by the present application is applied can be as Figure 6 shown. In this system architecture 400, the server cluster 410 is implemented by one or more servers. Optionally, in cooperation with other computing devices, such as devices such as a data storage, a router, and a load balancer. The server cluster 410 can use the data in the data storage system 250 or call the program code in the data storage system 250 to implement the steps of the video denoising method provided by the present application.

[0185] Users can operate their respective user devices (such as local device 401 and local device 402) to interact with the server cluster 410. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smartphone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.

[0186] The local device of each user can interact with the server cluster 410 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof. Specifically, the communication network can include a wireless network, a wired network, or a combination of a wireless network and a wired network, etc. The wireless network includes, but is not limited to: a fifth-generation mobile communication technology (5th-Generation, 5G) system, a long term evolution (LTE) system, a global system for mobile communication (GSM), or a code division multiple access (CDMA) network, a wideband code division multiple access (WCDMA) network, wireless fidelity (WiFi), Bluetooth, Zigbee, radio frequency identification (RFID), Long Range (Lora) wireless communication, near field communication (NFC), or any combination of one or more of them. The wired network can include a fiber optic communication network or a network composed of coaxial cables, etc.

[0187] Exemplarily, in an application scenario, any one of the servers in the server cluster 410 can obtain video data from a data storage system or other devices, such as terminals, PCs, etc., and then perform denoising processing on the video data through the video denoising method provided in this application to make each frame image in the video data clearer, improve the user experience, and send the denoised video data to the local device.

[0188] However, even with the advancement of technology, due to the randomness of the acquisition process and / or challenging sensing conditions, such as more noise in images captured in low-light scenarios, digital images are always affected by some inherent or external interference factors. These interference factors are random and can be modeled as a random variable, and its random fluctuations are called "noise". The noise in images captured by ordinary image sensors such as charge-coupled device (CCD) or complementary metal oxide semiconductor (CMOS) approximately follows a combination of Poisson distribution and Gaussian distribution, modeling signal-dependent and signal-independent noise sources respectively. Denoising refers to removing such random noise from noisy data without introducing artifacts or modifying the original image structure.

[0189] Generally, the photographing and imaging of terminal devices are limited by the hardware performance of the optical sensors of the terminal devices. Due to the imperfection of the acquisition process, the formation of digital images is always affected by different forms of noise and degradation, and image restoration algorithms must be used to restore the degraded input to a high-quality input. There are many image restoration methods, such as denoising, demosaicking, super-resolution, etc. Specifically, denoising is an essential task for the camera processing pipeline (e.g., in smartphones or video surveillance cameras), because as the first typical operation in the camera pipeline, the denoising quality will directly affect the output results of all subsequent tasks.

[0190] Some commonly used denoising algorithms are based on image processing techniques and utilize the statistical characteristics of the input data. For example, denoising can be performed through algorithms with low computational complexity such as local correlation, non-local correlation methods, or sparse algorithms. More complex algorithms such as NL-means or BM3D can produce higher-quality outputs, but the inference speed of these algorithms on general-purpose processors is very slow and they need to be implemented on specific hardware. Generally, in order to better generalize to input images at any noise level, multiple parameters need to be carefully adjusted manually one by one to explicitly control these denoising algorithms to obtain better denoising effects. Therefore, these denoising methods require a large amount of manual participation and rich manual debugging experience, and the denoising cost is high.

[0191] For video denoising, some commonly used methods are based on non-local imaging or motion estimation for denoising. Although these methods are effective, they usually require extremely high computational power to process multiple frames of input simultaneously. Or the spatio-temporal fusion of video information can be processed in a cyclic recursive manner, but the denoising quality is often unacceptable.

[0192] Generally, a neural network (such as a CNN) is based on a large number of trainable convolutional kernels, and the parameters of these convolutional kernels are optimized in a supervised manner by a task-specific loss function. Given enough data, a large number of parameters can automatically establish a mapping from a degraded noisy input to a restored denoised output during the training process. Another advantage of a standard feedforward CNN is its fast inference time because its basic operation (i.e., convolution) can be easily optimized on hardware. However, a CNN requires a large number of parameters to effectively and reliably solve problems, and once the complexity of the CNN decreases, the performance drops sharply. In addition, since a CNN requires a large amount of training data, the training cost is high.

[0193] Therefore, the present application provides a video processing method and a video denoising method, which can improve the denoising effect while achieving lightweight denoising, obtain clearer images, and further obtain video data with better image quality. The video processing method and the video denoising method provided by the present application can be applied to consumer products such as mobile phones, video surveillance, or TVs, or high-performance computing devices such as cloud products, improve the imaging quality of videos, reduce the denoising difficulty of the noise map through decorrelation transformation, and eliminate the noise existing in the video images through time-domain fusion, spatial denoising, and refinement, etc., to obtain a clean output and enhance the imaging quality of the video images.

[0194] First, the process of the video processing method will be introduced. Refer to Figure 7A .

[0195] 701. Obtain the current frame and the first fusion image.

[0196] Among them, the current frame can be a non-first frame in the video data, that is, any frame arranged after the first frame in a preset order. The information of at least one frame of image adjacent to the current frame in a preset order is fused in the first fusion image, such as one or more frames of images arranged before or after the current frame, etc.

[0197] The preset order can be the playing order of the video data, or the order arranged in chronological order, or the order opposite to the playing order, etc., and can be specifically adjusted according to the actual application scenario, which is not limited here.

[0198] For example, according to the playing order of the video, each frame of the video can be denoised to obtain a clearer image. During the process of denoising each frame of the image, a fusion image of one or more frames of images adjacent to each frame of the image can be selected to denoise each frame of the image, so as to combine the time-domain information in the video to denoise the image, improve the video denoising effect, and thus obtain a video with better image quality.

[0199] It should be understood that if the current frame is the first frame arranged in a preset order in the video data, when processing the first frame, there may be no first fused image, and the step of fusing the current frame and the first fused image does not need to be executed. If the current frame is the second frame, the first fused image can directly be the first frame.

[0200] 702. Extract features from the current frame to obtain first features, and extract features from the first fused image to obtain second features.

[0201] Among them, features can be extracted from the current frame and the first fused image respectively. For the convenience of understanding, the features extracted from the current frame are called first features, and the features extracted from the first fused image are called second features.

[0202] Specifically, a feature extraction network can be used to extract features from the image. The feature extraction network can be a CNN as described above, or other neural networks including one or more convolutions, etc.

[0203] 703. Determine a first fusion weight and a second fusion weight according to the first features and the second features.

[0204] After extracting features from the current frame and the first fused image, it is also necessary to fuse the first features and the second features. Before that, according to the features of the foreground and background respectively included in the first features and the second features, the weights occupied by each image when fusing the current frame and the first fused image can be determined respectively.

[0205] It can be understood that the first features can include the features of the foreground part and the background part in the current frame, and can be used to identify the positions of the foreground part and the background part in the current frame; the second features can include the features of the foreground part and the background part in the first fused image, and can be used to identify the positions of the foreground part and the background part in the first fused image. For example, in video data, as time changes, the position of the foreground in each frame image may be different, and the position of the foreground included in the current frame is usually more accurate. Therefore, it is necessary to distinguish the foreground part and the background part in the current frame for subsequent fusion. And because the first fused image incorporates the information of at least one frame of image arranged before the current frame, the noise it includes is usually less and the image quality is better.

[0206] In this embodiment, the features extracted from the first fused image can be used to distinguish the panoramic part and the background part in the first fused image, so that when subsequently fusing the current frame and the first fused image, the weights occupied by the foreground part and the background part can be determined more accurately, thereby smoothing the noise in the current frame and making the noise included in the fused image less.

[0207] Among them, the weight corresponding to the foreground in the current frame is not less than the weight occupied by the foreground in the first fused image, and the weight occupied by the background part in the current frame is not greater than the weight occupied by the background part in the first fused image. Specifically, the foreground part in the current frame can be determined through the first feature, and the foreground part in the first fused image can be determined through the second feature, and then a higher weight can be given to the foreground part in the current frame. For example, the weight of the foreground part in the current frame can be set to 0.8, and the weight of the background part in the first fused image can be set to 0.2. The background part in the current frame can also be determined through the first feature, and the background part in the first fused image can be determined through the second feature, and then a lower weight can be given to the background part in the current frame. For example, the weight of the background part in the current frame can be set to 0.4, and the weight of the background part in the first fused image can be set to 0.6.

[0208] 704. Fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image.

[0209] After obtaining the first fusion weight and the second fusion weight, the current frame and the first fused image can be fused according to the first fusion weight and the second fusion weight to obtain a second fused image.

[0210] Generally, the position of the foreground part in the current frame is more accurate, and the noise in the first fused image is less. Therefore, when fusing the current frame and the first fused image, more reference can be made to the foreground in the current frame and the background in the first fused image, so as to reduce the ghosting in the fused image and reduce the noise in the obtained fused image.

[0211] In the embodiment of the present application, by fusing the current frame and the first fused image, it is equivalent to combining the temporally related information between adjacent frames in the scene, smoothing the noise in the current frame, reducing the noise in the image, and obtaining a second fused image with less noise. And when fusing the current frame and the first fused image, more reference is made to the foreground part in the current frame and the background part in the first fused image. Since the objects in the video data may change positions over time, in the embodiment of the present application, the foreground included in the current frame and the background included in the first fused image can greatly reduce the ghosting in the second fused image and improve the quality of the finally obtained image.

[0212] Optionally, before fusing the current frame and the first fused image, the current frame can also be color-transformed by a color transformation matrix to obtain at least one chrominance component and at least one luminance component, which are referred to as the first chrominance component and the first luminance component for easy distinction here. The at least one first chrominance component and the at least one first luminance component can form a new current frame, that is, the current frame after color transformation. The first fused image is also color-transformed by the color transformation matrix to obtain at least one chrominance component and at least one luminance component, which are referred to as the second chrominance component and the second luminance component for easy distinction here. The at least one second chrominance component and the at least one second luminance component can form a new first fused image. When fusing the current frame and the first fused image, the new current frame and the new first fused image after color transformation can be fused to obtain the second fused image.

[0213] The color transformation performed on the current frame and the first fused image is equivalent to performing a color correlation removal transformation on the current frame and the first fused image, transforming the current frame and the first fused image from the pixel domain to a different domain to facilitate subsequent denoising operations.

[0214] Specifically, the aforementioned color matrix can be preset or obtained by training at least one convolutional kernel. For example, the method provided in this application can be implemented through a neural network. When updating the neural network with a large number of samples, the color transformation matrix can be updated simultaneously to obtain an updated color transformation matrix.

[0215] The difference from the video denoising method provided in this application includes that after obtaining the second fused image, the second fused image is color-transformed by an inverse color transformation matrix to obtain an updated second fused image. The inverse color transformation matrix is the inverse matrix of the color transformation matrix. In the implementation manner of this application, if a color correlation removal transformation is performed, an inverse color transformation can also be performed to restore the color in the image and obtain an image with color.

[0216] Optionally, before fusing the current frame and the first fused image, a decorrelation transform of the spatial frequency can also be performed on the current frame and the first fused image to facilitate subsequent image denoising. Specifically, a wavelet transform can be performed on the current frame and the first fused image to obtain at least one low-frequency component and at least one high-frequency component of the current frame and at least one low-frequency component and at least one high-frequency component of the first fused image. For ease of distinction, the low-frequency component and the high-frequency component of the current frame are respectively referred to as the first low-frequency component and the first high-frequency component, and at least one first low-frequency component and at least one first high-frequency component form a new current frame; the low-frequency component and the high-frequency component of the first fused image are respectively referred to as the second low-frequency component and the second high-frequency component, and at least one second low-frequency component and at least one second high-frequency component form a new first fused image, thereby separating the main structure and detail information in the image from the dimension of spatial frequency, so as to better denoise the main structure and details in the image respectively and improve the denoising effect.

[0217] Specifically, wavelet coefficients can be used to perform a wavelet transform on the current frame to obtain a first low-frequency component and a first high-frequency component. The wavelet coefficients are preset coefficients or obtained by training at least one convolution kernel. The first low-frequency component and the first high-frequency component form a new current frame, and a wavelet transform is performed on the first fused image to obtain a second low-frequency component and a second high-frequency component. The second low-frequency component and the second high-frequency component form a new first fused image.

[0218] If color transformation and wavelet transform need to be performed on the current frame and the first fused image, the wavelet transform can be performed after the color transformation, so as to achieve decorrelation of the color correlation and spatial frequency of the current frame and the first fused image, facilitate subsequent denoising, and obtain an image with better denoising effect.

[0219] The difference from the video denoising method provided in this application includes that if a wavelet transform is performed, after obtaining the second fused image, an inverse wavelet transform can also be performed on the second fused image. For example, an inverse wavelet transform is performed on the second fused image through inverse wavelet coefficients to obtain an updated second fused image. The inverse wavelet coefficients are the inverse matrix of the wavelet coefficients. Thus, the inverse wavelet transform can accurately restore the high-frequency component and the low-frequency component, which is equivalent to restoring the structure and details in the image, and obtaining a clearer second fused image.

[0220] In addition, it can also be understood that when processing the next frame of the current frame, that is, when the next frame is used as the new current frame, the second fused image can be used as the new first fused image to perform denoising processing on the new current frame, thereby realizing iterative denoising of video data, using the temporal correlation between adjacent frames in the video data for denoising, improving the denoising quality of the image, and obtaining a denoised image with better image quality.

[0221] To further reduce the noise in the second fused image, the second fused image can also be further denoised. Refer to Figure 7B , which is a schematic flowchart of a video denoising method provided by this application.

[0222] It should be noted that steps 701-704 in the embodiments of this application can refer to the relevant introductions in the foregoing Figure 7A and will not be elaborated here.

[0223] 705. Denoise the second fused image to obtain a denoised image.

[0224] Among them, after fusing the current frame and the first fused image, the noise in the obtained second fused image has decreased relative to the current frame. At this time, the second fused image can continue to be denoised to reduce the noise in the second fused image and obtain a denoised image with less noise.

[0225] Specifically, the second fused image can be filtered, such as FIR (finite impulse response) filtering, median filtering, Wiener filtering, etc., so as to reduce the noise in the second fused image and obtain a denoised image with better image quality.

[0226] Optionally, when denoising the second fused image, the first fusion weight, the second fusion weight, the current frame, and the first fused image can be combined for denoising. Therefore, when denoising, the information used for fusing the current frame and the first fused image can be reused to improve data utilization and make the subsequent denoising effect better.

[0227] Optionally, the variance of each pixel point in the second fused image can be calculated by combining the first fusion weight, the first fused image, and the current frame to obtain the variance of the fused image. Then, by fusing the variance of the fused image and the second fused image, a further denoised image can be obtained.

[0228] For example, the variance of each pixel point in the fused image can be calculated through a cyclic recursive formula to obtain the variance of the fused image, and the information of time-domain fusion can be reused to calculate the variance of each pixel point in the second fused image. Then, this variance is used in combination with the second fused image for denoising to obtain a denoised image with better image quality.

[0229] More specifically, features can be extracted from the second fused image to obtain a third feature, and features can be extracted from the variance of the fused image to obtain a sixth feature. Then, by fusing the third feature and the sixth feature, a denoised image can be obtained. In this embodiment, by fusing the variance of the fused image, the noise in the second fused image is further smoothed, and a denoised image with less noise is obtained.

[0230] Specifically, for example, the denoising step can be performed by a denoising CNN, which includes one or more convolutional layers and rectified linear unit (ReLU). By using the second fused graph and the fused graph variance as the input of the denoising CNN, a denoised image can be output. The denoising CNN can be trained using a large number of samples and can be used to extract features and fuse the extracted features.

[0231] In addition, when denoising, the current frame can also be combined for denoising. Specifically, for example, when using the denoising CNN for denoising, in addition to using the fused graph variance and the second fused image as the input of the denoising CNN, the current frame can also be used as the input of the denoising CNN to output a denoised image.

[0232] Therefore, in the embodiments of the present application, by iteratively updating each frame in the video, which is equivalent to utilizing the temporal correlation between the images in the video, the noise in each frame image can be smoothed, improving the denoising effect of each frame in the video data, thereby obtaining video data with better image quality and improving the user experience.

[0233] Optionally, if the current frame and the first fused image are de-colored before fusing the current frame and the first fused image, after denoising the second fused image, an inverse color transformation can be performed on the denoised image. Specifically, an inverse transformation can be performed on the denoised image through an inverse color transformation matrix, which is the inverse matrix of the color transformation matrix, that is, the product of the color matrix and the inverse color matrix results in an identity matrix. When updating at least one convolutional kernel, the product of the color matrix and the inverse color matrix being an identity matrix can be used as a constraint for updating, so that the inverse color transformation can accurately restore the color of the image and obtain a clearer denoised image.

[0234] It should be understood that if step 705 is not performed, an inverse color transformation can be directly performed on the second fused image, and the transformation method is similar to that of the above-mentioned denoised image, so as to restore the color of the second fused image and obtain a clear second fused image.

[0235] Optionally, if a wavelet transformation is performed on the current frame and the first fused image before fusing the current frame and the first fused image, after denoising the second fused image, an inverse wavelet transformation can be performed on the denoised image to obtain a denoised image related to the spatial frequency dimension. Specifically, an inverse wavelet transformation can be performed on the denoised image through inverse wavelet coefficients to obtain an updated denoised image, and the inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

[0236] It should be understood that if step 705 is not performed, the inverse wavelet transform can be directly performed on the second fused image, and the transformation method is similar to the transformation method of the denoised image described above, so as to restore the structure and details of the second fused image and obtain a clear second fused image.

[0237] Therefore, in the embodiments of the present application, through color removal transformation and decorrelation transformation in the spatial dimension, the color, high frequency, and low frequency of the image can be separated, so that noise can be filtered out more accurately, and a denoised image with a better denoising effect can be obtained.

[0238] 706. Fuse the second fused image and the denoised image to obtain an updated denoised image.

[0239] After obtaining the denoised image, in order to avoid excessive smoothing, the second fused image and the denoised image can be fused, so that the details in the denoised image can be enriched through the information included in the second fused image, so as to obtain a denoised image including more details and improve the quality of the denoised image.

[0240] Optionally, the specific method of fusing the second fused image and the denoised image may include: First, extract features from the second fused image to obtain a third feature; Subsequently, extract features from the denoised image to obtain a fourth feature; Determine a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature. The third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among them, in the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value are negatively correlated; According to the third fusion weight and the fourth fusion weight, fuse the third feature and the fourth feature to obtain an updated denoised image. Therefore, in the embodiments of the present application, the weight value can be determined according to the frequency of the pixel point. The higher the frequency, the lower the corresponding weight value, so that the noise included in the high-frequency information can be effectively smoothed, and the denoising effect can be achieved.

[0241] Therefore, in the embodiments of the present application, after denoising the second fused image to obtain a denoised image, the second fused image and the denoised image can also be fused, which can prevent the denoised image from being overly smoothed, use the second fused image to enrich the details in the denoised image, obtain a denoised image with richer details, improve the picture quality of the denoised image, and improve the user experience.

[0242] The foregoing has introduced the process of the method provided by the present application. For ease of understanding, the process of the video denoising method provided by the present application will be further introduced based on a more detailed application scenario.

[0243] Exemplarily, the process of another video denoising processing method provided by the present application can be referred to Figure 8 .

[0244] The embodiments of the present application can be understood as proposing a denoising method that cyclically uses a multi-stage processing method to remove noise in a video. The input video can be an image sequence collected by any camera or sensor, and each frame of the image can be identified by different time steps, such as represented as {0, 1, 2,..., t,...}.

[0245] First, due to carrying noise, the aforementioned current frame is represented as a noisy frame Noisy t , and the noisy frame Noisy t can be any frame in the video data. If the noisy frame Noisy t is not the first frame, a fused image that combines the previous frame of the image can be obtained, such as represented as Fused t-1 .

[0246] Then, decorrelation operations are performed on Noisy t and Fused t-1 , such as decorrelating colors, decorrelating spatial frequencies, etc., in order to facilitate subsequent noise smoothing. The input video image can first be converted to a color-decorrelated luminance-chrominance space using a color transformation, and then a wavelet transform is used for frequency transformation. These two transformations can reduce the difficulty of denoising, and the wavelet transform has the advantages of reducing the image resolution and reducing the computational complexity.

[0247] Temporal fusion refers to fusing the decorrelated Noisy t and Fused t-1 at different time periods to obtain a fused image Fused t after fusion. The fused image Fused t fuses the temporal correlation information between the current frame and the previous fused frame, thereby smoothing the noise. When processing the next frame, Fused t can be fused to smooth the noise, realizing iterative denoising of the video data. The temporal fusion step can detect the object information moving between the frames of the video. By recursively using the fused image of the previous frame (t - 1) and the noisy input of the current frame in a loop, the static background of the image gradually converges to the effect of multi-frame averaging, while the moving foreground comes from the current frame. This will greatly reduce the noise in the input image, obtain a fused image, and greatly reduce the difficulty of denoising.

[0248] After obtaining Fused t , spatial denoising can be performed. Filter the noise in Fused t to obtain a denoised image Denoised t-1 .

[0249] Then, combine Fused t with Denoised t-1Perform fine processing, that is, fuse Fused t The details included into Denoised t-1 , so as to avoid the loss of details caused by excessive smoothing of Denoised t-1 .

[0250] Subsequently, perform an inverse transformation on the refined Denoised t-1 , that is, the inverse transformation of the aforementioned decorrelation transformation, so as to restore the color, scale, etc. of the image, and obtain a new denoised image corresponding to the color or scale, etc. of Noisy t , such as denoted as Output t .

[0251] In the embodiment of the present application, in the fusion stage, the current frame at time (t) is temporally fused with the image fused in the previous time step (t - 1). The fusion stage utilizes the temporal correlation between consecutive frames to reduce noise in the best way. Subsequently, the denoising stage effectively removes all the residual noise on the fused image completely. The initial denoising in the fusion stage is beneficial to the denoising stage, making the denoising task easier.

[0252] However, denoising may produce imperfect outputs. Therefore, the present application adopts a refinement stage to fuse the current fused image and the denoised image. The purpose of refinement is to extract the image structure from the fused image rich in details but still having noise and add it to the denoised image that is noise-free but may be overly smoothed. Eventually, a final output image of higher quality is produced.

[0253] It should be noted that taking Figure 8 as an example, the processes of the video processing method and the video denoising method provided by the present application are similar. The difference is that in the video processing method, the denoising and refinement steps may not need to be executed, and an image with improved image quality can be obtained directly. The input of the subsequent inverse transformation step is directly Fused t . The present application exemplarily illustrates with the overall process of a video denoising method as an example.

[0254] For ease of understanding, each step in the process of the video denoising method provided by the present application will be introduced in detail below. The process of the video denoising method provided by the present application can be specifically divided into steps such as decorrelation transformation, temporal fusion, denoising, refinement, and inverse transformation. These steps will be introduced in more detail by way of example below.

[0255] First, for ease of understanding, an exemplary introduction to the noise that may be included in the image is provided.

[0256] The noise image of the t-th frame (Noisy t ) can be regarded as a clean image (Clean t) and random noise (Noise t ) added together: Noisy t =Clean t +Noise t , where the noise can be a random variable following a signal-dependent heteroscedastic zero-mean Gaussian distribution. For example, the variance distribution can be expressed as:

[0257] Noise t ~Gaussian(mean=0,variance=Shot t Clean t +Read t ).

[0258] Where Shot t represents shot noise, and Read t represents read noise. Shot noise and read noise are usually determined by the performance of the device for shooting the video, shooting parameters, or environmental information. For example, shot noise and read noise are determined according to exposure, gain, or ISO (International Organization for Standardization), etc.

[0259] Specifically, for example, the ISO corresponding to each frame of the image may be different. For example, the ISO value and the noise value usually have a positive correlation. The larger the ISO value, the greater the noise. The noise may also be related to the environmental brightness. For example, the higher the brightness, the greater the noise. The relationship between the noise parameters (such as shot noise and read noise) and the settings of the device for shooting the video is called the noise level function (NLF). Usually, the NLF of the device can be constructed using a calibration method.

[0260] Through the video denoising method provided by this application, shot noise or read noise, etc., can be effectively filtered out, so as to obtain a denoised image with better denoising effect. The detailed steps of denoising are introduced below.

[0261] I. Decorrelation transform

[0262] Among them, the decorrelation transform includes decorrelation transform of color and / or decorrelation transform of spatial frequency, etc. The decorrelation transform of color or the decorrelation transform of spatial frequency will be introduced separately below.

[0263] 1. Decorrelation transform of color

[0264] The decorrelation transform of color (abbreviated as color transform) can be understood as decorrelating the colors of the pixels of the image, so as to reduce the correlation between each pixel point in the color dimension, facilitating subsequent denoising processing.

[0265] Specifically, color transformation and inverse color transformation can be understood as forward and reverse linear decorrelation transformations of an image, which transform the image from the pixel domain to different domains, such as the luminance domain, the chrominance domain, etc., to facilitate subsequent denoising of the image.

[0266] More specifically, decorrelation transformation can be performed through YUV transformation to convert a color image, such as an RGB image, into a luminance component and a chrominance component. The luminance component is the luminance in Noisy t For example, this luminance can be the average value of the colors of the pixel points in Noisy t The chrominance component is the measure that describes the color itself of Noisy t For example, if the color of an image can be divided into three channels, each channel corresponds to a chrominance component. At the same time, by averaging the colors of all channels of each pixel point, a luminance component can be obtained.

[0267] As Figure 9 shown, the input image (i.e., Noisy t ) can be represented by four RGGB channels, that is, each pixel point has a color value in each channel dimension. The color values of each channel in the input image form a one-dimensional matrix, and then a color transformation matrix is used to perform point-by-point convolution on the four-dimensional matrix to obtain three chrominance components and a luminance component.

[0268] Color transformation can be understood as matrix multiplication. The length of the matrix is equal to the number of input color channels. For example, if the input image can be represented as a 4*4 matrix, color transformation can be achieved through point-by-point convolution.

[0269] In the embodiments of the present application, through decorrelation transformation of color, the input image can be converted into a luminance component and at least one chrominance component. The luminance component represents the luminance of the input image, and the chrominance component is used to describe each color of the input image. Since the luminance is usually the average color value, the noise is smoothed through the average color value, reducing the noise.

[0270] 2. Decorrelation transformation of spatial frequency

[0271] Decorrelation transformation of spatial frequency can be understood as performing decorrelation transformation on the pixel points of an image in the frequency dimension. Specifically, high-pass or low-pass filtering can be performed on the input image respectively to obtain high-frequency components and low-frequency components. Thus, the pixel points of the input image can be discretely distributed in separate spatial frequencies, realizing decorrelation of the spatial frequency of the input image, separating the effective part and the noise part of the image into different frequency components, making the denoising effect of subsequent denoising operations better and the denoising simpler.

[0272] Specifically, wavelet transform can be performed on the input image to achieve a de-spatial-frequency correlation transform. The wavelet transform can specifically use high-pass filters and low-pass filters to filter in the vertical and horizontal directions of the input image, that is, each color channel of the input image is decomposed into a low-frequency component and a high-frequency component (or called a low-frequency subband and a high-frequency subband). The size of each component can be equal to that of the input image or smaller than the input image. For example, the size of each component is half of the input image. Usually, the frequency-based representation separates the effective part and noise of the image into different frequency components, making denoising simpler.

[0273] For example, wavelet transform can be used to decompose the input image into four different frequency components, including a low-frequency component and three high-frequency components. The size of each component can be half of the input image. The elements within each subband are called coefficients. These four subbands represent low frequency (such as obtained by filtering the input image with a low-pass kernel) and high frequencies filtered in the vertical, horizontal, and diagonal directions (such as obtained by filtering the image with a directional high-pass kernel).

[0274] If the input image includes multiple channels, each channel can be represented as a one-dimensional image. Wavelet transform can be performed on each one-dimensional image to obtain a low-frequency component and three low-frequency components corresponding to each one-dimensional image. For example, if the input image is represented by four channels, wavelet transform can be performed on each channel, thus obtaining 16 frequency components, including 4 low-frequency components and 12 high-frequency components.

[0275] As Figure 10 shown, the input image is an H*W*C image, where H represents the length, W represents the width, and C represents the number of channels. First, determine the wavelet coefficients. Two one-dimensional forward decomposition filters of the selected wavelet family (how to understand, coefficients) are implemented (such as Harr wavelet). Then calculate the outer product of the paired wavelet coefficients to obtain the wavelet kernel. For example, four two-dimensional convolutional kernels are generated from the outer product of each pair of one-dimensional filters. These kernels can be calculated for the output in a convolution with a stride of 2. Then use the wavelet kernel to perform a convolution operation on the input image. The stride of the convolution operation is 2, and each channel corresponds to a low-frequency component and three high-frequency components.

[0276] Usually, the frequency characteristics corresponding to the image structure and noise are very different. After performing the de-frequency correlation transform, the noise is mainly concentrated in the low absolute value components in the high-frequency components. This is a powerful prior information for denoising. For example, soft thresholding and hard thresholding denoising methods can be used to set the low values of the high-frequency components to zero to filter a large amount of noise and achieve better denoising effects. In the embodiments of the present application, a non-linear model CNN can also be used for denoising to process the frequency components obtained after the decorrelation transform and achieve a better denoising effect.

[0277] 3. Combination of color - related transform removal and spatial - frequency - related transform removal

[0278] Among them, one of the color - related transform removal and the spatial - frequency - related transform removal can be selected to transform the input image, or both the color - related transform removal and the spatial - frequency - related transform removal can be performed on the input image. For example, after performing color - removal transformation to obtain a luminance component and three chrominance components, the one luminance component and three chrominance components (i.e., an image with four channels) are used as the input of wavelet transform, and one low - frequency component and three high - frequency components corresponding to each channel are output.

[0279] The color transformation is a color - related linear transformation similar to YUV transformation, and the wavelet transform is a spatial - frequency - related transformation. These transformations can obtain the optimal transformation parameters through learning, and can be perfectly reconstructed through a reversible loss function.

[0280] Therefore, in the embodiments of the present application, the chrominance component and the luminance component can be separated through color - related transform removal, thereby smoothing the noise in the image and achieving a denoising effect. Or through frequency - related transform removal, the high - frequency and low - frequency of each pixel point are discretely distributed, so as to facilitate subsequent simpler filtering and achieve a better noise filtering effect.

[0281] II. Temporal fusion

[0282] After performing the decorrelation transformation, the input of temporal fusion can include the decorrelated Noisy t and the fused image Fused t-1 . Temporal fusion can be performed by a temporal - fusion CNN. Specifically, features can be extracted from the decorrelated Noisy t and the fused image Fused t-1 respectively, and then the weights corresponding to the decorrelated Noisy t and the fused image Fused t-1 are determined according to the features, and then the decorrelated Noisy t and the fused image Fused t-1 are fused according to the weights. Generally, since the foreground in video data may be in a moving state and its position may be different in each frame, while the background may change little.

[0283] Therefore, when fusing the decorrelated Noisy t and the fused image Fused t-1 , a higher weight can be set for the foreground part of the current frame, i.e., the decorrelated Noisy t , such as for the decorrelated Noisy tThe foreground part in the image has a weight of 0.8, which is the Fused image after decorrelation transformation. t-1 The weight of the pixel corresponding to the foreground is set to 0.2. t The weight of the background part is not greater than the Fused after decorrelation transformation t-1 The weight of the corresponding pixel in . For example, the Noisy t The weight of the background part can be set to 0.4, and the Fused after decorrelation transformation t-1 The weight of the corresponding pixel in is set to 0.6.

[0284] Therefore, in the implementation manner of the present application, the fusion weights of each image can be determined by combining the foreground part in the current frame and the background part in the fused image, so that the foreground part in the fused image refers more to the foreground in the current frame, and the background part refers more to the fused image of the previous frame, so that the fused image has better performance in both the foreground and background parts, and can reduce the generation of ghost images.

[0285] In addition, the input of time domain fusion can also include noise variance, such as the noise variance is expressed as:

[0286] Var[Noisy t ]=Shot t ·Luminance(Noisy t )+Read t

[0287] Among them, Luminance is the Noisy after decorrelation t Brightness, such as the average of multiple color channels, Shot t represents shot noise, Read t Represents the read noise. The noise variance can be represented as a single-channel image, and represents an estimate of the noise variance in the input noise map. For the time-domain fusion CNN, the noise level can be determined so that the task of fusing images can naturally adapt to different noise levels.

[0288] like Figure 11 As shown, the fused CNN can be used to determine the Noisy after decorrelation by combining the noise variance t and Fused after decorrelation t-1 The corresponding weights can be obtained through a large number of sample training, and the noise variance and the decorrelated Noisy t and Fused after decorrelation t-1 As the input of the fused CNN, the decorrelated Noisy t and Fused after decorrelation t-1The corresponding weights, and the noise variance can be used to fuse the CNN to determine the noise level, so as to obtain the de-correlated Noisy t and the de-correlated Fused t-1 and assign the corresponding weights, that is, the fusion weights Weights Figure 11 as shown in t . Then, according to the corresponding weights of the de-correlated Noisy t and the de-correlated Fused t-1 output by the fusion CNN, fuse the de-correlated Noisy t and the de-correlated Fused t-1 , and the fused image is expressed as:

[0289] Fused t = Fused t-1 ·(1 - Weights t ) + Noisy t ·Weights t .

[0290] The fusion CNN can include multiple layers of convolutional operators and ReLU non-linear operators. The final output is two channels activated by sigmoid or softmax to ensure the convexity of the image fusion equation. The goal of the fusion stage is to utilize the inherent temporal correlation in natural videos to minimize the noise present in the images while avoiding temporal artifacts (such as ghosting), so as to retain as much structure and detail in the images as possible. In a static background, the output fused image is the time average that can best reduce the noise in the fused image, and the final predicted weight is close to zero. Conversely, for a moving foreground, the output weight is close to 1, and the fused image will be closer to the current noisy frame but contain less noise. In the fusion result, the static background can gradually converge to the ideal frame average result, while the dynamic foreground directly comes from the current noisy frame. Fusion can minimize noise while avoiding ghosting.

[0291] III. Denoising

[0292] Among them, denoising can be achieved through filters, CNNs, etc. That is, the denoising model mentioned in the embodiments of the present application can include filters or CNNs, etc. Exemplarily, the present application can be implemented through a CNN. For the convenience of distinction, the CNN used for denoising is called the denoising CNN, and the denoising CNN can include multiple convolutional layers and ReLU activation functions, etc.

[0293] Specifically, after the aforementioned de-correlation transformation, the noise level of the obtained fused image Fused t has decreased, and the fused image Fused tAs the input of the denoising CNN, the denoised image Denoised is output t .

[0294] In addition, optionally, Noisy t can also be used as the input of the denoising CNN, so as to output Denoised t .

[0295] Optionally, the variance of the fused image can also be calculated, and then both the variance of the fused image and Fused t are used as the input of the denoising CNN to output the denoised image.

[0296] For example, as Figure 12 shown, the variance of the fused image Var[Fused t , Noisy t and Fused t are all used as the input of the denoising CNN to output the denoised image Denoised t . For example, the weights Weights t mentioned above can be combined, that is, the weights corresponding to Noisy t and the weights corresponding to Fused t-1 are used to calculate the variance of each pixel point to obtain the variance of the fused image. For example, the variance of the fused image can also be expressed as:

[0297]

[0298] In this formula, the fusion weight Weights t is less than 1, which makes the variance of the fused image decrease as the number of denoised frames increases, that is, it means that the noise in the fused image is also getting smaller and smaller.

[0299] Therefore, in the embodiments of the present application, denoising can be performed through the denoising CNN. By inputting the variance of the fused image into the denoising CNN, the denoising CNN can know the noise level, and then achieve a better denoising effect.

[0300] IV. Refinement processing

[0301] Due to the defects of the denoising operation itself, the spatial denoising stage is very likely to cause over-smoothing. This phenomenon is more serious especially in the case of low signal-to-noise ratio or when the computational complexity constraint of the algorithm is particularly strict. After obtaining the denoised image, to avoid the loss of information in the denoised image caused by over-smoothing, the inherent limitations of the denoising algorithm can be overcome by fusing the information of the fused image and the denoised image. For example, Fused t and Denoised t can be continuously fused to fuse the details in Fused t into Denoised through the fusion operationt , making the details included in Denoised t richer and improving the image quality.

[0302] For example, as shown in Figure 13 , the refinement process can also be implemented by direct fusion or CNN. Here, taking the implementation of the refinement process using CNN as an example, for ease of distinction, this CNN is called the refinement CNN. Regarding Var[Fused t , Fused t and Denoised t are all used as inputs to the refinement CNN, and the weights corresponding to Fused t and Denoised t are output respectively. When determining the weights corresponding to Fused t and Denoised t respectively, the noise level corresponding to the variance of the fused image can be referred to, and corresponding weights are assigned to high-frequency and low-frequency noises.

[0303] Generally, in the refinement process, the weight value of high frequency is smaller, and the weight value of low frequency is higher. Generally, low frequency contains the main structure of the image, while high frequency contains the detailed information of the image. Therefore, in order to prevent the denoised image from being overly smoothed, the main structure of the image included in the low frequency can be referred to more, so that the structure in the image is clearer. After determining the weights corresponding to Fused t and Denoised t respectively, Fused t and Denoised t can be fused. For example, the refined output can be expressed as:

[0304]

[0305] is the weight corresponding to Denoised t , is the weight corresponding to Fused t .

[0306] It can be understood that the refinement weights can be used to extract high-frequency information from the fused image rich in details but still a bit noisy and transfer the high-frequency detailed information to the denoised image that is noise-free but may be overly smoothed. As shown in Figure 13 , even if the CNN has a very low complexity, this formula can provide high-quality results.

[0307] The form of the refined CNN can be diverse. In the embodiments of the present application, it can be composed of multiple layers of convolution and the RELU activation function. To ensure that the weight distribution of the refined formula is between [0,1], the final predicted weights are activated by sigmoid or softmax.

[0308] Therefore, in the embodiments of the present application, by fusing Fused t and Denoised t to avoid the over-smoothing of Denoised t a denoised image with richer details is obtained. Moreover, by assigning lower weights to high frequencies and higher weights to low frequencies, the effect of smoothing the noise included in the high-frequency information can be achieved, and a denoised image with a better denoising effect is obtained.

[0309] V. Inverse transformation

[0310] Among them, the inverse transformation is the inverse transformation of the aforementioned decorrelation transformation. If the color decorrelation transformation is performed previously, after obtaining the denoised image, an inverse transformation needs to be performed on the denoised image. If the wavelet transformation is performed for the aforementioned decorrelation transformation, an inverse wavelet transformation needs to be performed after obtaining the denoised image.

[0311] It should be understood that if the decorrelation transformation is not performed, there is no need to perform the inverse transformation.

[0312] In the video processing method provided by the present application, if the aforementioned denoising step is not performed, the inverse transformation can be directly performed on the obtained fused image Fused t Hereinafter, an example of performing an inverse transformation on the denoised image will be used for illustrative purposes. In some scenarios, it can also be replaced by performing an inverse transformation on Fused t The transformation methods are similar, and the only difference is the input image. The present application will not elaborate on this.

[0313] For example, as Figure 14 shown, before denoising, a color decorrelation operation is performed on the input image through color transformation to obtain the luminance component and the chrominance component. After subsequent denoising, a point-by-point convolution operation is performed on the denoised image using the inverse color transformation matrix, and a color image, that is, a new denoised image, is output. It is equivalent to separating the colors of each pixel point in the aforementioned decorrelation step, and restoring the color of the denoised image through the inverse transformation to obtain a denoised image with color.

[0314] Among them, the color transformation matrix and the inverse color transformation matrix are inverse matrices of each other, that is, the product of the color transformation matrix and the inverse color transformation matrix is the identity matrix, thereby ensuring that the color of the denoised image can be accurately restored. For example, the color matrix can be represented as C forward, since the color transformation is linear and defined by a matrix, the inverse transformation matrix C of the color transformation can be derived by inverting the matrix. inverse , so at the end of the algorithm process, the inverse transformation matrix can be used for point-by-point convolution to reconstruct the color information. Since the color transformations in both the forward and inverse directions are linear convolution operators, the matrix weights of the transformation can be learned in a data-driven manner to adapt to specific denoising tasks and input data. To ensure the invertibility of the matrix during the training process, C forward ·C inverse = I can be added as a constraint term, where I is the identity matrix, to ensure that the application of the forward and inverse operators can accurately reconstruct.

[0315] For another example, as Figure 15 shown, before denoising, the input image is wavelet-transformed using wavelet coefficients. After denoising, an inverse wavelet kernel is used to perform deconvolution operations on the denoised high-frequency and low-frequency components to obtain an output image with fused low-frequency and high-frequency components, that is, the new denoised image, restoring the main structure and details of the denoised image. The inverse wavelet kernel is obtained by taking the outer product of the inverse wavelet coefficients, and the inverse wavelet coefficients and the wavelet coefficients are inverse matrices of each other, that is, the product of the inverse wavelet coefficients and the wavelet coefficients is the identity matrix, thus ensuring the accurate restoration of the main structure and details of the image.

[0316] For example, the decorrelation spatial transformation is also a linear operation, so it can be implemented using a convolution operator. The wavelet coefficients can be represented as W forward , and the inverse wavelet coefficients can be represented as W inverse , W forward and W inverse are inverse matrices of each other, that is, W forward ·W inverse = I, where I is the identity matrix. Similar to the color transformation, the weights of the inverse transformation of the decorrelation spatial frequency transformation can also be learned in a data-driven manner. In the embodiments of the present application, when updating W inverse , W forward ·W inverse = I can be used as a constraint, so that the inverse wavelet transformation can accurately restore the high-frequency and low-frequency components to obtain a clearer denoised image.

[0317] If both the decorrelation color transformation and the wavelet transformation are performed before denoising, then during the inverse transformation, an inverse color transformation and an inverse wavelet transformation need to be performed, so that the denoised image can restore color and frequency.

[0318] Therefore, in the embodiments of the present application, after performing decorrelation transformation, temporal fusion, denoising, and refinement steps, a denoised image with excellent denoising effect can be obtained. It can be understood that by utilizing the correlation between adjacent frames in the video data in the time dimension, the noise in the image is reduced, and at the same time, ghosting can be avoided, resulting in an image with a better denoising effect.

[0319] In addition, to further improve the image quality, multi-scale processing can also be performed on the current frame and the first fused image. That is, after performing the decorrelation operation on the current frame and the first fused image, the image after the decorrelation operation is downsampled one or more times to obtain images of multiple scales, and then the images of each scale are respectively processed by the aforementioned steps 702-705, such as the aforementioned temporal fusion, denoising, and upsampling operations, etc. The images of each scale are iteratively denoised to obtain a denoised image with a better denoising effect. Each processing is performed by combining the fused image or denoised image obtained from the previous scale (i.e., a smaller scale) for the current iteration, so as to fuse the fused image with a smaller scale, which is equivalent to achieving the effect of smoothing the noise.

[0320] Specifically, the process of the multi-scale processing may include: performing at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image; performing denoising processing on at least one downsampled frame and at least one downsampled fused image to obtain a multi-scale fused image; fusing the denoised image and the multi-scale fused image to obtain an updated denoised image.

[0321] Exemplarily, taking the denoising process of an image of one scale (such as the first downsampled fusion image and the first downsampled frame of the same scale, which are images obtained by downsampling the current frame and the first fusion image once) as an example, according to the features extracted from the first downsampled frame and the features extracted from the first downsampled fusion image, determine the weight corresponding to the first downsampled frame to obtain the fifth fusion weight, and determine the weight corresponding to the first downsampled fusion image to obtain the sixth fusion weight. The first downsampled fusion image and the first downsampled frame have the same size, which can also be understood as the first downsampled fusion image and the first downsampled frame have the same number of downsampling times; and the methods for calculating the fifth fusion weight and the sixth fusion weight are similar to the methods for calculating the first fusion weight and the second fusion weight before, except that the scales of the images are different, which will not be elaborated here. Subsequently, determine the weight corresponding to the second downsampled fusion image according to the features extracted from the second downsampled fusion image to obtain the seventh fusion weight. The second downsampled fusion image incorporates the information of at least one downsampled frame with a scale smaller than the first downsampled frame; fuse the first downsampled frame, the first downsampled fusion image, and the second fusion image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain the third downsampled fusion image. The upsampled image of the third downsampled image is used to fuse with at least one downsampled frame with a scale larger than the first downsampled frame; denoise the third downsampled fusion image to obtain the first downsampled denoised image; upsample the first downsampled denoised image to obtain the upsampled denoised image. The upsampled image is used to combine with the fusion image of the same scale as the upsampled image for denoising to obtain an image with a scale larger than the first downsampled frame.

[0322] Specifically, for example, as Figure 16 shown, wavelet transform can be selected for downsampling. After decorrelation, Noisy t and Fused t-1 are respectively downsampled once or multiple times to obtain downsampled images of multiple scales. As Figure 16 shown, in the decorrelation operation, a wavelet transform is performed once, and an image with a size of H×W×C is downsampled to an image with a size of H / 2×W / 2×4C, that is, the downsampled frame or the downsampled fusion image obtained by the first downsampling; in the next downsampling, the image with a size of H / 2×W / 2×4C is downsampled to an image with a size of H / 4×W / 4×16C, that is, the downsampled frame or the downsampled fusion image obtained by two downsamplings, and so on for subsequent downsamplings.

[0323] In the processing of one of the scales, as Figure 16 shown in scale 2, temporal fusion can be performed on the image of this scale, that is, fuse the decorrelated Noisy t and Fused t-1The steps of temporal domain fusion for the image obtained after downsampling are similar to those of the temporal domain fusion in Step 2 described above, and will not be elaborated here.

[0324] Meanwhile, when performing temporal domain fusion on the images of each scale, it is also possible to fuse the upsampled images of the images that have undergone temporal domain fusion in the next scale process. For example, Figure 16 as shown in t Noisy and Fused t-1 after one downsampling can be upsampled during the processing of Scale 2 to obtain upsampled images. During the processing of Scale 1, when performing temporal domain fusion, in addition to fusing the Noisy t and Fused t-1 after decorrelation transform, it is also possible to fuse the upsampled images obtained during the processing of Scale 2. For example, the upsampled images and the Noisy t after decorrelation transform can be fused to obtain a new Noisy t . Or, the upsampled images can be used as the input to the fusion CNN to output the weights corresponding to the upsampled images, and then the upsampled images, Noisy t and Fused t-1 etc. can be fused according to the weights.

[0325] In the temporal domain fusion step of Scale 2, in addition to fusing the Noisy t and Fused t-1 after one downsampling, it is also possible to fuse the upsampled images obtained in Scale 3, and so on.

[0326] During the processing of Scale 2, denoising can also be performed. After performing temporal domain fusion to obtain the Scale 2 fused image, denoising is performed on this Scale 2 fused image. The denoising process is similar to the denoising process in Step 3 described above, except that the scale of the input image is different, and will not be elaborated here. After denoising, inverse wavelet transform is performed on the denoised image, which is equivalent to upsampling, and then the image obtained by inverse wavelet transform is used as the input for the denoising step of Scale 1 to obtain the denoised image.

[0327] Similarly, during the denoising process of Scale 2, it is also possible to denoise the image obtained by inverse wavelet transform after denoising in Scale 3. Taking the denoising process of Scale 1 as an example, the input can include the fused image Fused t output from the temporal domain fusion step in Scale 1 and the upsampled denoised image output after inverse wavelet transform in Scale 2. The upsampled denoised image, Fused t and the Noisy t can all be used as the input to the denoising CNN to output the denoised image. This is equivalent to fusing the upsampled denoised image, Fused tand as Noisy t , thus smoothing the noise in the image and obtaining a denoised image with less noise.

[0328] Therefore, in the embodiments of the present application, the current frame can be downsampled multiple times to obtain images of multiple scales, and then the images of multiple scales are iteratively processed. Each iterative processing process includes a time-domain fusion step and a denoising step. That is, the denoising of each scale of the image utilizes the correlation between adjacent frames in the video data to achieve denoising. Images with better denoising effects can be obtained at each scale, improving the denoising effect and obtaining a denoised image with less noise, clearer, and fewer ghosts. Moreover, denoising is completed at a lower resolution, and the denoising task is separated into multiple specific stages, which can greatly simplify the task, thereby generating high-quality output with the least number of operations. Through special processing stages such as time-domain fusion, spatial denoising, and refinement, the video correlation based on the temporal, spatial, and spatio-temporal dimensions of the video is significantly utilized. This design enables almost no damage to the image quality of the final output even when the complexity is greatly reduced.

[0329] The above has introduced the process of the video denoising method provided by the present application in detail. Next, based on the process of the foregoing video denoising method, the structure of the video denoising device that executes this process will be introduced.

[0330] Refer to Figure 17 , a schematic structural diagram of a video denoising device provided by the present application is as follows.

[0331] An acquisition module 1701, configured to acquire a current frame and a first fused image, where the current frame is any frame image after the first frame arranged in a preset order in the video data, and the first fused image includes information of at least one frame adjacent to the current frame in the video data according to a preset order;

[0332] A time-domain fusion module 1702, configured to extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature;

[0333] The time-domain fusion module 1702 is further configured to determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The weight of the foreground in the current frame is not less than the weight of the foreground in the first fused image, and the weight of the background in the current frame is not greater than the weight of the background in the first fused image;

[0334] The time-domain fusion module 1702 is further configured to fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image;

[0335] A denoising module 1703, configured to denoise the second fused image to obtain a denoised image.

[0336] In a possible implementation, the video denoising device may further include:

[0337] A refinement module 1704, configured to fuse the second fused image and the denoised image after denoising the second fused image to obtain an updated denoised image.

[0338] In a possible implementation, the denoising module 1703 is specifically configured to: extract features from the second fused image to obtain a third feature; extract features from the denoised image to obtain a fourth feature; determine a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature, where the third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value have a negative correlation; fuse the third feature and the fourth feature according to the third fusion weight and the fourth fusion weight to obtain an updated denoised image.

[0339] In a possible implementation, the video denoising device may further include: a color correlation transformation module 1705, configured to perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component before extracting features from the current frame and extracting features from the first fused image. The first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fused image to obtain a second chrominance component and a second luminance component. The second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel;

[0340] The time-domain fusion module 1702 is specifically configured to extract features from the new current frame to obtain a first feature, and extract features from the new first fused image to obtain a second feature.

[0341] In a possible implementation, the video denoising device may further include:

[0342] An inverse color correlation transformation module 1706, configured to perform color transformation on the denoised image through an inverse color transformation matrix after denoising the second fused image to obtain an updated denoised image. The inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0343] In a possible implementation, the video denoising device may further include:

[0344] A de-frequency-related transformation module 1707, configured to perform a wavelet transform on a current frame by using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component before the time-domain fusion module 1702 extracts features from the current frame and extracts features from a first fused image. The first low-frequency component and the first high-frequency component form a new current frame, and perform a wavelet transform on the first fused image to obtain a second low-frequency component and a second high-frequency component. The second low-frequency component and the second high-frequency component form a new first fused image. The wavelet coefficients are preset coefficients or obtained by training at least one convolution kernel.

[0345] In a possible implementation manner, the video denoising device may further include:

[0346] An inverse frequency-related transformation module 1708, configured to perform an inverse wavelet transform on a denoised image by using inverse wavelet coefficients to obtain an updated denoised image after the denoising module 1703 denoises a second fused image. The inverse wavelet coefficients are an inverse matrix of the wavelet coefficients.

[0347] In a possible implementation manner, the time-domain fusion module 1702 is specifically configured to: calculate shot noise and read noise according to shooting parameters of a device used for shooting video data; determine a first fusion weight corresponding to the current frame by combining the shot noise, the read noise, a first feature, and a second feature.

[0348] In a possible implementation manner, the time-domain fusion module 1702 is specifically configured to: calculate a noise variance of each pixel point in the current frame according to the shot noise and the read noise; extract a fifth feature from the noise variance; determine a first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature.

[0349] In a possible implementation manner, the denoising module 1703 is specifically configured to: denoise a second fused image by combining a first fusion weight, a second fusion weight, a first fused image, and the current frame to obtain a denoised image.

[0350] In a possible implementation manner, the denoising module 1703 is specifically configured to: combine a first fusion weight, a second fusion weight, a first fused image, and the current frame to calculate a variance of each pixel point in the second fused image to obtain a fused image variance; use the fused image variance and the second fused image as inputs of a denoising model, and output a denoised image, where the denoising model is used to remove noise in the input image.

[0351] In a possible implementation manner, the denoising module 1703 is specifically configured to use the current frame, the fused image variance, and the second fused image as inputs of a denoising model, and output a denoised image.

[0352] In a possible implementation manner, the video denoising device may further include: a downsampling module 1709;

[0353] The downsampling module 1709 is configured to perform at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image;

[0354] The denoising module 1703 is further configured to perform denoising processing on at least one downsampled frame and at least one downsampled fused image to obtain a multi-scale fused image;

[0355] The denoising module 1703 is further configured to fuse the denoised image and the multi-scale fused image to obtain an updated denoised image.

[0356] In a possible implementation manner, any denoising process in the process of the denoising module 1703 performing denoising processing on at least one downsampled frame and at least one downsampled fused image may include: determining the weight corresponding to the first downsampled frame according to the features extracted from the first downsampled frame and the features extracted from the first downsampled fused image to obtain a fifth fusion weight, and determining the weight corresponding to the first downsampled fused image to obtain a sixth fusion weight. The first downsampled frame is any one of the at least one downsampled frame, and the first downsampled fused image is a frame with the same scale as the first downsampled frame among the at least one downsampled fused images; determining the weight corresponding to the second downsampled fused image according to the features extracted from the second downsampled fused image to obtain a seventh fusion weight. The second downsampled fused image incorporates the information of the images with scales smaller than the first downsampled frame among the at least one downsampled frame; fusing the first downsampled frame, the first downsampled fused image, and the second fused image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain a third downsampled fused image. The upsampled image of the third downsampled image is used to fuse with the images with scales larger than the first downsampled frame among the at least one downsampled frame; denoising the third downsampled fused image to obtain a first downsampled denoised image; upsampling the first downsampled denoised image to obtain an upsampled denoised image. The upsampled image is used to combine with the fused image with the same scale as the upsampled image to perform denoising to obtain an image with a scale larger than the first downsampled frame.

[0357] In a possible implementation manner, the downsampling module 1709 is specifically configured to perform at least one wavelet transform on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image.

[0358] The present application further provides a video processing device for executing the foregoing Figure 7A corresponding method steps. Refer to Figure 18, a structural schematic diagram of a video processing device provided by this application is described as follows.

[0359] An acquisition module 1801, configured to acquire a current frame and a first fused image, where the current frame is any frame image after the first frame arranged in a preset order in video data, and the first fused image includes information of at least one frame adjacent to the current frame in the video data according to the preset order;

[0360] A time-domain fusion module 1802, configured to extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature;

[0361] The time-domain fusion module 1802 is further configured to determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image;

[0362] The time-domain fusion module 1802 is further configured to fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image.

[0363] In a possible implementation manner, the above device may further include:

[0364] A de-color correlation transform module 1803, configured to perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component before the time-domain fusion module 1802 extracts features from the current frame and extracts features from the first fused image. The first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fused image to obtain a second chrominance component and a second luminance component. The second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel;

[0365] The time-domain fusion module 1802 is specifically configured to extract features from the new current frame to obtain a first feature, and extract features from the new first fused image to obtain a second feature.

[0366] In a possible implementation manner, the above device may further include: an inverse color correlation transform module 1804, configured to perform color transformation on the second fused image through an inverse color transformation matrix to obtain an updated second fused image, and the inverse color transformation matrix is the inverse matrix of the color transformation matrix.

[0367] In a possible implementation, the above device may further include: a de-frequency correlation transformation module 1805, configured to perform wavelet transformation on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component before the time-domain fusion module 1802 extracts features from the current frame and extracts features from the first fused image. The first low-frequency component and the first high-frequency component form a new current frame, and perform wavelet transformation on the first fused image to obtain a second low-frequency component and a second high-frequency component. The second low-frequency component and the second high-frequency component form a new first fused image. The wavelet coefficients are preset coefficients or obtained by training at least one convolutional kernel.

[0368] In a possible implementation, the above device may further include: an inverse frequency correlation transformation module 1806, configured to perform inverse wavelet transformation on the second fused image through inverse wavelet coefficients to obtain an updated second fused image. The inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

[0369] In a possible implementation, the time-domain fusion module 1802 is specifically configured to: calculate shot noise and read noise according to the shooting parameters of the device used to shoot the video data; determine the first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature.

[0370] In a possible implementation, the time-domain fusion module 1802 is specifically configured to: calculate the noise variance of each pixel point in the current frame according to the shot noise and the read noise; extract a fifth feature from the noise variance; determine the first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature.

[0371] In a possible implementation, the above device may further include: a downsampling module 1807, configured to perform at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. At least one first downsampled image is obtained by performing at least one downsampling on the current frame, and at least one second downsampled image is obtained by performing at least one downsampling on the first fused image;

[0372] The time-domain fusion module 1802 is further configured to fuse each frame in the at least one downsampled frame and the downsampled fused image of the same scale to obtain a multi-scale fused image;

[0373] The time-domain fusion module 1802 is further configured to fuse the second fused image and the multi-scale fused image to obtain an updated second fused image.

[0374] In a possible implementation, the time-domain fusion module 1802 fuses any one of at least one downsampled frame and a downsampled fusion image of the same scale, which may include: determining the weight corresponding to the first downsampled frame based on the features extracted from the first downsampled frame and the features extracted from the first downsampled fusion image to obtain a fifth fusion weight, and determining the weight corresponding to the first downsampled fusion image to obtain a sixth fusion weight, where the first downsampled frame is any one of at least one downsampled frame, and the first downsampled fusion image is a frame in at least one downsampled fusion image with the same scale as the first downsampled frame; determining the weight corresponding to the second downsampled fusion image based on the features extracted from the second downsampled fusion image to obtain a seventh fusion weight, where the second downsampled fusion image incorporates the information of images in at least one downsampled frame with a scale smaller than the first downsampled frame; fusing the first downsampled frame, the first downsampled fusion image, and the second downsampled fusion image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight to obtain a third downsampled fusion image, and the upsampled image of the third downsampled image is used to fuse with images in at least one downsampled frame with a scale larger than the first downsampled frame.

[0375] In a possible implementation, the downsampling module 1807 is specifically configured to perform at least one wavelet transform on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image.

[0376] Please refer to Figure 19 , a schematic structural diagram of another video denoising device provided by the present application is described as follows.

[0377] The video denoising device may include a processor 1901 and a memory 1902. The processor 1901 and the memory 1902 are interconnected by a line. Among them, program instructions and data are stored in the memory 1902.

[0378] The memory 1902 stores the program instructions and data corresponding to the steps in the foregoing Figures 4 - 16 .

[0379] The processor 1901 is configured to execute the method steps performed by the video denoising device shown in any one of the foregoing Figures 4 - 16 embodiments.

[0380] Optionally, the video denoising device may further include a transceiver 1903 for receiving or sending data.

[0381] In an embodiment of the present application, a computer-readable storage medium is further provided. A program for generating the vehicle driving speed is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps in the method described in the foregoing Figures 4 - 14 embodiments.

[0382] Optionally, the foregoing Figure 19 The video denoising device shown in is a chip.

[0383] Please refer to Figure 20 The structural schematic diagram of another video processing device provided by this application is described as follows.

[0384] The video processing device may include a processor 2001 and a memory 2002. The processor 2001 and the memory 2002 are interconnected by a line. Among them, program instructions and data are stored in the memory 2002.

[0385] The program instructions and data corresponding to the steps in the foregoing Figure 7A are stored in the memory 2002.

[0386] The processor 2001 is used to execute the method steps executed by the video processing device shown in any one of the foregoing Figure 7A embodiments.

[0387] Optionally, the video processing device may further include a transceiver 2003 for receiving or sending data.

[0388] In an embodiment of this application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a program for generating a vehicle driving speed. When it runs on a computer, it causes the computer to execute the steps in the method described in the foregoing Figure 7A embodiment.

[0389] Optionally, the foregoing Figure 20 The video processing device shown in is a chip.

[0390] In an embodiment of this application, a video denoising device is further provided. The video denoising device may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the method steps executed by the video denoising device shown in any one of the foregoing Figures 4 - 14 embodiments.

[0391] In an embodiment of this application, a video processing device is further provided. The video processing device may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the method steps executed by the video processing device shown in any one of the foregoing Figure 7A embodiments.

[0392] The embodiment of the present application further provides a digital processing chip. Circuits for implementing the foregoing processor 1901 / 2001 or the functions of the processor 1901 / 2001 and one or more interfaces are integrated in the digital processing chip. When a memory is integrated in the digital processing chip, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When a memory is not integrated in the digital processing chip, it can be connected to an external memory through a communication interface. The digital processing chip implements the actions performed by the video denoising device or the video processing device in the foregoing embodiment according to the program code stored in the external memory.

[0393] The embodiment of the present application further provides a computer program product. When it runs on a computer, it causes the computer to execute the steps performed by the video denoising device or the video processing device in the method described in the foregoing Figures 4 - 16 embodiment as shown.

[0394] The video denoising device or the video processing device provided by the embodiment of the present application may be a chip. The chip may include: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit to cause the chip in the server to execute the Figures 4 - 16 video denoising method or the video processing method described in the embodiment as shown. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0395] Specifically, the foregoing processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0396] Exemplarily, please refer to Figure 21 , Figure 21 which is a schematic structural diagram of a chip provided by an embodiment of the present application. The chip may be embodied as a neural network processor NPU 210. The NPU 210 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are allocated by the Host CPU. The core part of the NPU is the arithmetic circuit 2103, and the arithmetic circuit 2103 is controlled by the controller 2104 to extract matrix data from the memory and perform multiplication operations.

[0397] In some implementations, the arithmetic circuit 2103 includes multiple processing units (process engine, PE) inside. In some implementations, the arithmetic circuit 2103 is a two-dimensional systolic array. The arithmetic circuit 2103 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2103 is a general matrix processor.

[0398] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of the matrix B from the weight memory 2102 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of the matrix A from the input memory 2101 and performs matrix operations with the matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 2108.

[0399] The unified memory 2106 is used to store input data and output data. The weight data is directly transferred through the direct memory access controller (DMAC) 2105, and the DMAC transfers it to the weight memory 2102. The input data is also transferred to the unified memory 2106 through the DMAC.

[0400] The bus interface unit (BIU) 2110 is used for the interaction between the AXI bus, the DMAC, and the instruction fetch buffer (IFB) 2109.

[0401] The bus interface unit 2110 (bus interface unit, BIU) is used for the instruction fetch buffer 2109 to obtain instructions from the external memory, and also for the storage unit access controller 2105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0402] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 2106, or transfer the weight data to the weight memory 2102, or transfer the input data to the input memory 2101.

[0403] The vector calculation unit 2107 includes multiple arithmetic processing units, which further process the output of the arithmetic circuit as needed, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as batch normalization, pixel-level summation, upsampling of the feature plane, etc.

[0404] In some implementations, the vector calculation unit 2107 can store the processed output vector in the unified memory 2106. For example, the vector calculation unit 2107 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 2103, such as linear interpolation of the feature plane extracted by the convolutional layer, and for another example, the vector of the accumulated value to generate the activation value. In some implementations, the vector calculation unit 2107 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 2103, such as for use in subsequent layers in the neural network.

[0405] The instruction fetch buffer 2109 connected to the controller 2104 is used to store the instructions used by the controller 2104;

[0406] The unified memory 2106, the input memory 2101, the weight memory 2102, and the fetch memory 2109 are all On-Chip memories. The external memory is private to the NPU hardware architecture.

[0407] Among them, the operations of each layer in the recurrent neural network can be executed by the arithmetic circuit 2103 or the vector computing unit 2107.

[0408] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above Figures 4 - 16 method.

[0409] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0410] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0411] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0412] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0413] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0414] Finally, it should be noted that the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A video denoising method, characterized in that, Including: Obtain a current frame and a first fused image, where the current frame is any frame image in the video data that is arranged after the first frame in a preset order, and the first fused image includes information of at least one frame adjacent to the current frame in the video data in the preset order; Extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature; Determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fused image. The weight corresponding to the foreground in the current frame is not less than the weight corresponding to the foreground in the first fused image, and the weight corresponding to the background in the current frame is not greater than the weight corresponding to the background in the first fused image; Fuse the current frame and the first fused image according to the first fusion weight and the second fusion weight to obtain a second fused image; Denoise the second fused image to obtain a denoised image; Perform at least one downsampling on the current frame and the first fused image to obtain at least one downsampled frame and at least one downsampled fused image, and the scales of the images obtained by each downsampling are different. The at least one downsampled frame is obtained by performing at least one downsampling on the current frame, and the at least one downsampled fused image is obtained by performing at least one downsampling on the first fused image; Perform denoising processing on the at least one downsampled frame and the at least one downsampled fused image to obtain a multi-scale fused image; Fuse the denoised image and the multi-scale fused image to obtain an updated denoised image.

2. The method according to claim 1, wherein After denoising the second fused image, the method further includes: Fuse the second fused image and the denoised image to obtain an updated denoised image.

3. The method according to claim 2, wherein The fusing the second fused image and the denoised image includes: Extract features from the second fused image to obtain a third feature; Extract features from the denoised image to obtain a fourth feature; Determine a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature. The third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value have a negative correlation; Fuse the third feature and the fourth feature according to the third fusion weight and the fourth fusion weight to obtain an updated denoised image.

4. The method according to claim 1, wherein Before extracting features from the current frame and extracting features from the first fused image, the method further includes: Perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component. The first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fused image to obtain a second chrominance component and a second luminance component. The second chrominance component and the second luminance component form a new first fused image. The color transformation matrix is a preset matrix or obtained by training at least one convolutional kernel; The extracting features from the current frame to obtain a first feature and extracting features from the first fused image to obtain a second feature includes: Extract features from the new current frame to obtain the first feature, and extract features from the new first fused image to obtain the second feature.

5. The method according to claim 4, wherein After denoising the second fused image, the method further includes: Perform color transformation on the denoised image through an inverse color transformation matrix to obtain an updated denoised image. The inverse color transformation matrix is the inverse matrix of the color transformation matrix.

6. The method according to any one of claims 1-5, characterized in that, Before extracting features from the current frame and extracting features from the first fused image, the method further includes: Perform wavelet transform on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component. The first low-frequency component and the first high-frequency component form a new current frame, and perform wavelet transform on the first fused image to obtain a second low-frequency component and a second high-frequency component. The second low-frequency component and the second high-frequency component form a new first fused image. The wavelet coefficients are preset coefficients or obtained by training at least one convolutional kernel.

7. The method according to claim 6, characterized in that, After denoising the second fused image, the method further includes: Perform inverse wavelet transform on the denoised image through inverse wavelet coefficients to obtain an updated denoised image. The inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

8. The method according to any one of claims 1-5, characterized in that The determining the first fusion weight corresponding to the current frame according to the first feature and the second feature includes: Calculate shot noise and read noise according to the shooting parameters of the device used to shoot the video data; Determine the first fusion weight corresponding to the current frame in combination with the shot noise, the read noise, the first feature and the second feature.

9. The method according to claim 8, characterized in that The combining the shot noise, the read noise, the first feature and the second feature to determine the first fusion weight corresponding to the current frame includes: Calculate the noise variance of each pixel point in the current frame according to the shot noise and the read noise; Extract a fifth feature from the noise variance; Determine the first fusion weight corresponding to the current frame in combination with the fifth feature, the first feature and the second feature.

10. The method according to any one of claims 1-5, characterized in that, The denoising the second fused image includes: Denoise the second fused image in combination with the first fusion weight, the second fusion weight, the first fused image and the current frame to obtain the denoised image.

11. The method according to claim 10, wherein The combining the first fusion weight, the second fusion weight, the first fused image and the current frame to denoise the second fused image includes: Calculate the variance of each pixel in the second fused image by combining the first fusion weight, the second fusion weight, the first fused image, and the current frame, to obtain the fused image variance; Use the fused image variance and the second fused image as the input of a denoising model to output the denoised image, where the denoising model is used to remove noise from the input image.

12. The method according to claim 11, wherein The step of using the fused image variance and the second fused image as the input of the denoising model includes: Use the current frame, the fused image variance, and the second fused image as the input of the denoising model to output the denoised image.

13. The method according to any one of claims 1 to 5, characterized in that, Any one of the denoising processes in the process of denoising the at least one downsampled frame and the at least one downsampled fused image includes: Determine the weight corresponding to the first downsampled frame, to obtain the fifth fusion weight, and determine the weight corresponding to the first downsampled fused image, to obtain the sixth fusion weight, according to the features extracted from the first downsampled frame and the features extracted from the first downsampled fused image, where the first downsampled frame is any one of the at least one downsampled frame, and the first downsampled fused image is a frame in the at least one downsampled fused image with the same scale as the first downsampled frame; Determine the weight corresponding to the second downsampled fused image, to obtain the seventh fusion weight, according to the features extracted from the second downsampled fused image, where the second downsampled fused image incorporates the information of the images in the at least one downsampled frame with a scale smaller than the first downsampled frame; Fuse the first downsampled frame, the first downsampled fused image, and the second fused image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight, to obtain a third downsampled fused image, and the upsampled image of the third downsampled fused image is used to fuse with the images in the at least one downsampled frame with a scale larger than the first downsampled frame; Denoise the third downsampled fused image to obtain a first downsampled denoised image; Upsample the first downsampled denoised image to obtain an upsampled denoised image, and the upsampled denoised image is used to combine with the fused image with the same scale as the upsampled denoised image for denoising, to obtain an image with a scale larger than the first downsampled frame.

14. The method according to claim 13, wherein The step of performing at least one downsampling on the current frame and the first fused image includes: Perform at least one wavelet transform on the current frame and the first fused image to obtain the at least one downsampled frame and the at least one downsampled fused image.

15. A video denoising device, characterized in that, Includes: An acquisition module, configured to acquire a current frame and a first fused image, where the current frame is any frame image in the video data arranged after the first frame in a preset order, and the first fused image includes the information of at least one frame adjacent to the current frame in the video data in the preset order; A time-domain fusion module, configured to extract features from the current frame to obtain a first feature, and extract features from the first fused image to obtain a second feature; The time-domain fusion module is further configured to determine a first fusion weight and a second fusion weight according to the first feature and the second feature. The first fusion weight includes the weight corresponding to the current frame, and the second fusion weight includes the weight corresponding to the first fusion image. The weight of the foreground in the current frame is not less than the weight of the foreground in the first fusion image, and the weight of the background in the current frame is not greater than the weight of the background in the first fusion image; The time-domain fusion module is further configured to fuse the current frame and the first fusion image according to the first fusion weight and the second fusion weight to obtain a second fusion image; The denoising module is configured to denoise the second fusion image to obtain a denoised image; The downsampling module is configured to perform at least one downsampling on the current frame and the first fusion image to obtain at least one downsampled frame and at least one downsampled fusion image, and the scales of the images obtained by each downsampling are different. The at least one downsampled frame is obtained by performing at least one downsampling on the current frame, and the at least one downsampled fusion image is obtained by performing at least one downsampling on the first fusion image; The denoising module is further configured to perform denoising processing on the at least one downsampled frame and the at least one downsampled fusion image to obtain a multi-scale fusion image; The denoising module is further configured to fuse the denoised image and the multi-scale fusion image to obtain an updated denoised image.

16. The device according to claim 15, wherein The apparatus further includes: The refinement module is configured to fuse the second fusion image and the denoised image after denoising the second fusion image to obtain an updated denoised image.

17. The device according to claim 16, characterized in that, The denoising module is specifically configured to: Extract features from the second fusion image to obtain a third feature; Extract features from the denoised image to obtain a fourth feature; Determine a third fusion weight and a fourth fusion weight according to the third feature and the fourth feature. The third fusion weight is the weight corresponding to the third feature, and the fourth fusion weight is the weight corresponding to the fourth feature. Among the third feature and the fourth feature, the frequency of each pixel point and the corresponding weight value have a negative correlation; Fuse the third feature and the fourth feature according to the third fusion weight and the fourth fusion weight to obtain an updated denoised image.

18. The apparatus according to claim 15, wherein The apparatus further includes: a color decorrelation transformation module, configured to, before extracting features from the current frame and extracting features from the first fusion image, perform color transformation on the current frame through a color transformation matrix to obtain a first chrominance component and a first luminance component. The first chrominance component and the first luminance component form a new current frame, and perform color transformation on the first fusion image to obtain a second chrominance component and a second luminance component. The second chrominance component and the second luminance component form a new first fusion image. The color transformation matrix is a preset matrix or obtained by training at least one convolution kernel; The time-domain fusion module is specifically configured to extract features from the new current frame to obtain the first feature, and extract features from the new first fusion image to obtain the second feature.

19. The device according to claim 18, characterized in that, The device further includes: An inverse color correlation transformation module, configured to perform color transformation on the denoised image through an inverse color transformation matrix after denoising the second fusion image, to obtain an updated denoised image, where the inverse color transformation matrix is the inverse matrix of the color transformation matrix.

20. The device according to any one of claims 15-17, characterized in that, The device further includes: A de-frequency correlation transformation module, configured to perform wavelet transformation on the current frame using wavelet coefficients to obtain a first low-frequency component and a first high-frequency component before the time-domain fusion module extracts features from the current frame and extracts features from the first fusion image, where the first low-frequency component and the first high-frequency component form a new current frame, and perform wavelet transformation on the first fusion image to obtain a second low-frequency component and a second high-frequency component, where the second low-frequency component and the second high-frequency component form a new first fusion image, and the wavelet coefficients are preset coefficients or obtained by training at least one convolution kernel.

21. The device according to claim 20, wherein The device further includes: An inverse frequency correlation transformation module, configured to perform inverse wavelet transformation on the denoised image through inverse wavelet coefficients after the denoising module denoises the second fusion image, to obtain an updated denoised image, where the inverse wavelet coefficients are the inverse matrix of the wavelet coefficients.

22. The device according to any one of claims 15-17, characterized in that, The time-domain fusion module is specifically configured to: Calculate shot noise and read noise according to the shooting parameters of the device used to shoot the video data; Determine a first fusion weight corresponding to the current frame by combining the shot noise, the read noise, the first feature, and the second feature.

23. The device according to claim 22, characterized in that, The time-domain fusion module is specifically configured to: Calculate the noise variance of each pixel point in the current frame according to the shot noise and the read noise; Extract a fifth feature from the noise variance; Determine a first fusion weight corresponding to the current frame by combining the fifth feature, the first feature, and the second feature.

24. The device according to any one of claims 15-17, characterized in that, The denoising module is specifically configured to: Denoise the second fusion image by combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame, to obtain the denoised image.

25. The device according to claim 24, characterized in that, The denoising module is specifically configured to: Calculate the variance of each pixel point in the second fusion image by combining the first fusion weight, the second fusion weight, the first fusion image, and the current frame, to obtain a fusion map variance; Use the fusion map variance and the second fusion image as inputs to a denoising model, and output the denoised image, where the denoising model is used to remove noise in the input image.

26. The device according to claim 25, wherein The denoising module is specifically configured to use the current frame, the fusion map variance, and the second fusion image as inputs to a denoising model, and output the denoised image.

27. The device according to any one of claims 15-17, characterized in that, Any one of the denoising processes during which the denoising module denoises the at least one downsampled frame and the at least one downsampled fusion image includes: Determine the weight corresponding to the first downsampled frame based on the features extracted from the first downsampled frame and the features extracted from the first downsampled fused image, to obtain a fifth fusion weight, and determine the weight corresponding to the first downsampled fused image, to obtain a sixth fusion weight. The first downsampled frame is any one of the at least one downsampled frame, and the first downsampled fused image is a frame in the at least one downsampled fused image having the same scale as the first downsampled frame; Determine the weight corresponding to the second downsampled fused image based on the features extracted from the second downsampled fused image, to obtain a seventh fusion weight. The information of the images in the at least one downsampled frame with scales smaller than the first downsampled frame is fused in the second downsampled fused image; Fuse the first downsampled frame, the first downsampled fused image, and the second fused image according to the fifth fusion weight, the sixth fusion weight, and the seventh fusion weight, to obtain a third downsampled fused image. The upsampled image of the third downsampled fused image is used to fuse with the images in the at least one downsampled frame with scales larger than the first downsampled frame; Denoise the third downsampled fused image to obtain a first downsampled denoised image; Upsample the first downsampled denoised image to obtain an upsampled denoised image. The upsampled denoised image is used to combine with a fused image having the same scale as the upsampled denoised image for denoising, to obtain an image with a scale larger than the first downsampled frame.

28. The apparatus according to claim 27, wherein The downsampling module is specifically configured to perform at least one wavelet transform on the current frame and the first fused image, to obtain the at least one downsampled frame and the at least one downsampled fused image.

29. A video denoising device, characterized in that, Comprising a processor, the processor is coupled to a memory, and the memory stores a program. When the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 14 is implemented.

30. A computer-readable storage medium stores a program, which when executed by a processing unit, executes the method according to any one of claims 1 to 14.

31. A video denoising device, characterized in that, Comprising a processing unit and a communication interface, the processing unit obtains program instructions through the communication interface. When the program instructions are executed by the processing unit, the method according to any one of claims 1 to 14 is implemented.

Citation Information

Patent Citations

  • Noise reduction method, terminal and storage medium

    CN111127347A