Video Denoising Method, Apparatus, Electronic Device, and Readable Storage Medium

By using the position information of video frames for motion estimation, the problems of poor video denoising effect, large calculation volume and high power consumption in the prior art are solved, and high quality and low power consumption video denoising processing is achieved.

CN115035456BActive Publication Date: 2025-06-20VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210769643.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-06-20
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing video denoising method is not effective under high noise conditions, and has large calculation volume and high power consumption, making it difficult to achieve real-time processing on mobile devices.

Method used

By obtaining the position information of the video frame to be denoised and the reference video frame, motion estimation is performed, the amount of motion change between pixels is obtained, and video denoising is performed based on this.

Benefits of technology

Improve the quality of video denoising, reduce the calculation amount and power consumption, and realize real-time video denoising processing on mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035456B_ABST
    Figure CN115035456B_ABST
Patent Text Reader

Abstract

The present application discloses a video denoising method, apparatus, electronic device and readable storage medium, belonging to the technical field of video processing. The method includes: obtaining the first pose information of a first video frame and the second pose information of a second video frame, where the first video frame is the video frame to be denoised and the second video frame is the reference video frame; performing motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; and performing denoising processing on the first video frame according to the first motion estimation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of video processing, and particularly relates to a video denoising method, device, electronic device, and readable storage medium. Background Art

[0002] Noise is one of the important factors affecting the quality of video images captured or displayed by electronic devices (such as digital imaging devices). Existing denoising methods can be mainly divided into two categories: image denoising and video (image sequence) denoising. Among them, video denoising can achieve better denoising effects.

[0003] Currently, the main methods of video denoising include video denoising methods based on three-dimensional filtering, video denoising methods based on block matching, denoising methods based on motion estimation and compensation, and video denoising methods based on deep learning. However, the above methods have poor effects on video denoising processing. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a video denoising method, device, electronic device, and readable storage medium, which can improve the effect of an electronic device in processing video noise.

[0005] In a first aspect, the embodiments of this application provide a video denoising method, which includes: obtaining the first pose information of a first video frame and the second pose information of a second video frame, where the first video frame is a video frame to be denoised, and the second video frame is a reference video frame; performing motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; and performing denoising processing on the first video frame according to the first motion estimation information.

[0006] In a second aspect, the embodiments of this application provide a video denoising device, which includes: an obtaining module and a processing module. The obtaining module is used to obtain the first pose information of a first video frame and the second pose information of a second video frame, where the first video frame is a video frame to be denoised, and the second video frame is a reference video frame; the processing module is used to perform motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information obtained by the obtaining module to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; and perform denoising processing on the first video frame according to the first motion estimation information.

[0007] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0011] In the embodiment of the present application, an electronic device obtains the first pose information of a first video frame and the second pose information of a second video frame. The first video frame is a video frame to be denoised, and the second video frame is a reference video frame. According to the first pose information and the second pose information, motion estimation is performed on the first video frame and the second video frame to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame. According to the first motion estimation information, denoising processing is performed on the first video frame. Since the electronic device can perform motion estimation on the video frame to be denoised and the reference video frame based on the pose information of the video frame to be denoised and the reference video frame, to obtain the motion change amount of pixels in the first video frame and pixels in the second video frame, that is, motion estimation information, and then the electronic device can perform denoising processing on the first video frame according to the motion change amount of pixels in the video frame. Therefore, the electronic device can make full use of the motion estimation information from the pose information, so that the electronic device can obtain a higher-quality processing result when denoising the video frame, and reduce the calculation amount generated by the electronic device when denoising the video frame, saving the power consumption of the electronic device. In this way, the electronic device can obtain a better video denoising effect. Description of the Drawings

[0012] Figure 1 is one of the schematic diagrams of a video denoising method provided by an embodiment of the present application;

[0013] Figure 2 is one of the example schematic diagrams of a video denoising method provided by an embodiment of the present application;

[0014] Figure 3 is the second of the example schematic diagrams of a video denoising method provided by an embodiment of the present application;

[0015] Figure 4 is the third of the schematic diagrams of a video denoising method provided by an embodiment of the present application;

[0016] Figure 5 It is the fourth schematic diagram of a video denoising method provided by an embodiment of the present application;

[0017] Figure 6 It is a schematic structural diagram of a video denoising method device provided by an embodiment of the present application;

[0018] Figure 7 It is one of the schematic hardware structures of an electronic device provided by an embodiment of the present application;

[0019] Figure 8 It is the second schematic hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.

[0022] Next, some terms / nouns involved in the embodiments of the present invention will be explained.

[0023] Two-dimensional: 2Dimension, 2D

[0024] Three-dimensional: 3Dimension, 3D

[0025] Six degrees of freedom: 6Degree of Freedom, 6DoF

[0026] Extended reality: extented Reality, XR

[0027] Scale-invariant feature transform: Scale-invariant feature transform, SIFT

[0028] Features from Accelerated Segment Test, FAST

[0029] Speeded-Up Robust Features, SURF

[0030] Inertial Measurement Unit, IMU

[0031] Convolutional Neural Network, CNN

[0032] U-shape Net, UNet

[0033] Currently, electronic devices (such as digital imaging devices) can be used to take pictures or videos, or provide subsequent intelligent tasks in the form of image or image sequence information, such as object detection, recognition, and tracking. Therefore, for electronic devices, good imaging quality is of self-evident importance. However, noise is one of the important factors affecting the picture quality of digital imaging devices. Imaging is challenged by noise in low-light environments such as at night, dawn, dusk, and in areas such as under bridges, tunnels, canyons, and caves.

[0034] Existing denoising methods can be mainly divided into two categories: image denoising and video (image sequence) denoising. Among them, image denoising is the basis of video denoising. Compared with image denoising, video denoising has additional temporal continuous information that can be used, thus enabling better denoising effects. Moreover, the image processing algorithm flow of existing digital imaging devices generally includes two parts: spatial domain (image) denoising and temporal domain denoising.

[0035] For image sequences, the data volume increases exponentially. How to efficiently utilize temporal data is the key to video denoising. In existing technologies, the main solutions for video denoising include: (1) methods based on 3D filtering; (2) methods based on block matching; (3) methods based on motion estimation and compensation; (4) methods based on deep learning.

[0036] For methods based on 3D filtering:

[0037] To address the motion issue in video denoising, the most straightforward approach is to extend from two dimensions to three dimensions. For example, 2D Gaussian filtering is extended to 3D Gaussian filtering, 2D wavelet is extended to 3D wavelet, and 2D bilateral filtering is extended to 3D bilateral filtering. Taking 3D bilateral filtering as an example: Bilateral filtering has a Gaussian kernel for detecting edges based on luminance information changes. Then, this Gaussian kernel can also be applied to the time domain. Therefore, at the same spatial position, if the luminance change is small, it is considered that no motion occurs; if the luminance change is large, it is considered that motion occurs. Since this detection is carried out in the form of a Gaussian kernel, the change detection is implicitly included in this time-domain Gaussian kernel.

[0038] For the block-matching-based method:

[0039] In an image sequence, the search for similar blocks can be carried out in adjacent frames. Since the change information is implicitly included in the matched similar blocks, after the similar block matching is completed, the selected matching blocks can be weighted averaged to filter out noise, where the weights depend on the distances of each block to the reference block. Before performing the weighted average, 3D DCT hard threshold noise suppression can be carried out on the clustering of each similar block to make the calculation of the weights more accurate. Since multiple 2D similar blocks form a 3D clustering, denoising can also be carried out on this 3D clustering in the wavelet domain. When the block-matching-based method matches similar blocks in the time domain, it preferentially selects blocks at the same spatial position. If a block can be matched at the same position, it indicates that this block is very likely not to have changed; if no block can be matched at the same position, then consider searching at other spatial positions in adjacent frames, and finally consider other spatial positions in the current frame. This spatio-temporal block-matching method completes the estimation of changes, enabling the noise suppression of the current block to consider the temporal information before and after. Moreover, in addition to 2D blocks, 3D blocks may be more suitable. Since a 3D block is a local spatio-temporal block. Therefore, the 3D block itself contains temporal motion information and can better identify the time-dependent characteristics related to motion. Since these multiple similar 3D blocks form a 4D clustering, noise reduction can be carried out based on these 4D clusters.

[0040] For the motion-estimation-compensation-based method:

[0041] For 3D pixel-domain filtering, transform-domain filtering, or 2D block or 3D block matching, the motion information is implicitly represented. The motion-estimation-compensation-based method suppresses noise through explicit motion information, that is, by tracking the matched similar blocks and filtering along the motion trajectory. Although this method uses 2D blocks, the motion information is explicitly estimated and utilized. In addition to the pixel domain, motion estimation can also be carried out in the frequency domain and wavelet domain. Moreover, in addition to similar blocks, optical flow is also often used for motion estimation.

[0042] For the deep-learning-based method:

[0043] With the development of deep learning, this method has also achieved good results in the field of video denoising. For temporal information, deep learning extracts useful temporal information through different network structures. Typical network structures include: deformable convolution, recurrent neural network, long short-term memory network, optical flow network based on deep learning, etc.

[0044] Currently, there are a large number of research results for the above various methods. Among them, many methods have also been applied to commercial products, such as video denoising based on block-matching motion estimation, video denoising based on lightweight neural networks, etc. However, no matter which technical solution is adopted, for digital imaging devices, a video denoising method with good denoising performance, small computational load, and low power consumption is still the goal to be pursued.

[0045] Although existing video denoising methods have achieved certain effects, there are still deficiencies. For methods based on 3D filtering, simple filter kernels cannot represent complex motions, and their denoising effects are generally poor. For methods based on block matching, when the noise is large, block matching itself will be severely interfered by noise, resulting in matching failures. In addition, the search for similar blocks in the 3D spatio-temporal domain brings a large amount of computational consumption. For methods based on motion estimation and compensation, the denoising effect depends on the accuracy of motion estimation. Similarly, in the presence of large noise, the performance of motion estimation will be greatly degraded, leading to failed denoising. For methods based on deep learning, due to the prior information from the supervised images, the denoising effect has been greatly improved. However, it has the problems of large computational load and high power consumption. The root cause of this problem also lies in the need to perform convolution on multiple frames of images before and after to collect useful motion information. Moreover, for different models of sensors, there are problems such as difficulty in collecting real training data sets and difficulty in modeling simulation data sets.

[0046] In summary, the common problems existing in the prior art for video denoising methods are: 1) The extraction of temporal information has poor effects under large noise conditions, resulting in a sharp decline in the denoising effect; 2) The computational load of the temporal information extraction module is large and the power consumption is high, and some methods cannot even be processed in real time on mobile devices.

[0047] The purpose of the embodiments of this application is to provide a video denoising method with excellent effects, small computational load, and low power consumption for an electronic device (mobile device) with a 6DoF pose estimation module. That is: based on the 6DoF pose information of the electronic device, an image motion estimation method is proposed for video denoising: based on the motion estimation from the pose information, by designing a new block matching strategy or motion coding network, while ensuring the denoising performance, the computational load and power consumption are reduced.

[0048] The video denoising method provided by the embodiments of the present application can be applied to an electronic device with a 6DoF pose information generation module and is applicable to scenarios where video denoising is required. Among them, devices with a 6DoF pose information generation module mainly include typical mobile devices: such as mobile phones, tablets, action cameras, watches with cameras, etc., robotic devices such as unmanned aerial vehicle devices, unmanned vehicle devices, unmanned ship devices, and also include mobile devices such as self-driving passenger cars and trucks.

[0049] The video denoising method provided by the embodiments of the present application can also be used for generating recorded videos and preview videos for users to enjoy, or photo synthesis of multiple frames of images, or robotic devices with only intelligent tasks such as target detection, recognition, and tracking, or devices such as XR helmets and XR glasses that require gesture interaction recognition and environmental perception.

[0050] Next, in conjunction with the accompanying drawings, through specific embodiments and their application scenarios, the video denoising method provided by the embodiments of the present application will be described in detail.

[0051] The embodiments of the present application provide a video denoising method. Figure 1 The flowchart of a video denoising method provided by the embodiments of the present application is shown. This method can be applied to an electronic device. As Figure 1 shown, the video denoising method provided by the embodiments of the present application may include the following steps 201 to 203.

[0052] Step 201: The electronic device obtains the first pose information of the first video frame and the second pose information of the second video frame.

[0053] In the embodiments of the present application, the first video frame is the video frame to be denoised, and the second video frame is the reference video frame.

[0054] In the embodiments of the present application, if the electronic device needs to perform denoising processing on the video, it can obtain the pose information corresponding to the video frame to be denoised in the target and the pose information corresponding to the reference video frame. Thus, the electronic device can perform denoising processing on the first video frame to be denoised according to the corresponding pose information of the video frame and the pose information corresponding to the reference video frame.

[0055] Optionally, in the embodiments of the present application, the electronic device can obtain a continuous image sequence from a digital image sensor through a video image sequence acquisition module, that is, the target video including the first video frame and the second video frame, so as to obtain the image sequence of the target video, and through a 6DoF pose acquisition module, obtain the 6DoF pose information of the electronic device corresponding to the target video, that is, obtain the 6DoF pose information corresponding to each frame of the target video.

[0056] Optionally, in the embodiments of the present application, the electronic device may acquire multiple video frames of the target video and 6DoF sequence information, and calculate the 6DoF information corresponding to each video frame.

[0057] It should be noted that the 6DoF pose acquisition module can be used to acquire the 6DoF pose information of the electronic device. Among them, the data sources for generating the 6DoF pose information include SLAM methods, UAV POS data, laser gyroscopes, etc.

[0058] Exemplarily, the electronic device may acquire the pose information corresponding to the target video through a monocular RGB camera, or multiple sensors such as IMU, ultrasonic, and barometer.

[0059] It should be noted that the pose information can provide the position information and attitude information of the electronic device in the physical space. The image changes caused by the video frames in the captured video are exactly due to the pose changes of the electronic device. Among them, the position information is represented as a three-axis translation position quantity t starting from the origin of the world coordinate system: (x, y, z), and the attitude information is represented as the rotation quantity of the electronic device itself. The typical rotation quantity can be represented in the form of Euler angles (for example: pitch angle, yaw angle, roll angle) as R: (α, β, γ), and can also be represented in forms such as rotation matrix, rotation vector, and quaternion. Generally, the generation time of the image, that is, the generation time of each video frame and the generation time of the pose information both correspond to their own frequencies, and the frequency values may be different.

[0060] Exemplarily, the target video can be 30fps or 60fps, while the pose information can be 10fps or 20fps. Therefore, the electronic device can obtain the pose information corresponding to the generation time of the video frames in the target video through interpolation.

[0061] It should be noted that since the pose information can be 100fps or 200fps, or even higher frequencies, when the pose data frame rate is greater than the target video frame rate, the electronic device can find the nearest neighbor pose data according to the image timestamp and pose timestamp of the target video. If the pose information is represented by relative change amounts, then multiple-frame integration is required to obtain the rotation change amount R and translation change amount t corresponding to the images of two video frames.

[0062] Optionally, in the embodiments of the present application, the electronic device may first cache multiple consecutive video frames in the target video to obtain the image information of the consecutive video frames, and calculate the corresponding pose information of the consecutive video frames. Thus, the electronic device can denoise the first video frame in the consecutive video frames according to the images of the consecutive video frames and the pose information corresponding to the images of the consecutive video frames, and output the first video frame.

[0063] Step 202: The electronic device performs motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information, and obtains first motion estimation information.

[0064] In the embodiment of the present application, the first motion estimation information includes the motion change amount of the pixels in the first video frame and the pixels in the second video frame.

[0065] Optionally, in the embodiment of the present application, the electronic device may perform motion estimation on the first video frame and the second video frame through a 2D image motion estimation module to obtain the motion change amount of the pixels in the first video frame and the pixels in the second video frame.

[0066] It should be noted that the above 2D image motion estimation module can be used to perform motion estimation in the 2D image domain on the video frame and its corresponding pose information to obtain the motion information of the pixels in the video frame.

[0067] In the embodiment of the present application, after the electronic device obtains the first pose information and the second pose information, it can perform motion estimation on the first image in two ways to obtain the first motion estimation information.

[0068] Optionally, in the embodiment of the present application, the electronic device may project the pose change of the electronic device in the 3D space onto the 2D image space by using an epipolar geometry model for motion estimation to obtain the first motion estimation information.

[0069] Optionally, in the embodiment of the present application, the first video frame includes at least one first pixel, and the second video frame includes at least one second pixel; the above step 202 may be specifically implemented by the following steps 202a and 202b.

[0070] Step 202a: The electronic device calculates a second target pixel that matches the first target pixel according to the first pose information, the second pose information, and the target internal parameters.

[0071] In the embodiment of the present application, the first target pixel is a pixel among at least one first pixel, the second target pixel is a pixel among at least one second pixel, and the target internal parameter is the camera internal parameter for collecting the first video frame and the second video frame.

[0072] Optionally, in the embodiment of the present application, the above first target pixel may be a pixel feature point of the first video frame, or may be a pixel feature vector of the first video frame.

[0073] Optionally, in the first method provided in the embodiment of the present application, the electronic device may use a feature point extraction operator to extract the pixel feature points of the first video frame and the second video frame and describe them in the form of a feature descriptor.

[0074] Exemplarily, the electronic device may use feature point extraction operators such as SIFT, FAST, and SURF to extract pixel feature points.

[0075] Optionally, in the second method provided in the embodiments of the present application, the electronic device may use a neural network to extract pixel feature vectors of the first video frame and the second video frame.

[0076] In the embodiments of the present application, the first target pixel is a pixel among at least one first pixel whose matching degree with at least one second pixel is greater than or equal to a preset threshold, and the second target feature pixel is a pixel among at least one second pixel whose matching degree with at least one first pixel is greater than or equal to a preset threshold.

[0077] Optionally, in the first method provided in the embodiments of the present application, the electronic device may use a distance operator such as the Hamming distance, and calculate a second target pixel that matches the first target pixel according to the first pose information, the second pose information, and the target internal parameters.

[0078] Optionally, in the second method provided in the embodiments of the present application, the electronic device may use a measure network to perform distance discrimination on at least one first pixel and at least one second pixel, so that the electronic device can calculate a second target pixel that matches the first target pixel according to the first pose information, the second pose information, and the target internal parameters, so as to complete the matching of the pixel feature vectors.

[0079] Step 202b: The electronic device calculates the target pixel displacement amount between the first target pixel and the second target pixel to obtain the first motion estimation information.

[0080] In the embodiments of the present application, the electronic device may calculate the displacement vector of the pixel position of the second target pixel corresponding to the first target pixel according to the first pose information, the second pose information, and the target internal parameters, so as to obtain the first motion estimation information.

[0081] In the embodiments of the present application, after calculating the second target pixel that matches the first target pixel, the electronic device may use a pair-wise geometric model to calculate the target pixel displacement amount between the first target pixel and the second target pixel, so as to obtain the first motion estimation information.

[0082] Optionally, in the first method provided in the embodiments of the present application, the electronic device may use a pair-wise geometric model to calculate the target pixel displacement amount between the pixel feature points of the first target pixel and the pixel feature points of the second target pixel, so as to obtain the first motion estimation information.

[0083] Optionally, in the second method provided in the embodiments of the present application, the electronic device may adopt an epipolar geometry model to obtain the target pixel displacement amount of the pixel feature vector of the first target pixel and the pixel feature vector of the second target pixel, so as to obtain the first motion estimation information, and the first motion estimation information is a motion field of W×H×2.

[0084] Wherein, W is the image width and H is the image height.

[0085] Exemplarily, it is assumed that through the extraction and feature matching of the pixel feature points and pixel feature vectors of the first video frame and the second video frame, the pixel positions corresponding to the same object feature points in the two video frames can be obtained, that is, p1(u1, v1) and p2(u2, v2). Then, the epipolar constraint can be obtained by adopting the epipolar geometry model as follows:

[0086]

[0087] Wherein, K is the camera internal parameter model, R and t are the rotation and translation amounts of the camera when taking two images, and (·)^ is the skew-symmetric symbol. The middle part of Equation (1) can be denoted in the form of the essential matrix E and the fundamental matrix F:

[0088] E = t^R, F = K^(-1)EK -T EK -1 (2)

[0089] It can be seen from Equation (1) that when the rotation amount R, translation amount t of the camera motion and the camera internal parameter K are known, the position relationship between two pixels can be obtained. That is, if the pixel p1 is known, the pixel p2 can be calculated, and thus the displacement amount (Δu, Δv) = (u2 - u1, v2 - v1) of the pixel on the 2D image can be obtained. The rotation amount R and translation amount t of the camera motion can be obtained from the pose information of the mobile device, and the camera internal parameter K can be obtained by calibration.

[0090] Step 203: The electronic device performs denoising processing on the first video frame according to the first motion estimation information.

[0091] Optionally, in the embodiments of the present application, the above step 203 may be specifically implemented by the following steps 203a1 to 203c1.

[0092] Step 203a1: The electronic device calculates a fourth pixel that matches a third pixel according to the first motion estimation information.

[0093] In the embodiments of the present application, the third pixel is a pixel in the first video frame, and the fourth pixel is a pixel in the second video frame.

[0094] Optionally, in the embodiments of the present application, the electronic device may calculate a pixel in the second video frame that matches a pixel in the first video frame according to the first motion estimation information.

[0095] Step 203b1: The electronic device searches for at least one fifth pixel that matches the third pixel in the first video frame based on a block matching strategy.

[0096] Optionally, in the embodiments of the present application, the electronic device may search at a fixed position near the reference pixel block corresponding to the pixel of the video frame to be denoised in the 2D spatial domain of the image. If a matching pixel block is found, further search is performed in this direction.

[0097] Exemplarily, as Figure 2 shown, if the right side is the video frame to be denoised and the solid block is the current reference pixel block to be processed. First, search and match can be performed at the position of the dashed box near the reference pixel block. If two matching pixel blocks above and below are detected, the electronic device can continue to search further in the up and down directions, that is, at the position of the dotted line block. The electronic device can preset two steps of search. In the first step, 8 positions are searched, and in the second step, at most 8 positions are searched. For ease of implementation, the positions of each searched pixel block are relatively fixed. Among them, the step size during the search may not adopt Figure 2 the side length of the pixel block shown, and a smaller side length can be adopted. Specifically, the embodiments of the present application do not impose any restrictions here.

[0098] Optionally, in the embodiments of the present application, in the 3D time domain, the position of the matching pixel block can be directly calculated according to the first motion estimation result.

[0099] Exemplarily, the electronic device may use the first three video frames and the last three video frames of the video frame to be denoised as the time domain frames to be searched. Then, the electronic device can perform spatial domain search based on the pixel to be denoised in the video frame to be denoised, so as to obtain 16 pixel blocks. And since each of the first three frames and the last three frames of the video frame to be denoised can provide 1 pixel block through motion information, that is, a total of 6 pixel blocks are provided. Thus, the electronic device can add the 16 pixel blocks, the 6 pixel blocks, and the 1 pixel block of the pixel to be denoised in the video frame to be denoised, so as to obtain at most 23 pixel blocks in the spatial domain and the time domain. Then, the electronic device can screen out the similar pixel blocks with higher matching degrees from the obtained 23 pixel blocks. Among them, the matching of two pixel blocks can be calculated using the Euclidean distance of pixel differences, and usually, more than a dozen similar pixel blocks can be matched.

[0100] Among them, the formula for the Euclidean distance of pixel differences is:

[0101] Step 203c1: The electronic device performs spatio-temporal filtering denoising processing on the third pixel in the first video frame according to the fourth pixel and at least one fifth pixel.

[0102] In an embodiment of the present application, the electronic device can perform spatio-temporal filtering denoising processing on a third pixel in a first video frame according to a fourth pixel in a second video frame that matches the third pixel, and at least one fifth pixel that is searched for in the first video frame and matches the third pixel based on a block matching strategy.

[0103] Optionally, in an embodiment of the present application, the electronic device can perform non-local mean spatio-temporal filtering denoising processing on the third pixel in the first video frame based on a spatio-temporal denoising module according to the fourth pixel and at least one fifth pixel.

[0104] It should be noted that the above spatio-temporal denoising module can be used to perform time-domain and space-domain denoising in combination with motion information to output a final denoising result.

[0105] Optionally, in an embodiment of the present application, the electronic device can perform denoising processing on the third pixel in the first video frame in a weighted average manner, or in a wavelet domain, etc.

[0106] Exemplarily, in an embodiment of the present application, a weighted average method can be used to perform denoising on the third pixel in the first video frame, that is, through the following formula:

[0107] Y(i) = ∑ j∈I w(i, j)v(j) (3)

[0108] Where I is a noisy image, pixel i ∈ I, Y(i) is the filtering result, v(j) is the pixel value at pixel j of image I, and w(i, j) is the weight of the pixel where pixel i and pixel j are located.

[0109] The calculation method of w(i, j) is:

[0110]

[0111] Where N i 、N j are rectangular pixels centered on pixel i and pixel j, h is a Gaussian weight factor, and a is the standard deviation of the Gaussian kernel.

[0112] Z(i) is a normalization factor, and its calculation method is:

[0113]

[0114] After the above processing by the electronic device, the filtering denoising result Y(i) can be obtained, and by performing this filtering on the third pixel in the first video frame, the final filtering denoising result of the first video frame can be obtained.

[0115] In the first method provided by the application embodiment, explicit motion estimation is performed on the first pose information of the first video frame and the second pose information of the second video frame of the electronic device, and the motion change amount of the pixels in the first video frame and the pixels in the second video frame is obtained, that is, the first motion estimation information. Thus, the electronic device can calculate the pixels in the second video frame that match the pixels in the first video frame according to the first motion estimation information, and based on the block matching strategy, calculate at least one pixel in the first video frame that matches the third video frame, so that the electronic device can perform denoising processing on the first video frame based on the matched pixels, improving the quality of the motion estimation of the electronic device. Therefore, the number of times and the computational amount of block matching of the electronic device are reduced, the time for video denoising is accelerated, the power consumption of the electronic device is saved, and the effect of the electronic device for performing denoising processing on the video is improved.

[0116] Optionally, in the embodiment of the present application, the above step 203 can be specifically implemented by the following steps 203a2 to 203d2.

[0117] Step 203a2: The electronic device encodes the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information.

[0118] Optionally, in the example of the present application, since the first motion estimation information is a two-channel 2D vector, the electronic device encodes the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information.

[0119] Exemplarily, the electronic device can preset to perform three encoding processes. Among them, after the first encoding process, the electronic device can obtain -dimensional feature vectors; after the second encoding process, -dimensional feature vectors are obtained; after the third encoding process, -dimensional feature vectors are obtained.

[0120] Step 203b2: The electronic device stacks the first video frame and the second video frame to obtain a video frame matrix.

[0121] In the embodiment of the present application, the video frame matrix includes the image information after stacking the first video frame and the second video frame.

[0122] Step 203c2: The electronic device encodes the video frame matrix based on a convolutional neural network to obtain encoded image information.

[0123] Optionally, in the embodiments of the present application, the electronic device stacks the first video frame and the second video to obtain a video frame matrix, and acquires the image information of the video frame matrix. Thus, the electronic device can use a convolutional neural network to perform encoding processing on the video frame matrix to obtain encoded image information.

[0124] Exemplarily, for two consecutive video frames, first stack them to obtain a video frame matrix of W×H×6 dimensions, and then also use a convolutional neural network to perform encoding processing to obtain encoded image information.

[0125] It should be noted that when the electronic device encodes the video frame matrix, the number of encoding times is consistent with the width, height dimensions and the encoded motion estimation information, and the channel dimension can be inconsistent.

[0126] Step 203d2: The electronic device fuses the encoded motion estimation information and the encoded image information to obtain target encoded information, and performs decoding processing on the target encoded information to obtain the fourth video frame.

[0127] In the embodiments of the present application, the fourth video frame is the video frame obtained after denoising the first video frame.

[0128] Optionally, in the embodiments of the present application, the electronic device can use a motion information fusion module to fuse the encoded motion estimation information and the encoded image information to obtain target encoded information. Thus, the electronic device can perform decoding processing on the target encoded information to obtain the video frame obtained after denoising the first video frame.

[0129] It should be noted that the motion information fusion module is used to convert the direct motion information obtained by the 2D image motion estimation module into an information mode that can be processed by the spatio-temporal denoising module.

[0130] Optionally, in the embodiments of the present application, the electronic device can fuse the encoded motion estimation information and the encoded image information together through channel-wise splicing and fusion operations.

[0131] It should be noted that the above-mentioned channel-wise splicing operation refers to stacking an a×b×c1-dimensional feature vector and an a×b×c2-dimensional feature vector along the channel dimension to obtain an a×b×(c1 + c2)-dimensional feature vector.

[0132] Exemplarily, the electronic device can perform three encoding processes on the video frame matrix. Among them, after the first encoding process, the electronic device obtains a - dimensional feature vector, so that after fusing the encoded motion estimation information and the encoded image information, an a - dimensional feature vector is obtained; after the second encoding process, an dimensional feature vectors, and after fusing and processing the encoded motion estimation information and the encoded image information, dimensional feature vectors are obtained; after the third encoding process, dimensional feature vectors are obtained, and after fusing and processing the encoded motion estimation information and the encoded image information, dimensional feature vectors are obtained, which are the target encoding information. Then, the electronic device can perform decoding processing on the target encoding information, that is, dimensional feature vectors, that is, perform decoding processing using a convolutional network. Similar to UNet, skip connections can be used to prevent the loss of detailed information. Since the electronic device uses element-wise addition operations when using skip connections, the size of the feature vectors to be encoded needs to be consistent with the decoding processing.

[0133] Exemplarily, taking the above target encoding information as dimensional feature vectors as an example for illustration, starting from dimensional feature vectors for the first decoding process, dimensional feature vectors are obtained. After element-wise addition, the dimension of the feature vectors remains unchanged; after the second decoding process, dimensional feature vectors are obtained; after the third decoding process, a fourth video frame with dimensions W×H×3 is obtained.

[0134] Exemplarily, Figure 3 shows a flowchart of a second video denoising method provided by an embodiment of the present application.

[0135] In the second method provided by the embodiment of the present application, the electronic device uses a deep learning neural network, so that the motion estimation information from the pose information can be more fully utilized to achieve higher-quality motion estimation. In addition, in the network structure, there is no need for a dedicated motion estimation module, and only the obtained motion estimation information is required, thus reducing the computational amount of the electronic device and saving the power consumption of the electronic device. Therefore, the electronic device can obtain a better video denoising effect.

[0136] It should be noted that the training dataset required for the neural video denoising network in the embodiment of the present application can be composed of a simulated dataset with added noise or a real dataset collected by a special device. The loss function can use the Euclidean distance loss function or the perceptual distance loss function.

[0137] It should be noted that the video denoising method provided by the embodiment of the present application can also be used in other occasions that require motion estimation, such as video coding and decoding, video deblurring, action recognition, target tracking, video super-resolution, video frame interpolation, video segmentation, etc.

[0138] Optionally, the video denoising method provided by the embodiments of the present application further includes the following steps 301 and 302, and the above step 203 can be implemented by the following step 303.

[0139] Step 301: The electronic device obtains at least one third pose information of at least one third video frame.

[0140] In the embodiments of the present application, each third video frame corresponds to one third pose information.

[0141] In the embodiments of the present application, the electronic device can also perform denoising processing on multiple video frames of the target video.

[0142] Optionally, in the embodiments of the present application, the electronic device can obtain the first pose information of the first video frame and obtain at least one third pose information of at least one third video frame, so that the electronic device can perform motion estimation on the first video frame and at least one third video frame, and obtain the motion change amount of the pixels in the first video frame respectively and the pixels in one of the at least one third video frames.

[0143] Optionally, in the embodiments of the present application, the at least one third video frame is multiple video frames adjacent to the first video frame in the target video.

[0144] Optionally, in the embodiments of the present application, the at least one third video frame can be multiple video frames in the target video that are before the first video frame; or, the at least one third video frame can be multiple video frames in the target video that are after the first video frame; or, the at least one third video frame can be multiple video frames in the target video that are before the first video frame and multiple video frames in the target video that are after the first video frame.

[0145] It should be noted that the information similarity between the video frames adjacent to the first video frame in the target video and the first video is relatively high.

[0146] Step 302: The electronic device performs motion estimation on the first video frame and at least one third video frame according to the at least one third pose information and the first pose information, and obtains at least one second motion estimation information.

[0147] In the embodiments of the present application, the electronic device can perform motion estimation according to the first pose information and one of the at least one third pose information respectively, and obtain at least one second motion estimation information.

[0148] It should be noted that the process of the electronic device performing motion estimation according to the first pose information and one of the at least one third pose information can refer to the above steps, and will not be elaborated here.

[0149] Step 303: The electronic device performs denoising processing on the first video frame according to the first motion estimation information and at least one second motion estimation information.

[0150] In the embodiments of the present application, after obtaining the first motion estimation information and at least one second motion estimation information, the electronic device can perform denoising processing on the first video frame in two ways.

[0151] It should be noted that the implementation process of the first implementation method among the above two methods can specifically refer to the above steps 203a1 to 203c1, and the implementation process of the second implementation method among the above two methods can specifically refer to the above steps 203a2 to 203d2, which will not be elaborated here.

[0152] The embodiments of the present application provide a video denoising method. The electronic device obtains the first pose information of the first video frame and the second pose information of the second video frame. The first video frame is the video frame to be denoised, and the second video frame is the reference video frame; performs motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information to obtain the first motion estimation information, where the first motion estimation information includes the motion change amount of the pixels in the first video frame and the pixels in the second video frame; performs denoising processing on the first video frame according to the first motion estimation information. Since the electronic device can perform motion estimation on the video frame to be denoised and the reference video frame based on the pose information of the video frame to be denoised and the reference video frame, to obtain the motion change amount of the pixels in the first video frame and the pixels in the second video frame, that is, the motion estimation information, and then the electronic device can perform denoising processing on the first video frame according to the motion change amount of the pixels in the video frame. Therefore, the electronic device can make full use of the motion estimation information from the pose information, so that the electronic device can obtain a higher-quality processing result when denoising the video frame, and reduce the calculation amount generated by the electronic device when denoising the video frame, saving the power consumption of the electronic device. Thus, the electronic device can obtain a better video denoising effect.

[0153] The video denoising method provided by the embodiments of the present application can be implemented through the following two embodiments:

[0154] Embodiment 1

[0155] Figure 4 Shows the flowchart of the block-matching spatio-temporal video denoising method combining pose information provided by Embodiment 1 of the present application. As Figure 1 shown, the video denoising method provided by Embodiment 1 of the present application can include the following steps 11 to 14.

[0156] Step 11: The electronic device obtains the image of the target video and the 6DoF sequence information, and calculates the 6DoF information corresponding to each frame of the image.

[0157] Step 12: The electronic device uses the epipolar geometry model to perform 2D image pixel motion estimation.

[0158] Optionally, in the embodiments of the present application, the electronic device may project the pose change of the mobile device in the 3D space onto the 2D image space using the epipolar geometry model for pose estimation, and the above step 12 may be specifically implemented through the following steps 12a to 12c.

[0159] Step 12a: The electronic device uses a feature point extraction operator, such as SIFT, FAST, SURF, etc., to extract image feature points and describe them in the form of feature descriptors.

[0160] Step 12b: The electronic device uses a distance operator such as the Hamming distance to perform feature point matching on two consecutive frames of images.

[0161] Step 12c: The electronic device uses the epipolar geometry model to calculate the displacement amount (Δu, Δv) of the pixel (u, v) in the image space, that is, the motion estimation information is obtained.

[0162] Step 13: The electronic device executes a fast block matching strategy based on the motion estimation result to complete the fusion of the motion estimation information from the pose.

[0163] Step 14: The electronic device uses the matched image blocks to perform spatio-temporal non-local mean filtering for denoising.

[0164] Embodiment 2

[0165] Figure 5 The flowchart of the neural network video denoising method combined with pose information provided in Embodiment 2 of the present application is shown. As Figure 5 shown, the video denoising method provided in Embodiment 2 of the present application may include the following steps 21 to 24.

[0166] Step 21: The electronic device obtains the image of the target video and the 6DoF sequence information, and calculates the 6DoF information corresponding to each frame of the image.

[0167] Step 22: The electronic device uses the epipolar geometry model to perform 2D image pixel motion estimation.

[0168] Optionally, in the embodiments of the present application, the electronic device may use the epipolar geometry model to perform 2D image pixel motion estimation, and the above step 22 may be specifically implemented through the following steps 22a to 22c.

[0169] Step 22a: The electronic device uses a CNN to extract image feature vectors.

[0170] Step 22b: The electronic device uses the measurement network to perform distance discrimination on different image blocks, thereby completing block-based feature point matching.

[0171] Step 22c: The electronic device uses the epipolar geometry model to calculate the displacement of pixels in the image space and obtains motion estimation information.

[0172] Step 23: The electronic device constructs a motion estimation information encoding network.

[0173] Step 24: The electronic device fuses the motion estimation information and uses UNet for spatio-temporal video denoising.

[0174] It should be noted that in this embodiment, the electronic device can not only process two consecutive frames of images, but also process multiple consecutive frames of images. If it is multiple frames of images, for Step 23, the electronic device needs to stack the motion estimation information of multiple adjacent frames; for Step 24, multiple frames of images need to be stacked.

[0175] In the video denoising method provided in the embodiment of the present application, the execution subject can be a video denoising device. In the embodiment of the present application, taking the video denoising device executing the video denoising method as an example, the video denoising device provided in the embodiment of the present application is described.

[0176] Figure 6 Fig. shows a possible structural schematic diagram of the video denoising device involved in the embodiment of the present application. As Figure 6 shown, the video denoising device 60 may include: an acquisition module 61 and a processing module 62.

[0177] Among them, the acquisition module 61 is used to acquire the first pose information of the first video frame and the second pose information of the second video frame. The first video frame is the video frame to be denoised, and the second video frame is the reference video frame. The processing module 62 is used to perform motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information acquired by the acquisition module 61, and obtain first motion estimation information. The first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; and perform denoising processing on the first video frame according to the first motion estimation information.

[0178] An embodiment of the present application provides a video denoising device. Since the electronic device can perform motion estimation on the video frame to be denoised and the reference video frame based on the pose information of the video frame to be denoised and the reference video frame, obtain the motion change amount of the pixels in the first video frame and the pixels in the second video frame, that is, motion estimation information, and then the electronic device can perform denoising processing on the first video frame according to the motion change amount of the pixels in the video frame. Therefore, the electronic device can make full use of the motion estimation information from the pose information, so that the electronic device can obtain a higher-quality processing result when denoising the video frame, reduce the amount of calculation generated by the electronic device when denoising the video frame, save the power consumption of the electronic device, and thus the electronic device can obtain a better video denoising effect.

[0179] In a possible implementation manner, the first video frame includes at least one first pixel, and the second video frame includes at least one second pixel; the processor module 62 is specifically configured to calculate a second target pixel that matches a first target pixel according to the first pose information, the second pose information, and the target internal parameter, where the first target pixel is a pixel among at least one first pixel, the second target pixel is a pixel among at least one second pixel, and the target internal parameter is the camera internal parameter for collecting the first video frame and the second video frame; calculate the target pixel displacement amount between the first target pixel and the second target pixel to obtain the first motion estimation information.

[0180] In a possible implementation manner, the processing module 62 is specifically configured to calculate a fourth pixel that matches a third pixel according to the first motion estimation information, where the third pixel is a pixel in the first video frame and the fourth pixel is a pixel in the second video frame; search for at least one fifth pixel that matches the third pixel in the first video frame based on a block matching strategy; perform spatio-temporal filtering denoising processing on the third pixel in the first video frame according to the fourth pixel and the at least one fifth pixel.

[0181] In a possible implementation manner, the processing module 62 is specifically configured to perform encoding processing on the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information; stack the first video frame and the second video frame to obtain a video frame matrix; perform encoding processing on the video frame matrix based on a convolutional neural network to obtain encoded image information; perform fusion processing on the encoded motion estimation information and the encoded image information to obtain target encoded information, and perform decoding processing on the target encoded information to obtain a fourth video frame, where the fourth video frame is the video frame obtained after denoising the first video frame.

[0182] In a possible implementation manner, the processing module 62 is further configured to obtain at least one third pose information of at least one third video frame, where each third video frame corresponds to one third pose information; perform motion estimation on the first video frame and the at least one third video frame according to the at least one third pose information and the first pose information, so as to obtain at least one second motion estimation information. Specifically, the processing module 62 is configured to perform denoising processing on the first video frame according to the first motion estimation information and the at least one second motion estimation information.

[0183] The video denoising device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0184] The video denoising device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0185] The video denoising device provided in the embodiments of the present application can implement Figures 1 to 3 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0186] Optionally, as Figure 7 shown, the embodiments of the present application further provide an electronic device 900, including a processor 901 and a memory 902. A program or instruction that can run on the processor 901 is stored on the memory 902. When the program or instruction is executed by the processor 901, it implements each step of the above-mentioned video denoising method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0187] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0188] Figure 8 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0189] The electronic device 100 includes, but is not limited to: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, etc.

[0190] Those skilled in the art can understand that the electronic device 100 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 8 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0191] Among them, the processor 110 is used to obtain the first pose information of the first video frame and the second pose information of the second video frame. The first video frame is the video frame to be denoised, and the second video frame is the reference video frame; according to the first pose information and the second pose information, perform motion estimation on the first video frame and the second video frame to obtain first motion estimation information, and the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; according to the first motion estimation information, perform denoising processing on the first video frame.

[0192] The embodiments of the present application provide an electronic device. The electronic device can perform motion estimation on the video frame to be denoised and the reference video frame based on the pose information of the video frame to be denoised and the reference video frame, so as to obtain the motion change amount of pixels in the first video frame and pixels in the second video frame, that is, motion estimation information. Then, the electronic device can perform denoising processing on the first video frame according to the motion change amount of pixels in the video frame. Therefore, the electronic device can make full use of the motion estimation information from the pose information, so that the electronic device can obtain a higher-quality processing result when denoising the video frame, and reduce the calculation amount generated by the electronic device when denoising the video frame, saving the power consumption of the electronic device. In this way, the electronic device can obtain a better video denoising effect.

[0193] Optionally, in the embodiments of the present application, the first video frame includes at least one first pixel, and the second video frame includes at least one second pixel; specifically, the processor 110 is configured to calculate a second target pixel that matches a first target pixel according to the first pose information, the second pose information, and the target internal parameters, where the first target pixel is a pixel among the at least one first pixel, the second target pixel is a pixel among the at least one second pixel, and the target internal parameters are the camera internal parameters for collecting the first video frame and the second video frame; calculate the target pixel displacement amount between the first target pixel and the second target pixel to obtain the first motion estimation information.

[0194] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to calculate a fourth pixel that matches a third pixel according to the first motion estimation information, where the third pixel is a pixel in the first video frame and the fourth pixel is a pixel in the second video frame; search for at least one fifth pixel that matches the third pixel in the first video frame based on a block matching strategy; perform spatio-temporal filtering denoising processing on the third pixel in the first video frame according to the fourth pixel and the at least one fifth pixel.

[0195] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to perform encoding processing on the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information; stack the first video frame and the second video frame to obtain a video frame matrix; perform encoding processing on the video frame matrix based on a convolutional neural network to obtain encoded image information; perform fusion processing on the encoded motion estimation information and the encoded image information to obtain target encoded information, and perform decoding processing on the target encoded information to obtain a fourth video frame, where the fourth video frame is the video frame obtained after denoising processing of the first video frame.

[0196] Optionally, in the embodiments of the present application, the processor 110 is further configured to obtain at least one third pose information of at least one third video frame, where each third video frame corresponds to one third pose information; perform motion estimation on the first video frame and the at least one third video frame according to the at least one third pose information and the first pose information to obtain at least one second motion estimation information; specifically, the processor 110 is configured to perform denoising processing on the first video frame according to the first motion estimation information and the at least one second motion estimation information.

[0197] It should be understood that in the embodiments of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes the image data of static pictures or videos obtained by an image capturing device (such as a camera) in the video capture mode or the image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also referred to as a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated herein.

[0198] The memory 109 can be used to store software programs and various data. The memory 109 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0199] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110 either.

[0200] The embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned video denoising method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0201] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0202] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned video denoising method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0203] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0204] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned video denoising method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0205] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0206] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0207] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the purpose of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A video denoising method, characterized in that, Including: Obtain the first pose information of the first video frame and the second pose information of the second video frame, where the first video frame is the video frame to be denoised, and the second video frame is the reference video frame; Perform motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; Perform denoising processing on the first video frame according to the first motion estimation information; The first video frame includes at least one first pixel, and the second video frame includes at least one second pixel; The performing motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information to obtain first motion estimation information includes: Calculate a second target pixel that matches a first target pixel according to the first pose information, the second pose information, and the target internal parameters, where the first target pixel is a pixel among the at least one first pixel, the second target pixel is a pixel among the at least one second pixel, and the target internal parameters are the camera internal parameters for collecting the first video frame and the second video frame; Calculate the target pixel displacement amount between the first target pixel and the second target pixel to obtain the first motion estimation information.

2. The method according to claim 1, characterized in that, The performing denoising processing on the first video frame according to the first motion estimation information includes: Calculate a fourth pixel that matches a third pixel according to the first motion estimation information, where the third pixel is a pixel in the first video frame, and the fourth pixel is a pixel in the second video frame; Search for at least one fifth pixel that matches the third pixel in the first video frame based on a block matching strategy; Perform spatio-temporal filtering denoising processing on the third pixel in the first video frame according to the fourth pixel and the at least one fifth pixel.

3. The method according to claim 1, characterized in that, The performing denoising processing on the first video frame according to the first motion estimation information includes: Perform encoding processing on the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information; Stack the first video frame and the second video frame to obtain a video frame matrix; Perform encoding processing on the video frame matrix based on a convolutional neural network to obtain encoded image information; Fuse the encoded motion estimation information and the encoded image information to obtain target encoded information, and perform decoding processing on the target encoded information to obtain a fourth video frame, where the fourth video frame is the video frame obtained after denoising the first video frame.

4. The method according to claim 1, characterized in that, The method further includes: Obtain at least one third pose information of at least one third video frame, and each of the at least one third video frame corresponds to one of the third pose information; Perform motion estimation on the first video frame and the at least one third video frame according to the at least one third pose information and the first pose information to obtain at least one second motion estimation information; The performing denoising processing on the first video frame according to the first motion estimation information includes: Denoise the first video frame according to the first motion estimation information and the at least one second motion estimation information.

5. A video denoising device, characterized in that, The device includes: an acquisition module and a processing module; The acquisition module is configured to acquire the first pose information of the first video frame and the second pose information of the second video frame, where the first video frame is the video frame to be denoised, and the second video frame is the reference video frame; The processing module is configured to perform motion estimation on the first video frame and the second video frame according to the first pose information and the second pose information acquired by the acquisition module, to obtain first motion estimation information, where the first motion estimation information includes the motion change amount of pixels in the first video frame and pixels in the second video frame; and denoise the first video frame according to the first motion estimation information; The first video frame includes at least one first pixel, and the second video frame includes at least one second pixel; Specifically, the processing module is configured to calculate a second target pixel matching a first target pixel according to the first pose information, the second pose information, and the target internal parameters, where the first target pixel is a pixel among the at least one first pixel, the second target pixel is a pixel among the at least one second pixel, and the target internal parameters are the camera internal parameters for acquiring the first video frame and the second video frame; and calculate the target pixel displacement amount between the first target pixel and the second target pixel to obtain the first motion estimation information.

6. The device according to claim 5, characterized in that, Specifically, the processing module is configured to calculate a fourth pixel matching a third pixel according to the first motion estimation information, where the third pixel is a pixel in the first video frame and the fourth pixel is a pixel in the second video frame; search for at least one fifth pixel matching the third pixel in the first video frame based on a block matching strategy; and perform spatio-temporal filtering denoising processing on the third pixel in the first video frame according to the fourth pixel and the at least one fifth pixel.

7. The device according to claim 5, characterized in that, Specifically, the processing module is configured to perform encoding processing on the first motion estimation information based on a convolutional neural network to obtain encoded motion estimation information; stack the first video frame and the second video frame to obtain a video frame matrix; perform encoding processing on the video frame matrix based on a convolutional neural network to obtain encoded image information; and perform fusion processing on the encoded motion estimation information and the encoded image information to obtain target encoded information, and perform decoding processing on the target encoded information to obtain a fourth video frame, where the fourth video frame is the video frame obtained by denoising the first video frame.

8. An electronic device, characterized in that, It includes a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the video denoising method according to any one of claims 1 to 4 are implemented.

9. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by the processor, the steps of the video denoising method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Video image denoising method and device, electronic equipment and storage medium

    CN111652814A

  • Image denoising processing method and device, storage medium and electronic equipment

    CN114066771A