Video processing method, device, terminal and computer-readable storage medium
By generating noise level maps and registering video frames, the video frames are denoised, which solves the problem of high noise in video images under low light conditions, and improves image quality and reduces noise.
Patent Information
- Application Number
- CN202210302332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-24
AI Technical Summary
Under low light or short exposure time conditions, video images contain a lot of noise, which makes it difficult for the prior art to effectively reduce noise, affecting image quality.
By acquiring multiple video frames arranged in chronological order, a noise level map corresponding to the target video frame is generated, and other video frames are registered to the target video frame. The target video frame is denoised based on the noise level map and the registered video frame. The noise reduction frame reduces noise by calculating the difference map between the video frames.
Effective noise reduction and detail enhancement of video frames is achieved, noise reduction, image details and edges are restored, and motion blur and shadowing are avoided.
Smart Images

Figure CN114792291B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method, an apparatus, a terminal, and a computer-readable storage medium for processing a video. Background Art
[0002] In video processing, noise reduction is a key step in improving image quality. Under insufficient light intensity conditions, such as low light or short exposure time, a large amount of noise is contained in the captured image. Therefore, how to reduce noise in a video has become an urgent problem to be solved. Summary of the Invention
[0003] In view of this, the main object of the present invention is to provide a method, an apparatus, a terminal, and a computer-readable storage medium for processing a video.
[0004] To achieve the above object, the technical solution of the present invention is implemented as follows: A method for processing a video includes the following steps: obtaining N video frames I 1 , I 2 , …, I N , where N is a natural number and N≥2; generating a noise level map Noise corresponding to the video frame I N , generating a video frame I j registered to I N video frame I′ j , where j = 1, 2,..., N-1; the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N have the same length and the same width, and the same pixel coordinate system is established for the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N ; the denoised frame corresponding to the video frame I N is D j (x) = |I′ j (x) - I N (x)|, where α is a constant, where x is any coordinate in the pixel coordinate system, I′ j (x), I N (x) and Noise(x) are respectively I′ j , I N and the pixel values of the pixels corresponding to the coordinate x in Noise.
[0005] As an improvement of an embodiment of the present invention, the "generating video frame I" N corresponding noise level map Noise" specifically includes: Noise = Ψ 1 (|I N -φ(I N )|), where the function φ() is an edge-preserving smoothing function, and the function Ψ 1 () is a Gaussian smoothing function.
[0006] As an improvement of an embodiment of the present invention, the function φ() is a guided filtering function.
[0007] As an improvement of an embodiment of the present invention, the "generating video frame I" j registered to I N of video frame I′ j " specifically includes: generating video frame I j to video frame I N spatial transformation T j→N , and based on the spatial transformation T j→N generating video frame I j registered to I N of video frame I′ j .
[0008] As an improvement of an embodiment of the present invention, the "generating video frame I j to video frame I N spatial transformation T j→N " specifically includes: based on the ORB algorithm, obtaining a set of feature points P j composed of multiple FAST feature pixels from video frame I j , and obtaining a set of feature points P N composed of multiple FAST feature pixels from video frame I N ; obtaining the BRIEF descriptor of each feature point in the set of feature points P j , and obtaining the BRIEF descriptor of each feature point in the set of feature points P N ; pairing each element E j in the set of feature points P j with the element E N in the set of feature points P N to obtain multiple paired point pairs, and the distance of each point pair is F(E j ,E N ), and satisfying: F(E j ,E N ) = min(F(E j , E)|E ∈ set of feature points P N), F(x1, x2) is the Hamming distance between the BRIEF descriptor of x1 and the BRIEF descriptor of x2; a preset proportion of the point pairs with the shortest distances are selected from multiple point pairs to form multiple target point pairs, and the spatial transformation T is calculated based on the positional relationship between the two paired elements in each target point j→N , where 0 < preset proportion ≤ 1.
[0009] As an improvement of the embodiment of the present invention, the "generating video frame I j registered to I N of the video frame I' j " specifically includes: generating a video frame I based on the optical flow method, feature point method or neural network method j registered to I N of the video frame I' j .
[0010] As an improvement of the embodiment of the present invention, it further includes the following steps: generating a denoised frame corresponding detail enhancement map detail layer map w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants, Ψ 2 () is a Gaussian smoothing function, and are respectively and the pixel values of the pixels corresponding to the coordinate x in
[0011] The embodiment of the present invention provides a video processing device, including the following modules: an information acquisition module, configured to acquire N video frames I arranged in chronological order 1 , I 2 , …, I N , N is a natural number, N ≥ 2; a noise map generation module, configured to generate a noise level map Noise of the video frame I N , generate a video frame I j registered to I N of the video frame I' j , j = 1, 2, ..., N - 1; the lengths and widths of the noise level map Noise, I' 1 , I' 2 , …, I' N-1 and I N are equal, and for the noise level map Noise, I' 1 , I' 2 , …, I' N-1 and I NSet up the same pixel coordinate system; a processing module for video frame I N The corresponding denoised frame is D j (x) = |I′ j (x) - I N (x)|, where α is a constant, and x is any coordinate in the pixel coordinate system I′ j (x), I N (x) and Noise(x) are respectively I′ j 、I N and the pixel values of the pixels corresponding to the coordinate x in Noise; an enhancement module for generating a detail enhancement map corresponding to the denoised frame The corresponding detail enhancement map Detail layer map w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants Ψ 2 () is a Gaussian smoothing function and are respectively and The pixel values of the pixels corresponding to the coordinate x in..
[0012] An embodiment of the present invention provides a terminal, including: a memory for storing a computer program; a processor for implementing the steps of the above processing method when executing the computer program
[0013] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above processing method are implemented
[0014] The video processing method, device, terminal and computer-readable storage medium provided by the embodiments of the present invention have the following advantages: The embodiments of the present invention disclose a video processing method, device, terminal and computer-readable storage medium. The processing method includes the following steps: obtaining a plurality of video frames arranged in chronological order, generating a noise level map corresponding to the target video frame, generating matching video frames of all other video frames registered to the target video frame, and performing denoising processing on the target video frame based on the matching video frames, the noise level map and the target video frame. Performing enhancement processing on the denoised video frame based on the noise level map. This processing method can perform denoising and enhancement processing on the target video frame BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flowchart of the processing method provided by an embodiment of the present invention
[0016] Figure 2A 、 Figure 2B 、 Figure 3A 、 Figure 3B 、 Figure 3C 、 Figure 4A 、 Figure 4B 、 Figure 4C 、 Figure 5A 、 Figure 5B and Figure 5C are the experimental result diagrams of the processing method. Specific implementation manners
[0017] The present invention will be described in detail below in conjunction with the embodiments shown in the drawings. However, this embodiment does not limit the present invention, and any structural, method, or functional transformation made by those of ordinary skill in the art based on this embodiment is included in the protection scope of the present invention.
[0018] The following description and the drawings fully illustrate the specific embodiments herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another element, and do not require or imply any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a structure, device or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the structure, device or equipment including the said element. The various embodiments herein are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0019] The terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. in this text indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing this text and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In the description of this text, unless otherwise specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0020] Embodiment 1 of the present invention provides a method for processing images, as Figure 1 shown, including the following steps:
[0021] Step 101: Obtain N video frames I arranged in chronological order 1 , I 2 , …, I N , where N is a natural number and N≥2;
[0022] Here, assuming that the length of each frame of the picture in the video is H and the width is W, the video frame can be expressed as I∈R H*W , R H*W is a two-dimensional real number space. Optionally, the generation time of I 1 <the generation time of I 2 <... <the generation time of I N .
[0023] Step 102: Generate a noise level map Noise corresponding to the video frame I N , generate a video frame I j registered to I N of the video frame I′ j , j = 1, 2,..., N - 1; the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N have the same length and width, and the same pixel coordinate system is established for the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N ;
[0024] Here, for the convenience of description, in the noise level map Noise, I′1 , I′ 2 , …, I′ N-1 and I N Set up the same pixel coordinate system, that is, for each pixel, obtain its position in the length direction and then obtain its position in the width direction. Thus, the position in the length direction and the position in the width direction constitute the coordinates of the pixel, thereby forming a pixel coordinate system. For example, in these images, the coordinates of the top-leftmost pixel are (0, 0). In the rightward direction, the position of the width gradually increases, and in the downward direction, the position of the length gradually increases. It can be understood that for the coordinate x (including the position in the width direction and the position in the length direction), there is a corresponding pixel in the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N respectively.
[0025] Here, Figure 2A and Figure 2B show the effect of image noise level estimation. Figure 2A is the video frame I N , Figure 2B is its corresponding noise level map Noise. It can be seen from Figure 2A and Figure 2B that Figure 2A the darker the area, Figure 2B the higher the predicted noise level in Figure 2B , which is consistent with the theory. The noise level of an image under underexposure is greater than that under normal exposure. At the same time, it can be seen from
[0026] Step 103: The denoised frame corresponding to the video frame I N is D j (x) = |I′ j (x) - I N (x)|, where α is a constant, and x is any coordinate in the pixel coordinate system. I′ j (x), I N (x), and Noise(x) are the pixel values of the pixels corresponding to the coordinate x in I′ j , I N and Noise respectively. Here, is the pixel value of the pixel corresponding to the coordinate x in the denoised frame , and I′ j(x) is the pixel value of the pixel corresponding to the coordinate x in the video frame I′ j in I, N (x) is the pixel value of the pixel corresponding to the coordinate x in the video frame I N and Noise(x) is the pixel value of the pixel corresponding to the coordinate x in the noise level map Noise.
[0027] Here, x is any coordinate in the pixel coordinate system, that is, for any coordinate, it is calculated that and then the denoised frame can be obtained
[0028] Here, when specifically implementing this step, the following sub-steps may be included:
[0029] Step 1: Based on the formula D j (x) = |I′ j (x) - I N (x)|, calculate the difference map D N between the video frames I j and I′ j . It can be understood that D j (x) is the pixel value of the pixel corresponding to the coordinate x in the video frame D j ;
[0030] Step 2: Based on the formula calculate the motion mask image M j of the video frame I′ N relative to the video frame I′ j . It can be understood that usually during video shooting, in addition to the global motion between frames caused by the movement of the camera, the objects being photographed will also have local motion, such as: cars moving on the road in a street view video, the beating of the heart during human perspective projection. If these local motions are not processed, it will introduce local trailing and blurring phenomena in the time-domain denoising result. Here, the image is represented as the sum of the signal S and the noise n, I(x) = S(x) + n(x). When the difference between two images exceeds the fluctuation of the noise at a certain pixel point, we should consider that they belong to different signals and should not average these regions during time-domain denoising. Therefore, with the help of the inter-frame difference map D j and the noise level map Noise of the target frame, estimate the motion mask M j of the j-th frame:
[0031] Among them, the area with a pixel value of 0 in the motion mask M j represents the moving area, and the area with a pixel value of 1 represents the stationary area; since the smoothed noise map is used as the noise level map, the noise level map is multiplied by the coefficient α as the threshold at different pixel points.
[0032] In the long-term work of the inventors, it is found that in the field of natural images, the noise in a video can usually be modeled as Gaussian white noise with a mean of zero. In the field of medical images, the noise model is more complex. For example, in X-ray imaging, the noise is usually modeled as a mixture of Gaussian white noise and Poisson noise. It can be understood that the noise models of natural images and medical images are different, but there is a commonality in the noise of videos: that is, the noise in the video frames appears randomly, and the noise between frames is independent of each other. When conducting a comparative experiment, the inventors conduct an experiment on time-domain noise reduction. Time-domain noise reduction takes advantage of the strong correlation of the image signals in different frames, while the noise is random and has no correlation, to reduce the noise; time-domain noise reduction can restore the details and edges of the image while removing the noise; however, when the object being photographed moves, simple time weighting will cause the image to be blurred or have ghosting. The processing method in the embodiments of the present invention can well solve these problems.
[0033] Here, Figure 3A shows the video frame I without noise reduction processing N , Figure 3B shows the noise reduction effect diagram of the traditional method, Figure 3C shows the noise reduction frame During the shooting of this video, the human body turns from the frontal position to the lateral position. From Figure 3B it can be seen that due to the movement of the object being photographed, if the processing of step 103 is not performed, there is ghosting in the stomach area on the noise reduction image, and at the same time the spine is blurred. From Figure 3C it can be seen that after step 103, the spine looks clearer, the ghosting in the stomach area also disappears, and at the same time the noise of the image is also reduced, which shows that the combination of local motion self-recognition and noise reduction can reduce image blurring and ghosting.
[0034] The effect of the operation "generate the video frame I j registered to I N of the video frame I' j " in step 102 on noise reduction is compared as Figure 4A , Figure 4B and Figure 4C shown. Figure 4A shows the video frame I without noise reduction processing N , Figure 4B shows the noise reduction result without using the operation "generate the video frame I j registered to I N of the video frame I' j " in step 102, while Figure 4C shows the use of the operation "generate the video frame I j registered to I N of the video frame I' j”Denoising result. During the shooting of this video, the camera moves from top to bottom. From Figure 4B the result, we can see that due to the movement of the camera, if the operation "generate video frame I j registered to I N video frame I' of j " is not used, the spine in the middle of the denoised image is still blurred. From Figure 4C it can be seen that after using the operation "generate video frame I j registered to I N video frame I' of j ", the spine has restored a clear shape, and at the same time, the noise in the image has also decreased. This shows that after the operation "generate video frame I j registered to I N video frame I' of j " and step 103, the phenomenon of motion blur and smear can be further reduced.
[0035] In this embodiment, the "generate video frame I N corresponding noise level map Noise" specifically includes: Noise = Ψ 1 (|I N -φ(I N )|), where the function φ() is an edge-preserving smoothing function, and the function Ψ 1 () is a Gaussian smoothing function.
[0036] Here, the Gaussian smoothing function makes the entire input image transition uniformly and smoothly, removes details, and filters out noise.
[0037] It can be understood that the image is a noise map containing the noise in video frame I N . The Gaussian smoothing function Ψ 1 () smooths the absolute value of (that is, takes the absolute value of the pixel value of each pixel in ) to obtain an estimate Noise of the noise level map at different pixel points.
[0038] In this embodiment, the function φ() is a guided filter function. Here, this guided filter function is specifically disclosed in the following paper: Kaiming He, Jian Sun, Xiaoou Tang. Guided Image Filtering[J]. IEEE Transactions on Pattern Analysis & Machine Intelligence, 2013, 35(6): 1397 - 1409. This guided filter function belongs to edge-preserving filtering, can perform image smoothing processing, and has a good edge-preserving effect.
[0039] In this embodiment, the "generating video frame I" j registered to I N of video frame I' j " specifically includes: generating video frame I j to video frame I N of spatial transformation T j→n , based on the spatial transformation T j→N generating video frame I j registered to I N of video frame I' j .
[0040] Here, for video frames I 1 , I 2 , …, I N-1 respectively generate corresponding video frames I' 1 , I' 2 , …, I' N-1 . It can be understood that during the shooting of a video, the movement of the camera may cause the object being photographed to move within the field of view. Aligning the first N - 1 video frames to the target video frame helps reduce motion blur during the noise reduction process. Let the image pixel point space of image I j be ω j , and the image pixel point space of image I N be ω N . The goal of image alignment is to find a transformation T j→N : ω j →ω N such that the pixels in the transformed image I' j and I N correspond one by one, where the transformed image can be written as: I' j = T j→N (I j ).
[0041] The "generating spatial transformation T j from video frame I N to video frame I" j→N specifically includes:
[0042] Based on the ORB (Oriented Fast and Rotated Brief) algorithm, obtain a set of feature points P j composed of multiple FAST feature pixels from video frame I j , and obtain a set of feature points P N composed of multiple FAST feature pixels from video frame I N; Here, this step can be: (1) Select a pixel P from the picture and first set its brightness value to I P ; (2) Set a suitable threshold t. (3) Consider a discretized Bresenham circle with a radius equal to 3 pixels centered on this pixel P, and there are 16 pixels on the boundary of this circle; (4) Now, if there are n consecutive pixel points on this circle of 16 pixels, and their pixel values are either all greater than I P +t or all less than I P -t, then the pixel P is a FAST feature pixel.
[0043] Obtain the BRIEF descriptors of each feature point in the feature point set P j ; Obtain the BRIEF descriptors of each feature point in the feature point set P N ; Here, for the feature point set P j and P N calculate the BRIEF descriptors of the feature points. The BRIEF descriptor is a binary descriptor (optionally, a 128-bit binary string, etc.). Its calculation method is to randomly select 128 point pairs around the feature point. For the two points in each point pair, if the pixel value of the previous point is greater than that of the latter point, take 1, otherwise take 0.
[0044] Pair each element E j in the feature point set P j with the element E N in the feature point set P N to obtain multiple paired point pairs, and the distance of each point pair is F(E j , E N ), and satisfy: F(E j , E N ) = min(F(E j , E) | E ∈ the feature point set P N ), where F(x1, x2) is the Hamming distance between the BRIEF descriptors of x1 and x2; Select a preset proportion of the point pairs with the shortest distances from the multiple point pairs to form multiple target point pairs, and calculate the spatial transformation T j→N according to the positional relationship of the two paired elements in each target point, where 0 < preset proportion ≤ 1.
[0045] Optionally, the preset proportion = 90%.
[0046] In this embodiment, the "generate the video frame I j registered to I N 's video frame I′ j " specifically includes: Based on the optical flow method, feature point method, or neural network method, generate the video frame I jRegistered to I N of the video frame I' j .
[0047] Here, the optical flow is the instantaneous velocity of the pixel motion of a spatial moving object on the observation imaging plane. The optical flow method is a method that uses the change of pixels in the time domain in an image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, so as to calculate the motion information of the object between adjacent frames. It usually defines the instantaneous change rate of the gray level at a specific coordinate point on the two-dimensional image plane as the optical flow vector.
[0048] Here, the basic idea of the feature point method is as follows: First, extract many features from the image, then perform feature matching between the images, so as to obtain many well-matched points, and then perform image registration based on these points.
[0049] Here, a neural network can be pre-trained. This neural network can generate a registration network. Then, based on this neural network, generate the video frame I j Registered to I N of the video frame I' j .
[0050] In this embodiment, the following steps are further included:
[0051] Step 104: Generate a denoised frame The corresponding detail enhancement map Detail layer map w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants, Ψ 2 () is a Gaussian smoothing function, and are respectively and the pixel values of the pixels corresponding to the coordinate x in
[0052] Here, first, extract the detail layer map of the denoised frame That is, use the Gaussian smoothing function Ψ 2 () to extract the low-frequency map of the denoised map Use the denoised frame Subtract the low-frequency map to obtain the detail layer map
[0053] After that, calculate the enhancement weight of the detail layer map according to the noise level map Noise. The greater the pixel value in Noise, the greater the noise. Therefore, we use the noise level map Noise to distinguish For the details and noise within, the threshold th is used to define the regions with high noise within Noise, obtaining the detail enhancement coefficient map w, where w(x) = β * (1 - clip(Noise(x), th) / th). Here, β is the enhancement intensity, and the clip operation sets the intensity at the pixel points in Noise where the pixel value > th to th. As shown in the above formula, the enhancement coefficient calculated at the pixels in Noise with a noise level greater than th is 0, while the enhancement coefficient calculated at the pixels with a noise level of 0 is β. In other cases, the greater the noise, the smaller the enhancement coefficient.
[0054] Finally, multiply by w(x) and the denoised frame and add them together to obtain the image with enhanced details From the calculation of w(x), it can be seen that the areas with greater noise in the detail map are enhanced to a lesser extent. When the noise level is greater than a certain threshold th, the high-frequency components at that pixel are not enhanced. Thus, the purpose of adaptively enhancing details is achieved by means of the noise level map N(x).
[0055] Here, in order to further improve the visual effect of the video image, the denoised video frames are subjected to detail enhancement. Detail enhancement mainly strengthens the high-frequency information within the image. Some textures and edges within the image belong to high-frequency information. Enhancing this part of the signal helps to improve the sharpness and contrast of the image. In addition, a part of the high-frequency information is also contained in the noise. Adaptive enhancement based on the noise level will not amplify the noise while enhancing the details, nor will it weaken the effect of the previous image denoising.
[0056] Figure 5A shows the denoised frame Figure 5B shows the denoising effect diagram without being processed by step 104, while Figure 5C shows the denoising effect diagram after being processed by step 104, Figure 5B and Figure 5C uses the same enhancement intensity. It can be seen from Figure 5B that although the local contrast of the image is improved in Figure 5B , the noise is also enhanced synchronously, and the overall noise level of the entire image is significantly higher than that of the image before denoising. After being processed by step 104, Figure 5C only improves the contrast of the image edge details, and at the same time does not amplify the noise. The overall noise level of the entire image looks the same as the input denoised frame is consistent.
[0057] Embodiment 2 of the present invention provides a video processing device, including the following modules:
[0058] An information acquisition module, configured to acquire N video frames I arranged in chronological order 1 , I2 ,…,I N , where N is a natural number and N≥2;
[0059] A noise map generation module for generating the video frame I N The corresponding noise level map Noise, generating the video frame I j Registered to I N The video frame I′ j , j = 1, 2,..., N - 1; the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N have the same length and the same width, and set the same pixel coordinate system for the noise level map Noise, I′ 1 , I′ 2 , …, I′ N-1 and I N ;
[0060] A processing module for the video frame I N The corresponding denoised frame is D j (x) = |I′ j (x) - I N (x)|, where α is a constant, and x is any coordinate in the pixel coordinate system, I′ j (x), I N (x) and Noise(x) are respectively I′ j , I N and the pixel values of the pixels corresponding to the coordinate x in Noise;
[0061] An enhancement module for generating the detail enhancement map corresponding to the denoised frame Detail layer map w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants, Ψ 2 () is a Gaussian smoothing function, and are respectively and the pixel values of the pixels corresponding to the coordinate x in
[0062] Embodiment 3 of the present invention provides a terminal, including:
[0063] A memory for storing a computer program;
[0064] A processor for implementing the steps of the processing method in the first embodiment when executing the computer program.
[0065] Embodiment 3 of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the processing method in the first embodiment are implemented.
[0066] It should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0067] The series of detailed descriptions listed above are only specific descriptions of the feasible embodiments of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent embodiments or modifications made without departing from the technical spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for processing a video, characterized in that, it includes the following steps: Obtain N video frames I arranged in chronological order 1 , I 2 , …, I N , where N is a natural number and N ≥ 2; Generate video frame I N Generate video frame I corresponding to the noise level map Noise j Register to I N Video frame I j ′, j = 1, 2, ..., N - 1; the noise level map Noise, I′ 1 、I′ 2 、…、I′ N-1 and I N have the same length and the same width, and set the same pixel coordinate system for the noise level map Noise, I′ 1 、I′ 2 、…、I′ N-1 and I N ; Video frame I N The corresponding noise-reduced frame is D j (x) = |I j '(x) - I N (x)|, where α is a constant, and x is the pixel coordinate Any coordinate in the coordinate system, I j (x), I N (x) and Noise(x) are respectively I j ′, I N and the pixel values of the pixels corresponding to the coordinate x in Noise; Generate a noise-reduced frame The corresponding detail-enhanced image Detail layer image w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants Ψ 2 () is a Gaussian smoothing function and are respectively and the pixel values of the pixels corresponding to the coordinate x in 2. The processing method according to claim 1, characterized in that, The described "generating video frame I N The corresponding noise level map Noise" specifically includes: Noise = Ψ 1 (|I N - φ(I N )|), where the function φ() is an edge-preserving smoothing function, and the function Ψ 1 () is a Gaussian smoothing function.
3. The processing method according to claim 2, characterized in that: The function φ() is a guided filter function.
4. The processing method according to claim 1, characterized in that, The said "generating video frame I j registered to I N of video frame I j '" specifically includes: Generate video frame I j to video frame I N spatial transformation T j→N Based on the spatial transformation T j→N Generate video frame I j registered to I N video frame I j '.
5. The processing method according to claim 4, characterized in that, The "generation of video frame I j to video frame I N spatial transformation T j→N " specifically includes: Based on the ORB algorithm, multiple FAST feature pixels are obtained from video frame I j to form a feature point set P j , and multiple FAST feature pixels are obtained from video frame I N to form a feature point set P N ; Obtain the feature point set P j and the BRIEF descriptor of each feature point in it, and obtain the feature point set P N and the BRIEF descriptor of each feature point in it; The feature point set P j Each element E j With the feature point set P N The element E in N Pairing obtains multiple paired points, and the distance between each point pair is F(E j ,E N ), and satisfy: F(E j ,E N )=min(F(E j , E)|E∈feature point set P N ), F(x1, x2) is the Hamming distance between the BRIEF descriptor of x1 and the BRIEF descriptor of x2; a preset ratio of point pairs with the shortest distance is selected from multiple point pairs to form multiple target point pairs, and the spatial transformation T is calculated according to the positional relationship between the two paired elements in each target point. j→N , where 0<preset ratio≤1.
6. The processing method according to claim 1, characterized in that, The "generated video frame I j registered to I N of the video frame I j '" specifically includes: Generate video frame I based on optical flow method, feature point method or neural network method j Register to I N of video frame I j ′.
7. A video processing device, characterized in that, it includes the following modules: An information acquisition module for acquiring N video frames I arranged in chronological order 1 , I 2 , …, I N , where N is a natural number and N ≥ 2; A noise map generation module for generating video frame I N The corresponding noise level map Noise, generating video frame I j Registered to I N Of video frame I j ′, j = 1, 2,..., N - 1; The noise level map Noise, I′ 1 、I′ 2 、…、I′ N-1 And I N Have the same length and the same width, and set the same pixel coordinate system for the noise level map Noise, I′ 1 、I′ 2 、…、I′ N-1 And I N Establish the same pixel coordinate system; Processing module for video frame I N The corresponding denoised frame is D j (x) = |I j '(x) - I N (x)|, where α is a constant and x is in the pixel coordinate system Any coordinate, I j (x), I N (x) and Noise(x) are respectively I j ′, I N and the pixel values of the pixels corresponding to the coordinate x in Noise; Enhancement module for generating noise-reduced frames Corresponding detail enhancement map Detail layer map w(x) = β * (1 - clip(Noise(x), th) / th), where β and th are constants Ψ 2 () is a Gaussian smoothing function and are respectively and the pixel values of the pixels corresponding to the coordinate x in 8. A terminal, characterized in that, it includes: a memory for storing a computer program; a processor for implementing the steps of the processing method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN112950502A
Apparatus for removing a noise of image
KR101558532B1
Image registration method and terminal
WO2017107700A1