Video noise reduction method and device, electronic equipment and storage medium
By acquiring and aligning the first image of the video to be denoised and the second image after denoising of the previous frame, multi-scale denoising processing is performed, which solves the problem of poor video denoising effect in the existing technology and achieves better denoising effect and reduction of ghosting.
Patent Information
- Application Number
- CN202311245974.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-09-25
AI Technical Summary
Existing video noise reduction solutions cannot effectively improve the noise reduction effect. Single-frame solutions cause noise flickering, while multi-frame solutions are prone to ghosting in motion scenes.
By acquiring the first image of the video to be denoised and the second image of the previous frame after denoising, and performing alignment processing, multi-scale denoising is performed. The multi-scale denoising model is used to improve the denoising effect and reduce ghosting.
It alleviates and avoids ghosting issues caused by multi-frame video noise reduction, effectively improving video noise reduction performance.
Smart Images

Figure CN119788789B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a video noise reduction method and device, electronic equipment and storage medium. BACKGROUND
[0002] Image noise is a random variation in intensity or color information in an image (not present in the subject itself), usually a manifestation of electronic noise. It is generally produced by the sensor and circuitry of a scanner or digital camera, and can also be caused by film grain or the unavoidable shot noise in an ideal photodetector. Image noise is an unwanted byproduct of the image capture process, which brings errors and additional information to the image.
[0003] A video is composed of multiple continuous video frames, and a video frame is an image, so video noise reduction includes noise reduction of images in the video in units of frames. Image noise reduction refers to a technology that uses a certain method to maintain the original texture and details of the image as much as possible while suppressing or eliminating noise points in the image, so as to improve the visual quality of the image. The difference between video noise reduction and single-frame image noise reduction is that video can use the correlation between video frames in the time domain and / or spatial domain.
[0004] In related technologies, there are mainly two schemes when processing video noise reduction: one is a single-frame scheme, each frame is independently reduced by various methods, and the information of other frames is not associated; the other is a multi-frame scheme, which usually fuses the noise reduction result of the previous frame with the current frame in various ways to obtain the noise reduction result of the current frame, and the previous frame and the current frame have the same resolution.
[0005] However, the single-frame scheme does not consider the information in the time domain, resulting in noise flicker in the noise-reduced video, which seriously affects the noise reduction effect; the multi-frame scheme is prone to ghosting if the shooting time of the previous frame and the current frame is inconsistent, which affects the noise reduction effect if there is motion in the shooting scene. Therefore, the video noise reduction scheme provided in related technologies cannot effectively improve the noise reduction effect. SUMMARY
[0006] The purpose of the embodiments of the present application is to provide a video noise reduction method, device, electronic equipment and storage medium to solve the technical problem that the video noise reduction scheme provided in related technologies cannot effectively improve the noise reduction effect.
[0007] In a first aspect, the embodiments of the present application provide a video noise reduction method, comprising:
[0008] obtaining a first image to be de-noised and a second image after de-noising in a video to be de-noised, the first image being any one frame image in the video to be de-noised, and the second image being a previous frame image of the first image, and the size of the second image being not greater than that of the first image;
[0009] aligning the first image and the second image to obtain an aligned image corresponding to the second image;
[0010] performing multi-scale de-noising on the first image and the aligned image to obtain a target image after de-noising of the first image, and replacing the first image in the video to be de-noised with the target image.
[0011] In a second aspect, an embodiment of the present application provides a video de-noising device, comprising:
[0012] an obtaining module, configured to obtain a first image to be de-noised and a second image after de-noising in a video to be de-noised, the first image being any one frame image in the video to be de-noised, and the second image being a previous frame image of the first image;
[0013] an aligning module, configured to align the first image and the second image to obtain an aligned image corresponding to the second image;
[0014] a first de-noising module, configured to perform multi-scale de-noising on the first image and the aligned image to obtain a target image after de-noising of the first image, and replace the first image in the video to be de-noised with the target image.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements steps in the video de-noising method of any one of the above aspects when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement steps in the video de-noising method of any one of the above aspects.
[0017] The embodiment of the present application provides a video noise reduction method and device, electronic equipment and storage medium, the method obtains a first image to be reduced in any frame to be reduced in a to-be-reduced video and a second image after noise reduction in a previous frame of the first image, wherein the size of the second image is not greater than the size of the first image, then multi-scale noise reduction processing is carried out according to the first image and the aligned image, so that the first image is reduced according to the image information of the second image after noise reduction, the ghost problem caused by multi-frame video noise reduction can be relieved and avoided, and the noise reduction effect of the video is effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a flowchart of a video noise reduction method provided by the embodiment of the present application;
[0019] Figure 2 is a network structure diagram of a multi-scale noise reduction model provided by the embodiment of the present application;
[0020] Figure 3 is another flowchart of a video noise reduction method provided by the embodiment of the present application;
[0021] Figure 4 is a structural diagram of a video noise reduction device provided by the embodiment of the present application;
[0022] Figure 5 is another structural diagram of a video noise reduction device provided by the embodiment of the present application;
[0023] Figure 6 is a structural diagram of an electronic equipment provided by the embodiment of the present application;
[0024] Figure 7 is another structural diagram of an electronic equipment provided by the embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0026] It should be understood that each step recorded in the method embodiment of the present disclosure can be executed in different order and / or in parallel. In addition, the method embodiment can include additional steps and / or omit the execution of the shown steps. The scope of the present disclosure is not limited in this respect.
[0027] As used herein, the term "includes" and its variants are to be read to be analogous to an open-ended term such as "comprises," "includes," or "contains" that is to say, the term "includes" when used in this specification denotes "including, but not limited to." The term "based on" is to be read as "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms have analogous meanings.
[0028] In the related art, there are mainly two schemes when processing the video noise reduction: one is a single frame scheme, each frame is independently reduced by various methods, and the information of other frames is not associated; the other is a multi-frame scheme, which usually fuses the noise reduction result of the previous frame with the current frame in various ways to obtain the noise reduction result of the current frame, and the previous frame and the current frame have the same resolution.
[0029] However, the single frame scheme does not consider the information in the time domain, resulting in the problem of noise flicker in the noise-reduced video, which seriously affects the noise reduction effect; the multi-frame scheme is inconsistent in time if the previous frame and the current frame are shot, and if there is motion in the shooting scene, ghosting is easy to occur, which affects the noise reduction effect. Therefore, the video noise reduction scheme provided in the related art cannot effectively improve the noise reduction effect.
[0030] To solve the technical problems in the related art, the embodiments of the present application provide a video noise reduction method, which can be applied to an electronic device. When recording an object through a camera of the electronic device, due to the limitations of hardware and environment, the obtained video may have noise problems such as noise points, patches, and ghosting that affect the visual effect. At this time, the original video collected by the camera in the electronic device can be processed for noise reduction to obtain a noise-reduced video.
[0031] The electronic device can be mobile or fixed, for example, the electronic device can be a mobile phone, a camera, a camcorder, a vehicle, a tablet personal computer (TPC), a media player, a smart television, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a smart watch, an augmented reality (AR) / virtual reality (VR), a wearable device (WD), a game console, etc. with a video noise reduction function. The specific type of the electronic device is not limited in the embodiments of the present application.
[0032] Specifically, please refer to Figure 1 , Figure 1is a flowchart of a video noise reduction method provided by an embodiment of the present application, and the method comprises steps 101 to 103.
[0033] In step 101, a first image to be reduced in noise and a second image after noise reduction are obtained in a video to be reduced in noise, the first image is any one frame image of the video to be reduced in noise, and the second image is a previous frame image of the first image, and the size of the second image is not greater than the size of the first image.
[0034] In the present embodiment, the second image provided by the present embodiment is an image after noise reduction processing, and the specific noise reduction manner can be a conventional noise reduction algorithm or an artificial intelligence noise reduction algorithm using a trained convolutional neural network, and is not specifically limited here.
[0035] In the present embodiment, the first image to be reduced in noise provided by the present embodiment can be an RGB three-channel color image after image signal processing (ISP), or can be an original RAW image without ISP processing. In the case where the first image to be reduced in noise is a three-channel color image after ISP processing, the video noise reduction method provided by the present embodiment can be used as a post-processing scheme to optimize a video already taken by a user and improve noise reduction effect and inter-frame consistency. When the first image to be reduced in noise is an original RAW image without ISP processing, the video noise reduction method provided by the present embodiment can be used as a noise reduction module in mobile terminal video ISP to perform real-time noise reduction processing on the original RAW image.
[0036] In step 102, the first image and the second image are aligned to obtain an aligned image corresponding to the second image.
[0037] In the present embodiment, before the step of aligning the first image and the second image to obtain the aligned image corresponding to the second image provided by the present embodiment, the video noise reduction method provided by the present embodiment can further comprise: transforming the first image and the second image to obtain first and second transformed images of the same size.
[0038] In the present embodiment, the transformation processing provided by the present embodiment is mainly to transform the first image and / or the second image, and specifically, the transformation processing provided by the present embodiment includes but is not limited to size transformation, color space transformation, down-sampling, etc., as long as it can improve the processing efficiency of the alignment processing of the first image and the second image, and is not specifically limited here.
[0039] Specifically, after the first image and the second image are transformed to obtain the first transformed image and the second transformed image of the same size, the step of aligning the first image and the second image to obtain the aligned image corresponding to the second image provided by the embodiment can be: aligning the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image. By aligning the first transformed image and the second transformed image of the same size, the processing efficiency of the alignment processing can be effectively improved. At the same time, by aligning the second transformed image after noise reduction to the first transformed image, the information after noise reduction in the second transformed image after noise reduction can be conveniently introduced, and the ghost problem caused by the misalignment of pixels due to the motion between the front and rear frames can be reduced.
[0040] Optionally, since the size of the second image provided by the embodiment is not greater than the size of the first image, the size of the second image provided by the embodiment can be smaller than the size of the first image, and the size of the second image can also be equal to the size of the first image. In some embodiments, in the case where the size of the second image is smaller than the size of the first image, the sizes of the first image and the second image need to be unified first to reduce the size of the first image to the size of the second image, and then the alignment processing is performed. In this case, since the size of the second image after noise reduction is small and the resolution is low, the calculation amount of the alignment processing of the first image and the second image can be effectively reduced. In other embodiments, in the case where the size of the second image is equal to the size of the first image, the alignment processing can be directly performed. In this case, since the size and the resolution of the second image after noise reduction are not reduced, the first image and the second image are directly aligned, the information after noise reduction in the high-resolution second image can be effectively utilized, and the noise reduction effect of the subsequent noise reduction processing of the first image can be improved.
[0041] In some embodiments, the first image and the second image can also be transformed by the embodiment. Optionally, the transformation processing provided by the embodiment can include down-sampling processing. The step of transforming the first image and the second image to obtain the first transformed image and the second transformed image of the same size provided by the embodiment can be: performing down-sampling processing on the first image and the second image of the first size to obtain the first transformed image and the second transformed image of the second size, and the second size is not greater than the first size. By simultaneously performing down-sampling processing on the first image and the second image, the calculation amount of the alignment processing of the first image and the second image can be effectively reduced, thereby improving the noise reduction efficiency of the video noise reduction.
[0042] It should be noted that the transformation processing provided in this embodiment can also include up-sampling processing, that is, the first image and the second image can be simultaneously subjected to up-sampling processing to obtain the first transformed image and the second transformed image of the same third size, which is larger than the first size. In this way, the information after noise reduction in the high-resolution second transformed image can be effectively utilized, and the noise reduction effect of subsequent noise reduction processing on the first image can be further improved.
[0043] In addition, the transformation processing provided in this embodiment can not only transform the size of the image, but also transform and select the color space. For example, the first image to be denoised and the second image after noise reduction are three-channel color images, and the first image to be denoised and the second image after noise reduction are respectively subjected to alignment processing on the luminance channel image, and the aligned image input into the multi-scale denoising model is also the luminance image after alignment processing. In this way, by using only a single-channel luminance image to participate in the calculation, the calculation amount required by the noise reduction process can be further reduced while improving the noise reduction effect, which is beneficial to the deployment of mobile terminals.
[0044] As an optional embodiment, in order to reduce the interference of the noise existing in the first image on the alignment processing process, the alignment accuracy is improved. In this embodiment, the first image can be subjected to noise reduction processing before the first image and the second image are subjected to alignment processing, so as to reduce the noise existing in the first image. Specifically, before the step of performing alignment processing on the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image, the video denoising method provided in this embodiment can further include: performing noise reduction processing on the first transformed image to obtain a first denoised image; and performing motion estimation processing on the first denoised image and the second transformed image to obtain a motion estimation of the second transformed image relative to the first denoised image.
[0045] In this embodiment, in the process of performing noise reduction processing on the first transformed image, a low-precision noise reduction algorithm can be used to perform noise reduction processing on the first transformed image, so as to reduce the calculation amount of the noise reduction process; or a high-precision noise reduction algorithm can be used to perform noise reduction processing on the first transformed image, so as to improve the estimation accuracy of subsequent motion estimation and effectively reduce the interference of noise in the first image on the alignment processing process. Specifically, the precision of the noise reduction algorithm used by this embodiment to perform noise reduction processing on the first transformed image can be set according to actual application requirements, and is not specifically limited here.
[0046] In this embodiment, after obtaining the first denoised image and the motion estimation of the second transformed image relative to the first denoised image, the step of aligning the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image can be: aligning the first denoised image and the second transformed image according to the motion estimation to obtain the aligned image corresponding to the second transformed image. By aligning the second transformed image to the first denoised image, the embodiment can facilitate subsequent introduction of the previous frame denoised information, and reduce the ghosting problem caused by misalignment of pixels due to motion between the previous frame and the current frame.
[0047] In this embodiment, the alignment processing can include rotation, translation and other operations. The motion estimation processing provided by the embodiment is mainly to estimate the motion of the second transformed image relative to the first denoised image, so as to generate corresponding alignment parameters, facilitate the alignment process of the first denoised image and the second transformed image, and effectively improve the alignment accuracy of the alignment processing.
[0048] In step 103, the first image and the aligned image are subjected to multi-scale denoising processing to obtain a target image after denoising of the first image, and the target image replaces the first image in the to-be-denoised video.
[0049] In this embodiment, by performing multi-scale denoising processing on the first image and the aligned image with different scales, the embodiment can effectively reduce the ghosting problem caused by motion between the previous frame and the current frame while improving the consistency of the denoising effect of the previous frame and the current frame, thereby achieving the purpose of improving the denoising effect of video denoising.
[0050] Optionally, the embodiment can use a pre-trained multi-scale denoising model to perform multi-scale denoising processing on the first image and the aligned image to obtain the target image after denoising of the first image. Specifically, the step of performing multi-scale denoising processing on the first image and the aligned image to obtain the target image after denoising of the first image can be: inputting the first image and the aligned image into the trained multi-scale denoising model to obtain the target image after denoising of the first image.
[0051] Specifically, before the step of performing multi-scale denoising processing on the first image and the aligned image to obtain the target image after denoising of the first image, the video denoising method provided in the embodiment can further include: obtaining a training data set and a multi-scale denoising model to be trained, the training data set including a first training image after denoising of a first frame in a training video, a second training image of an arbitrary frame without denoising, and a labeled image after denoising of the second training image, the first training image including training images of multiple different sizes; inputting the first training image and the second training image into the multi-scale denoising model to be trained for denoising training to obtain a training denoising image corresponding to the second training image; and adjusting model parameters of the multi-scale denoising model to be trained until convergence according to a loss between the training denoising image and the labeled image, to obtain a trained multi-scale denoising model.
[0052] The manner of determining that the multi-scale denoising model to be trained is trained to convergence can be that, in a case where a training number of the multi-scale denoising model to be trained reaches a preset training number, it is determined that the model parameters of the multi-scale denoising model to be trained are converged. Specifically, the preset training number provided in the embodiment can be 50000 or 100000 or the like, and the specific value can be set according to actual application requirements, which is not specifically limited here.
[0053] In some embodiments, referring to Figure 2 , Figure 2 is a network structure diagram of the multi-scale denoising model provided in the embodiment, as shown in Figure 2 , the multi-scale denoising model provided in the embodiment can include multiple convolutional layers of different resolutions, and can include a full-resolution convolutional layer, a 1 / 4-resolution convolutional layer, a 1 / 16-resolution convolutional layer, and a 1 / 64-resolution convolutional layer, as shown in Figure 2 . Specifically, the process of inputting the first image and the aligned image of different sizes into the multi-scale denoising model for denoising processing provided in the embodiment can be that: the first image to be denoised is input into the full-resolution convolutional layer, the denoised aligned image is input into the smallest resolution convolutional layer of the multi-scale denoising model, so that the first image is gradually down-sampled to 1 / 64 of the full resolution, and the first image is image super-resolution reconstructed according to the denoised information of the aligned image at the 1 / 64-resolution convolutional layer to gradually up-sample to the original full resolution, so as to output the target image after denoising of the first image.
[0054] It should be noted that the lowest resolution of the multi-scale denoising model provided in the embodiment is not limited to 1 / 64 of the full resolution (i.e., the resolution of the second transformed image after the size transformation process), but can also be any resolution. Moreover, the position of the aligned image input into the multi-scale denoising model is also not limited to the lowest resolution of the model, but can also be any resolution convolutional layer in the multi-scale denoising model. For example, in Figure 2 , when the aligned image is 1 / 64 of the full resolution, and the input into the 1 / 16 or 1 / 4 or full resolution convolutional layer of the multi-scale denoising model, the aligned image can be first up-sampled to the resolution of the corresponding convolutional layer, and then the stepwise down-sampling and stepwise up-sampling processes are performed simultaneously with the first image. In this way, the denoising information in the denoised aligned image can be more effectively utilized to perform denoising processing on the first image to be denoised, so as to further improve the denoising effect on the first image, and at the same time improve the consistency of the denoising effect of the front and rear frames, while also reducing the ghosting problem caused by the motion of the front and rear frames.
[0055] As an optional embodiment, since the embodiment requires the first frame video frame image in the video to be denoised after denoising processing as the second image, when the first frame video frame image in the video to be denoised is not processed by denoising, and the first image is also selected as the first frame video frame image in the video to be denoised, the embodiment can directly perform denoising processing on the first frame video frame image in the video to be denoised, then take the first frame video frame image after denoising processing as the second image, and perform denoising processing by using the video denoising method provided in the embodiment, so as to realize the denoising processing on the first frame video frame image (i.e., the first image). Specifically, before the step of obtaining the first image to be denoised and the second image after denoising in the video to be denoised provided in the embodiment, the video denoising method provided in the embodiment can further include: when the first image is the first frame video frame image in the video to be denoised, performing denoising processing on the first image to obtain a second denoised image; and determining the second denoised image as the second image.
[0056] At this point, by using the video denoising method provided in the embodiment, the denoising processing on any one frame video frame image in the video to be denoised can be completed, so as to realize the denoising processing on the entire video to be denoised, and obtain the target video after denoising of the video to be denoised.
[0057] In order to better illustrate the video denoising method provided in the embodiment, please refer to Figure 3 , Figure 3 is another flowchart of the video denoising method provided in the embodiment. As Figure 3As shown, first, the first frame video frame image of the to-be-noise-reduced video after noise reduction processing is taken as a second image, then the second image and the first image needing noise reduction processing are transformed to obtain first and second transformed images with the same image size. Then, the first transformed image is subjected to noise reduction processing to obtain a first noise-reduced image. The first noise-reduced image and the second transformed image are subjected to motion estimation processing to obtain a motion estimation of the second transformed image relative to the first noise-reduced image, and corresponding alignment parameters are generated according to the motion estimation to align the first noise-reduced image and the second transformed image through the alignment parameters to obtain an aligned image corresponding to the second transformed image. After that, the noise-reduced aligned image and the first image to be noise-reduced are simultaneously input into a multi-scale noise reduction model to perform noise reduction processing on the first image to obtain a target image after noise reduction processing of the first image. Then, the target image is continuously taken as a second image, and the image of the next frame of the original first image is taken as a first image to continuously perform noise reduction processing on the images to be noise-reduced in the to-be-noise-reduced video until noise reduction processing on the entire to-be-noise-reduced video is completed.
[0058] To sum up, the embodiment of the present application provides a video noise reduction method, which comprises obtaining a first image to be noise-reduced and a second image after noise reduction in a to-be-noise-reduced video, the first image being any one frame video frame image in the to-be-noise-reduced video, the second image being a previous frame image of the first image, the size of the second image being not greater than the size of the first image, performing alignment processing on the first image and the second image to obtain an aligned image corresponding to the second image, performing multi-scale noise reduction processing on the first image and the aligned image to obtain a target image after noise reduction of the first image, and replacing the first image in the to-be-noise-reduced video with the target image. By using the embodiment of the present application, the ghost problem caused by multi-frame video noise reduction can be alleviated and avoided, and the noise reduction effect on the video can be effectively improved.
[0059] According to the method described in the above embodiment, the present embodiment will be further described from the perspective of a video noise reduction device. The video noise reduction device can be implemented as an independent entity or integrated in an electronic device, such as a terminal, which can include a mobile phone, a tablet computer, etc.
[0060] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of a video noise reduction device provided by the embodiment of the present application, as Figure 4 shown, the video noise reduction device 400 provided by the embodiment of the present application comprises an acquisition module 401, an alignment module 402, and a first noise reduction module 403.
[0061] The acquisition module 401 is configured to acquire a first image to be noise-reduced and a second image after noise reduction in a to-be-noise-reduced video, the first image being any one frame video frame image in the to-be-noise-reduced video, and the second image being a previous frame image of the first image.
[0062] The alignment module 402 is configured to perform alignment processing on the first image and the second image to obtain an aligned image corresponding to the second image.
[0063] The first noise reduction module 403 is configured to perform multi-scale noise reduction processing on the first image and the aligned image to obtain a target image in which the first image is replaced by the first noise-reduced image in the to-be-noise-reduced video.
[0064] In some embodiments, the first noise reduction module 403 is further configured to input the first image and the aligned image into the trained multi-scale noise reduction model to obtain the target image in which the first image is replaced by the first noise-reduced image.
[0065] In some embodiments, please refer to Figure 5 , Figure 5 is another structural schematic diagram of the video noise reduction device provided by the embodiment of the present application, as shown in Figure 5 The video noise reduction device provided by the embodiment of the present application can further include a transformation module 404, a second noise reduction module 405, a determination module 406, and a training module 407.
[0066] The transformation module 404 is configured to perform transformation processing on the first image and the second image to obtain first and second transformed images of the same size.
[0067] Specifically, the alignment module 402 is further configured to perform alignment processing on the first transformed image and the second transformed image to obtain an aligned image corresponding to the second transformed image.
[0068] The second noise reduction module 405 is configured to perform noise reduction processing on the first image to obtain a second noise-reduced image, in the case where the first image is a first frame video image in the to-be-noise-reduced video.
[0069] The determination module 406 is configured to determine the second noise-reduced image as the second image.
[0070] The training module 407 is configured to obtain a training data set and a to-be-trained multi-scale noise reduction model, the training data set including a first training image of a first frame in a training video after noise reduction processing, a second training image of an arbitrary frame without noise reduction processing, and a labeled image of the second training image after noise reduction processing, the first training image including training images of multiple different sizes; inputting the first training image and the second training image into the to-be-trained multi-scale noise reduction model for noise reduction training to obtain a training noise-reduced image corresponding to the second training image; and adjusting model parameters of the to-be-trained multi-scale noise reduction model according to a loss between the training noise-reduced image and the labeled image until convergence, to obtain a trained multi-scale noise reduction model.
[0071] In some embodiments, the transformation module 404 provided in the embodiments can include a downsampling processing, and the transformation module 404 is further configured to perform the downsampling processing on the first image and the second image of the first size to obtain the first transformed image and the second transformed image of a second size, and the second size is not greater than the first size.
[0072] In some embodiments, the alignment module 402 provided in the embodiments is further configured to perform the noise reduction processing on the first transformed image to obtain a first noise-reduced image, perform the motion estimation processing on the first noise-reduced image and the second transformed image to obtain a motion estimation of the second transformed image relative to the first noise-reduced image, and perform the alignment processing on the first noise-reduced image and the second transformed image according to the motion estimation to obtain an aligned image corresponding to the second transformed image.
[0073] In implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined as one or more entities, and the implementation of each of the above modules and / or units can refer to the implementation of the above method embodiments, and the beneficial effects achieved by the implementation can refer to the beneficial effects achieved by the above method embodiments, which will not be repeated here.
[0074] In addition, please refer to Figure 6 , Figure 6 is a structural schematic diagram of an electronic device provided in the embodiments of the present application, which can be a mobile terminal such as a smart phone, a tablet computer, or the like. As shown in Figure 6 , the electronic device 600 includes a processor 601 and a memory 602. The processor 601 is electrically connected to the memory 602.
[0075] The processor 601 is the control center of the electronic device 600, and connects each part of the electronic device 600 through various interfaces and lines. The processor 601 executes various functions of the electronic device 600 and processes data by running or loading an application program stored in the memory 602 and calling data stored in the memory 602, thereby monitoring the entire electronic device 600.
[0076] In the embodiments, the processor 601 in the electronic device 600 loads the instructions corresponding to the processes of one or more application programs into the memory 602, and runs the application program stored in the memory 602 by the processor 601, thereby implementing any step in the video noise reduction method provided in the above embodiments.
[0077] The electronic device 600 can implement the steps in any embodiment of the video noise reduction method provided in the embodiments of the present application, and thus can achieve the beneficial effects achieved by any video noise reduction method provided in the embodiments of the present application. Details can be found in the above embodiments, which will not be repeated here.
[0078] Please refer to Figure 7 , Figure 7 is another structural schematic diagram of an electronic device provided by an embodiment of the present application, as shown in Figure 7 , Figure 7 shows a specific structural block diagram of an electronic device provided by an embodiment of the present application, which can be used to implement the video noise reduction method provided in the above embodiments. The electronic device 700 can be a mobile terminal such as a smart phone or a notebook computer, etc.
[0079] The RF circuit 710 is used for receiving and sending electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 710 can include various existing circuit elements for performing these functions, for example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, etc. The RF circuit 710 can communicate with various networks such as the Internet, an intranet, a wireless network, or communicate with other devices through a wireless network. The wireless network can include a cellular phone network, a wireless local area network or a metropolitan area network. The wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g and / or 802.11n standards), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for mail, instant messaging and short messages, and any other suitable communication protocol, and even can include those protocols which are not yet developed.
[0080] The memory 720 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video noise reduction method described above in the embodiments, and the processor 780 can execute various functions and video noise reduction by running the software programs and modules stored in the memory 720.
[0081] The memory 720 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 720 can further include a memory disposed remotely from the processor 780, which can be connected to the electronic device 700 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0082] The input unit 730 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control. Specifically, the input unit 730 can include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display screen or touchpad, can collect user touch operations (such as user operations using fingers, styluses, etc. any suitable object or accessory on or near the touch-sensitive surface 731) on or near it, and drive the corresponding connection device according to the pre-set program. Optionally, the touch-sensitive surface 731 can include two parts of touch detection device and touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, and converts it into touch coordinates, and sends it to the processor 780, and can receive the command from the processor 780 and execute it. In addition, the touch-sensitive surface 731 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 can also include other input devices 732. Specifically, the other input devices 732 can include one or more of a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc.
[0083] The display unit 740 can be used to display information input by a user or provided to the user, as well as various graphical user interfaces of the electronic device 700, which can be composed of graphics, text, icons, video, and any combination thereof. The display unit 740 can include a display panel 741, which can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like, optionally. Further, the touch-sensitive surface 731 can cover the display panel 741, and when the touch-sensitive surface 731 detects a touch operation thereon or adjacent thereto, transmit to the processor 780 to determine the type of touch event, and then the processor 780 provides corresponding visual output on the display panel 741 according to the type of touch event. Although in the figure, the touch-sensitive surface 731 and the display panel 741 are implemented as two independent components to realize the input and output functions, in some embodiments, the touch-sensitive surface 731 and the display panel 741 can be integrated to realize the input and output functions.
[0084] The electronic device 700 can further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor can include an ambient light sensor that can adjust the brightness of the display panel 741 according to the brightness of ambient light, and a proximity sensor that can generate an interrupt when the cover is closed or turned off. As one of the motion sensors, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and when at rest, can detect the magnitude and direction of gravity, which can be used for applications such as identifying the posture of the mobile phone (such as switching between landscape and portrait screens, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometers, tapping), and the like; as well as other sensors that the electronic device 700 can be configured, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, and the like, which will not be described here.
[0085] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the electronic device 700. The audio circuit 760 can convert received audio data into an electrical signal, transmit the electrical signal to the speaker 761, and convert the electrical signal into a sound signal by the speaker 761 for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and converted into audio data, and then output to the processor 780 for processing, and then transmitted to another terminal via the RF circuit 710, or output to the memory 720 for further processing. The audio circuit 760 can also include an earphone jack to provide communication between an external earphone and the electronic device 700.
[0086] The electronic device 700 can help the user to receive requests, send information, etc. through the transmission module 770 (e.g., a Wi-Fi module), which provides the user with wireless broadband Internet access. Although the transmission module 770 is shown in the figure, it can be understood that it does not belong to the necessary components of the electronic device 700, and can be omitted as needed without changing the essence of the application.
[0087] The processor 780 is the control center of the electronic device 700, which connects various parts of the entire mobile phone through various interfaces and lines, executes various functions of the electronic device 700 and processes data by running or executing software programs and / or modules stored in the memory 720 and calling data stored in the memory 720, thereby monitoring the entire electronic device. Optionally, the processor 780 can include one or more processing cores; in some embodiments, the processor 780 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 780.
[0088] The electronic device 700 further includes a power supply 790 (such as a battery) for supplying power to various components, and in some embodiments, the power supply can be logically connected to the processor 780 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 790 can also include one or more direct or alternating power sources, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and any other components.
[0089] Although not shown, the electronic device 700 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be described here. In this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors to implement any of the steps in the video noise reduction method provided by the above embodiments.
[0090] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined as the same or several entities, and the specific implementation of each of the above modules can be referred to the method embodiments described above, which will not be described here.
[0091] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware by the instructions, and the instructions can be stored in a computer readable storage medium and loaded and executed by a processor. Therefore, the embodiments of the present application provide a storage medium, wherein a plurality of instructions are stored, and the instructions can be executed by a processor to implement any step in the video noise reduction method provided by the above embodiments.
[0092] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0093] Since the instructions stored in the storage medium can execute the steps in any embodiment of the video noise reduction method provided by the embodiments of the present application, the beneficial effects of any video noise reduction method provided by the embodiments of the present application can be achieved, which are described in detail in the above embodiments and will not be repeated here.
[0094] The above describes in detail a video noise reduction method, device, electronic equipment and storage medium provided by the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples. The above embodiment is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the principles of the present application, the specific implementation manners and application ranges can be changed. In summary, the content of the specification should not be understood as a limitation of the present application. Moreover, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered as the protection scope of the present application.
Claims
1. A video noise reduction method, characterized in that, include: Obtain a first image to be denoised and a second image after denoising from the video to be denoised. The first image is any frame of the video to be denoised, and the second image is the frame preceding the first image. The size of the second image is not greater than the size of the first image. The first image and the second image are aligned to obtain the aligned image corresponding to the second image. The first image and the aligned image are input into a trained multi-scale denoising model to obtain a target image after denoising the first image, and the target image replaces the first image in the video to be denoised. The multi-scale denoising model is used to downsample the first image and, with reference to the denoised information of the aligned image, to perform super-resolution reconstruction on the downsampled first image to obtain the denoised target image of the first image.
2. The method as described in claim 1, characterized in that, Before the step of aligning the first image and the second image to obtain the aligned image corresponding to the second image, the method further includes: The first image and the second image are transformed to obtain a first transformed image and a second transformed image of the same size. The step of aligning the first image and the second image to obtain the aligned image corresponding to the second image includes: Alignment processing is performed on the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image.
3. The method as described in claim 2, characterized in that, The transformation process includes downsampling; The transformation process of the first image and the second image to obtain a first transformed image and a second transformed image of the same size includes: Downsampling is performed on the first image and the second image of the first size to obtain a first transformed image and a second transformed image of the second size, wherein the second size is not larger than the first size.
4. The method as described in claim 2, characterized in that, Before the step of aligning the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image, the method further includes: The first transformed image is subjected to noise reduction processing to obtain the first denoised image; Motion estimation processing is performed on the first denoised image and the second transformed image to obtain the motion estimate of the second transformed image relative to the first denoised image; The step of aligning the first transformed image and the second transformed image to obtain the aligned image corresponding to the second transformed image includes: Based on the motion estimation, the first denoised image and the second transformed image are aligned to obtain the aligned image corresponding to the second transformed image.
5. The method as described in claim 1, characterized in that, Before the step of acquiring the first image to be denoised and the second image after denoising in the video to be denoised, the method further includes: If the first image is the first frame of the video to be denoised, the first image is denoised to obtain the second denoised image. The second denoised image is determined as the second image.
6. The method as described in claim 1, characterized in that, Before the step of performing multi-scale denoising processing on the first image and the aligned image to obtain the denoised target image of the first image, the method further includes: Obtain a training dataset and a multi-scale denoising model to be trained. The training dataset includes a first training image with denoising processing of the first frame in the training video, a second training image without denoising processing of any frame, and a labeled image of the second training image with denoising processing. The first training image includes training images of various sizes. The first training image and the second training image are input into the multi-scale denoising model to be trained for denoising training, and the training denoised image corresponding to the second training image is obtained. Based on the loss between the training denoised image and the labeled image, the model parameters of the multi-scale denoising model to be trained are adjusted until convergence, thus obtaining the trained multi-scale denoising model.
7. A video noise reduction device, characterized in that, include: The acquisition module is used to acquire a first image to be denoised and a second image after denoising in the video to be denoised, wherein the first image is any frame of the video to be denoised, and the second image is the frame preceding the first image; An alignment module is used to align the first image and the second image to obtain an aligned image corresponding to the second image. The first noise reduction module is used to input the first image and the aligned image into a trained multi-scale noise reduction model to obtain a target image after noise reduction of the first image, and replace the first image in the video to be denoised with the target image; wherein, the multi-scale noise reduction model is used to downsample the first image, and with reference to the information after noise reduction of the aligned image, to perform super-resolution reconstruction on the downsampled first image to obtain the target image after noise reduction of the first image.
8. The apparatus as claimed in claim 7, characterized in that, The video noise reduction device also includes: A transformation module is used to transform the first image and the second image to obtain a first transformed image and a second transformed image of the same size. The alignment module is further configured to perform alignment processing on the first transformed image and the second transformed image to obtain an aligned image corresponding to the second transformed image.
9. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Video noise reduction method, intelligent terminal and computer readable storage medium
CN114612312A
Video denoising method, video processing method and device
CN115063301A
Video noise reduction method and device, electronic equipment and storage medium
CN116664410A