Video denoising method and device
Through multi-chain recurrent neural network, the RAW format video frames are aligned with double-visual alignment, which solves the problem of poor denoising effect of RAW videos and achieves efficient denoising and video quality improvement.
Patent Information
- Application Number
- CN202311659549.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively denoise noise in RAW format videos, affecting the video quality.
The multi-chain recurrent neural network model is used to perform double alignment of video frames. By obtaining the subset of adjacent video frames in the sequence frame set, double-branched explicit state alignment and hidden state convolution are performed, and denoised video is finally reconstructed.
It improves the video denoising effect, reduces the denoising time, improves the video quality and denoising efficiency, and enhances the utilization of video timing information.
Smart Images

Figure CN120107093A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video denoising method and device. Background Art
[0002] With the development of science and technology, electronic devices support more and more functions. For example, you can shoot RAW format videos through smartphones and drones. More and more photographers are starting to use RAW format to shoot videos to get higher quality. However, due to the working principle of camera sensors, the noise of RAW format videos is more obvious. Therefore, denoising RAW format videos has become a focus of attention. Summary of the invention
[0003] The present disclosure provides a video denoising method and device to improve the video denoising effect while reducing the denoising time, which can improve the quality of the acquired video while improving the denoising efficiency. The technical solution of the present disclosure is as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a video denoising method is provided, comprising:
[0005] Obtain a sequence frame set corresponding to the first unprocessed RAW video;
[0006] Inputting any subset of sequence frames in the set of sequence frames into any recurrent neural network unit in a target multi-chain recurrent neural network (RNN) model, wherein the subset of sequence frames includes at least two adjacent video frames;
[0007] Using any of the recurrent neural network units to perform explicit and implicit dual alignment processing on the sequence frame subset to obtain hidden state representation information output by any of the recurrent neural network units;
[0008] The hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements.
[0009] According to some embodiments, the step of performing explicit and implicit dual alignment processing on the sequence frame subset using any one of the recurrent neural network units to obtain hidden state representation information output by any one of the recurrent neural network units includes:
[0010] Control any one of the recurrent neural network units to perform dual-branch explicit state alignment processing on the sequence frame subset based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the sequence frame subset;
[0011] Perform hidden state convolution based on weighted convolution on the explicit state alignment result to obtain hidden state representation information output by any recurrent neural network unit.
[0012] According to some embodiments, the controlling any one of the recurrent neural network units to perform dual-branch explicit state alignment processing on the sequence frame subset based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the sequence frame subset includes:
[0013] Control any one of the recurrent neural network units to obtain the single-channel data optical flow information corresponding to the sequence frame subset;
[0014] Control any one of the recurrent neural network units to perform Bell conversion processing on the single-channel data corresponding to the sequence frame subset to obtain the four-channel RGGB data corresponding to the sequence frame subset;
[0015] Obtaining feature information corresponding to the four-channel RGGB data;
[0016] Acquire an offset according to the optical flow information and the feature information;
[0017] The offset, the optical flow information and the feature information are input into a deformable convolutional network to obtain a display state alignment result corresponding to the sequence frame subset.
[0018] According to some embodiments, performing hidden state convolution based on weighted convolution on the visible state alignment result to obtain hidden state representation information output by any recurrent neural network unit includes:
[0019] According to the display state alignment result, an offset network is used to obtain an offset of a pixel position of any channel;
[0020] Using a weight network to obtain the weight corresponding to the pixel position of any channel;
[0021] According to the offset of the pixel position of any channel and the weight, the hidden state representation information output by any recurrent neural network unit is obtained.
[0022] According to some embodiments, obtaining a set of sequence frames corresponding to the first unprocessed RAW video includes:
[0023] Processing the first unprocessed RAW video using an optical flow algorithm to obtain a processed first unprocessed RAW video;
[0024] Key frame extraction is performed on the processed first unprocessed RAW video to obtain a sequence frame set corresponding to the first unprocessed RAW video.
[0025] According to some embodiments, the step of inputting any subset of sequence frames in the set of sequence frames into any recurrent neural network unit in the target multi-chain recurrent neural network comprises:
[0026] Obtaining sequence frame information corresponding to any sequence frame subset in the sequence frame set;
[0027] In a case where the sequence frame information indicates that the sequence frame subset is the first subset of the sequence frame set, inputting the sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network;
[0028] or
[0029] When the sequence frame information indicates that the sequence frame subset is not the first subset of the sequence frame set, obtaining a denoising result of the same scale corresponding to the sequence frame subset;
[0030] Feature fusion is performed based on the denoising result and the sequence frame subset to obtain a fused sequence frame subset, and the fused sequence frame subset is input into any recurrent neural network unit in the target multi-chain recurrent neural network.
[0031] According to some embodiments, reconstructing the hidden state representation information to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements includes:
[0032] Fusing the hidden state representation information with all hidden state representation information output by the previous coding layer to obtain fused hidden state representation information;
[0033] The output result of the topmost chain in the target multi-chain recurrent neural network is reorganized according to the input order of the sequence frame set to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirement.
[0034] According to some embodiments, inputting a subset of sequence frames in the set of sequence frames into any recurrent neural network unit in a target multi-chain recurrent neural network comprises:
[0035] Down-sampling the sequence frame subset to obtain at least one sequence frame subset different in size from the sequence frame subset;
[0036] The at least one sequence frame subset is input into the recurrent neural network unit corresponding to any sequence frame subset in the target multi-chain recurrent neural network.
[0037] According to some embodiments, the method further comprises:
[0038] Acquire a training sample set, wherein the training sample set includes a first training sample subset collected by a first camera device and a second training sample subset collected by a second camera device, the first camera device and the second camera device correspond to the same shooting conditions and different noise levels;
[0039] The initial multi-chain recurrent neural network is trained using the training sample set to obtain the target multi-chain recurrent neural network.
[0040] According to a second aspect of an embodiment of the present disclosure, a video denoising device is provided, comprising:
[0041] A set acquisition unit, used for acquiring a sequence frame set corresponding to the first unprocessed RAW video;
[0042] A subset input unit, used for inputting any sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the sequence frame subset includes at least two adjacent video frames;
[0043] An information acquisition unit, used to perform explicit and implicit dual alignment processing on the sequence frame subset using any of the recurrent neural network units to obtain hidden state representation information output by any of the recurrent neural network units;
[0044] The video acquisition unit is used to reconstruct the hidden state representation information and acquire a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets the noise requirement.
[0045] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0046] processor;
[0047] a memory for storing instructions executable by the processor;
[0048] The processor is configured to execute the instructions to implement the video denoising method described in any one of the aforementioned aspects.
[0049] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the video denoising method described in any one of the preceding aspects.
[0050] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements any one of the methods described in the preceding aspects when executed by a processor.
[0051] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0052] In some or related embodiments, a sequence frame set corresponding to a first unprocessed RAW video is obtained; any sequence frame subset in the sequence frame set is input into any recurrent neural network unit in a target multi-chain recurrent neural network model, wherein the sequence frame subset includes at least two adjacent video frames; any recurrent neural network unit is used to perform explicit and implicit dual alignment processing on the sequence frame subset to obtain hidden state representation information output by any recurrent neural network unit; the hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements. Therefore, explicit and implicit dual alignment processing can be performed to improve the adequacy of video timing information utilization, reduce the impact of video frame deformation on denoising results, improve explicit and implicit dual alignment processing effects, use multi-frame joint denoising to reduce the situation where single-frame denoising results in poor denoising quality, and process a subset of sequence frames to improve video denoising effects while reducing denoising duration, and improve the quality of the acquired video while improving denoising efficiency.
[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0055] Figure 1 is a flow chart of a video denoising method according to an exemplary embodiment;
[0056] Figure 2 is a flow chart of a video denoising method according to an exemplary embodiment;
[0057] Figure 3 is a schematic diagram of a framework of an explicit and implicit dual alignment structure according to an exemplary embodiment;
[0058] Figure 4 is a flow chart of a dual-branch alignment method based on deformable convolution according to an exemplary embodiment;
[0059] Figure 5 is a flow chart of a dual-branch alignment method based on deformable convolution according to an exemplary embodiment;
[0060] Figure 6 is a flow chart of a video denoising method according to an exemplary embodiment;
[0061] Figure 7 is a block diagram of a video denoising device according to an exemplary embodiment;
[0062] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0063] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0064] The embodiments of the present disclosure provide a video denoising method and device. In some embodiments, the video denoising method and information processing method, communication method and other terms can be interchangeable, the video denoising device and information processing device, communication device and other terms can be interchangeable, and the information processing system and communication system and other terms can be interchangeable.
[0065] The embodiments of the present disclosure are not exhaustive, but are only illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined, for example, some or all of the steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0066] In each embodiment of the present disclosure, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between the embodiments are consistent and can be referenced to each other, and the technical features in different embodiments can be combined to form a new embodiment based on their internal logical relationships.
[0067] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0068] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular form, such as "a", "an", "the", "above", "said", "aforementioned", "this", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun after the article may be understood as a singular expression or a plural expression.
[0069] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0070] In some embodiments, the terms “at least one,” “one or more,” “a plurality of,” “multiple,” etc. may be used interchangeably.
[0071] In some embodiments, "at least one of A and B", "A and / or B", "A in one case, B in another case", "in response to one case A, in response to another case B", etc., may include the following technical solutions according to the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). When there are more branches such as A, B, C, etc., the above is also similar.
[0072] In some embodiments, the recording method of "A or B" may include the following technical solutions according to the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). When there are more branches such as A, B, C, etc., the above is also similar.
[0073] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects, and do not constitute restrictions on the position, order, priority, quantity or content of the description objects. The statement of the description object refers to the description in the context of the claims or embodiments, and should not constitute unnecessary restrictions due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields", and the "first" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number, and can be one or more. Taking the "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes may be the same or different. For example, if the description object is "device", then the "first device" and the "second device" may be the same device or different devices, and their types may be the same or different. For another example, if the description object is "information", then the "first information" and the "second information" may be the same information or different information, and their contents may be the same or different.
[0074] In some embodiments, “including A”, “comprising A”, “used to indicate A”, and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0075] In some embodiments, terms such as "in response to ...", "in response to determining ...", "in the case of ...", "at the time of ...", "when ...", "if ...", "if ...", etc. can be used interchangeably.
[0076] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "no more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0077] In some embodiments, "terminal" or "terminal device" can be referred to as "user equipment (UE)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0078] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0079] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0080] Figure 1 is a flow chart of a video denoising method according to an exemplary embodiment. Figure 1 As shown, the video denoising method can be used in a video denoising scenario, comprising the following steps:
[0081] In step S11, a sequence frame set corresponding to a first unprocessed RAW video is obtained;
[0082] According to some embodiments, unprocessed RAW video may refer to, for example, a video in RAW format, such as an unprocessed RAW video shot by a camera supporting the RAW format, or a video acquired by a camera supporting the RAW format, such as a video transmitted by a camera supporting the RAW format through wireless transmission.
[0083] In some embodiments, the first unprocessed RAW video refers to a video to be subjected to video denoising. The disclosed embodiments do not limit the method for obtaining the first unprocessed RAW video. For example, it may be acquired by an electronic device through a camera device that supports the RAW format on the electronic device, or it may be acquired through a server. Among them, the first unprocessed RAW video does not specifically refer to a fixed video. For example, when the gain corresponding to the first unprocessed RAW video changes, the first unprocessed RAW video may also change accordingly. For example, when the acquisition time point of the first unprocessed RAW video changes, the first unprocessed RAW video may also change accordingly.
[0084] According to some embodiments, the sequence frame set may be, for example, a collection of at least two video frames, and the at least two video frames may be, for example, arranged in sequence. The sequence frame set does not specifically refer to a fixed set. For example, when the number of frames corresponding to the sequence frame set changes, the sequence frame set may also change accordingly. For example, when the first unprocessed RAW video changes, the sequence frame set may also change accordingly.
[0085] According to some embodiments, when an electronic device executes a video denoising method, a set of sequence frames corresponding to a first unprocessed RAW video may be obtained.
[0086] In step S12, any subset of sequence frames in the sequence frame set is input into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the subset of sequence frames includes at least two adjacent video frames;
[0087] According to some embodiments, the sequence frame subset may be, for example, a set including at least two video frames, the at least two video frames may be, for example, adjacent video frames, and the sequence frame subset may be a subset included in the sequence frame set. The sequence frame subset does not specifically refer to a fixed set. For example, when the video frame corresponding to the sequence frame subset changes, the sequence frame set may also change accordingly. For example, when the number of video frames corresponding to the sequence frame subset changes, the sequence frame set may also change accordingly.
[0088] In some embodiments, the recurrent neural network unit may be, for example, a processing unit included in a recurrent neural network, wherein a chain may correspond to a recurrent neural network unit. The recurrent neural network unit may, for example, perform explicit and implicit double alignment processing on a sequence frame.
[0089] According to some embodiments, when a sequence frame set is obtained, any sequence frame subset in the sequence frame set can be input into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the sequence frame subset includes at least two adjacent video frames.
[0090] In step S13, any recurrent neural network unit is used to perform explicit and implicit double alignment processing on the sequence frame subset to obtain the hidden state representation information output by any recurrent neural network unit;
[0091] According to some embodiments, explicit and implicit dual alignment may include explicit state alignment and hidden state alignment. Hidden state representation information may refer to information corresponding to a subset of sequence frames output by any recurrent neural network unit. For example, it may be information output by any recurrent neural network unit after completing hidden state alignment for a subset of sequence frames.
[0092] According to some embodiments, any recurrent neural network unit can be used to perform explicit and implicit double alignment processing on a subset of sequence frames to obtain hidden state representation information output by any recurrent neural network unit.
[0093] In one embodiment of the present disclosure, for example, explicit state alignment processing may be performed first, and then hidden state alignment processing may be performed after feature fusion of information output by the explicit state alignment operation, wherein the feature fusion may be, for example, SwinTransformer feature fusion.
[0094] In step S14, the hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets the noise requirement.
[0095] According to some embodiments, the second unprocessed RAW video may be, for example, a video obtained after denoising. The second unprocessed RAW video does not specifically refer to a fixed video. For example, when the explicit and implicit dual alignment process changes or the reconstruction process changes, the second unprocessed RAW video may also change accordingly. For example, when the first unprocessed RAW video changes, the second unprocessed RAW video may also change accordingly.
[0096] In some embodiments, the hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements.
[0097] In some or related embodiments, a sequence frame set corresponding to a first unprocessed RAW video is obtained; any sequence frame subset in the sequence frame set is input into any recurrent neural network unit in a target multi-chain recurrent neural network model, wherein the sequence frame subset includes at least two adjacent video frames; any recurrent neural network unit is used to perform explicit and implicit dual alignment processing on the sequence frame subset to obtain hidden state representation information output by any recurrent neural network unit; the hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirements. Therefore, explicit and implicit dual alignment processing can be performed to improve the adequacy of video timing information utilization, reduce the impact of video frame deformation on denoising results, improve the explicit and implicit dual alignment processing effect, use multi-frame joint denoising to reduce the situation where single-frame denoising results in poor denoising quality, balance the denoising intensity, retain real details while efficiently denoising, avoid excessive smoothing image quality loss, improve video denoising effect, and improve the quality of the acquired video. In addition, since the sequence frame subset includes at least two adjacent video frames, motion change information can be highlighted and the accuracy of video denoising can be improved.
[0098] Figure 2 is a flow chart of a video denoising method according to an exemplary embodiment. Figure 2 As shown, the video denoising method can be used in a RAW video denoising scenario, and includes the following steps:
[0099] In step S21, a sequence frame set corresponding to a first unprocessed RAW video is obtained;
[0100] The specific process is as above and will not be repeated here.
[0101] According to some embodiments, the method of the embodiments of the present disclosure may further include, for example, inverse ISP processing. For example, the inverse ISP model may be used to predict or optimize in the RGB video to obtain a RAW video corresponding to the RGB video. Among them, the inverse ISP model may be obtained, for example, by using an unsupervised or self-supervised Consistency Loss training method. The use of inverse ISP processing can provide better quality input videos for other task modules. Secondly, the network structure and loss function can be designed by adding image prior knowledge, a multi-branch or encoding-decoding structure can be used, and an inverse ISP model can be generated using GAN or the like to enhance the restoration effect of the inverse ISP.
[0102] According to some embodiments, Figure 3 is a schematic diagram of a framework of an explicit and implicit dual alignment structure according to an exemplary embodiment. Figure 3 As shown in the figure, the RNN structure can accept noisy videos and complete the linkage with historical denoising results, and use the reuse module to reuse the denoising results of historical video frames to enhance the denoising effect; for example, the input large-scale few-channel input can be converted into a small-scale multi-channel hidden state, thereby integrating the rich features of historical frames and the guiding information of future frames; finally, the hidden state is reconstructed into a noise-free RAW video frame through the reconstruction generator.
[0103] According to some embodiments, the first unprocessed RAW video is processed using an optical flow algorithm to obtain a processed first unprocessed RAW video;
[0104] The processed first unprocessed RAW video is subjected to key frame extraction to obtain a sequence frame set corresponding to the first unprocessed RAW video. Therefore, the optical flow algorithm can be used to process the video, motion information can be enhanced, and key frame extraction can be performed on the video, which can reduce the processing redundancy of the network model and improve the efficiency of video denoising.
[0105] The key frame extraction may be performed, for example, based on the video content or based on the video theme, which is not limited in the embodiments of the present disclosure.
[0106] In step S22, any subset of sequence frames in the sequence frame set is input into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the subset of sequence frames includes at least two adjacent video frames;
[0107] The specific process is as above and will not be repeated here.
[0108] According to some embodiments, the video frames corresponding to the sequence frame subset may be, for example, 2 frames or 3 frames. Different sequence frame subsets may correspond to the same number of frames or different numbers of frames. Therefore, the processing accuracy of each recurrent neural network unit can be increased, more information can be provided for a recurrent neural network unit, and the accuracy of the recurrent neural network in determining the temporal dependency relationship can be improved.
[0109] According to some embodiments, the method specifically comprises:
[0110] Obtaining sequence frame information corresponding to any sequence frame subset in the sequence frame set;
[0111] In a case where the sequence frame information indicates that the sequence frame subset is the first subset of the sequence frame set, inputting the sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network;
[0112] or
[0113] When the sequence frame information indicates that the sequence frame subset is not the first subset of the sequence frame set, obtaining a denoising result of the same scale corresponding to the sequence frame subset;
[0114] According to the denoising results and the sequence frame subset, feature fusion is performed to obtain the fused sequence frame subset, and the fused sequence frame subset is input into any recurrent neural network unit in the target multi-chain recurrent neural network. Therefore, the historical denoising results can be forwarded, which can improve the video denoising effect.
[0115] According to some embodiments, inputting a subset of sequence frames in a set of sequence frames into any recurrent neural network unit in a target multi-chain recurrent neural network comprises:
[0116] Down-sampling the sequence frame subset to obtain at least one sequence frame subset having a size different from the sequence frame subset;
[0117] At least one subset of sequence frames is input into the recurrent neural network unit corresponding to any subset of sequence frames in the target multi-chain recurrent neural network. Therefore, the multi-scale invariance of the video can be achieved and the denoising effect of the video can be improved.
[0118] According to some embodiments, downsampling the sequence frame subset may be, for example, downsampling to half or one quarter of the original size of the sequence frame subset. This is not limited in the embodiments of the present disclosure.
[0119] In step S23, any recurrent neural network unit is controlled to perform dual-branch explicit state alignment processing on the sequence frame subset based on the deformable convolutional network to obtain the explicit state alignment result corresponding to the sequence frame subset;
[0120] The specific process is as above and will not be repeated here.
[0121] According to some embodiments, controlling any recurrent neural network unit to perform dual-branch explicit state alignment processing on a subset of sequence frames based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the subset of sequence frames includes:
[0122] Control any recurrent neural network unit to obtain single-channel data optical flow information corresponding to a subset of sequence frames;
[0123] Control any recurrent neural network unit to perform Bell conversion processing on the single-channel data corresponding to the sequence frame subset to obtain the four-channel RGGB data corresponding to the sequence frame subset;
[0124] Obtain feature information corresponding to four-channel RGGB data;
[0125] Obtain the offset based on the optical flow information and feature information;
[0126] The offset, optical flow information and feature information are input into the deformable convolutional network to obtain the display state alignment result corresponding to the subset of sequence frames. Therefore, the single-channel data can be processed into four channels, which can enrich the feature information extraction, reduce the situation where the feature information is insufficient and the video denoising cannot be performed, and improve the video denoising effect.
[0127] According to some embodiments, Figure 4 is a flowchart of a dual-branch alignment method based on deformable convolution according to an exemplary embodiment. Figure 4 As shown in the figure, the upper part uses a single channel input to extract the optical flow information. Since the adjacent pixels of the single channel data are adjacent, the optical flow information can be improved. t-1,t+1 The accuracy of the acquisition; the lower part uses Bell conversion to create four-channel RGGB data for single-channel data. The feature extraction network extracts features and upsamples them, which can increase the information density of the input and improve the extracted feature information. Therefore, the output of the dual branches can be used to obtain an alignment offset that is more accurate than the optical flow [o ′ t-1→t ,o ′ t+1→t ], the optimized offset and feature information with higher information density are combined and input into the deformable convolutional network to obtain the explicit state alignment result.
[0128] In step S24, the hidden state convolution based on weighted convolution is performed on the visible state alignment result to obtain the hidden state representation information output by any recurrent neural network unit;
[0129] The specific process is as above and will not be repeated here.
[0130] According to some embodiments, performing hidden state convolution based on weighted convolution on the explicit state alignment result to obtain hidden state representation information output by any recurrent neural network unit includes:
[0131] According to the display state alignment result, an offset network is used to obtain the offset of the pixel position of any channel;
[0132] Use a weight network to obtain the weight corresponding to the pixel position of any channel;
[0133] According to the offset and weight of the pixel position of any channel, the hidden state representation information output by any recurrent neural network unit is obtained. Therefore, the accuracy of obtaining the hidden state representation information can be improved by obtaining the hidden state representation information according to the offset and weight, and each frame of video is reduced from a single channel to RGGB four-channel data through the Bell matrix. When adjacent pixels are not adjacent in the original single-channel data, only the optical flow network is used to calculate the optical flow of the input multiple frames. Due to the low information density, the amount of feature information fused is insufficient, resulting in a poor denoising effect. This can improve the video denoising effect.
[0134] According to some embodiments, Figure 5 is a flowchart of a dual-branch alignment method based on deformable convolution according to an exemplary embodiment. Figure 5 As shown, for example, an offset network can be used to obtain several input data points that may correspond to each data point in the hidden state according to a weight sampling method; at the same time, a weight network can also be used to obtain the weights of the corresponding data points. For example, the sum of the weights of the points with corresponding relationships is 1. The full feature map composed of the corresponding data points is multiplied and summed with the weight map in the channel dimension to obtain the aligned fused hidden state representation information. The data point can be, for example, a pixel point in a video frame.
[0135] For example, the following formula can be used to obtain hidden state representation information:
[0136]
[0137]
[0138]
[0139]
[0140] Among them, s t Indicates the selected point position obtained by the offset network om t Alignment results in the displayed state The sampling is done above and then the channels are connected;
[0141] h′ t Indicates that the hidden state h tand t The results are connected at the corresponding channel positions.
[0142] Represents the weight graph wm generated by the weight network t and h′ t The result of multiplying is the principle of weighted average.
[0143] h t Indicates that The results are added up in the corresponding channel range in the channel direction and connected to obtain the final hidden state representation information h of this unit. t .
[0144] In step S25, the hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets the noise requirement.
[0145] The specific process is as above and will not be repeated here.
[0146] According to some embodiments, reconstructing the hidden state representation information to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets the noise requirement includes:
[0147] Fusing the hidden state representation information with all the hidden state representation information output by the previous encoding layer to obtain the fused hidden state representation information;
[0148] The output result of the top chain in the target multi-chain recurrent neural network is reorganized according to the input order of the sequence frame set to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirement. Therefore, all the hidden state representation information output by the previous coding layer can be fused, which can improve the denoising effect and the visual quality of the video.
[0149] According to some embodiments, the encoded hidden state representation information output by the target multi-chain recurrent neural network is used as a condition and input into the Swin Transformer decoder. The Swin Transformer is initialized with the hidden state representation information, and at each encoder decoder layer, the encoded hidden state representation information is fused with the hidden state representation information output by the previous layer. At each layer of the Swin Transformer decoder, the pixel information of the image is predicted through the MLP head, which can reconstruct the input image sequence while improving its visual quality.
[0150] According to some embodiments, Figure 6 is a flow chart of a video denoising method according to an exemplary embodiment. Figure 6As shown, the input RAW video can be input into RNN in multiple frames, and then input into the multi-chain RNN network in a unit manner, wherein the RNN unit of each chain can be used to perform explicit and implicit dual alignment processing on multiple frames of RAW video in parallel, which can improve the efficiency of explicit and implicit dual alignment, and can complete the explicit and implicit dual alignment task of multiple scales in the same period of time. In the alignment process, for example, the optical flow information and multi-channel feature information can be extracted from the explicit state, and then the alignment and fusion operation is performed through the deformable convolution network to obtain the feature information after the fusion of the input multiple frames. Then, the historical denoised hidden state can be fused and aligned with the noisy input features by weighted convolution to achieve the purpose of obtaining multi-channel denoising features. Finally, the noise-free frame can be reconstructed by the reconstruction module based on SwinTransformer, and the multi-scale denoised frames can be reconstructed into the original scale video by multi-level linkage, and the denoising result reuse module can be used to pass the denoising information down, and the output of the top chain can be collected and reorganized in the order of input to obtain the denoised RAW video.
[0151] According to some embodiments, the method further comprises:
[0152] Acquire a training sample set, wherein the training sample set includes a first training sample subset collected by a first camera device and a second training sample subset collected by a second camera device, the first camera device and the second camera device correspond to the same shooting conditions and different noise levels;
[0153] The initial multi-chain recurrent neural network is trained using a training sample set to obtain a target multi-chain recurrent neural network. Therefore, the model can be trained using RAW videos with different noise levels, which can improve the denoising ability of the multi-chain recurrent neural network under various noise conditions, improve the generalization ability of the multi-chain recurrent neural network, and improve the robustness and denoising accuracy of the multi-chain recurrent neural network.
[0154] For example, both the first camera and the second camera can capture videos in RAW format. For example, the shooting conditions corresponding to the first camera and the second camera are the same, for example, the white balance, exposure parameters and scene positions are the same. Among them, the noise levels corresponding to the first camera and the second camera are different, for example, they can be adjusted by adjusting the gain and / or brightness gain ISO. Among them, for example, detailed metadata METADATA metadata of the first camera and the second camera can be recorded to mark different data. Among them, for example, the shooting scene can also be selected, for example, representative static and motion training scenes can be selected to improve the training effect of the model. When the shot video is obtained, the shot video can be enhanced in the spatial domain or time domain to improve the richness of the training sample set, and the robustness of the multi-chain recurrent neural network can be improved.
[0155] Among them, the video input to the recurrent neural network can be, for example, a sequence of frames of a RAW video. For example, the difference between adjacent input video frames can highlight motion changes. The optical flow algorithm can also be used to process the RAW video to enhance the motion information of the RAW video. In addition, key frames in the RAW video can be extracted as input to reduce redundancy. The input of multimodal channels such as motion and scenes can be increased to improve the applicability and denoising accuracy of the recurrent neural network. Finally, data enhancement methods such as rotation and cropping can be used to increase sample diversity.
[0156] According to some embodiments, during the recurrent neural network training process, and when the RAW video is input into the recurrent neural network, the RAW video can be input into multiple recurrent neural units of the recurrent neural network, which can reduce the number of timing expansion steps, speed up the training speed and reasoning process of the model, reduce the situation where the number of parameters is large due to the use of multiple recurrent neural networks, reduce the number of parameters, and improve the nonlinear representation ability of the recurrent neural network, reduce the model training time and improve the model training efficiency.
[0157] In some embodiments, during the training of the recurrent neural network, for example, the historical denoising results and the current input can be fused to enhance the RNN model's ability to model dependencies between time steps. The previous state can more directly affect the current time step, and the feedback of the previous feature can also provide additional contextual information for the current RNN unit, helping to extract richer feature representations. That is, an attention mechanism can be provided, which can focus the current input on relevant features through the previous feedback. In addition, feature fusion can also increase the nonlinear expression ability of the RNN unit and enhance its ability to model video time series.
[0158] In some embodiments, the entire model is trained end-to-end, the RNN encoder is responsible for the decoder to encode the timing information of the input sequence, and the Swin Transformer decoder fully utilizes its powerful generation ability to output high-quality images. Compared with using RNN or Swin Transformer alone, this joint model can better integrate timing information and improve the quality of generated videos.
[0159] Among them, in one embodiment of the present disclosure, Swin Transformer can be composed of Multi-head Self-Attention and MLP similar to the standard Transformer. Shifted windows are used inside each block to divide the feature map into multiple windows, and these windows are connected in a shifted manner to enhance the sensitivity to local information. Swin Transformer Blocks are connected hierarchically, that is, the input of each Block comes from the output of all Blocks in the previous layer, which can enhance the fusion of multi-scale information. In each Stage, Swin Transformer can gradually reduce the resolution of the feature map and increase the number of channels to expand the receptive field. When processing data, Swin Transformer first patches and embeds the input image, and then performs multi-level processing through multiple Swin Transformer Blocks to gradually increase the receptive field and reduce the resolution. Among them, regularization strategies such as DropPath can also be used during the training process. Through this new structural design and staged multi-scale processing, the advantages of Transformer are retained, and the modeling capabilities of local and multi-scale information can also be enhanced.
[0160] In some or related embodiments, by controlling any recurrent neural network unit, a dual-branch explicit state alignment process is performed on a subset of sequence frames based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the subset of sequence frames; a hidden state convolution based on weighted convolution is performed on the explicit state alignment result to obtain the hidden state representation information output by any recurrent neural network unit. Two branches can be used to process optical flow information and feature information respectively, which can improve the accuracy of optical flow information determination while improving the richness of multi-channel data features, and can improve
[0161] The upper limit of the alignment effect of deformable convolution can improve the video denoising effect.
[0162] Figure 7 is a block diagram of a video denoising device according to an exemplary embodiment. Figure 7 , the device comprises:
[0163] A set acquisition unit 701 is used to acquire a sequence frame set corresponding to a first unprocessed RAW video;
[0164] A subset input unit 702, used to input any sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the sequence frame subset includes at least two adjacent video frames;
[0165] An information acquisition unit 703 is used to perform explicit and implicit dual alignment processing on a subset of sequence frames using any recurrent neural network unit to obtain hidden state representation information output by any recurrent neural network unit;
[0166] The video acquisition unit 704 is used to reconstruct the hidden state representation information and acquire a second unprocessed RAW video corresponding to the first unprocessed RAW video and having a video noise that meets the noise requirement.
[0167] According to some embodiments, the information acquisition unit 703 is used to perform explicit and implicit dual alignment processing on a subset of sequence frames using any recurrent neural network unit to obtain the hidden state representation information output by any recurrent neural network unit, specifically for:
[0168] Control any recurrent neural network unit to perform dual-branch explicit state alignment processing on the sequence frame subset based on the deformable convolutional network to obtain the explicit state alignment result corresponding to the sequence frame subset;
[0169] The hidden state convolution based on weighted convolution is performed on the explicit state alignment result to obtain the hidden state representation information output by any recurrent neural network unit.
[0170] According to some embodiments, the information acquisition unit 703 is used to control any recurrent neural network unit to perform dual-branch explicit state alignment processing on a subset of sequence frames based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the subset of sequence frames, specifically for:
[0171] Control any recurrent neural network unit to obtain single-channel data optical flow information corresponding to a subset of sequence frames;
[0172] Control any recurrent neural network unit to perform Bell conversion processing on the single-channel data corresponding to the sequence frame subset to obtain the four-channel RGGB data corresponding to the sequence frame subset;
[0173] Obtain feature information corresponding to four-channel RGGB data;
[0174] Obtain the offset based on the optical flow information and feature information;
[0175] The offset, optical flow information and feature information are input into the deformable convolutional network to obtain the display state alignment result corresponding to the subset of sequence frames.
[0176] According to some embodiments, the information acquisition unit 703 is used to perform hidden state convolution based on weighted convolution on the explicit state alignment result, and obtain the hidden state representation information output by any recurrent neural network unit, specifically for:
[0177] Obtain the offset of the pixel position of any channel according to the display state alignment result and the offset network;
[0178] Use a weight network to obtain the weight corresponding to the pixel position of any channel;
[0179] According to the offset and weight of the pixel position of any channel, the hidden state representation information of any recurrent neural network unit output is obtained.
[0180] According to some embodiments, the set acquisition unit 701 is used to acquire a sequence frame set corresponding to the first unprocessed RAW video, specifically to:
[0181] Using an optical flow algorithm to process the first unprocessed RAW video to obtain a processed first unprocessed RAW video;
[0182] Key frame extraction is performed on the processed first unprocessed RAW video to obtain a sequence frame set corresponding to the first unprocessed RAW video.
[0183] According to some embodiments, the subset input unit 702 is used to input a subset of sequence frames in any sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network, specifically for:
[0184] Obtaining sequence frame information corresponding to any sequence frame subset in the sequence frame set;
[0185] In a case where the sequence frame information indicates that the sequence frame subset is the first subset of the sequence frame set, inputting the sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network;
[0186] or
[0187] When the sequence frame information indicates that the sequence frame subset is not the first subset of the sequence frame set, obtaining a denoising result of the same scale corresponding to the sequence frame subset;
[0188] Feature fusion is performed based on the denoising result and the sequence frame subset to obtain a fused sequence frame subset, and the fused sequence frame subset is input into any recurrent neural network unit in the target multi-chain recurrent neural network.
[0189] According to some embodiments, the video acquisition unit 704 is used to reconstruct the hidden state representation information, and when acquiring the second unprocessed RAW video corresponding to the first unprocessed RAW video and the video noise meets the noise requirement, specifically:
[0190] Fusing the hidden state representation information with all the hidden state representation information output by the previous encoding layer to obtain the fused hidden state representation information;
[0191] The output result of the topmost chain in the target multi-chain recurrent neural network is obtained, and the output results are reorganized according to the input order of the sequence frame set to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirement.
[0192] According to some embodiments, the subset input unit 702 is used to input a subset of sequence frames in a sequence frame set into any recurrent neural network unit in a target multi-chain recurrent neural network, specifically for:
[0193] Down-sampling the sequence frame subset to obtain at least one sequence frame subset having a size different from the sequence frame subset;
[0194] At least one subset of sequence frames is input into the recurrent neural network unit corresponding to any subset of sequence frames in the target multi-chain recurrent neural network.
[0195] According to some embodiments, the video acquisition unit 704 is further configured to:
[0196] Acquire a training sample set, wherein the training sample set includes a first training sample subset collected by a first camera device and a second training sample subset collected by a second camera device, the first camera device and the second camera device correspond to the same shooting conditions and different noise levels;
[0197] The initial multi-chain recurrent neural network is trained using a training sample set to obtain a target multi-chain recurrent neural network.
[0198] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0199] In some or related embodiments, a set acquisition unit is used to acquire a sequence frame set corresponding to a first unprocessed RAW video; a subset input unit is used to input any sequence frame subset in the sequence frame set into any recurrent neural network unit in a target multi-chain recurrent neural network, wherein the sequence frame subset includes at least two adjacent video frames; an information acquisition unit is used to perform explicit and implicit dual alignment processing on the sequence frame subset using any recurrent neural network unit to obtain the hidden state representation information output by any recurrent neural network unit; a video acquisition unit is used to reconstruct the hidden state representation information to acquire a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirement. Therefore, explicit and implicit dual alignment processing can be performed to improve the adequacy of video timing information utilization, reduce the influence of video frame deformation on denoising results, improve the explicit and implicit dual alignment processing effect, use multi-frame joint denoising to reduce the situation where single-frame denoising results in poor denoising quality, improve the video denoising effect, and improve the quality of the acquired video.
[0200] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable electronic devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0201] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0202] A number of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other electronic devices through a computer network such as the Internet and / or various telecommunication networks.
[0203] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as video denoising or video denoising methods. For example, in some embodiments, video denoising or video denoising methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the video denoising or video denoising method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform video denoising or a video denoising method in any other suitable manner (eg, by means of firmware).
[0204] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0205] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0206] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0207] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0208] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0209] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0210] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0211] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A video denoising method, It is characterized in that include: Obtain a sequence frame set corresponding to the first unprocessed RAW video; Inputting any subset of sequence frames in the set of sequence frames into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the subset of sequence frames includes at least two adjacent video frames; Using any of the recurrent neural network units to perform explicit and implicit dual alignment processing on the sequence frame subset to obtain hidden state representation information output by any of the recurrent neural network units; The hidden state representation information is reconstructed to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements.
2. The method according to claim 1, It is characterized in that The step of using any of the recurrent neural network units to perform explicit and implicit dual alignment processing on the sequence frame subset to obtain the hidden state representation information output by any of the recurrent neural network units includes: Control any one of the recurrent neural network units to perform dual-branch explicit state alignment processing on the sequence frame subset based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the sequence frame subset; Perform hidden state convolution based on weighted convolution on the explicit state alignment result to obtain hidden state representation information output by any recurrent neural network unit.
3. The method according to claim 2, It is characterized in that The controlling any one of the recurrent neural network units to perform dual-branch explicit state alignment processing on the sequence frame subset based on a deformable convolutional network to obtain an explicit state alignment result corresponding to the sequence frame subset includes: Control any one of the recurrent neural network units to obtain the single-channel data optical flow information corresponding to the sequence frame subset; Control any one of the recurrent neural network units to perform Bell conversion processing on the single-channel data corresponding to the sequence frame subset to obtain the four-channel RGGB data corresponding to the sequence frame subset; Obtaining feature information corresponding to the four-channel RGGB data; Acquire an offset according to the optical flow information and the feature information; The offset, the optical flow information and the feature information are input into a deformable convolutional network to obtain a display state alignment result corresponding to the sequence frame subset.
4. The method according to claim 2, It is characterized in that The step of performing hidden state convolution based on weighted convolution on the visible state alignment result to obtain hidden state representation information output by any recurrent neural network unit includes: According to the display state alignment result, an offset network is used to obtain an offset of a pixel position of any channel; Using a weight network to obtain the weight corresponding to the pixel position of any channel; According to the offset of the pixel position of any channel and the weight, the hidden state representation information output by any recurrent neural network unit is obtained.
5. The method according to claim 1, It is characterized in that The step of obtaining a sequence frame set corresponding to the first unprocessed RAW video includes: Processing the first unprocessed RAW video using an optical flow algorithm to obtain a processed first unprocessed RAW video; Key frame extraction is performed on the processed first unprocessed RAW video to obtain a sequence frame set corresponding to the first unprocessed RAW video.
6. The method according to claim 1, It is characterized in that The step of inputting any subset of sequence frames in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network comprises: Obtaining sequence frame information corresponding to any sequence frame subset in the sequence frame set; In a case where the sequence frame information indicates that the sequence frame subset is the first subset of the sequence frame set, inputting the sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network; or When the sequence frame information indicates that the sequence frame subset is not the first subset of the sequence frame set, obtaining a denoising result of the same scale corresponding to the sequence frame subset; Feature fusion is performed based on the denoising result and the sequence frame subset to obtain a fused sequence frame subset, and the fused sequence frame subset is input into any recurrent neural network unit in the target multi-chain recurrent neural network.
7. The method according to claim 1, It is characterized in that The step of reconstructing the hidden state representation information to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets noise requirements includes: Fusing the hidden state representation information with all hidden state representation information output by the previous coding layer to obtain fused hidden state representation information; The output result of the topmost chain in the target multi-chain recurrent neural network is reorganized according to the input order of the sequence frame set to obtain a second unprocessed RAW video corresponding to the first unprocessed RAW video and whose video noise meets the noise requirement.
8. The method according to claim 1, It is characterized in that The step of inputting a subset of the sequence frames in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network comprises: Down-sampling the sequence frame subset to obtain at least one sequence frame subset different in size from the sequence frame subset; The at least one sequence frame subset is input into the recurrent neural network unit corresponding to any sequence frame subset in the target multi-chain recurrent neural network.
9. The method according to claim 1, It is characterized in that The method further comprises: Acquire a training sample set, wherein the training sample set includes a first training sample subset collected by a first camera device and a second training sample subset collected by a second camera device, the first camera device and the second camera device correspond to the same shooting conditions and different noise levels; The initial multi-chain recurrent neural network is trained using the training sample set to obtain the target multi-chain recurrent neural network.
10. A video denoising device, It is characterized in that include: A set acquisition unit, used for acquiring a sequence frame set corresponding to the first unprocessed RAW video; A subset input unit, used for inputting any sequence frame subset in the sequence frame set into any recurrent neural network unit in the target multi-chain recurrent neural network, wherein the sequence frame subset includes at least two adjacent video frames; An information acquisition unit, used to perform explicit and implicit dual alignment processing on the sequence frame subset using any of the recurrent neural network units to obtain hidden state representation information output by any of the recurrent neural network units; The video acquisition unit is used to reconstruct the hidden state representation information and acquire a second unprocessed RAW video corresponding to the first unprocessed RAW video and having video noise that meets the noise requirement.
11. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method as claimed in any one of claims 1 to 9. 12 . A storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to claim 1 .