Video denoising model training method and device, equipment and storage medium

Through iterative training and motion information adjustment of the video denoising model, the problem of random noise interference in the monitoring video is solved, the stability and clarity of the video are improved, and the information continuity is ensured.

CN120450995APending Publication Date: 2025-08-08WEBANK (CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566598.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art has random noise interference in surveillance videos, resulting in a decrease in video clarity and stability, supervision methods lead to a drag, and self-supervision methods lead to a loss of key details.

Method used

The video denoising model is used for iterative training, and preliminary denoising is performed by fusion of noise-added features and image features, and the motion information of multiple other denoising images is used to adjust the pixel points to optimize the model parameters.

Benefits of technology

It improves the stability and clarity of the denoised video, reduces the possibility of key information loss, and enhances the information continuity between video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450995A_ABST
    Figure CN120450995A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video denoising model training method and device, equipment and a storage medium, and can be applied to the technical field of artificial intelligence. In the method, a denoising model is adopted to perform preliminary denoising on a sample video frame based on a noise adding feature and an image feature; the denoising model can more fully learn the difference between the noise adding feature and the image feature, and the accuracy of preliminary denoising of the sample video frame is improved, so that the stability of the denoised video is improved; and adjusting a plurality of pixel points in the preliminary de-noised image based on the motion information between the plurality of other de-noised images and the preliminary de-noised image, so that the de-noising model learns the motion trail of an object between the other de-noised images and the preliminary de-noised image, and therefore, when the video contains the moving object, the de-noising model can learn the motion trail of the object between the other de-noised images and the preliminary de-noised image. The features of the same object in a plurality of video frames can still be accurately associated, the continuity of information interaction between different video frames is improved, and the possibility of key information loss is reduced, so that the definition and the stability of the de-noised video are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, and storage medium for a video denoising model. Background Art

[0002] In the field of video surveillance, due to reasons such as dim light at night or infrared imaging, surveillance videos are prone to severe random noise, which interferes with target recognition and behavior analysis in surveillance videos, thereby affecting the recognition results of surveillance videos.

[0003] In related technologies, supervised or self-supervised methods are usually used to denoise surveillance videos. Supervised methods often denoise surveillance videos based on temporal filtering, which uses the inter-frame redundancy of video frames to average out random noise. However, when faced with frames with large relative motion, this method can cause the surveillance video to be discontinuous in the temporal dimension, which can easily produce ghosting and reduce the clarity of the surveillance video. Self-supervised methods generally use blind spot networks to denoise surveillance videos, filling in the blind spots of video frames through contextual reasoning. However, the filled information over-smoothes the video frames, resulting in the loss of key details of the video frames, causing the surveillance video to flicker and reducing its stability.

[0004] Therefore, there is an urgent need for a video denoising method that can improve the stability and clarity of surveillance videos. Summary of the Invention

[0005] Embodiments of the present invention provide a method, apparatus, device, and storage medium for training a video denoising model, which are used to improve the stability and consistency of denoised videos.

[0006] In one aspect, an embodiment of the present application provides a method for training a video denoising model, the method comprising:

[0007] The denoising model to be trained is iteratively trained using multiple sample video frames until a stopping condition is met, thereby obtaining a trained target denoising model. Each round of iteration includes the following steps:

[0008] Using the denoising model used in this round, feature encoding is performed on the sample video frame, and the corresponding image features are fused with the preset noise to obtain noise-added features; based on the noise-added features and the image features, preliminary denoising is performed on the sample video frame to obtain a corresponding preliminary denoised image;

[0009] Adjusting multiple pixels in the preliminary denoised image based on motion information between multiple other denoised images and the preliminary denoised image to obtain a target video frame; the multiple other denoised images are preliminary denoised images of other video frames adjacent to the sample video frame;

[0010] A model loss value is obtained based on the target video frame and the sample video frame, and parameters of the denoising model used in this round are adjusted according to the model loss value.

[0011] Optionally, performing preliminary denoising on the sample video frame based on the noise feature and the image feature to obtain a corresponding preliminary denoised image includes:

[0012] generating denoising prompt information of the sample video frame based on the noise addition feature and the image feature;

[0013] Based on the denoising prompt information, preliminary denoising is performed on the sample video frame to obtain a corresponding preliminary denoised image.

[0014] Optionally, performing preliminary denoising on the sample video frame based on the denoising prompt information to obtain a corresponding preliminary denoised image includes:

[0015] generating an attention feature of the sample video frame based on the image feature of the sample video frame;

[0016] Fusing the attention feature and the denoising prompt information to obtain a fusion weight;

[0017] Based on the fusion weights, the sample video frames are adjusted to obtain corresponding preliminary denoised images.

[0018] Optionally, adjusting a plurality of pixels in the preliminary denoised image based on motion information between a plurality of other denoised images and the preliminary denoised image to obtain a target video frame includes:

[0019] For each pixel in the preliminary denoised image, the following operations are performed: obtaining motion information associated with the pixel, the motion information including denoised pixels corresponding to the same object as the pixel in each of the other denoised images; updating the preliminary pixel value of the pixel to a second pixel value based on the obtained first pixel value of at least one denoised pixel;

[0020] A target video frame is obtained based on the second pixel value of each pixel in the preliminary denoised image.

[0021] Optionally, updating the preliminary pixel value of the pixel point to a second pixel value based on the obtained first pixel value of the at least one denoised pixel point includes:

[0022] Obtaining a weight of the preliminary pixel value and a weight of each of the at least one first pixel value based on an association relationship between the preliminary pixel value and the at least one first pixel value obtained;

[0023] Based on the weight of the preliminary pixel value and the respective weights of the at least one first pixel value, a weighted sum is performed on the preliminary pixel value and the at least one first pixel value to obtain the second pixel value.

[0024] Optionally, obtaining a model loss value based on the target video frame and the sample video frame includes:

[0025] For each pixel point in the preliminary denoised image, respectively performing: determining a first position offset between the pixel point and at least one denoised pixel point;

[0026] For each pixel in the sample video frame, respectively performing the following steps: obtaining third position information of the pixel in the sample video frame, and fourth position information of an original pixel corresponding to the same object as the pixel in each of the other video frames; and determining a second position offset between the third position information of the pixel and at least one associated fourth position information.

[0027] A model loss value is obtained based on the obtained at least one first position offset and the obtained at least one second position offset.

[0028] Optionally, obtaining a model loss value based on the obtained at least one first position offset and at least one second position offset includes:

[0029] Obtaining a reconstruction loss value based on a degree of difference between the target video frame and the preliminary denoised image;

[0030] obtaining an offset loss value based on the obtained at least one first position offset and the obtained at least one second position offset;

[0031] A model loss value is obtained based on the reconstruction loss value and the offset loss value.

[0032] In one aspect, an embodiment of the present application provides a video denoising method, the method comprising:

[0033] For multiple original video frames in the video to be processed, respectively: inputting one original video frame into a target denoising model for denoising to obtain a corresponding denoised video frame; wherein the target denoising model is obtained by training using a video denoising model training method;

[0034] Based on the obtained multiple denoised video frames, a target denoised video is obtained.

[0035] In one aspect, an embodiment of the present application provides a device for training a video denoising model, the device comprising:

[0036] A denoising module is configured to perform feature encoding on the sample video frame using the denoising model used in this round, and fuse the corresponding image features with the preset noise to obtain a noise-added feature; based on the noise-added feature and the image features, perform preliminary denoising on the sample video frame to obtain a corresponding preliminary denoised image;

[0037] an adjustment module, configured to adjust a plurality of pixels in the preliminary denoised image based on motion information between a plurality of other denoised images and the preliminary denoised image, to obtain a target video frame; the plurality of other denoised images being preliminary denoised images of other video frames adjacent to the sample video frame;

[0038] A correction module is used to obtain a model loss value based on the target video frame and the sample video frame, and adjust the parameters of the denoising model used in this round according to the model loss value.

[0039] In one aspect, an embodiment of the present application provides a video denoising apparatus, the apparatus comprising:

[0040] An input module is configured to execute, for each of the plurality of original video frames in the video to be processed, the following steps: inputting an original video frame into a target denoising model for denoising, thereby obtaining a corresponding denoised video frame; wherein the target denoising model is obtained by training using the training device for the video denoising model;

[0041] The acquisition module is used to obtain a target denoised video based on the obtained multiple denoised video frames.

[0042] In one aspect, an embodiment of the present application provides a computer device, comprising:

[0043] a memory for storing program instructions;

[0044] The processor is used to call the program instructions stored in the memory and execute the steps of the above-mentioned video denoising model training method according to the obtained program.

[0045] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the program is run on the computer device, the computer executes the steps of the above-mentioned video denoising model training method.

[0046] On the one hand, an embodiment of the present application provides a computer program product, including a computer program stored on a computer-readable storage medium, wherein the computer program includes program instructions, and when the program instructions are executed by a computer device, the computer device executes the steps of the above-mentioned video denoising model training method.

[0047] In an embodiment of the present application, a denoising model is used to perform preliminary denoising on sample video frames based on noise features and image features, so that the denoising model can more fully learn the difference between noise features and image features, improve the accuracy of preliminary denoising of sample video frames, and thus improve the stability of the denoised video; based on the motion information between multiple other denoised images and the preliminary denoised image, multiple pixels in the preliminary denoised image are adjusted, so that the denoising model learns the motion trajectory of an object between other denoised images and the preliminary denoised image. Therefore, when the video contains a moving object, it can still accurately associate the features of the same object in multiple video frames, thereby improving the continuity of information interaction between different video frames, reducing the possibility of loss of key information, and thus improving the clarity and stability of the denoised video. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0049] Figure 1 A schematic diagram of the structure of a system architecture provided in an embodiment of the present application;

[0050] Figure 2 A flowchart of a method for training a video denoising model provided in an embodiment of the present application;

[0051] Figure 3 A schematic diagram of the structure of a sample video frame preprocessing provided in an embodiment of the present application;

[0052] Figure 4A A schematic diagram of a structure for preliminary denoising of a sample video frame provided in an embodiment of the present application;

[0053] Figure 4B A schematic diagram of a structure for preliminary denoising of a sample video frame provided in an embodiment of the present application;

[0054] Figure 5A A flowchart of a method for training a video denoising model provided in an embodiment of the present application;

[0055] Figure 5B A schematic diagram of the structure of pixel position information provided in an embodiment of the present application;

[0056] Figure 5C A schematic diagram of the structure of pixel position offset information provided in an embodiment of the present application;

[0057] Figure 6A schematic diagram of the structure of a video denoising model training method provided in an embodiment of the present application;

[0058] Figure 7 A schematic diagram of the structure of a video denoising model provided in an embodiment of the present application;

[0059] Figure 8 A flowchart of a video denoising method provided in an embodiment of the present application;

[0060] Figure 9 A schematic diagram of the structure of a video denoising model training device provided in an embodiment of the present application;

[0061] Figure 10 A schematic diagram of the structure of a video denoising device provided in an embodiment of the present application;

[0062] Figure 11 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0064] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0065] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0066] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0067] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

[0069] Some of the terms used in the examples of this application are explained below to facilitate understanding by those skilled in the art.

[0070] Explicit optical flow calculation: By explicitly calculating the motion vectors (i.e., optical flow) of pixels or feature points between consecutive frames, these vectors are used to guide image alignment. Its core is to directly model geometric transformations such as translation, rotation, and deformation.

[0071] Implicit alignment: This involves automatically learning alignment relationships between frames or data through models or algorithms, without explicitly computing geometric transformations (such as optical flow and affine transformations) or feature matching. The core idea is to leverage the representational learning capabilities of deep learning and other technologies to directly model associations at the pixel, feature, or semantic level, thereby implicitly completing the alignment task.

[0072] The following is a brief introduction to the system architecture diagram applicable to the technical solution of the embodiment of the present application. It should be noted that the process introduced below is only used to illustrate the embodiment of the present application and is not a limitation.

[0073] refer to Figure 1 , which is a system architecture diagram applicable to an embodiment of the present application. The system architecture includes at least a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. This application does not specifically limit the number of terminal devices 101 and servers 102.

[0074] The terminal device 101 is pre-installed with an application that has a video denoising function, which can be a client application, a web application, a mini-program application, etc. The terminal device 101 can be a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart home appliance, an intelligent voice interaction device, an intelligent vehicle-mounted device, etc., but is not limited thereto.

[0075] Server 102 is the background server of the application. Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, but is not limited to these.

[0076] It should be noted that the method in the embodiment of the present application can be executed by the terminal device 101 or the server 102 alone, or can be executed by the terminal device 101 and the server 102 together.

[0077] In the embodiment of the present application, the terminal device 101 and the server 102 can be directly or indirectly connected to each other through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless Fidelity (WIFI) network. Of course, other possible networks are also possible, and the embodiment of the present application does not limit this.

[0078] The following is based on Figure 1 The system architecture diagram shown in FIG. 1 shows a process of a method for training a video denoising model. The process of the method can be Figure 1 The terminal device 101 shown in FIG. 1 may also be executed by the server 102, or the terminal device 101 and the server 102 may interact to execute the execution. Figure 2 As shown, the following steps are included:

[0079] The denoising model to be trained is iteratively trained using multiple sample video frames until a stopping condition is met, thereby obtaining a trained target denoising model. Each round of iterative process includes steps 201 to 203:

[0080] In step 201, the sample video frame is feature-encoded by the denoising model used in this round, and the corresponding image features are fused with the preset noise to obtain the noise-added features; based on the noise-added features and the image features, the sample video frame is preliminarily denoised to obtain the corresponding preliminary denoised image.

[0081] Specifically, the denoising model includes: a prompt information generation module, a preliminary denoising module and a temporal consistency module.

[0082] After inputting a sample video frame into the denoising model to be trained, the sample video frame is downsampled before feature encoding to obtain multiple corresponding sampling subgraphs. This involves dividing the sample video frame into multiple blocks, and each block is further divided into multiple windows. Each downsampling of the sample video frame involves selecting a window from the multiple windows. After multiple downsampling steps, the multiple identical windows are aggregated into a sampling subgraph, resulting in multiple sampling subgraphs associated with the sample video frame.

[0083] For example, refer to Figure 3 , a sample video frame is divided into 4 blocks, each block is further divided into 4 windows, and a window set 301 is obtained. The window set 301 includes 16 windows.

[0084] One downsampling is to select a window from the 16 windows of the sample video frame, downsample the sample video frame 16 times, and merge the obtained identical windows into the same sampling sub-graph 302. After 16 downsamplings, 4 sampling sub-graphs can be obtained, namely sampling sub-graph 302 to sampling sub-graph 305.

[0085] In the embodiment of the present application, by downsampling the sample video frames, the size of the sample video frames is reduced, the computational complexity of the subsequent denoising model on the sample video frames is reduced, and the denoising efficiency of the sample video frames is improved.

[0086] In addition, the prompt information generation module performs feature encoding on multiple sampled sub-images of the sample video frame to obtain image features. Forward diffusion is performed on the image features to obtain a preset noise. The forward diffusion process gradually adds noise to the image features until the image features become the preset noise, where the preset noise is pure noise. The preset noise is then spliced onto the image features to obtain a noisy feature.

[0087] by Figure 3 Taking sampling sub-image 302 as an example, the prompt information generation module performs a convolution calculation on the pixel values of each local image region (each local image region includes 9 pixels) in sampling sub-image 302 using a convolution kernel of a fixed size (e.g., 3×3) to obtain the corresponding local image features. In parallel, maximum pooling is performed on the pixel values of each pixel in the sampling sub-image to capture the global image features of sampling sub-image 302. The global features of sampling sub-image 302 are then fused with multiple local features to obtain the sub-image features of sampling sub-image 302.

[0088] Similarly, the prompt information generation module performs the same operation on the sampling subgraphs 303 to 305 to obtain the subgraph features of the sampling subgraphs 303 to 305 .

[0089] The sample video frames are subjected to maximum pooling in parallel to capture the global image features of the sample video frames, and the global image features of the sample video frames and the sub-image features of the sampling sub-images 302 to 305 are fused to obtain the image features of a sample video frame.

[0090] In the prompt information generation module, random noise is added to the image features of the sample video frames over K time steps. This noise covers the image information, such as color, brightness, texture, and edge information, until it is converted to the preset noise. The preset noise is then added to the end of the image features to generate the noisy features.

[0091] In an embodiment of the present application, a convolution kernel of a fixed size is used to obtain the local image features of a sampling sub-image, and the local image features are fused with the global image features of the sampling sub-image to obtain the sub-image features of the sampling sub-image; the sub-image features of multiple sampling sub-images and the global image features of the sample video frame are fused to obtain the image features of the sample video frame, thereby fully extracting the key information in the sample video frame, alleviating the problem of easy loss of key information of the sample video frame in the prior art, and improving the accuracy and stability of the denoised video.

[0092] In some embodiments, denoising prompt information of a sample video frame is generated based on the noise addition feature and the image feature; and based on the denoising prompt information, the sample video frame is preliminarily denoised to obtain a corresponding preliminarily denoised image.

[0093] Specifically, the noise-added features are obtained by splicing the image features and the preset noise; the prompt information generation module learns the difference between the image features and the part with the preset noise added in the noise-added features; based on the learned difference, the image features of the sample video frame are generatively optimized to repair the missing parts and the noise parts in the image features to obtain cleaner and more realistic denoising prompt information, and the denoising prompt information is subsequently used as a denoising condition to guide the denoising process.

[0094] The prompt information generation module includes: a diffusion model, specifically, a standard diffusion model (Denoising Diffusion Probabilistic Models, DDPM), an accelerated sampling diffusion model (Pseudo Linear Multistep Methods, PLMS), a latent space diffusion model, etc. This application does not make any specific limitation on this.

[0095] The calculation formula for obtaining denoising prompt information based on the noise addition feature and image feature is shown in the following formula (1):

[0096] P gen =G θ (P init |I noisy)……………………(1)

[0097] Among them, P gen Indicates denoising prompt information, G θ represents the parameters of the diffusion model; P init Represents image features; I noisy Represents the noise feature.

[0098] In practical applications, the denoising prompt information P gen It can be a clean and noise-free reference image, or it can be a reference image feature or other forms of information that characterizes a clean and noise-free reference image, which is not specifically limited in this application; the image feature P init It is obtained by extracting the image information such as color, brightness, texture, edge information, etc. of the sample video frame; the noise feature I noisy It is the image feature P init The preset noise is obtained by splicing the preset noise based on the above, wherein the preset noise can be in the form of a pure noise image or other forms.

[0099] Optionally, the prompt information generation module may also use a variational autoencoder (VAE), such as a denoising autoencoder (DAE), to generate denoised prompt information based on the noise features and the image features.

[0100] In an embodiment of the present application, the noise-added features obtained based on the image features and preset noise of the sample video frame are a priori cognition of the sample video frame, and their function is to guide the subsequent optimization process of the image features based on the noise-added features using a diffusion model or a variational autoencoder, so that the obtained denoising prompt information can retain key information such as the main contours, object edges and texture directions of the sample video frame, thereby improving the clarity of the denoising of the sample video frame.

[0101] In some embodiments, based on the image features of the sample video frame, an attention feature of the sample video frame is generated; the attention feature and the denoising prompt information are fused to obtain a fusion weight; based on the fusion weight, the sample video frame is adjusted to obtain a corresponding preliminary denoised image.

[0102] Specifically, in the preliminary denoising module of the denoising model, reference Figure 4AImage features are obtained by extracting image information such as color, brightness, texture, and edge information from sample video frames. The size of the image features is h×w×C. Global average pooling is performed on the obtained image features to obtain multiple attention features corresponding to the sample video frames, where h, w, and C are all positive integers greater than or equal to 1; the size of an attention feature is 1×1×C. Attention features represent the image information that requires special attention in the sample video frames. For example, the image information of objects or people that require special attention in the sample video frames of surveillance videos, where h represents the number of pixels in the vertical direction of a sample video frame, w represents the number of pixels in the horizontal direction of a sample video frame, and C represents the number of channels per pixel. For example, a pixel in a grayscale image has only one channel, i.e., a grayscale image is a single-channel image; a pixel in an RGB image has three channels, including red, blue, and green.

[0103] The attention features are concatenated and fused with the denoising hints, and the fusion result is fed into a linear layer to obtain the corresponding fusion weight for the sample video frame. The fusion weight for each attention feature is 1×1×C. The fusion weight assigns greater weight to image information that requires special attention. Based on the fusion weights, the corresponding network weight layer is generated. During this process, the convolution kernel size, the number of input channels, and the stride of the convolution kernel in the network weight layer are set according to the size of the sample video frame and the size of the sample subgraph of the sample video frame. The number of output channels is also set as needed, resulting in a convolution kernel weight shape of (output channels, input channels, kernel height, kernel width). The sample video frame is fed into the network weight layer with multiple convolution kernels, and the initial denoised image corresponding to the sample video frame is output. The number of input channels is the number of channels per pixel in the sample video frame.

[0104] For example, refer to Figure 4B Taking the sample video frame as an RGB image (i.e., C = 3) as an example, image feature 401 has a size of 100×100×3. Average pooling is performed on image feature 401 to obtain 100 attention features 403, each of which has a size of 1×1×3. For each attention feature, attention feature 403 is fused with denoising hint information 402 of the same size to obtain a fusion weight 404, each of which has a size of 1×1×3. Based on fusion weight 404, the convolution kernel size of the network weight layer is set to 3×3, and the weight shape of the convolution kernel of the network weight layer is set to (16, 3, 3, 3). At the same time, at least one convolution kernel with a different number of output channels is set for the network weight layer, such as a convolution kernel with a weight shape of (32, 3, 3, 3) or a convolution kernel with a weight shape of (128, 3, 3, 3), so that the network weight layer can learn diverse features.

[0105] The image feature 402 is input into the set network weight layer, and each convolution kernel in the network weight layer slides in the feature of the image feature 402, and local information such as image edge information and image texture information is captured through the weight. Through the combination of different convolution kernels in the network weight layer, a variety of feature maps are generated based on the image feature 401. Based on the diverse feature maps, a preliminary denoised image 405 corresponding to the sample video frame is obtained by fusion.

[0106] In the embodiment of the present application, the image features of the sample video frame are preliminarily denoised in combination with the denoising prompt information, guiding the denoising process to restore clearer texture and edge information to the sample video frame, reducing the possibility of losing key information of the sample video frame, and improving the stability of video denoising.

[0107] Step 202 : Based on motion information between multiple other denoised images and the preliminary denoised image, multiple pixels in the preliminary denoised image are adjusted to obtain a target video frame; the multiple other denoised images are preliminary denoised images of other video frames adjacent to the sample video frame.

[0108] Specifically, the temporal consistency module of the denoising model determines the pixel values of a pixel in the preliminary denoised image and the denoised pixel values of the denoised pixel based on the motion information between the pixel and the denoised pixels of multiple other denoised images that correspond to the same object. The pixels in the preliminary denoised image are then adjusted to obtain the target video frame.

[0109] The motion information between the multiple other denoised images and the preliminary denoised image refers to the dynamic change characteristics of an object in the continuous preliminary denoised images in the time dimension, wherein an object is composed of multiple pixels.

[0110] In an embodiment of the present application, for each pixel in the preliminary denoised image, the pixel value of the denoised pixel for the same object in multiple other denoised images is used to adjust the pixel, thereby improving the coherence and stability of the denoised video.

[0111] In some embodiments, for each pixel point in the preliminary denoised image, the following operations are performed: motion information associated with a pixel point is obtained, and the motion information includes: denoised pixel points corresponding to the same object as the pixel point in each other denoised image; based on the obtained first pixel value of at least one denoised pixel point, the preliminary pixel value of the pixel point is updated to a second pixel value; based on the second pixel value of each pixel point in the preliminary denoised image, a target video frame is obtained.

[0112] Specifically, the other denoised images refer to preliminary denoised images of other video frames adjacent to the sample video frame.

[0113] In some embodiments, the present invention uses at least the following method to obtain the second pixel value:

[0114] Based on the correlation between the preliminary pixel value and at least one first pixel value, a weight of the preliminary pixel value and a weight of each of the at least one first pixel value are obtained; then, based on the weight of the preliminary pixel value and a weight of each of the at least one first pixel value, a weighted sum of the preliminary pixel value and the at least one first pixel value is performed to obtain a second pixel value.

[0115] In the specific implementation, taking the number of other denoised images as N, the following steps are included: Figure 5A As shown:

[0116] Step 501: Acquire a preliminary denoised image and N other denoised images adjacent to the preliminary denoised image, where N is greater than or equal to 1.

[0117] Step 502 : For each pixel point of the preliminary denoised image, obtain first position information ρ0=(x0, y0) of the pixel point in the preliminary denoised image.

[0118] The first position information is: the position coordinates of the pixel points in the preliminary denoised image. For example, see Figure 5B , setting the upper left corner of the preliminary denoised image as the coordinate origin, the position coordinate of the pixel point 5011 in the preliminary denoised image is ρ0 = (500, 500).

[0119] Step 503: Use the video denoising model to be trained to predict the denoised pixel corresponding to the same object as the pixel in N other denoised images, as well as the first position offset Δρ of the denoised pixel relative to the pixel. i =(Δx i , Δy i ), obtain the second position information ρ corresponding to the denoised pixel based on the first position offset i =(x0+Δx i ,y0+Δy i ).

[0120] Among them, ρ i Indicates second position information of a denoised pixel corresponding to the same object as a pixel in the i-th other denoised image among the N other denoised images.

[0121] The first position offset is the difference between the position coordinates of a pixel point in the preliminary denoised image and the position coordinates of denoised pixels corresponding to the same object in N other denoised images.

[0122] Take N=8 as an example, see Figure 5BFor the pixel 5011 in the preliminary denoised image, the other eight denoised pixels corresponding to the same object are: denoised pixel 5012, denoised pixel 5013, denoised pixel 5014, denoised pixel 5015, denoised pixel 5016, denoised pixel 5017, denoised pixel 5018, and denoised pixel 5019.

[0123] See also Figure 5C Relative to pixel 5011, the first position offset 5022 of the denoised pixel 5012 is Δρ1=(-6, -5), the first position offset 5023 of the denoised pixel 5013 is Δρ2=(-4, -4), the first position offset 5024 of the denoised pixel 5014 is Δρ3=(-3, -4), the first position offset 5025 of the denoised pixel 5015 is Δρ4=(-1, -1), the first position offset 5026 of the denoised pixel 5016 is Δρ5=(2, 1), the first position offset 5027 of the denoised pixel 5017 is Δρ6=(3, 4), the first position offset 5028 of the denoised pixel 5018 is Δρ7=(5, 4), and the first position offset 5029 of the denoised pixel 5019 is Δρ8=(5, 5).

[0124] Based on the position coordinates of pixel point 5011 and the first position offset of the denoised pixel point relative to pixel point 5011, the position coordinates of each denoised pixel point (i.e., the second position information of each denoised pixel point) are determined, specifically including: the position coordinates of the denoised pixel point 5012 are ρ1 = (494, 495), the position coordinates of the denoised pixel point 5013 are ρ2 = (496, 496), the position coordinates of the denoised pixel point 5014 are ρ3 = (497, 496), the position coordinates of the denoised pixel point 5015 are ρ4 = (499, 499), the position coordinates of the denoised pixel point 5016 are ρ5 = (502, 501), the position coordinates of the denoised pixel point 5017 are ρ6 = (503, 504), the position coordinates of the denoised pixel point 5018 are ρ7 = (505, 504), and the position coordinates of the denoised pixel point 5019 are ρ8 = (505, 505). According to the position coordinates of each denoised pixel point, the first pixel value of the denoised pixel point is obtained from other denoised images where the pixel point is located.

[0125] Step 504 : The preliminary pixel value of a pixel and the first pixel values of the other N denoised pixels are combined into a vector X, and the weights of the respective pixel values in the vector X are calculated.

[0126] Specifically, the self-attention formula is used to calculate the weights of the initial pixel value and the N first pixel values in the vector X. The specific formula is shown in the following formula (2):

[0127]

[0128] Where, set Q = XW Q , K=XW K , V=XW V ;W Q 、W K 、W V are the weights learned by the self-attention network for the query vector Q, key vector K, and value vector V, respectively. x Represents the weight value of vector X (i.e. the weight of each pixel value); d k Represents the feature dimension.

[0129] For example, vector X is composed of a pixel in the initial denoised image and N other denoised pixels. Take N = 8 as an example, see Figure 5B , the pixel value of pixel 5011 in the preliminary denoised image is x1, and according to the second position information of each of the eight denoised pixels, the pixel values at the second position information are obtained, and the following are respectively obtained: the pixel value of denoised pixel 5012 is x2, the pixel value of denoised pixel 5013 is x3, the pixel value of denoised pixel 5014 is x4, the pixel value of denoised pixel 5015 is x5, the pixel value of denoised pixel 5016 is x6, the pixel value of denoised pixel 5017 is x7, the pixel value of denoised pixel 5018 is x8, and the pixel value of denoised pixel 5019 is x9, then the vector X = [x1, x2, x3, x4, x5, x6, x7, x8, x9], and the corresponding weight values are

[0130] Step 505 : Based on the weights of the pixel values in the vector X, perform weighted summation on the pixel values to obtain a second pixel value.

[0131] Specifically, the second pixel value is the adjusted pixel value of a pixel point in the preliminary denoised image.

[0132] For example, refer to Figure 6 When N=8, for the preliminary denoised image 602 , and the first four other denoised images 601 and the last four other denoised images 603 adjacent to the preliminary denoised image 602 .

[0133] For each pixel in the preliminary denoised image 602, the temporal consistency module in the denoising model is used to predict the first position information of the central pixel in the preliminary denoised image 602, i.e., the position coordinate (500, 500). The temporal consistency module is also used to predict the denoised pixel corresponding to the same object as the pixel in each of the first four other denoised images 601, as well as the first position offset of the denoised pixel relative to the first position information. For example, the first position offsets of the denoised pixels in the first four other denoised images 601 are (5, 5), (5, 4), (3, 4), and (2, 1), respectively.

[0134] Similarly, the temporal consistency module is used to predict the denoised pixels corresponding to the same object as the pixel in the next four other denoised images 603, as well as the first position offset of the denoised pixel relative to the first position information. For example, the first position offsets of the corresponding denoised pixels in the next four other denoised images 603 are (-1, -1), (-3, -4), (-4, -4), and (-6, -5), respectively.

[0135] For each denoised pixel, second position information of the denoised pixel is obtained based on the obtained first position information and the first position offset associated with the denoised pixel. Based on the one piece of first position information and the eight pieces of second position information, a position set 604 is obtained. Position set 604 includes: (505, 505), (505, 504), (503, 504), (502, 501), (500, 500), (499, 499), (497, 496), (496, 496), and (494, 495).

[0136] Based on the first position information (500,500), a preliminary pixel value of a pixel at the corresponding position is obtained. Based on the second position information of each of the eight denoised pixels (i.e., (505,505), (505,504), (503,504), (502,501), (499,499), (497,496), (496,496), and (494,495)), the first pixel values of the eight denoised pixels at the corresponding positions in their respective preliminary denoised images are obtained. According to the preliminary pixel value and the eight first pixel values, the weight values 605 corresponding to each of the one pixel and the eight denoised pixels are calculated using the self-attention formula of the temporal consistency module.

[0137] Based on the weight value 605, the preliminary pixel value of a pixel and the first pixel values of each of the eight denoised pixels are weightedly summed to obtain the second pixel value of the pixel. Based on the second pixel values of multiple pixels in the preliminary denoised image 602, a target video frame 606 is obtained.

[0138] In this embodiment of the present application, multiple denoised images are implicitly aligned with the preliminary denoised image based on the motion information between pixels in the multiple other images and the preliminary denoised image. This implicit alignment, through self-attention calculation of each pixel value in the preliminary denoised image, enables pixel-level alignment of moving objects without the need for explicit optical flow calculations, improving the accuracy of information exchange between different sample video frames.

[0139] Step 203: Obtain a model loss value based on the target video frame and the sample video frame, and adjust the parameters of the denoising model used in this round according to the model loss value.

[0140] In an embodiment of the present application, a denoising model is used to perform preliminary denoising on sample video frames based on noise features and image features, so that the denoising model can more fully learn the difference between noise features and image features, improve the accuracy of preliminary denoising of sample video frames, and thus improve the stability of the denoised video; based on the motion information between multiple other denoised images and the preliminary denoised image, multiple pixel points in the preliminary denoised image are adjusted, so that the denoising model learns the motion trajectory of an object between other denoised images and the preliminary denoised image. Therefore, when the video contains a moving object, it can still accurately associate the features of the same object in multiple video frames, thereby improving the continuity of information interaction between different video frames, reducing the possibility of loss of key information, and thus improving the clarity and stability of the denoised video.

[0141] In some embodiments, for each pixel in the preliminary denoised image, the following are performed: determining a first position offset between a pixel and at least one denoised pixel; for each pixel in the sample video frame, the following are performed: obtaining third position information of a pixel in the sample video frame, and fourth position information of the original pixel of the same object corresponding to a pixel in each other video frame in the corresponding other video frame; determining a second position offset between the third position information of a pixel and at least one associated fourth position information; and obtaining a model loss value based on the obtained at least one first position offset and at least one second position offset.

[0142] Specifically, the temporal consistency module predicts the first position offset of a pixel corresponding to the same object in each other denoised image relative to the pixel based on the first position information of the pixel in the preliminary denoised image.

[0143] For each pixel point in the sample video frame, the third position information of the pixel point in the sample video frame is obtained by displaying optical flow calculation or feature matching method, as well as the fourth position information of the original pixel point of the same object corresponding to the pixel point in each other video frame in the corresponding other video frame. Based on the third position information of the pixel point and the fourth position information of the original pixel point, the second position offset of the original pixel point relative to a pixel point in the sample video frame is determined.

[0144] For example, a sample video frame has Y pixels, where Y is greater than 1. Accordingly, the initial denoised image also has Y pixels. For the pixel at the center of the initial denoised image, the first position information for this pixel is (500, 500), and its first position offset is recorded as (0, 0). The temporal consistency module predicts the first position offsets of the denoised pixels corresponding to the same object in the other eight denoised images, resulting in (5, 5), (5, 4), (3, 4), (2, 1), (-1, -1), (-3, -4), (-4, -4), and (-6, -5). Thus, a pixel has nine associated first position offsets. For a preliminary denoised image with Y pixels, 9*Y first position offsets are ultimately obtained.

[0145] For the pixel point at the center position in the sample video frame, the first position information is (500,500), and its own second position offset is recorded as (0,0). The second position offsets of the original pixel points of the same object corresponding to the pixel point at the center position in 8 other video frames are calculated by displaying the optical flow, and the results are (5,4), (4,4), (3,4), (1,1), (-1,-2), (-3,-3), (-5,-4), and (-5,-5). In this way, one pixel point has 9 associated second position offsets. For a sample video frame with Y pixels, 9*Y second position offsets are finally obtained.

[0146] The 9*Y first position offsets are compared with the corresponding 9*Y second position offsets to obtain the offset loss value between the preliminary denoised image and the sample video frame, and the offset loss value is used as the model loss value of the denoising model.

[0147] In an embodiment of the present application, a first position offset is predicted by a denoising model, a second position offset is obtained by display optical flow calculation or feature matching, and a model loss value is obtained based on the first position offset and the second position offset. The difference between the input and output of the denoising model is determined from the pixel dimension, and the parameters of the denoising model in each round of training are corrected according to the model loss value that characterizes the difference, thereby improving the accuracy of the denoising model and improving the stability and clarity of the denoised video.

[0148] In some embodiments, a reconstruction loss value is obtained based on the degree of difference between the target video frame and the preliminary denoised image; an offset loss value is obtained based on at least one first position offset and at least one second position offset obtained; and a model loss value is obtained based on the reconstruction loss value and the offset loss value.

[0149] Specifically, the reconstruction loss value represents the difference between the overall pixel distribution and structure of the target video frame and the preliminary denoised image.

[0150] The calculation formula for obtaining the reconstruction loss value based on the degree of difference between the target video frame and the preliminary denoised image is shown in the following formula (3):

[0151]

[0152] Among them, L rec Represents the reconstruction loss value; represents the preliminary denoised image; Indicates the target video frame corresponding to the preliminary denoised image.

[0153] For example, for Figure 7 The pixel value of each pixel in the preliminary denoised image 703 in , is the pixel value of each pixel in the target video frame 705. rec is the difference between the preliminary denoised image 703 and the target video frame 705 in the pixel value dimension.

[0154] A calculation formula for the offset loss value is obtained based on at least one first position offset and a second position offset, as shown in the following formula (4):

[0155]

[0156] Among them, L offset Indicates the offset loss value; is the second position offset; is the first position offset.

[0157] The first position offset, the second position offset and the offset loss value have been introduced in the previous text with examples and will not be repeated here.

[0158] The calculation formula of the model loss value is shown in the following formula (5):

[0159] L tc =L rec +L offset …………………(5)

[0160] Among them, L tc Represents the model loss value; L recRepresents the reconstruction loss value; L offset Indicates the offset loss value.

[0161] In order to better explain the embodiment of the present application, the following introduces a method for training a video denoising model provided by the embodiment of the present application in combination with the network architecture of the video denoising model. The process of the method is as follows: Figure 1 The server execution shown includes the following modules, such as Figure 7 As shown:

[0162] The video denoising model includes: a prompt information generation module 702 , a preliminary denoising module 704 and a temporal consistency module 706 .

[0163] First, a sample video frame 701 is input into a prompt information generation module 702. In the prompt information generation module 702, feature encoding is performed on the sample video frame based on a diffusion model to obtain corresponding image features. Noise is gradually added to the image features to obtain a preset noise. The preset noise is fused with the image features to obtain a noise-added feature. Based on the noise-added feature and the image features, denoised prompt information for the sample video frame is generated.

[0164] Specifically, after generating the denoising prompt information for the sample video frame, a sampling subgraph is selected from the multiple sampling subgraphs of the sample video frame as a reference sampling subgraph, and the remaining multiple sampling subgraphs are used as labels to implement self-supervised optimization, thereby obtaining the diffusion loss value of the diffusion model in the prompt information generation module 702. The diffusion model is then reversely corrected based on the diffusion loss value to improve the accuracy of the denoising prompt information. The specific calculation formula for the diffusion loss value is shown in the following formula (6):

[0165]

[0166] Where L represents the diffusion loss value; M represents the number of sample video frames in the training set; N represents the number of sampled sub-graphs in a sample video frame; f(I t1 ) represents a reference sampling sub-graph of a sample video frame; I t2 , I tj Represents the 1st sampling sub-image and the j-1th sampling sub-image in a sample video frame excluding the reference sampling sub-image.

[0167] In the embodiment of the present application, the diffusion loss value of the diffusion model is calculated to optimize the parameters of the diffusion model, thereby improving the accuracy of the denoising prompt information and thus improving the denoising ability of the video denoising model.

[0168] The denoising prompt information and the sample video frame 701 are input into the preliminary denoising module 704. In the preliminary denoising module 704, an attention mechanism is used to perform preliminary denoising on the sample video frame based on the denoising prompt information, and a corresponding preliminary denoised image 703 is obtained.

[0169] The preliminary denoised image 703 and multiple other denoised images are input into the temporal consistency module 706 through a local sliding time window. In the temporal consistency module 706, the following operations are performed on each pixel in the preliminary denoised image 703 using a self-attention mechanism:

[0170] The first position information of a pixel in the preliminary denoised image 703 is obtained, and the denoised pixel corresponding to the pixel in multiple other denoised images of the same object is predicted. A first position offset relative to the first position information in the preliminary denoised image 703 is used, and based on the multiple first position offsets, the second position information of each denoised pixel in the multiple other denoised images is obtained. A preliminary pixel value of the pixel is determined based on the first position information in the preliminary denoised image 703, and first pixel values of the multiple denoised pixels are determined based on the respective second position information of the denoised pixels in the multiple other denoised images. The self-attention mechanism in the temporal consistency module 706 is used to calculate weights corresponding to the preliminary pixel value and the multiple first pixel values based on the correlation between the preliminary pixel value and the multiple first pixel values. Based on the corresponding weights, the preliminary pixel value and the multiple first pixel values are weighted and summed to obtain a second pixel value. The second pixel value is the pixel value of the pixel in the preliminary denoised image after adjustment by the temporal consistency module 706. Based on the obtained second pixel values of the multiple pixels, a target video frame is obtained and output.

[0171] In an embodiment of the present application, multiple other denoised images are input into a temporal consistency module together with the preliminary denoised image, so that the temporal consistency module uses implicitly aligned self-attention to calculate each pixel value in the preliminary denoised image. This allows pixel-level alignment of moving objects without the need for explicit optical flow calculations, which is equivalent to using time-domain filtering to eliminate any remaining slight flicker, thereby improving the accuracy of information interaction between different frames and the stability of the denoised video. The use of a local sliding time window mechanism can control the scope of attention, and only focuses on a number of adjacent frames of each frame (for example, 4 frames before and after each frame), rather than the entire video. This local sliding time window prevents distant frames from interfering with the current preliminary denoised image, improving the anti-interference ability of the denoising process while reducing computational complexity.

[0172] The present application also proposes a video denoising method, the process of which can be as follows: Figure 1 The terminal device 101 shown in FIG. 1 may also be executed by the server 102, or the terminal device 101 and the server 102 may interact to execute the execution. Figure 8 As shown, the following steps are included:

[0173] Step 801 , for a plurality of original video frames in a video to be processed, respectively execute: input an original video frame into a target denoising model for denoising, and obtain a corresponding denoised video frame.

[0174] The target denoising model is obtained by training using the denoising model training method.

[0175] Specifically, in the process of applying the target denoising model, the temporal consistency module can correct the output results of the preliminary denoising module, reduce the discontinuity of the original video frame boundaries caused by segmenting the video to be processed, and smooth the output results between frames through simple post-processing such as time domain filtering.

[0176] Step 802: Obtain a target denoised video based on the obtained multiple denoised video frames.

[0177] In the embodiment of the present application, since the target denoising model incorporates denoising hint information, it can more completely reconstruct the key details covered by noise in the original video frame, effectively avoiding the loss of details in the original video frame, restoring the clear edges and textures of the original video frame, and improving the clarity of the target denoised video. At the same time, since the temporal consistency module in the target denoising model uses an implicit alignment method, combined with other video frames adjacent to the original video frame, each pixel is adjusted, thereby reducing the problem of instability between frames in the target denoised video. Compared with the existing method of completely independent denoising frame by frame, the consistency and stability of the target denoised video in terms of color, brightness, etc. are improved, making the target recognition and behavior analysis results of surveillance videos more reliable in practical applications.

[0178] Based on the same technical concept, the present application embodiment provides a structural diagram of a training device for a video denoising model, such as Figure 9 As shown, the training device 900 of the video denoising model includes:

[0179] Denoising module 901 is configured to perform feature encoding on the sample video frame using the denoising model used in this round, and fuse the corresponding image features with the preset noise to obtain a noise-added feature; and perform preliminary denoising on the sample video frame based on the noise-added feature and the image features to obtain a corresponding preliminary denoised image;

[0180] an adjustment module 902 for adjusting a plurality of pixels in the preliminary denoised image based on motion information between a plurality of other denoised images and the preliminary denoised image to obtain a target video frame; the plurality of other denoised images being preliminary denoised images of other video frames adjacent to the sample video frame;

[0181] The correction module 903 is used to obtain a model loss value based on the target video frame and the sample video frame, and adjust the parameters of the denoising model used in this round according to the model loss value.

[0182] Optionally, the denoising module 901 is specifically configured to:

[0183] generating denoising prompt information of the sample video frame based on the noise addition feature and the image feature;

[0184] Based on the denoising prompt information, preliminary denoising is performed on the sample video frame to obtain a corresponding preliminary denoised image.

[0185] Optionally, the denoising module 901 is specifically configured to:

[0186] generating an attention feature of the sample video frame based on the image feature of the sample video frame;

[0187] Fusing the attention feature and the denoising prompt information to obtain a fusion weight;

[0188] Based on the fusion weights, the sample video frames are adjusted to obtain corresponding preliminary denoised images.

[0189] Optionally, the adjustment module 902 is specifically configured to:

[0190] For each pixel in the preliminary denoised image, the following operations are performed: obtaining motion information associated with the pixel, the motion information including denoised pixels corresponding to the same object as the pixel in each of the other denoised images; updating the preliminary pixel value of the pixel to a second pixel value based on the obtained first pixel value of at least one denoised pixel;

[0191] A target video frame is obtained based on the second pixel value of each pixel in the preliminary denoised image.

[0192] Optionally, the adjustment module 902 is specifically configured to:

[0193] Obtaining a weight of the preliminary pixel value and a weight of each of the at least one first pixel value based on an association relationship between the preliminary pixel value and the at least one first pixel value;

[0194] Based on the weight of the preliminary pixel value and the weight of each of the at least one denoised pixel points, a weighted sum is performed on the preliminary pixel value and the at least one first pixel value to obtain the second pixel value.

[0195] Optionally, the correction module 903 is specifically configured to:

[0196] For each pixel point in the preliminary denoised image, respectively performing: determining a first position offset between a pixel point and at least one associated denoised pixel point;

[0197] For each pixel in the sample video frame, respectively performing the following steps: obtaining third position information of the pixel in the sample video frame, and fourth position information of an original pixel corresponding to the same object as the pixel in each of the other video frames; and determining a second position offset between the third position information of the pixel and at least one associated fourth position information.

[0198] A model loss value is obtained based on the obtained at least one first position offset and the obtained at least one second position offset.

[0199] Optionally, the correction module 903 is specifically configured to:

[0200] Obtaining a reconstruction loss value based on a degree of difference between the target video frame and the preliminary denoised image;

[0201] obtaining an offset loss value based on the obtained at least one first position offset and the obtained at least one second position offset;

[0202] A model loss value is obtained based on the reconstruction loss value and the offset loss value.

[0203] In an embodiment of the present application, a denoising model is used to perform preliminary denoising on sample video frames based on noise features and image features, so that the denoising model can more fully learn the difference between noise features and image features, improve the accuracy of preliminary denoising on sample video frames, and thus improve the stability of the denoised video; based on the motion information between multiple other denoised images and the preliminary denoised image, multiple pixels in the preliminary denoised image are adjusted, so that the denoising model learns the motion trajectory of an object between other denoised images and the preliminary denoised image. Therefore, when the video contains a moving object, it can still accurately associate the features of the same object in multiple video frames, thereby improving the continuity of information interaction between different video frames, reducing the possibility of loss of key information, and thus improving the clarity and stability of the denoised video.

[0204] Based on the same technical concept, the embodiment of the present application provides a structural diagram of a video denoising device, such as Figure 10 As shown, the video denoising device 1000 includes:

[0205] The input module 1001 is configured to execute, for each of the plurality of original video frames in the video to be processed, the following steps: inputting an original video frame into a target denoising model for denoising, thereby obtaining a corresponding denoised video frame; wherein the target denoising model is obtained by training using the aforementioned video denoising model training device;

[0206] The acquisition module 1002 is configured to acquire a target denoised video based on the acquired multiple denoised video frames.

[0207] Based on the same technical concept, the embodiment of the present application provides a computer device, which can be Figure 1 The server shown, such as Figure 11 As shown, it includes at least one processor 1101 and a memory 1102 connected to the at least one processor. The specific link medium between the processor 1101 and the memory 1102 is not limited in the embodiment of the present application. Figure 11 For example, the processor 1101 and the memory 1102 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, and the like.

[0208] In an embodiment of the present application, the memory 1102 stores instructions executed by at least one processor 1101, and the at least one processor 1101 can execute the steps of the above-mentioned video denoising model training method by executing the instructions stored in the memory 1102.

[0209] Among them, the processor 1101 is the control center of the computer device, which can use various interfaces and lines to connect various parts of the computer device, and realize the training of the video denoising model by running or executing instructions stored in the memory 1102 and calling data stored in the memory 1102. Optionally, the processor 1101 may include one or more processing modules. The processor 1101 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1101. In some embodiments, the processor 1101 and the memory 1102 may be implemented on the same chip. In some embodiments, they may also be implemented separately on independent chips.

[0210] The processor 1101 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0211] Memory 1102 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. Memory 1102 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. Memory 1102 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer device, but is not limited thereto. The memory 1102 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0212] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program that can be executed by a computer device. When the program runs on the computer device, the computer device executes the steps of the above-mentioned video denoising model training method.

[0213] Based on the same inventive concept, an embodiment of the present application provides a computer program product, including a computer program stored on a computer-readable storage medium, wherein the computer program includes program instructions. When the program instructions are executed by a computer device, the computer device executes the steps of the above-mentioned video denoising model training method.

[0214] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0215] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0216] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0217] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0218] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A video denoising model training method, characterized in that: include: The denoising model to be trained is iteratively trained using multiple sample video frames until a stopping condition is met, thereby obtaining a trained target denoising model. Each round of iteration includes the following steps: Using the denoising model used in this round, feature encoding is performed on the sample video frame, and the corresponding image features are fused with the preset noise to obtain noise-added features; based on the noise-added features and the image features, preliminary denoising is performed on the sample video frame to obtain a corresponding preliminary denoised image; Adjusting multiple pixels in the preliminary denoised image based on motion information between multiple other denoised images and the preliminary denoised image to obtain a target video frame; the multiple other denoised images are preliminary denoised images of other video frames adjacent to the sample video frame; A model loss value is obtained based on the target video frame and the sample video frame, and parameters of the denoising model used in this round are adjusted according to the model loss value.

2. The method according to claim 1, wherein The performing preliminary denoising on the sample video frame based on the noise addition feature and the image feature to obtain a corresponding preliminary denoised image includes: generating denoising prompt information of the sample video frame based on the noise addition feature and the image feature; Based on the denoising prompt information, preliminary denoising is performed on the sample video frame to obtain a corresponding preliminary denoised image.

3. The method according to claim 2, wherein The performing preliminary denoising on the sample video frame based on the denoising prompt information to obtain a corresponding preliminary denoised image includes: generating an attention feature of the sample video frame based on the image feature of the sample video frame; Fusing the attention feature and the denoising prompt information to obtain a fusion weight; Based on the fusion weights, the sample video frames are adjusted to obtain corresponding preliminary denoised images.

4. The method according to claim 1, wherein The step of adjusting a plurality of pixels in the preliminary denoised image based on motion information between the plurality of other denoised images and the preliminary denoised image to obtain a target video frame includes: For each pixel in the preliminary denoised image, the following operations are performed: obtaining motion information associated with the pixel, the motion information including denoised pixels corresponding to the same object as the pixel in each of the other denoised images; updating the preliminary pixel value of the pixel to a second pixel value based on the obtained first pixel value of at least one denoised pixel; A target video frame is obtained based on the second pixel value of each pixel in the preliminary denoised image.

5. The method according to claim 4, wherein The updating of the preliminary pixel value of the pixel point to a second pixel value based on the obtained first pixel value of each of the at least one denoised pixel points comprises: Obtaining a weight of the preliminary pixel value and a weight of each of the at least one first pixel value based on an association relationship between the preliminary pixel value and the at least one first pixel value obtained; Based on the weight of the preliminary pixel value and the respective weights of the at least one first pixel value, a weighted sum is performed on the preliminary pixel value and the at least one first pixel value to obtain the second pixel value.

6. The method according to claim 4, wherein The obtaining of a model loss value based on the target video frame and the sample video frame includes: For each pixel point in the preliminary denoised image, respectively performing: determining a first position offset between the pixel point and at least one denoised pixel point; For each pixel in the sample video frame, respectively performing the following steps: obtaining third position information of the pixel in the sample video frame, and fourth position information of an original pixel corresponding to the same object as the pixel in each of the other video frames; and determining a second position offset between the third position information of the pixel and at least one associated fourth position information. A model loss value is obtained based on the obtained at least one first position offset and the obtained at least one second position offset.

7. The method according to claim 6, wherein The obtaining of the model loss value based on the obtained at least one first position offset and the obtained at least one second position offset comprises: Obtaining a reconstruction loss value based on a degree of difference between the target video frame and the preliminary denoised image; obtaining an offset loss value based on the obtained at least one first position offset and the obtained at least one second position offset; A model loss value is obtained based on the reconstruction loss value and the offset loss value.

8. A video denoising method, characterized in that: include: For multiple original video frames in the video to be processed, respectively: input an original video frame into a target denoising model for denoising to obtain a corresponding denoised video frame; wherein the target denoising model is trained using the method described in any one of claims 1 to 7; Based on the obtained multiple denoised video frames, a target denoised video is obtained.

9. A computer device, characterized in that: include: a memory for storing program instructions; A processor is configured to call the program instructions stored in the memory and execute the steps of the method according to any one of claims 1 to 7 according to the obtained program.

10. A computer-readable storage medium, characterized in that It stores a computer program that can be executed by a computer device. When the program is run on the computer device, the computer device executes the steps of the method according to any one of claims 1 to 7.