Event camera and Transform-UNet combined video denoising method in low-light environment

By combining the video denoising method of event camera and Transformer-UNet architecture, the problem of insufficient utilization of time domain information in low-light environments is solved, and efficient video denoising effect and visual quality optimization are achieved.

CN120047345APending Publication Date: 2025-05-27XIAN UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510117027.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the time domain information during video denoising in low-light environments, resulting in poor denoising effect.

Method used

Combining the high-temporal resolution characteristics of event cameras and the deep time domain analysis capabilities of Transformer-UNet architecture, an efficient method of video denoising is achieved through multi-layer network processing and loss function optimization.

Benefits of technology

It significantly improves the video noise removal effect, optimizes the overall visual quality of the video, and enhances the adaptability and interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047345A_ABST
    Figure CN120047345A_ABST
Patent Text Reader

Abstract

The invention discloses an event camera and Transform-UNet combined video denoising method in a low-light environment. The method comprises the following steps: firstly, dividing a video event frame data set into a training data set and a test data set; then reconstructing an event frame by using a UNet network, designing an Encodex network based on the UNet to perform down-sampling operation on a low-illumination video frame in a DID data set, designing an Encodey network based on Transform and the UNet, and further enhancing detail and texture information in a video through a fusion module in combination with a self-attention mechanism; and finally, enabling the gradual brightness enhancement task and the fusion task of the low-illumination video to reach an optimal balance state in the training process. According to the method, the limitation that time domain information is not fully utilized in the denoising process in the prior art is solved, so that the video denoising effect is remarkably improved, and the overall visual quality of the video is greatly optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer digital image processing, and particularly relates to a video denoising method combining an event camera and Transformer-UNet in a low-light environment. Background Art

[0002] With the rapid development of digital imaging technology, videos have become an indispensable information dissemination medium in daily life. However, during the acquisition, transmission, and reception of videos, the interference of noise is always present. Especially in environments with insufficient light, the low signal-to-noise ratio makes the noise problem particularly obvious. These noises not only seriously erode the clarity of the video but also have a profound impact on subsequent video processing tasks such as image segmentation and target recognition. Therefore, in order to improve the subjective and objective quality of video images, enhance the compression efficiency of video images, and reduce the bandwidth required for transmission, it is particularly crucial to perform denoising processing on videos.

[0003] In existing video denoising technologies, there are generally two major challenges: First, the utilization of temporal information is insufficient, resulting in unsatisfactory denoising effects; Second, the computational cost of the temporal information extraction module is high, consuming a large amount of energy, and it is difficult to achieve real-time processing. As an emerging imaging technology, event cameras have shown unique advantages in the processing of dynamic scenes due to their excellent temporal resolution and sensitivity to dynamic changes. The Transformer technology, which initially shone in the field of natural language processing, has also begun to stand out in the field of computer vision in recent years, especially in video processing tasks. Its self-attention mechanism can effectively capture the long-range dependencies between video frames, which is crucial for the video denoising task.

[0004] In addition, the UNet architecture, as a widely praised deep learning model in the field of medical image segmentation, has shown excellent performance in the field of image processing due to its unique symmetric U-shaped structure and effective feature fusion strategy. Through its encoder-decoder structure and skip connections, UNet can achieve precise image segmentation and restoration while maintaining image details. This structure not only enhances the network's ability to learn image features but also improves the accuracy of the model in dealing with small objects and fine structures in images.

[0005] Therefore, it has become an urgent task to develop an efficient and practical video denoising model that combines Transformer-UNet and event cameras. This combination can not only make full use of the high temporal resolution characteristics of event cameras but also leverage the powerful feature fusion ability of UNet and the advantages of Transformer in dealing with long-range dependencies to jointly improve the video denoising effect and optimize the overall visual quality of the video. Summary of the Invention

[0006] The objective of the present invention is to provide a video denoising method that combines an event camera with Transformer-UNet in a low-light environment, which solves the limitation in the prior art that time-domain information is not fully utilized during the denoising process, thereby significantly improving the video denoising effect and greatly optimizing the overall visual quality of the video.

[0007] The technical solution adopted by the present invention is a video denoising method that combines an event camera with Transformer-UNet in a low-light environment, which is specifically implemented according to the following steps:

[0008] Step 1: Divide the video event frame dataset into a training dataset and a test dataset;

[0009] Step 2: Use the UNet network to reconstruct the event frames to enhance the detailed information of the event frames;

[0010] Step 3: Design an Encoderx network based on UNet to perform downsampling operations on the low-light video frames in the DID dataset, design an Encodery network based on Transformer and UNet, combine the self-attention mechanism, and further enhance the detailed and texture information in the video through a fusion module;

[0011] Step 4: Use the SCNN and TCNN networks to analyze the video sequence obtained in Step 3;

[0012] Step 5: Perform upsampling through the Decoder to generate the denoising result;

[0013] Step 6: Design a loss function, and through multiple rounds of training, make the low-light video progressive brightness enhancement task and the fusion task obtained in Step 5 reach the best balance state during the training process.

[0014] The present invention is also characterized in that

[0015] Step 1 is specifically implemented according to the following steps:

[0016] Step 1.1: Synthesize the DID dataset into a video information stream, and then input the low-light video into V2E to obtain a video event stream. Among them, the input frame_rate is 30, the slowmotion_factor is 1, ensuring that the event camera maintains the original frame rate when performing video frame interpolation, the video output size is (346, 260), and the sigma threshold is 0.01;

[0017] The event camera is described as:

[0018]

[0019] Where I i and I i-1 are absolute light intensities with timestamps t i , y i at coordinates (x i and t i-1 respectively. The parameter a is the gain of the logarithmic amplifier, and b is the offset to prevent log(0). Then, the logarithmic amplified signal Ω is used to determine by a comparator whether it is strong enough to generate an output, which is described as:

[0020]

[0021] where θ is the threshold of the comparator. If the absolute value of Ω exceeds the preset threshold θ, the comparator will generate an event according to the direction of the luminance gradient and determine its polarity. When the luminance increases, the comparator will output a positive event; when the luminance decreases, a negative event is output. Finally, after an arbitration mechanism designed to reduce the data volume, the arbitration circuit outputs a quaternion e i (x i , y i , t i , p i );

[0022] The output signal of the event-based camera is modified to:

[0023]

[0024] where m is the event number in the event stream, and refer to real events and noise respectively;

[0025] Step 1.2. Reshape the DID dataset into picture frames of size (256, 256). Through normalization, limit the numerical range of the input DID dataset to [0, 1] to eliminate the influence of different dimensions and ensure the balance of each feature during the model training process;

[0026] Step 1.3. According to different low-light scenarios of the dataset, divide the processed normalized data into a training dataset and a test dataset to ensure the diversity and representativeness of model training and evaluation.

[0027] Step 2 is specifically implemented according to the following steps:

[0028] Step 2.1. Preprocessing of the event frames obtained in Step 1.1. First, pass through two convolutional blocks. Each block consists of a convolutional layer with a kernel size of (3,3), a BatchNorm2d layer, and a LeakyReLU activation function. These operations together form the DoubleConv structure. This process expands the number of input channels from 3 to 64, enabling the capture of feature details from shallow to deep layers and providing a rich feature basis for subsequent image reconstruction.

[0029] Step 2.2. Build the Encoder. The first layer of the Encoder is a DoubleConv structure, which contains two consecutive convolutional operations. Each convolutional layer is followed by batch normalization BatchNorm2d and the LeakyReLU activation function. This step extracts the initial features of the image and enhances the feature representation ability. After being processed by DoubleConv, the data undergoes downsampling through the max pooling MaxPooling operation, reducing the spatial dimension of the image while increasing the number of feature channels, preparing for capturing more abstract features. The data continues to pass through a combination of multiple DoubleConv and pooling layers, gradually delving deeper into the network. Each layer increases the number of feature channels while reducing the spatial resolution, achieving the gradual abstraction of features. Finally, the Encoder outputs a high-dimensional feature map, which encodes the deep feature information of the input image and provides rich context for the subsequent fusion process.

[0030] Step 2.3. Build the Decoder. At the end of the downsampling path, i.e., the "bottleneck" layer, the network gradually restores the spatial dimension of the image through the upsampling module. The decoder receives two parameters x1 and x2, where x1 is the low-resolution feature map passed down from the encoder, and x2 is the high-resolution feature map passed down through the skip connection from the corresponding encoder layer. x1 is first upsampled through the transposed layer in the up layer of the UNet network to increase its spatial dimension. Calculate the differences in height and width between x1 and x2, and pad x1 accordingly to ensure that their spatial dimensions are consistent. Concatenate the adjusted x1 and x2 in the channel dimension to fuse the feature information from different levels. The fused feature map is adjusted in terms of channel size through DoubleConv. Through the above operations, high-quality event frame reconstruction is achieved.

[0031] Step 3 is specifically implemented according to the following steps:

[0032] Step 3.1: Establish Encoder Encoderx. The data received by the Encoderx network is a low-light image. First, it passes through multiple convolutional layers with a convolutional kernel of (3, 3) and a stride of 1 to extract deep features and enhance the network's ability to abstract features. Then, it goes through a grouped convolutional layer with a convolutional kernel of (3, 3) and a stride of 2 to reduce the spatial resolution of the feature map while keeping the number of feature channels unchanged, thus completing the downsampling UNetDown process.

[0033] Step 3.2: Establish Encoder Encodery. The data received by the Encodery network is the reconstructed event frame x output by the UNet network and the low-light image y. First, x and y are input into the self-attention module. By learning the feature weights of the event frame x, the importance of each pixel point in the low-light image y is dynamically evaluated. Then, according to the calculated attention weights, the feature map is weighted to highlight significant features. Through the self-attention mechanism, the network can pay more attention to the dynamic changes and important features in the image. Its mathematical expression is:

[0034] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix of the low-light image y respectively. Attention weights are assigned by calculating the similarity between the query and the key, and then the value matrix is weighted. In addition, Through this mechanism, the information that was not fully attended to in the event frame x can be further refined and enhanced through the low-light image y. This complementarity and refinement of information enable the final event frame to retain key dynamic information while also showing richer details and better visual effects.

[0035] Step 3.3: After being processed by the Encoderx and Encodery networks, the self-attention processed feature a1 and the original low-light image y are respectively downsampled through the UNetDown module to further reduce the spatial dimension of the feature map and prepare for feature fusion.

[0036] Step 3.4: After being processed multiple times in Steps 3.2 to 3.3, the downsampled feature maps are sent to the FusionBlock fusion module. First, the features are concatenated in the channel dimension, and then the concatenated result passes through two convolutional layers with a convolutional kernel of (3, 3) and a stride of 1. After each convolutional layer, a CBAM module, instance normalization, and a ReLU activation function are connected to introduce non-linearity and enhance the feature expression ability.

[0037] Step 3.5: The CBAM module integrates channel attention and spatial attention to achieve a comprehensive adjustment of the feature map.

[0038] Step 3.5 is specifically implemented according to the following steps:

[0039] Step 3.5.1: First, multiply the input feature map x by the channel attention weights obtained by calculating through the ChannelAttentionModule to strengthen the important feature channels;

[0040] Step 3.5.2: Then, pass the channel-weighted feature map to the SpatialAttention module, calculate the obtained spatial attention weights, and multiply them with the feature map to highlight the important spatial positions;

[0041] Step 3.5.3: The finally output feature map out is the result after double attention adjustment of channels and space. Such a feature map can more effectively express the key information in the image and provide richer information for the subsequent network layers;

[0042] Step 3.5.4: The ChannelAttentionModule focuses on adjusting the weights in the channel dimension of the feature map, enabling the network to pay more attention to the important feature channels. Use nn.AdaptiveAvgPool2d and nn.AdaptiveMaxPool2d to perform global average pooling and max pooling on the input feature map x respectively to obtain the global average value and maximum value of the feature map. The pooled features pass through a shared multi-layer perceptron, which contains a ReLU activation function to learn the correlations between channels and generate channel attention weights. Use the Sigmoid function to convert the output of the MLP into channel attention weights. The MLP consists of two nn.Conv2d layers. Add the results of processing the global average feature and the maximum feature through the MLP, and then obtain the final channel attention weights through the Sigmoid function, and multiply them with the original feature map x channel by channel to achieve channel weighting;

[0043] Step 3.5.5: The SpatialAttention module focuses on the spatial dimension of the feature map, aiming to enhance the network's attention to spatial positions. First, perform global average and max pooling on the input feature map x to obtain the feature descriptions in the spatial dimension. Concatenate the average feature and the maximum feature obtained by pooling in the channel dimension to form a new feature description. The concatenated features pass through a convolutional layer nn.Conv2d, which is used to learn the correlations between spatial positions. Then still use the Sigmoid function to convert the output of the convolutional layer into spatial attention weights. Multiply the spatial attention weights with the original feature map x element by element to achieve spatial weighting.

[0044] Step 4 is specifically implemented according to the following steps:

[0045] Step 4.1: By stitching together the matched patches, these artificial frames contain patches similar to the original frames but with different noise realizations, providing additional information for denoising.

[0046] Step 4.2: Construct the spatial domain denoising network SCNN: The input of the spatial domain denoising network SCNN is Patch-CraftFrames. The spatial domain denoising network SCNN consists of multiple blocks. Each block includes a SepConv layer followed by a ReLU activation function. The middle blocks have batch normalization BN between SepConv and ReLU. Among them, the SepConv layer consists of three convolutional filters: conv_vh, conv_f, and conv_n. Each filter works on a sub-dimension and treats the remaining dimensions as independent tensors. The input and output are both five-dimensional tensors with sizes n in ×f in ×c×v×h and n out ×f out ×c×v×h. [v, h] is the frame size, c is the number of color layers, f is the patch size, n is the number of neighbors used. The Conv_vh filter applies 2D convolution with a kernel size of m×m, taking dimension c as the input channel and dimensions n and f as independent dimensions. The Conv_f filter applies 2D convolution with a 1×1 kernel, taking dimensions c and f as the input channels and dimension n as an independent dimension. The Conv_n filter applies 2D convolution with a 1×1 kernel, taking dimension n as the input channel and dimensions c and f as independent dimensions. Each SepConv layer reduces the number of neighbors by half, that is This network works in the residual domain, predicting the noise n, and the output frame is obtained by subtracting the predicted noise from the noisy frame.

[0047] Step 4.3: Construct the temporal domain denoising network TCNN: The first part of TCNN is Tf3D, which consists of T t blocks. Each block contains a 3D convolution with a 3×3×3 kernel followed by a Leaky ReLU activation function. The second part is Tf2D, which contains a 2D convolution with a 3×3 kernel followed by a Leaky ReLU activation function. The 3D kernel of Tf3D does not apply padding in the time dimension and applies zero padding in the space dimension. The kernel of Tf2D also applies zero padding. TCNN works in a sliding window manner. The input is 2T t +1 frames. Each TCNN input frame is the concatenation of the input and output of SCNN along the color dimension, and then it is concatenated with the adjacent frame y 1 downsampled by Encodery in the second dimension as the input frame. Similar to SCNN, TCNN works in the residual domain, predicting the noise z t , by subtracting from the partially denoised frame Subtract the predicted noise from it to obtain the output frame

[0048] Step 4.1 is specifically implemented according to the following steps:

[0049] Step 4.1.1: First, extract all possible overlapping patches from the currently processed video frame. The size of these patches is For each extracted patch, the algorithm searches for the most similar patch within the spatio-temporal window defined around it. The size of this window is B×B×(2T S +1), where B represents the size of the spatial axis and 2T S +1 represents the size of the time window, including T S frames forward and T S frames backward. Use the L2 norm as the distance metric to evaluate the similarity between patches;

[0050] Step 4.1.2: The found nearest neighbor patches are used to construct Patch-Craft Frames. During the construction process, only the central part of the patch is used, and the size is

[0051] Step 4.1.3: When constructing Patch-Craft Frames, f groups are created. Each group contains n + 1 frames. The first frame in each group is composed of the first nearest neighbor patch stitched together, the second frame is composed of the second nearest neighbor patch stitched together, and so on until the nth nearest neighbor patch;

[0052] Step 4.1.4: Different groups are constructed by using different patch offsets. The first group uses patches without offsets, that is, the offset is [0,0], the second group uses patches with an offset of [0,1], and so on until the offset is

[0053]

[0054] Step 5 is specifically implemented according to the following steps:

[0055] Build a decoder Decoder. The decoder receives the denoised video frames generated in step 4 The upsampling layer of the decoder consists of two key parts. The first part contains a transposed convolutional layer with a convolutional kernel size of (2,2), a stride set to 2, and two two-dimensional convolutional layers with convolutional kernel sizes of (3,3) and a stride of 1. The second part consists of two-dimensional convolution, instance normalization Instance Normalization, and the Leaky ReLU activation function. These components together introduce non-linear changes and enhance the expressive power of the model.

[0056] Step 6 is implemented according to the following steps:

[0057] Step 6.1: During the training phase, the Adam optimizer is used and the mean square error (MSE) is used as the performance evaluation indicator. The mathematical expression is as follows:

[0058] Among them, K represents the total number of samples in the test set, x K and Respectively represent the denoised video frames output by the model and the corresponding true values;

[0059] Step 6.2: After multiple rounds of iterative training, the model achieves an optimal balance during the training process, thereby being able to produce denoised frames with close to high-level visual quality.

[0060] The beneficial effect of the present invention is that the video denoising method combining event camera and Transformer-UNet in low-light environment proposes an innovative video denoising method by combining the high temporal resolution of event camera and the deep temporal analysis capability of Transformer-UNet architecture. This method not only improves the denoising accuracy and optimizes the video processing flow, but also enhances the adaptability and interpretability of the model, so that it can provide high-quality denoising effects under a variety of video content and noise conditions, bringing significant technological progress to the field of video analysis and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is the overall design idea of ​​the video denoising method combining event camera and Transformer-UNet in low-light environment of the present invention;

[0062] Figure 2 It is a design diagram of the SCNN module in the video denoising method combining the event camera and Transformer-UNet in a low-light environment of the present invention;

[0063] Figure 3 This is a design diagram of the TCNN module in the video denoising method combining an event camera and Transformer-UNet in a low-light environment in the present invention. DETAILED DESCRIPTION

[0064] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] like Figure 1As shown in the figure, the present invention first performs enhanced preprocessing on the event frames. The event frames are reconstructed by UNet to improve the detail information and texture information of the event frames. This provides better event frame image quality for the subsequent fusion tasks in Encodery. Secondly, by designing an Encoderx network based on UNet, downsampling operations are performed on the low-light video frames to extract key features. At the same time, an Encodery network based on Transformer and UNet is designed, combined with the self-attention mechanism, to further enhance the feature information of the low-light video frames. Then, through the fusion module, the detail and texture information in the video is further enhanced to ensure the richness and authenticity of the video content even in low-light environments. Next, the SCNN and TCNN networks are used to perform in-depth feature analysis on the video sequence. Through the joint processing in the spatio-temporal domain, noise is effectively removed and important information is retained. Finally, upsampling is performed through the Decoder to generate high-quality denoising results that conform to the characteristics of human visual perception. At the same time, a loss function is designed to guide the training process. Through multiple loop trainings, the low-light video progressive brightness enhancement task and the fusion task reach the best balance state during the training process.

[0066] The overall design idea of the video denoising method combining an event camera and Transformer-UNet in a low-light environment according to the present invention is as Figure 1 shown, and it is specifically implemented according to the following steps:

[0067] Step 1: Use V2E to simulate the event camera to synthesize the video event frames corresponding to the DID dataset, and divide the dataset to form a training dataset and a test dataset;

[0068] Step 1 is specifically implemented according to the following steps:

[0069] Step 1.1: Synthesize the video information stream of the DID dataset, and then input the low-light video into V2E to obtain the video event stream. Among them, the input frame_rate is 30, the slowmotion_factor is 1, ensuring that the event camera maintains the original frame rate when performing video frame interpolation, the video output size is (346, 260), and the sigma threshold is 0.01;

[0070] The event camera simulates the perception mechanism of the visual sub-pathways of humans and non-human primates, abandons the traditional integral imaging mechanism, and detects the brightness difference on a logarithmic scale. Therefore, it is extremely sensitive to changes in brightness differences. Described as:

[0071]

[0072] where I i and I i-1 are at coordinates (x i , y i) are respectively timestamped with t i and t i-1 of the absolute light intensity. The parameter a is the gain of the logarithmic amplifier, and b is the offset to prevent log(0). Then, the logarithmic amplified signal Ω is used to determine by a comparator whether it is strong enough to generate an output, which is described as:

[0073]

[0074] where θ is the threshold of the comparator. If the absolute value of Ω exceeds the preset threshold θ, the comparator will generate an event according to the direction of the luminance gradient and determine its polarity. When the luminance increases, the comparator will output a positive event; when the luminance decreases, a negative event is output. Finally, after an arbitration mechanism designed to reduce the data volume, the arbitration circuit outputs a quaternion e including the coordinate position, timestamp, and polarity i (x i , y i , t i , p i );

[0075] The above assumption is based on the premise that the output event stream is noise-free. However, due to hardware defects, the event stream will inevitably be mixed with random noise N. Therefore, the output signal of the event-based camera is modified to:

[0076] where m is the event number in the event stream, and respectively refer to the real event and the noise;

[0077] Step 1.2: Reshape the DID dataset into picture frames of size (256, 256). Through normalization, limit the numerical range of the input DID dataset to [0, 1] to eliminate the influence brought by different dimensions and ensure the balance of each feature during the model training process;

[0078] Step 1.3: According to different low-light scenarios of the dataset, divide the processed normalized data into a training dataset and a test dataset to ensure the diversity and representativeness of model training and evaluation.

[0079] Step 2: Use the UNet network to reconstruct the event frames and enhance the detailed information of the event frames;

[0080] Step 2 is specifically implemented according to the following steps:

[0081] Step 2.1. Preprocessing of the event frames obtained in Step 1.1. First, it passes through two convolutional blocks. Each block consists of a convolutional layer with a kernel size of (3,3), a BatchNorm2d layer, and a LeakyReLU activation function. These operations together form the DoubleConv structure. This process expands the number of input channels from 3 to 64, enabling the capture of feature details from shallow to deep layers and providing a rich feature basis for subsequent image reconstruction.

[0082] Step 2.2. Build the Encoder. The first layer of the Encoder is a DoubleConv structure, which contains two consecutive convolutional operations. Each convolutional layer is followed by batch normalization BatchNorm2d and the LeakyReLU activation function. This step extracts the initial features of the image and enhances the feature expression ability. After being processed by DoubleConv, the data undergoes downsampling through the MaxPooling operation, reducing the spatial dimension of the image while increasing the number of feature channels, preparing for capturing more abstract features. The data continues to pass through a combination of multiple DoubleConv and pooling layers, gradually delving deeper into the network. Each layer increases the number of feature channels while reducing the spatial resolution, achieving layer-by-layer abstraction of features. Finally, the Encoder outputs a high-dimensional feature map, which encodes the deep feature information of the input image and provides rich context for the subsequent fusion process.

[0083] Step 2.3. Build the Decoder. At the end of the downsampling path, i.e., the "bottleneck" layer, the network gradually restores the spatial dimension of the image through the upsampling module. The decoder receives two parameters x1 and x2, where x1 is the low-resolution feature map passed down from the encoder, and x2 is the high-resolution feature map passed down through the skip connection from the corresponding encoder layer. x1 is first upsampled through the transposed layer in the up layer of the UNet network to increase its spatial dimension. After upsampling, the spatial dimension of x1 may not match that of x2. Therefore, calculate the differences in height and width between x1 and x2, and pad x1 accordingly (F.pad) to ensure that their spatial dimensions are consistent. Then, concatenate the adjusted x1 and x2 in the channel dimension (torch.cat) to fuse the feature information from different levels. The fused feature map is adjusted in terms of channel size through DoubleConv. Through the above operations, high-quality event frame reconstruction is achieved.

[0084] Step 3: Design an Encoderx network based on UNet to downsample the low-light video frames in the DID dataset, extract key features, and design an Encoderry network based on Transformer and UNet. Combined with the self-attention mechanism, the event frame can enhance the feature information of the low-light video frame. Then, the fusion module is used to further enhance the details and texture information in the video.

[0085] Step 3 is implemented as follows:

[0086] Step 3.1, establish encoder Encoderx: The data received by the Encoderx network is a low-light image (the low-light frame in step 1.2). It first passes through multiple convolution layers with a convolution kernel of (3, 3) and a step size of 1 to extract deep features and deepen the network's ability to abstract features. Then, it passes through a group convolution layer with a convolution kernel of (3, 3) and a step size of 2 to reduce the spatial resolution of the feature map while ensuring that the number of feature channels remains unchanged, completing the downsampling UNetDown process;

[0087] Step 3.2, establish encoder Encoder: The data received by the Encoder network is the reconstructed event frame x and the low-light image y output by the UNet network. First, x and y are input into the self-attention module. By learning the feature weights of the event frame x, the importance of each pixel in the low-light image y is dynamically evaluated. Then, the feature map is weighted according to the calculated attention weight to highlight the significant features. Through the self-attention mechanism, the network can pay more attention to the dynamic changes and important features in the image. Its mathematical expression is:

[0088] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix of the low-light image y, respectively. The attention weights are assigned by calculating the similarity between the query and the key, and then the weighted value matrix. In addition, Through this mechanism, the information that was not sufficiently focused on in the event frame x can be further refined and enhanced through the low-light image y. This complementary and refined information enables the final event frame to retain key dynamic information while also showing richer details and better visual effects.

[0089] Step 3.3, after being processed by the Encoderx and Encodery networks, the self-attention processed feature a1 and the original low-light image y are downsampled through the UNetDown module (the same as the UNetDown module in step 3.1) to further reduce the spatial dimension of the feature map in preparation for feature fusion;

[0090] Step 3.4: After being processed through multiple rounds of Step 3.2 to Step 3.3, the downsampled feature map obtained is fed into the FusionBlock fusion module. First, the features are concatenated in the channel dimension, and then the concatenated result passes through two convolutional layers with a convolutional kernel of (3, 3) and a stride of 1. After each convolutional layer, a CBAM module, instance normalization, and ReLU activation function are connected to introduce non-linearity and enhance the feature expression ability.

[0091] Step 3.5: The CBAM module integrates channel attention and spatial attention to achieve a comprehensive adjustment of the feature map.

[0092] Step 3.5 is specifically implemented according to the following steps:

[0093] Step 3.5.1: First, multiply the input feature map x by the channel attention weights calculated by the ChannelAttentionModule to strengthen the important feature channels.

[0094] Step 3.5.2: Then, pass the channel-weighted feature map to the SpatialAttention module, calculate the obtained spatial attention weights, and multiply them with the feature map to highlight the important spatial positions.

[0095] Step 3.5.3: The finally output feature map out is the result after double attention adjustment in the channel and spatial dimensions. Such a feature map can express the key information in the image more effectively and provide richer information for the subsequent network layers.

[0096] Step 3.5.4: The ChannelAttentionModule focuses on adjusting the weights in the channel dimension of the feature map, enabling the network to pay more attention to important feature channels. Use nn.AdaptiveAvgPool2d and nn.AdaptiveMaxPool2d to perform global average pooling and max pooling on the input feature map x respectively to obtain the global average value and maximum value of the feature map. The pooled features pass through a shared multi-layer perceptron, which contains a ReLU activation function to learn the correlations between channels and generate channel attention weights. Use the Sigmoid function to convert the output of the MLP into channel attention weights. The MLP consists of two nn.Conv2d layers. Add the results of processing the global average feature and the maximum feature through the MLP, and then obtain the final channel attention weights through the Sigmoid function, and multiply them with the original feature map x channel by channel to achieve channel weighting.

[0097] Step 3.5.5: The SpatialAttention module focuses on the spatial dimension of the feature map, aiming to enhance the network's attention to spatial positions. First, global average and max pooling are performed on the input feature map x to obtain feature descriptions in the spatial dimension. The average feature and max feature obtained by pooling are concatenated in the channel dimension to form a new feature description. The concatenated feature passes through a convolutional layer nn.Conv2d, which is used to learn the correlation between spatial positions. Then, the Sigmoid function is still used to convert the output of the convolutional layer into spatial attention weights. The spatial attention weights are multiplied element-wise with the original feature map x to achieve spatial weighting.

[0098] Step 4: Use the SCNN and TCNN networks to perform in-depth feature analysis on the video sequence obtained in Step 3;

[0099] Step 4 is specifically implemented according to the following steps:

[0100] Step 4.1: Construct frames similar to the Ground truth by stitching matching patches. These artificial frames contain patches similar to the original frames but with different noise realizations, providing additional information for denoising.

[0101] Step 4.2: Construct the spatial denoising network SCNN: The input of the spatial denoising network SCNN is Patch-CraftFrames. The spatial denoising network SCNN consists of multiple blocks. Each block includes a SepConv layer followed by a ReLU activation function. The middle blocks have batch normalization BN between the SepConv and ReLU. Among them, the SepConv layer consists of three convolutional filters: conv_vh, conv_f, and conv_n. Each filter works on a sub-dimension and treats the remaining dimensions as independent tensors. The input and output are both five-dimensional tensors with sizes of n in ×f in ×c×v×h and n out ×f out ×c×v×h. [v, h] is the frame size, c is the number of color layers, f is the patch size, n is the number of neighbors used. The Conv_vh filter applies 2D convolution with a kernel size of m×m, taking dimension c as the input channel and dimensions n and f as independent dimensions. The Conv_f filter applies 2D convolution with a 1×1 kernel, taking dimensions c and f as the input channels and dimension n as an independent dimension. The Conv_n filter applies 2D convolution with a 1×1 kernel, taking dimension n as the input channel and dimensions c and f as independent dimensions. Each SepConv layer reduces the number of neighbors by half, that is This network works in the residual domain, predicts the noise n, and obtains the output frame by subtracting the predicted noise from the noise frame.

[0102] Step 4.3, Construct the Temporal Denoising Network TCNN: Although the S-CNN processes the information of adjacent frames, it does not guarantee temporal continuity. Therefore, it is necessary to build the TCNN network. The first part of the TCNN is Tf3D (Temporal Filter3D), which consists of T t blocks, each block contains 3D convolution and a 3×3×3 kernel followed by a Leaky ReLU activation function. The second part is Tf2D (Temporal Filter2D), which contains 2D convolution and a 3×3 kernel followed by a Leaky ReLU activation function. The 3D kernel of Tf3D does not apply padding in the temporal dimension and applies zero padding in the spatial dimension. The kernel of Tf2D also applies zero padding. The TCNN works in a sliding window manner. The input is 2T t +1 frames. Each TCNN input frame is the concatenation of the input of the SCNN and the output of the SCNN along the color dimension, and then it is concatenated with the adjacent frame y 1 downsampled by Encodery in the second dimension as the input frame. Similar to the SCNN, the TCNN works in the residual domain and predicts the noise z t by subtracting the predicted noise from the partially denoised frame to obtain the output frame

[0103] Step 4.1 is specifically implemented according to the following steps:

[0104] Step 4.1.1, First, extract all possible overlapping patches from the currently processed video frames. The size of these patches is For each extracted patch, the algorithm searches for the most similar patch within the spatio-temporal window defined around it. The size of this window is B×B×(2T S +1), where B represents the size of the spatial axis and 2T S +1 represents the size of the time window, including T S frames forward and T S frames backward. The L2 norm (Euclidean distance) is used as the distance metric to evaluate the similarity between patches;

[0105] Step 4.1.2, The found nearest neighbor patches are used to construct Patch-Craft Frames. During the construction process, only the central part of the patches is used, and the size is

[0106] Step 4.1.3, When constructing Patch-Craft Frames, f groups are created, each group contains n + 1 frames. The first frame in each group is composed of the first nearest neighbor patch concatenated, the second frame is composed of the second nearest neighbor patch concatenated, and so on until the nth nearest neighbor patch;

[0107] Step 4.1.4: Different groups are constructed by using different patch offsets. The first group uses patches with no offset, i.e., the offset is [0,0], the second group uses patches with an offset of [0,1], and so on until the offset is

[0108]

[0109] Among them, formula (5) calculates the fractional feature map of each patch through convolution operations, which helps the network learn the similarity between patches. Formula (6) calculates the mean squared difference of all patches to obtain a simplified fractional feature map, which more directly reflects the overall difference between the Patch-Craft frame and the original frame.

[0110] Step 5: Upsample through the Decoder to generate high-quality denoising results that conform to the characteristics of human visual perception;

[0111] Step 5 is specifically implemented according to the following steps:

[0112] Build a decoder Decoder, and the decoder receives the denoised video frames generated in step 4 The upsampling layer of the decoder consists of two key parts. The first part contains a transposed convolutional layer with a convolutional kernel size of (2,2) and a stride of 2, and two two-dimensional convolutional layers with a convolutional kernel size of (3,3) and a stride of 1. This configuration is designed to expand the spatial dimension through the transposed convolutional layer while reducing the number of channels. However, the transposed convolutional layer will introduce a checkerboard effect. Therefore, two two-dimensional convolutional layers are used to effectively prevent this problem. The second part consists of a two-dimensional convolution, instance normalization, and a Leaky ReLU activation function. These components together introduce non-linear changes and enhance the expressive ability of the model.

[0113] Step 6: Design a loss function to guide the training process of the video denoising algorithm combining an event camera and Transformer-UNet. Through multiple rounds of training, the low-light video progressive brightness enhancement task and the fusion task obtained in step 5 reach the best balance state during the training process.

[0114] Step 6 is specifically implemented according to the following steps:

[0115] Step 6.1: In the training phase, the Adam optimizer is used to generate denoised video frames that are as close to the ground truth as possible. The mean square error (MSE) is used as the performance evaluation indicator. The mathematical expression is as follows:

[0116]

[0117] Among them, K represents the total number of samples in the test set, x K and Respectively represent the denoised video frames output by the model and the corresponding true values;

[0118] Step 6.2: After multiple rounds of iterative training, the model achieves an optimal balance during the training process, thereby being able to produce denoised frames with close to high-level visual quality.

[0119] Example 1

[0120] The video denoising method of the present invention combining event camera and Transformer-UNet in low-light environment has the following general design ideas: Figure 1 As shown, the specific implementation steps are as follows:

[0121] Step 1: Use the V2E simulated event camera to synthesize the video event frames corresponding to the DID dataset, and divide the dataset into a training dataset and a test dataset;

[0122] Step 2: Use the UNet network to reconstruct the event frame and enhance the detailed information of the event frame;

[0123] Step 3: Design an Encoderx network based on UNet to downsample the low-light video frames in the DID dataset, extract key features, and design an Encoderry network based on Transformer and UNet. Combined with the self-attention mechanism, the event frame can enhance the feature information of the low-light video frame. Then, the fusion module is used to further enhance the details and texture information in the video.

[0124] Step 4: Use SCNN and TCNN networks to perform in-depth feature analysis on the video sequence obtained in step 3;

[0125] Step 5: Upsample through Decoder to generate high-quality denoising results that conform to human visual perception characteristics;

[0126] Step 6: Design a loss function to guide the training process of the video denoising algorithm based on the combination of event camera and Transformer-UNet. Through multiple cycles of training, the low-light video progressive brightness enhancement task and fusion task obtained in step 5 can reach the optimal balance during the training process.

[0127] Example 2

[0128] The video denoising method combining an event camera with Transformer-UNet in a low-light environment according to the present invention is specifically implemented according to the following steps:

[0129] Step 1: Divide the video event frame dataset into a training dataset and a test dataset;

[0130] Step 1 is specifically implemented according to the following steps:

[0131] Step 1.1: Synthesize the DID dataset into a video information stream, and then input the low-light video into V2E to obtain a video event stream. Among them, the input frame_rate is 30, the slowmotion_factor is 1, to ensure that the event camera maintains the original frame rate when doing video frame interpolation, the video output size is (346, 260), and the sigma threshold is 0.01;

[0132] The event camera is described as:

[0133]

[0134] where I i and I i-1 are the absolute light intensities with timestamps t i , y i at coordinates (x i and t i-1 respectively. The parameter a is the gain of the logarithmic amplifier, and b is the offset to prevent log(0). Then, the logarithmic amplified signal Ω is used to judge whether it is strong enough to generate an output through a comparator, which is described as:

[0135]

[0136] where θ is the threshold of the comparator. If the absolute value of Ω exceeds the preset threshold θ, the comparator will generate an event according to the direction of the brightness gradient and determine its polarity. When the brightness increases, the comparator will output a positive event; when the brightness decreases, a negative event is output. Finally, after the arbitration circuit, after the arbitration mechanism aiming to reduce the data volume, a quaternion e i (x i , y i , t i , p i ) including the coordinate position, timestamp, and polarity is output;

[0137] The output signal of the event-based camera is modified to:

[0138]

[0139] Among them, m is the event number in the event stream, and respectively refer to the real event and noise;

[0140] Step 1.2: Reshape the DID dataset into picture frames of size (256, 256). Through normalization, limit the numerical range of the input DID dataset within [0, 1] to eliminate the influence brought by different dimensions and ensure the balance of each feature during the model training process;

[0141] Step 1.3: According to different low-light scenarios of the dataset, divide the processed normalized data into a training dataset and a test dataset to ensure the diversity and representativeness of model training and evaluation.

[0142] Step 2: Use the UNet network to reconstruct the event frame and enhance the detailed information of the event frame;

[0143] Step 3: Design the Encoderx network based on UNet to perform downsampling operations on the low-light video frames in the DID dataset. Design the Encodery network based on Transformer and UNet, combine the self-attention mechanism, and further enhance the detailed and texture information in the video through the fusion module;

[0144] Step 4: Use the SCNN and TCNN networks to analyze the video sequence obtained in Step 3;

[0145] Step 5: Perform upsampling through the Decoder to generate the denoising result;

[0146] Step 6: Design a loss function. Through multiple rounds of training, make the low-light video progressive brightness enhancement task and the fusion task obtained in Step 5 reach the best balance state during the training process.

[0147] Example 3

[0148] The video denoising method combining an event camera and Transformer-UNet in a low-light environment of the present invention is specifically implemented according to the following steps:

[0149] Step 1: Divide the video event frame dataset into a training dataset and a test dataset;

[0150] Step 1 is specifically implemented according to the following steps:

[0151] Step 1.1. Synthesize the DID dataset into a video information stream, and then input the low-light video into V2E to obtain a video event stream. Among them, the input frame_rate is 30, the slowmotion_factor is 1, ensuring that the event camera maintains the original frame rate during video frame interpolation. The video output size is (346, 260), and the sigma threshold is 0.01;

[0152] The event camera is described as:

[0153]

[0154] Where I i and I i-1 are the absolute light intensities with timestamps t i , y i at coordinates (x i , y i-1 ), respectively. The parameter a is the gain of the logarithmic amplifier, and b is the offset to prevent log(0). Then, the logarithmic amplified signal Ω is used to determine whether it is strong enough to generate an output through a comparator, which is described as:

[0155]

[0156] Where θ is the threshold of the comparator. If the absolute value of Ω exceeds the preset threshold θ, the comparator will generate an event according to the direction of the brightness gradient and determine its polarity. When the brightness increases, the comparator will output a positive event; when the brightness decreases, a negative event will be output. Finally, after the arbitration circuit, following an arbitration mechanism aimed at reducing the data volume, it outputs a quaternion e i (x i , y i , t i , p i );

[0157] The output signal of the event-based camera is modified to:

[0158]

[0159] Where m is the event number in the event stream, and refer to the real event and noise, respectively;

[0160] Step 1.2. Reshape the DID dataset into picture frames of size (256, 256). Through normalization, limit the numerical range of the input DID dataset to [0, 1], eliminate the influence brought by different dimensions, and ensure the balance of each feature during the model training process;

[0161] Step 1.3: Divide the processed normalized data into a training dataset and a test dataset according to different low-light scenarios in the dataset to ensure the diversity and representativeness of model training and evaluation.

[0162] Step 2: Use the UNet network to reconstruct the event frames and enhance the detailed information of the event frames;

[0163] Step 2 is specifically implemented according to the following steps:

[0164] Step 2.1: The preprocessing of the event frames obtained in Step 1.1 first passes through two convolutional blocks. Each block consists of a convolutional layer with a kernel size of (3,3), a BatchNorm2d layer, and a LeakyReLU activation function. These operations together constitute the DoubleConv structure. This process expands the number of input channels from 3 to 64, achieving the capture of feature details from shallow to deep layers and providing a rich feature basis for subsequent image reconstruction.

[0165] Step 2.2: Build an encoder Encoder. The first layer of the encoder Encoder is a DoubleConv structure, which contains two consecutive convolutional operations. Each convolutional layer is followed by a batch normalization BatchNorm2d and a LeakyReLU activation function. This step extracts the preliminary features of the image and enhances the expression ability of the features. After being processed by DoubleConv, the data undergoes downsampling through a max pooling MaxPooling operation, reducing the spatial dimension of the image while increasing the number of feature channels, preparing for capturing more abstract features. The data continues to pass through a combination of multiple DoubleConv and pooling layers, gradually delving deeper into the network. Each layer increases the number of feature channels while reducing the spatial resolution, achieving the gradual abstraction of features. Finally, Encoderx outputs a high-dimensional feature map, which encodes the deep feature information of the input image and provides rich context for the subsequent fusion process.

[0166] Step 2.3: Build a decoder (Decoder). At the end of the downsampling path, i.e., the "bottleneck" layer, the network gradually restores the spatial dimension of the image through an upsampling module. This decoder receives two parameters, x1 and x2, where x1 is the low-resolution feature map passed down from the encoder, and x2 is the high-resolution feature map passed down through a skip connection from the corresponding encoder layer. First, x1 is upsampled through the transposed layer in the up layer of the UNet network to increase its spatial dimension. Calculate the differences in height and width between x1 and x2, and pad x1 accordingly to ensure that their spatial dimensions are consistent. Then, the adjusted x1 and x2 are concatenated in the channel dimension to fuse the feature information from different levels. The fused feature map is adjusted in terms of channel size through DoubleConv. Through the above operations, high-quality event frame reconstruction is achieved.

[0167] Step 3: Design an Encoderx network based on UNet to perform downsampling operations on low-light video frames in the DID dataset. Design an Encodery network based on Transformer and UNet, combined with the self-attention mechanism, and further enhance the detail and texture information in the video through a fusion module.

[0168] Step 4: Use SCNN and TCNN networks to analyze the video sequence obtained in Step 3.

[0169] Step 5: Perform upsampling through the Decoder to generate a denoising result.

[0170] Step 6: Design a loss function, and through multiple rounds of training, make the progressive brightness enhancement task and the fusion task of the low-light video obtained in Step 5 reach the best balance state during the training process.

[0171] Example 4

[0172] The video denoising method combining an event camera and Transformer-UNet in a low-light environment according to the present invention is specifically implemented according to the following steps:

[0173] Step 1: Divide the video event frame dataset into a training dataset and a test dataset.

[0174] Step 2: Use the UNet network to reconstruct the event frames and enhance the detail information of the event frames.

[0175] Step 3: Design an Encoderx network based on UNet to perform downsampling operations on low-light video frames in the DID dataset. Design an Encodery network based on Transformer and UNet, combined with the self-attention mechanism, and further enhance the detail and texture information in the video through a fusion module.

[0176] Step 3 is specifically implemented according to the following steps:

[0177] Step 3.1: Establish Encoderx. The data received by the Encoderx network is a low-light image. First, it passes through multiple convolutional layers with a convolutional kernel of (3, 3) and a stride of 1 to extract deep features, enhancing the network's ability to abstract features. Then, it passes through a grouped convolutional layer with a convolutional kernel of (3, 3) and a stride of 2 to reduce the spatial resolution of the feature map while keeping the number of feature channels unchanged, completing the downsampling UNetDown process.

[0178] Step 3.2: Establish Encodery. The data received by the Encodery network is the reconstructed event frame x output by the UNet network and the low-light image y. First, x and y are input into the self-attention module. By learning the feature weights of the event frame x, the importance of each pixel point in the low-light image y is dynamically evaluated. Then, according to the calculated attention weights, the feature map is weighted to highlight significant features. Through the self-attention mechanism, the network can pay more attention to the dynamic changes and important features in the image. Its mathematical expression is:

[0179]

[0180] Among them, Q, K, and V respectively represent the query matrix, key matrix, and value matrix of the low-light image y. Attention weights are assigned by calculating the similarity between the query and the key, and then the value matrix is weighted. Additionally, Through this mechanism, the information that was not fully attended to in the event frame x can be further refined and enhanced through the low-light image y. This complementary and refined information enables the final event frame to retain key dynamic information while also showing more abundant details and better visual effects.

[0181] Step 3.3: After being processed by the Encoderx and Encodery networks, the self-attention processed feature a1 and the original low-light image y are respectively downsampled through the UNetDown module to further reduce the spatial dimension of the feature map, preparing for feature fusion.

[0182] Step 3.4: After being processed multiple times in Steps 3.2 to 3.3, the downsampled feature maps are fed into the FusionBlock fusion module. First, the features are concatenated in the channel dimension, and then the concatenated result passes through two convolutional layers with a convolutional kernel of (3, 3) and a stride of 1. After each convolutional layer, a CBAM module, instance normalization, and ReLU activation function are connected to introduce non-linearity and enhance the feature expression ability.

[0183] Step 3.5: The CBAM module integrates channel attention and spatial attention to achieve a comprehensive adjustment of the feature map.

[0184] Step 4: Analyze the video sequence obtained in Step 3 using the SCNN and TCNN networks;

[0185] Step 5: Upsample through the Decoder to generate the denoising result;

[0186] Step 6: Design a loss function and, through multiple rounds of training, enable the low-light video progressive brightness enhancement task and the fusion task obtained in Step 5 to reach the optimal balance state during the training process.

[0187] Example 5

[0188] The video denoising method combining an event camera and Transformer-UNet in a low-light environment of the present invention is specifically implemented according to the following steps:

[0189] Step 1: Divide the video event frame dataset into a training dataset and a test dataset;

[0190] Step 2: Use the UNet network to reconstruct the event frames and enhance the detailed information of the event frames;

[0191] Step 3: Design an Encoderx network based on UNet to perform downsampling operations on the low-light video frames in the DID dataset, design an Encodery network based on Transformer and UNet, and combine the self-attention mechanism to further enhance the detailed and texture information in the video through the fusion module;

[0192] Step 4: Analyze the video sequence obtained in Step 3 using the SCNN and TCNN networks;

[0193] Step 4 is specifically implemented according to the following steps:

[0194] Step 4.1: By stitching together matching patches, these artificial frames contain patches similar to the original frames but with different noise realizations, providing additional information for denoising.

[0195] Step 4.2, construct the spatial denoising network SCNN: The input of the spatial denoising network SCNN is Patch-CraftFrames. The spatial denoising network SCNN consists of multiple blocks. Each block includes a SepConv layer followed by a ReLU activation function. The middle blocks have batch normalization BN between SepConv and ReLU. Among them, the SepConv layer consists of three convolutional filters: conv_vh, conv_f, and conv_n. Each filter works on a sub-dimension and treats the remaining dimensions as independent tensors. The input and output are both five-dimensional tensors with sizes n in ×f in ×c×v×h and n out ×f out ×c×v×h. [v, h] is the frame size, c is the number of color layers, f is the patch size, and n is the number of neighbors used. The Conv_vh filter applies 2D convolution with a kernel size of m×m, taking dimension c as the input channel and dimensions n and f as independent dimensions. The Conv_f filter applies 2D convolution with a 1×1 kernel, taking dimensions c and f as the input channels and dimension n as an independent dimension. The Conv_n filter applies 2D convolution with a 1×1 kernel, taking dimension n as the input channel and dimensions c and f as independent dimensions. Each SepConv layer reduces the number of neighbors by half, that is This network works in the residual domain, predicting the noise n, and the output frame is obtained by subtracting the predicted noise from the noisy frame

[0196] Step 4.3, construct the temporal denoising network TCNN: The first part of TCNN is Tf3D, which consists of T t blocks. Each block contains 3D convolution and a 3×3×3 kernel followed by a Leaky ReLU activation function. The second part is Tf2D, which contains 2D convolution and a 3×3 kernel followed by a Leaky ReLU activation function. The 3D kernel of Tf3D does not apply padding in the time dimension and applies zero padding in the space dimension. The kernel of Tf2D also applies zero padding. TCNN works in a sliding window manner. The input is 2T t +1 frames. Each TCNN input frame is the concatenation of the input and output of SCNN along the color dimension, and then it is concatenated with the adjacent frame y 1 downsampled by Encodery in the second dimension as the input frame. Similar to SCNN, TCNN works in the residual domain, predicting the noise z t , and the output frame is obtained by subtracting the predicted noise from the partially denoised frame

[0197] Step 5, perform upsampling through the Decoder to generate the denoising result; ​

[0198] Step 6: Design a loss function and, through multiple cycles of training, make the low-light video progressive brightness enhancement task and fusion task obtained in step 5 reach the best balance during the training process.

[0199] Example 6

[0200] The video denoising method combining an event camera and Transformer-UNet in a low-light environment of the present invention is specifically implemented according to the following steps:

[0201] Step 1: Divide the video event frame dataset into a training dataset and a test dataset;

[0202] Step 2: Use the UNet network to reconstruct the event frame and enhance the detailed information of the event frame;

[0203] Step 3: Design an Encoderx network based on UNet to downsample the low-light video frames in the DID dataset, design an Encoderry network based on Transformer and UNet, combine the self-attention mechanism, and further enhance the details and texture information in the video through the fusion module;

[0204] Step 4: Use SCNN and TCNN networks to analyze the video sequence obtained in step 3;

[0205] Step 5: Upsample through Decoder to generate denoising results;

[0206] Step 6: Design a loss function and, through multiple cycles of training, achieve the best balance between the low-light video progressive brightness enhancement task and the fusion task during the training process.

[0207] Step 6 is implemented according to the following steps:

[0208] Step 6.1: During the training phase, the Adam optimizer is used and the mean square error (MSE) is used as the performance evaluation indicator. The mathematical expression is as follows:

[0209]

[0210] Among them, K represents the total number of samples in the test set, x K and Respectively represent the denoised video frames output by the model and the corresponding true values;

[0211] Step 6.2: After multiple rounds of iterative training, the model achieves an optimal balance during the training process, thereby being able to produce denoised frames with close to high-level visual quality.

[0212] The video denoising method combining an event camera with Transformer-UNet in low-light environments according to the present invention is specifically designed for low-light environments. By integrating the high temporal resolution characteristics of the event camera with the powerful temporal information processing ability of the Transformer-UNet architecture, it effectively overcomes the limitation of the prior art in not fully utilizing temporal information during the denoising process, thereby significantly improving the video denoising effect and greatly optimizing the overall visual quality of the video. The combination of this technology can not only improve the denoising accuracy and optimize the video processing flow, but also enhance the adaptability and interpretability of the model, enabling it to provide high-quality denoising effects under various video contents and noise conditions, bringing significant technological progress to the field of video analysis and processing.

[0213] The present invention significantly overcomes the problem of video denoising in low-light environments by innovatively integrating the event frames synthesized by the event camera. The algorithm not only effectively denoises the low-light video frames, but also fully considers the temporal information of the video frames during the processing, ensuring the superiority of the denoising effect and being closer to the visual perception habits of humans.

Claims

1. A video denoising method combining event camera and Transformer-UNet in low-light environment, characterized in that: Follow the steps below to implement it: Step 1: Divide the video event frame dataset into a training dataset and a test dataset; Step 2: Use the UNet network to reconstruct the event frame and enhance the detailed information of the event frame; Step 3: Design an Encoderx network based on UNet to downsample the low-light video frames in the DID dataset, design an Encoderry network based on Transformer and UNet, combine the self-attention mechanism, and further enhance the details and texture information in the video through the fusion module; Step 4: Use SCNN and TCNN networks to analyze the video sequence obtained in step 3; Step 5: Upsample through Decoder to generate denoising results; Step 6: Design a loss function and, through multiple cycles of training, make the low-light video progressive brightness enhancement task and fusion task obtained in step 5 reach the best balance during the training process.

2. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 1, characterized in that: The step 1 is specifically implemented according to the following steps: Step 1.1: Synthesize the DID data set into a video information stream, and then input the low-light video into V2E to obtain a video event stream. The input frame_rate is 30, the slowmotion_factor is 1, and the event camera maintains the original frame rate when inserting video frames. The video output size is (346, 260), and the sigma threshold is 0.

01. The event camera is described as: Among them I i and I i-1 is at the coordinate (x i ,y i ) are respectively marked with a time stamp t i and t i-1 The absolute light intensity is given by the logarithmic amplifier, the parameter a is the gain of the logarithmic amplifier, and b is to prevent the offset of log(0). Then, the logarithmic amplified signal Ω is used by the comparator to determine whether it is strong enough to generate an output, which is described as: Among them, θ is the threshold of the comparator. If the absolute value of Ω exceeds the preset threshold θ, the comparator will generate an event according to the direction of the brightness gradient and determine its polarity. When the brightness increases, the comparator will output a positive event; when the brightness decreases, it will output a negative event. Finally, the arbitration circuit outputs a quaternion e including coordinate position, timestamp and polarity after an arbitration mechanism designed to reduce the amount of data. i (x i ,y i ,t i ,p i ); The output signal of the event-based camera is modified to: Where m is the event number in the event stream, and They refer to real events and noise, respectively; Step 1.2: Reshape the DID dataset into a picture frame of size (256, 256). Through normalization, limit the numerical range of the input DID dataset to [0, 1] to eliminate the influence of different dimensions and ensure the balance of various features during model training. Step 1.3: Divide the processed normalized data into training data sets and test data sets according to the different dark light scenes of the data sets to ensure the diversity and representativeness of model training and evaluation.

3. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 2, characterized in that: The step 2 is specifically implemented according to the following steps: Step 2.1: The preprocessing of the event frame obtained in step 1.1 first passes through two convolution blocks. Each block consists of a convolution layer with a kernel size of (3,3), a BatchNorm2d layer, and a LeakyReLU activation function. These operations together form a DoubleConv structure. This process expands the number of input channels from 3 to 64, achieving feature detail capture from shallow to deep layers, and providing a rich feature basis for subsequent image reconstruction. Step 2.2, establish the encoder. The first layer of the encoder is a DoubleConv structure, which contains two consecutive convolution operations. Each convolution layer is followed by batch normalization BatchNorm2d and LeakyReLU activation function. This step extracts the preliminary features of the image and enhances the expressiveness of the features. After DoubleConv processing, the data is downsampled by the maximum pooling MaxPooling operation to reduce the spatial dimension of the image and increase the number of feature channels to prepare for capturing more abstract features. The data continues to pass through a combination of multiple DoubleConv and pooling layers, gradually going deeper into the network. Each layer reduces the spatial resolution while increasing the number of feature channels to achieve layer-by-layer abstraction of features. Finally, Encoderx outputs a high-dimensional feature map, which encodes the deep feature information of the input image and provides rich context for the subsequent fusion process. Step 2.3, establish a decoder Decoder. At the end of the downsampling path, that is, the "bottleneck" layer, the network gradually restores the spatial dimension of the image through the upsampling module. The decoder receives two parameters x1 and x2, where x1 is the low-resolution feature map passed down from the encoder, and x2 is the high-resolution feature map passed down from the corresponding encoder layer through the jump connection. x1 is first upsampled through the transpose layer in the up layer of the UNet network to increase its spatial dimension. The difference in height and width between x1 and x2 is calculated, and x1 is padded accordingly to ensure that the spatial dimensions of the two are consistent. The adjusted x1 and x2 are spliced ​​in the channel dimension to fuse feature information from different levels. The fused feature map is adjusted by DoubleConv to adjust the channel size. Through the above operations, high-quality event frame reconstruction is achieved.

4. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 3, characterized in that: The step 3 is specifically implemented according to the following steps: Step 3.1, establish encoder Encoderx: The data received by the Encoderx network is a low-light image. It first passes through multiple convolution layers with a convolution kernel of (3, 3) and a step size of 1 to extract deep features and deepen the network's ability to abstract features. Then, it passes through a group convolution layer with a convolution kernel of (3, 3) and a step size of 2 to reduce the spatial resolution of the feature map while ensuring that the number of feature channels remains unchanged, completing the downsampling UNetDown process; Step 3.2, establish encoder Encoder: The data received by the Encoder network is the reconstructed event frame x and the low-light image y output by the UNet network. First, x and y are input into the self-attention module. By learning the feature weights of the event frame x, the importance of each pixel in the low-light image y is dynamically evaluated. Then, the feature map is weighted according to the calculated attention weight to highlight the significant features. Through the self-attention mechanism, the network can pay more attention to the dynamic changes and important features in the image. Its mathematical expression is: Among them, Q, K, and V represent the query matrix, key matrix, and value matrix of the low-light image y, respectively. The attention weights are assigned by calculating the similarity between the query and the key, and then the weighted value matrix. In addition, Through this mechanism, the information that was not sufficiently focused on in the event frame x can be further refined and enhanced through the low-light image y. This complementary and refined information enables the final event frame to retain key dynamic information while also showing richer details and better visual effects. Step 3.3, after being processed by the Encoderx and Encodery networks, the self-attention processed feature a1 and the original low-light image y are downsampled through the UNetDown module respectively to further reduce the spatial dimension of the feature map in preparation for feature fusion; Step 3.4: After multiple processing steps 3.2 to 3.3, the obtained downsampled feature map is sent to the FusionBlock fusion module. First, the features are concat spliced ​​in the channel dimension, and then the spliced ​​results are passed through two convolutional layers with a convolution kernel of (3, 3) and a step size of 1. After each convolutional layer, the CBAM module, instance normalization and ReLU activation function are connected to introduce nonlinear characteristics and enhance feature expression capabilities. Step 3.5, the CBAM module integrates channel attention and spatial attention to achieve comprehensive adjustment of feature maps.

5. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 4, characterized in that: The step 3.5 is specifically implemented according to the following steps: Step 3.5.

1. First, multiply the channel attention weight calculated by ChannelAttentionModule by the input feature map x by x to strengthen the important feature channels. Step 3.5.2, then, pass the channel-weighted feature map to the SpatialAttention module, calculate the spatial attention weight, and multiply it with the feature map to highlight the important spatial positions; Step 3.5.3, the final output feature map out is the result of channel and spatial dual attention adjustment. Such a feature map can more effectively express the key information in the image and provide richer information for subsequent network layers; Step 3.5.4, ChannelAttentionModule focuses on adjusting the weights on the channel dimension of the feature map so that the network can pay more attention to important feature channels. nn.AdaptiveAvgPool2d and nn.AdaptiveMaxPool2d are used to perform global average pooling and maximum pooling on the input feature map x to obtain the global average and maximum values ​​of the feature map respectively. The pooled features pass through a shared multi-layer perceptron, which contains a ReLU activation function to learn the correlation between channels and generate channel attention weights. The output of MLP is converted into channel attention weights using the Sigmoid function. MLP consists of two nn.Conv2d layers. The global average feature and the maximum feature are added after processing by MLP, and then the final channel attention weight is obtained by the Sigmoid function, which is multiplied by the original feature map x channel by channel to achieve channel weighting. Step 3.5.5, the SpatialAttention module focuses on the spatial dimension of the feature map, aiming to enhance the network's attention to spatial positions. First, global averaging and maximum pooling are performed on the input feature map x to obtain a feature description in the spatial dimension. The average features and maximum features obtained by pooling are concatenated in the channel dimension to form a new feature description. The concatenated features are passed through a convolutional layer nn.Conv2d, which is used to learn the correlation between spatial positions. Then, the Sigmoid function is still used to convert the output of the convolutional layer into spatial attention weights, and the spatial attention weights are multiplied element-by-element with the original feature map x to achieve spatial weighting.

6. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 5, characterized in that: The step 4 is specifically implemented according to the following steps: Step 4.1: By splicing matched patches, these artificial frames contain patches similar to the original frames, but with different noise realizations, providing additional information for denoising. Step 4.2, build the spatial denoising network SCNN: The input of the spatial denoising network SCNN is Patch-Craft Frames. The spatial denoising network SCNN consists of multiple blocks, each of which includes a SepConv layer followed by a ReLU activation function. The intermediate block has a batch normalization BN between SepConv and ReLU. Among them, the SepConv layer consists of three convolutional filters: conv_vh, conv_f and conv_n. Each filter works on a sub-dimension and treats the remaining dimensions as independent tensors. The input and output are both five-dimensional tensors with sizes n and n respectively. in ×f in ×c×v×h and n out ×f out ×c×v×h, [v,h] is the frame size, c is the number of color layers, f is the patch size, n is the number of neighbors used, Conv_vh filters apply 2D convolutions with a kernel size of m×m, taking dimension c as input channels and dimensions n and f as independent dimensions, Conv_f filters apply 2D convolutions with a 1×1 kernel, taking dimensions c and f as input channels and dimension n as independent dimensions, Conv_n filters apply 2D convolutions with a 1×1 kernel, taking dimension n as input channels and dimensions c and f as independent dimensions, and each SepConv layer reduces the number of neighbors by half, i.e. The network works in the residual domain, predicts the noise n, and obtains the output frame by subtracting the predicted noise from the noisy frame Step 4.3: Construct the time domain denoising network TCNN: The first part of TCNN is Tf3D, which is composed of T t The first part is composed of blocks, each of which contains 3D convolution and 3×3×3 kernel followed by Leaky ReLU activation function. The second part is Tf2D, which contains 2D convolution and 3×3 kernel followed by Leaky ReLU activation function. The 3D kernel of Tf3D does not apply padding in the time dimension, but applies zero padding in space. The kernel of Tf2D also applies zero padding. TCNN works in a sliding window manner, and the input is 2T t +1 frame, each TCNN input frame is the concatenation of the SCNN input and SCNN output along the color dimension, and then it is concatenated with the adjacent frame y1 obtained by downsampling the Encoder in the second dimension as the input frame. Similar to SCNN, TCNN works in the residual domain and predicts the noise z t , by removing noise from the part of the frame Subtract the predicted noise from the output frame 7. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 6, characterized in that: The step 4.1 is specifically implemented according to the following steps: Step 4.1.

1. First, extract all possible overlapping patches from the currently processed video frame. The size of these patches is For each extracted patch, the algorithm searches for the most similar patch within a spatiotemporal window defined around it, the size of this window is B×B×(2T S +1), where B represents the size of the spatial axis, 2T S +1 indicates the size of the time window, including T S Frame forward and T S Frame backward, using the L2 norm as the distance metric to evaluate the similarity between patches; Step 4.1.2, find the nearest neighbor patches used to construct Patch-Craft Frames. During the construction process, only the central part of the patch is used, with a size of Step 4.1.3: When constructing Patch-Craft Frames, f groups are created, each containing n+1 frames. The first frame in each group is concatenated from the first nearest neighbor patch, the second frame is concatenated from the second nearest neighbor patch, and so on, until the nth nearest neighbor patch. Step 4.1.

4. Different groups are constructed by using different patch offsets. The first group uses patches with no offset, i.e., offset [0,0], the second group uses patches with offset [0,1], and so on, until the offset is 8. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 7, characterized in that: The step 5 is specifically implemented according to the following steps: Establish a decoder Decoder, the decoder receives the denoised video frame generated in step 4 The upsampling layer of the decoder consists of two key parts. The first part contains a transposed convolution layer with a convolution kernel size of (2,2) and a stride of 2, and two two-dimensional convolution layers with a convolution kernel size of (3,3) and a stride of 1. The second part consists of two-dimensional convolution, instance normalization, and leaky ReLU activation function. These components together introduce nonlinear changes and enhance the expressiveness of the model.

9. The video denoising method combining event camera and Transformer-UNet in low light environment according to claim 8, characterized in that: The step 6 is specifically implemented according to the following steps: Step 6.1: During the training phase, the Adam optimizer is used and the mean square error (MSE) is used as the performance evaluation indicator. The mathematical expression is as follows: Among them, K represents the total number of samples in the test set, x K and Respectively represent the denoised video frames output by the model and the corresponding true values; Step 6.2: After multiple rounds of iterative training, the model achieves an optimal balance during the training process, thereby being able to produce denoised frames with close to high-level visual quality.

Citation Information

Cited By

  • Method and device for generating clear enhanced image by fusing image and event under low illumination

    CN120339122A

  • Low illumination enhancement method based on event and image bidirectional collaborative guidance

    CN120852257A

  • Low-light enhancement method based on event and image bidirectional collaborative guidance

    CN120852257B

  • Digital PCR image real-time denoising method and system based on gradient information enhancement

    CN120953116A