Low-illumination event stream enhancement method based on conditional diffusion model
By applying a denoising framework based on the conditional diffusion model in low-illumination event stream processing, the noise problem of low-illumination event stream is solved, significantly improving the signal-to-noise ratio and target information integrity.
Patent Information
- Application Number
- CN202510226163.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The event stream captured by the event camera in low-illumination scenarios is susceptible to noise interference, resulting in a reduced signal-to-noise ratio and it is difficult to preserve the integrity of the scene target information.
A denoising framework based on the conditional diffusion model is adopted to perform diffusion denoising processing on low-illumination event streams. The specific steps include preprocessing the data set, stacking the event stream time-series every 5 consecutive video frames, adding noise in the forward process, denoising the network through noise in the reverse process, and finally outputting the denoised image.
It significantly improves the signal-to-noise ratio of the event stream after denoising, retains more complete scene target information, and enhances the richness of texture details and visual characteristics.
Smart Images

Figure CN120070244A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer video processing, and particularly relates to a method for enhancing low-light event streams based on a conditional diffusion model. Background Art
[0002] Benefiting from significant advantages such as high dynamic range, low latency, and high temporal resolution, event cameras can capture the structural features of moving targets in low-light scenarios. Since the output of an event camera is an asynchronous event stream rather than a traditional continuous frame image, a series of challenges are faced in processing and analyzing these event streams, and one of them is the noise problem. To address the above issues, the present invention proposes a denoising framework based on a diffusion model, aiming to perform diffusion denoising on low-light event streams, thereby improving the signal-to-noise ratio of the denoised event streams and retaining more complete scene target information. Summary of the Invention
[0003] The object of the present invention is to provide a method for enhancing low-light event streams based on a conditional diffusion model, which can retain the integrity of scene target information as much as possible and enhance the richness of texture details and visual features.
[0004] The technical solution adopted by the present invention is a method for enhancing low-light event streams based on a conditional diffusion model, which is specifically implemented according to the following steps:
[0005] First, preprocess the dataset. Stack every 5 consecutive video frames of the event stream under normal light conditions in time series. For the low-light event stream, perform the same operation. In the forward process, sequentially add noise to the stacked sequences of each event stream under normal light conditions. In the reverse process, sequentially perform denoising operations on the stacked sequences of each low-light event stream, and finally output the denoised image.
[0006] The characteristics of the present invention also lie in that
[0007] It is specifically implemented according to the following steps:
[0008] Step 1: Use all the frames of the event streams in different scenarios under normal light conditions and all the frames of the low-light event stream corresponding to each scenario as the dataset. First, preprocess the dataset. Stack every 5 consecutive video frames of the event stream in each scenario under normal light conditions in time series to obtain a stacked sequence. For all the frames of the low-light event stream in each scenario, perform the same operation. In the forward process, concatenate the stacked sequences of the event streams under normal light conditions and the corresponding stacked sequences of the low-light event streams along the channel dimension to strengthen the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0009] Step 2: Design a noise estimation network for the forward diffusion process. Gradually add Gaussian noise to the Y estimate of the concatenated sequence along the channel dimension of the stacked sequence of event streams under normal light conditions and the corresponding stacked sequence of low-light event streams at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can process temporal and spatial information in video data.
[0010] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence with noise at step t, and the second part is the stacked sequence of low-light event streams. They will be used as the conditional input of the reverse diffusion model. Concatenate these two parts along the channel dimension as the conditional input of the above noise estimation network, fit the current posterior mean and variance, and finally calculate and output the denoised image based on the mean, log variance, and random noise predicted by this noise estimation network.
[0011] Step 1 is specifically implemented according to the following steps:
[0012] Step 1.1: Respectively give a stacked sequence of normal light event streams and the corresponding stacked sequence of low-light event streams Concatenate them along the channel dimension to obtain a sequence
[0013] Step 1.2: Add noise ε to the data at each step t from 0 to T t , and this process is expressed as:
[0014]
[0015] where is the sequence after adding noise ε to the data at the current step t from the previous step t - 1 ; β t is the scale parameter of the noise, set as a small positive number; I is the identity matrix; t
[0016] Step 1.3: Gradually add noise, and the distribution of the data gradually approaches the standard Gaussian distribution N(0,Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, and a series of noise-polluted data of the normal illumination event stream video frames is obtained
[0017] The result generated by the forward process is expressed as:
[0018]
[0019] where is the data at the current step t from the previous step t - 1 Add noise ε t to the sequence; is a parameter related to β t used to represent the degree of noise accumulation, and I is the identity matrix.
[0020] Step 2 is specifically implemented according to the following steps:
[0021] Step 2.1: Adopt the 3D U-Net architecture as the noise estimation network;
[0022] Step 2.2: Use the concatenated sequence containing noise in the t-step process as the input of the noise estimation network and output the predicted noise, denoted as
[0023] Step 2.3: Starting from the training objective of the noise estimation network, design the loss function during the training process of the noise estimation network.
[0024] The loss function during the training process of the noise estimation network in Step 2.3 is denoted as:
[0025]
[0026] Use ||.|| 2 as the loss function to measure the difference between the predicted noise and the true noise.
[0027] Step 3 is specifically implemented according to the following steps:
[0028] Step 3.1: Estimate and remove the noise added in the forward process;
[0029] Step 3.2: After the reverse process is completed for all time steps, the denoised image is finally obtained.
[0030] Step 3.1 is specifically implemented according to the following steps:
[0031] At each time step t, the diffusion model proposed in this paper will estimate the distribution of the original data from the noisy data through learning, and this estimation process is denoted as:
[0032]
[0033] where, μ θ and ∑ θ are the mean and variance parameterized by the model;
[0034] The diffusion model gradually recovers from time step T to time step 0 by gradually denoising, that is:
[0035]
[0036] Step 3.2 is specifically implemented according to the following steps:
[0037] At each step, by means of sampling, using the mean and variance estimated in the previous step, from generate This process starts from the initial noise and gradually recovers to through the reverse diffusion steps, that is, new data samples are generated. When the reverse process completes all time steps, the finally obtained is the finally denoised image;
[0038] Each step of the reverse process is described in the following form:
[0039]
[0040] where α t and β t are noise parameters in the forward process, is the cumulative noise parameter, ε t is the noise estimated by the model, σ t is the standard deviation of each step, used to control the denoising intensity, Z t is an additional noise term to ensure the randomness of the generation process. When predicting the noise ε t at the t-th step, the method is the same as that in Step 2.2 described above.
[0041] The beneficial effect of the present invention is that for the low-light event stream enhancement method based on the conditional diffusion model, the dataset is first preprocessed. Artificially, every 5 consecutive video frames of the event stream under normal light conditions are stacked in time series. For the low-light event stream, the same operation is also performed; in the forward process, noise is sequentially added to the stacked sequences of each event stream under normal light conditions; in the reverse process, denoising operations are sequentially performed on the stacked sequences of each low-light event stream, and finally the denoised image is output. That is, a diffusion model is adopted as the denoising framework. In the forward process, the event stream data is gradually transformed into a noise distribution by constructing a Markov chain, and a 3D U-Net architecture is constructed to estimate the noise added by the forward diffusion; in the reverse reconstruction process, combined with the above-mentioned constructed noise estimation network, clear event features are accurately restored from the noisy event stream data, thereby significantly improving the quality and reliability of the event stream data. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the algorithm block diagram of the low-light event stream enhancement method based on the conditional diffusion model of the present invention;
[0043] Figure 2These are the consecutive 5-frame visual effect diagrams of the low-light event stream enhancement method based on the conditional diffusion model of the present invention. Detailed implementation manners
[0044] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.
[0045] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 , and is specifically implemented according to the following steps:
[0046] First, preprocess the data set. Artificially stack the video frames of every consecutive 5 frames of the event stream under normal light conditions in time series. For the low-light event stream, perform the same operation. In the forward process, add noise to the stacked sequence of the event stream under normal light conditions one by one. In the reverse process, perform denoising operations on the stacked sequence of the low-light event stream one by one, and finally output the denoised image.
[0047] Specifically, it is implemented according to the following steps:
[0048] Step 1: Use all the frames of the event streams in different scenarios under normal light conditions and all the frames of the low-light event stream corresponding to each scenario as the data set. First, preprocess the data set, stack the video frames of every consecutive 5 frames of the event stream in each scenario under normal light conditions in time series to obtain a stacked sequence. For all the frames of the low-light event stream in each scenario, perform the same operation. In the forward process, splice the stacked sequence of the event stream under normal light conditions and the stacked sequence of the corresponding low-light event stream along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each spliced sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise;
[0049] Step 1 is specifically implemented according to the following steps:
[0050] Step 1.1: Respectively give a stacked sequence of the event stream under normal light and the stacked sequence of the corresponding low-light event stream Splice them along the channel dimension to obtain the sequence
[0051] Step 1.2: Add noise ε to the data at each step t from 0 to T t , and this process is expressed as:
[0052]
[0053] where is the data at the current step t adding noise ε to the data at the previous step t-1 adding noise εt The sequence after; β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix;
[0054] Step 1.3, Gradually add noise, and the data distribution gradually approaches the standard Gaussian distribution N(0,Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, and a series of noise-polluted data of the video frames of the normal illumination event stream are obtained
[0055] At each step of the forward process, a certain amount of noise is accumulated. This cumulative effect causes the original data to gradually lose its structure and information and finally be completely covered by noise. The result generated by the forward process is expressed as:
[0056]
[0057] where, is the data at the previous step t - 1 when moving forward at the current step t with the added noise ε t The sequence after; is a parameter related to β t used to represent the degree of noise accumulation, and I is the identity matrix.
[0058] Step 2, Design a noise estimation network for the forward diffusion process. Gradually add Gaussian noise to the Y estimation of the concatenated sequence along the channel dimension of the stacked sequence of the event stream under normal light conditions and the corresponding stacked sequence of the low-light event stream at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can handle temporal and spatial information in video data.
[0059] Step 2 is specifically implemented according to the following steps:
[0060] Step 2.1, Adopt a 3D U-Net architecture as the noise estimation network;
[0061] Step 2.2, Use the concatenated sequence containing noise during the t-step process as the input of the noise estimation network and output the predicted noise, expressed as
[0062] Step 2.3, Starting from the training objective of the noise estimation network, make the noise prediction output by the network as close as possible to the actually added noise, and design the loss function during the training process of the noise estimation network.
[0063] The loss function during the training process of the noise estimation network in Step 2.3 is expressed as:
[0064]
[0065] Use ||.|| 2 as the loss function to measure the difference between the predicted noise and the true noise.
[0066] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at step t, and the second part is the stacked sequence of low-light event streams. They will be used as the conditional input of the reverse diffusion model. Concatenate these two parts along the channel dimension as the conditional input of the above noise estimation network, fit the current posterior mean and variance, and finally calculate and output the denoised image according to the mean, log variance, and random noise predicted by this noise estimation network.
[0067] Step 3 is specifically implemented according to the following steps:
[0068] In the reverse process, use the above-trained noise estimation network, and use the Y of the stacked sequence of every 5 video frames of the low-light event stream as the conditional input Remove noise step by step.
[0069] Step 3.1: Estimate and remove the noise added in the forward process;
[0070] Step 3.1 is specifically implemented according to the following steps:
[0071] At each time step t, the diffusion model proposed in this paper will learn to estimate the distribution of the original data from the noisy data This estimation process is expressed as:
[0072]
[0073] where μ θ and ∑ θ are the mean and variance parameterized by the model;
[0074] The diffusion model gradually recovers from time step T to time step 0 by gradually denoising, that is:
[0075]
[0076] Step 3.2: After all time steps are completed in the reverse process, finally obtain the denoised image.
[0077] Step 3.2 is specifically implemented according to the following steps:
[0078] At each step, by sampling, using the mean and variance estimated in the previous step, from generate This process starts from the initial noise Start by gradually restoring through the reverse diffusion step to That is, new data samples are generated, and this process reflects the path of the data gradually restoring from the Gaussian noise distribution to the original data distribution. When the reverse process completes all time steps, the finally obtained is the finally denoised image;
[0079] Each step of the reverse process is described in the following form:
[0080]
[0081] where α t and β t are the noise parameters in the forward process, is the cumulative noise parameter, ε t is the noise estimated by the model, σ t is the standard deviation of each step, used to control the denoising intensity, Z t is the additional noise term to ensure the randomness of the generation process. When predicting the noise ε t at the t-th step, the method is the same as that in step 2.2 described above.
[0082] Example 1
[0083] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 and is specifically implemented according to the following steps:
[0084] First, preprocess the dataset. Artificially stack the event streams under normal light conditions every 5 consecutive video frames in time series, and perform the same operation for the low-light event streams; in the forward process, sequentially add noise to the stacked sequences of each event stream under normal light conditions; in the reverse process, sequentially perform denoising operations on the stacked sequences of each low-light event stream, and finally output the denoised image.
[0085] Example 2
[0086] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 and is specifically implemented according to the following steps:
[0087] First, preprocess the dataset. Artificially stack the event streams under normal light conditions every 5 consecutive video frames in time series, and perform the same operation for the low-light event streams; in the forward process, sequentially add noise to the stacked sequences of each event stream under normal light conditions; in the reverse process, sequentially perform denoising operations on the stacked sequences of each low-light event stream, and finally output the denoised image.
[0088] Specifically, it is implemented according to the following steps:
[0089] Step 1: Use all the frames of the event streams of different scenarios under normal light conditions and all the frames of the low-light event streams of each corresponding scenario as a dataset. First, preprocess the dataset. Stack the video frames of each event stream of each scenario under normal light conditions in sequence every 5 consecutive frames to obtain a stacked sequence. For all the frames of the low-light event stream of each scenario, perform the same operation. In the forward process, concatenate the stacked sequences of the event streams under normal light conditions and the corresponding stacked sequences of the low-light event streams along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0090] Step 2: Design a noise estimation network for the forward diffusion process. Gradually add Gaussian noise to the Y estimate of the sequence obtained by concatenating the stacked sequences of the event streams under normal light conditions and the corresponding stacked sequences of the low-light event streams along the channel dimension at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure, which can process the temporal and spatial information in video data.
[0091] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at the t-th step, and the second part is the stacked sequence of the low-light event stream. They will be used as the conditional input of the reverse diffusion model. Concatenate these two parts along the channel dimension as the conditional input of the above noise estimation network, fit the current posterior mean and variance, and finally calculate and output the denoised image based on the mean, log variance, and random noise predicted by this noise estimation network.
[0092] Example 3
[0093] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 , and is specifically implemented according to the following steps:
[0094] First, preprocess the dataset. Manually stack the video frames of the event stream under normal light conditions in sequence every 5 consecutive frames. For the low-light event stream, perform the same operation. In the forward process, add noise to the stacked sequence of each event stream under normal light conditions in sequence. In the reverse process, perform denoising operations on the stacked sequence of each low-light event stream in sequence, and finally output the denoised image.
[0095] Specifically, it is implemented according to the following steps:
[0096] Step 1: Take all the frames of the event streams of different scenarios under normal light conditions and all the frames of the low-light event streams of each corresponding scenario as a dataset. First, preprocess the dataset. Stack every 5 consecutive video frames of the event stream of each scenario under normal light conditions in time series to obtain a stacked sequence. For all the frames of the low-light event stream of each scenario, perform the same operation. In the forward process, concatenate the stacked sequence of the event stream under normal light conditions with the stacked sequence of the corresponding low-light event stream along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0097] Step 1 is specifically implemented according to the following steps:
[0098] Step 1.1: Respectively give a stacked sequence of a normal-light event stream and the stacked sequence of the corresponding low-light event stream Concatenate them along the channel dimension to obtain a sequence
[0099] Step 1.2: At each step t from 0 to T, add noise ε to the data t , and this process is expressed as:
[0100]
[0101] where is the sequence after adding noise ε to the data from the previous step t - 1 at the current step t; β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix; t is the sequence after adding noise ε to the data
[0102] Step 1.3: Gradually add noise, and the distribution of the data gradually approaches the standard Gaussian distribution N(0, Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, and a series of noise-polluted data of the video frames of the normal-illuminance event stream is obtained
[0103] At each step of the forward process, a certain amount of noise is accumulated. This cumulative effect makes the original data gradually lose its structure and information and is finally completely covered by noise. The result generated by the forward process is expressed as:
[0104]
[0105] where is the sequence after adding noise ε to the data Add noise ε t The resulting sequence; is a parameter related to β t used to represent the degree of noise accumulation, and I is the identity matrix.
[0106] Step 2: Design a noise estimation network for the forward diffusion process. Gradually add Gaussian noise to the Y estimation of the concatenated sequence along the channel dimension of the stacked sequence of the event stream under normal light conditions and the corresponding stacked sequence of the low-light event stream at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can process temporal and spatial information in video data.
[0107] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at the t-th step, and the second part is the stacked sequence of the low-light event stream. They will be used as the conditional input of the reverse diffusion model. Concatenate these two parts along the channel dimension as the conditional input of the above noise estimation network, fit the current posterior mean and variance, and finally calculate and output the denoised image based on the mean, log variance, and random noise predicted by this noise estimation network.
[0108] Example 4
[0109] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 and is specifically implemented according to the following steps:
[0110] First, preprocess the dataset. Manually stack the event streams under normal light conditions in time series for every 5 consecutive video frames, and perform the same operation for the low-light event stream; in the forward process, add noise to the stacked sequence of the event stream under normal light conditions in turn; in the reverse process, perform denoising operations on the stacked sequence of the low-light event stream in turn, and finally output the denoised image.
[0111] Specifically, it is implemented according to the following steps:
[0112] Step 1: Use all the frames of the event streams in different scenarios under normal light conditions and all the frames of the corresponding low-light event stream in each scenario as the dataset. First, preprocess the dataset by stacking the event streams in each scenario under normal light conditions in time series for every 5 consecutive video frames to obtain a stacked sequence; perform the same operation for all the frames of the low-light event stream in each scenario. In the forward process, concatenate the stacked sequence of the event stream under normal light conditions and the corresponding stacked sequence of the low-light event stream along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0113] Step 1 is specifically implemented according to the following steps:
[0114] Step 1.1: Respectively give a stacked sequence of normal light event streams and the corresponding stacked sequence of low-light event streams Concatenate them along the channel dimension to obtain a sequence
[0115] Step 1.2: At each step t from 0 to T, add noise ε to the data t , and this process is expressed as:
[0116]
[0117] where, is the data at the current step t for the previous step t - 1 after adding noise ε t ; β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix;
[0118] Step 1.3: Gradually add noise, and the distribution of the data gradually approaches the standard Gaussian distribution N(0, Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, and a series of noise-polluted data of the video frames of the normal illumination event stream is obtained
[0119] At each step of the forward process, a certain amount of noise is accumulated. This cumulative effect causes the original data to gradually lose its structure and information and is finally completely covered by noise. The result generated by the forward process is expressed as:
[0120]
[0121] where, is the data at the current step t for the previous step t - 1 after adding noise ε t ; is a parameter related to β t used to represent the degree of noise accumulation, and I is the identity matrix.
[0122] Step 2: Design a noise estimation network for the forward diffusion process, and gradually add Gaussian noise to the Y estimation of the sequence obtained by concatenating the stacked sequence of event streams under normal light conditions and the corresponding stacked sequence of low-light event streams along the channel dimension at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can process the temporal and spatial information in video data.
[0123] Step 2 is specifically implemented according to the following steps:
[0124] Step 2.1: Adopt a 3D U-Net architecture as the noise estimation network;
[0125] Step 2.2: Use the concatenated sequence containing noise in the t-th step as the input of the noise estimation network and output the predicted noise, denoted as
[0126] Step 2.3: Starting from the training objective of the noise estimation network, design the loss function in the training process of the noise estimation network to make the noise prediction output by the network as close as possible to the actually added noise.
[0127] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at the t-th step, and the second part is the stacked sequence of low-light event streams. They will be used as the conditional input of the reverse diffusion model. These two parts are concatenated along the channel dimension as the conditional input of the above noise estimation network to fit the current posterior mean and variance. According to the mean, log variance, and random noise predicted by this noise estimation network, the denoised image is finally calculated and output.
[0128] Example 5
[0129] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 , and is specifically implemented according to the following steps:
[0130] First, preprocess the dataset. Artificially stack the event streams under normal light conditions in time series for every 5 consecutive video frames. The same operation is performed for the low-light event streams. In the forward process, noise is added to the stacked sequence of each event stream under normal light conditions in turn. In the reverse process, denoising operations are performed on the stacked sequence of each low-light event stream in turn, and finally the denoised image is output.
[0131] Specifically, it is implemented according to the following steps:
[0132] Step 1: Take all the frames of the event streams of different scenarios under normal light conditions and all the frames of the low-light event streams of each corresponding scenario as a dataset. First, preprocess the dataset. Stack every 5 consecutive video frames of the event stream of each scenario under normal light conditions in time series to obtain a stacked sequence. For all the frames of the low-light event stream of each scenario, perform the same operation. In the forward process, concatenate the stacked sequence of the event stream under normal light conditions with the stacked sequence of the corresponding low-light event stream along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0133] Step 1 is specifically implemented according to the following steps:
[0134] Step 1.1: Given a stacked sequence of the event stream under normal light and the stacked sequence of the corresponding low-light event stream concatenate them along the channel dimension to obtain a sequence
[0135] Step 1.2: Add noise ε to the data at each step t from 0 to T. t This process is expressed as:
[0136]
[0137] where is the sequence after adding noise ε to the data at the previous step t - 1 at the current step t; β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix. t
[0138] Step 1.3: Gradually add noise, and the distribution of the data gradually approaches the standard Gaussian distribution N(0, Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, obtaining a series of noise-polluted data of the video frames of the normal illumination event stream
[0139] At each step of the forward process, a certain amount of noise is accumulated. This cumulative effect causes the original data to gradually lose its structure and information and finally be completely covered by noise. The result generated by the forward process is expressed as:
[0140]
[0141] where is the data Add noise ε t The sequence after that; is a parameter related to β t used to represent the degree of noise accumulation, and I is the identity matrix.
[0142] Step 2: Design a noise estimation network for the forward diffusion process, and gradually add Gaussian noise to the Y estimation of the stacked sequence of event streams under normal light conditions and the stacked sequence of corresponding low-light event streams concatenated along the channel dimension at multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can process temporal and spatial information in video data.
[0143] Step 2 is specifically implemented according to the following steps:
[0144] Step 2.1: Adopt a 3D U-Net architecture as the noise estimation network;
[0145] Step 2.2: Use the concatenated sequence containing noise in the t-step process as the input of the noise estimation network and output the predicted noise, denoted as
[0146] Step 2.3: Starting from the training objective of the noise estimation network, make the noise prediction output by the network as close as possible to the actually added noise, and design the loss function in the training process of the noise estimation network.
[0147] The loss function in the training process of the noise estimation network in Step 2.3 is expressed as:
[0148]
[0149] Use ||.|| 2 as the loss function to measure the difference between the predicted noise and the true noise.
[0150] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at the t-th step, and the second part is the stacked sequence of low-light event streams. They will be used as the conditional input of the reverse diffusion model. Concatenate these two parts along the channel dimension as the conditional input of the above noise estimation network, fit the current posterior mean and variance, and finally calculate and output the denoised image according to the mean, log variance, and random noise predicted by this noise estimation network.
[0151] Example 6
[0152] The low-light event stream enhancement method based on the conditional diffusion model of the present invention combines Figure 1 , Figure 2 , and is specifically implemented according to the following steps:
[0153] First, preprocess the dataset. Artificially stack the event streams under normal light conditions in chronological order for every 5 consecutive video frames. For the low-light event streams, perform the same operation. During the forward process, sequentially add noise to the stacked sequences of each event stream under normal light conditions. During the reverse process, sequentially perform denoising operations on the stacked sequences of each low-light event stream, and finally output the denoised images.
[0154] Specifically, it is implemented according to the following steps:
[0155] Step 1: Use all the frames of the event streams in different scenarios under normal light conditions and all the frames of the corresponding low-light event streams in each scenario as the dataset. First, preprocess the dataset by stacking the event streams in each scenario under normal light conditions in chronological order for every 5 consecutive video frames to obtain stacked sequences. For all the frames of the low-light event streams in each scenario, perform the same operation. During the forward process, concatenate the stacked sequences of the event streams under normal light conditions and the corresponding stacked sequences of the low-light event streams along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each concatenated sequence at multiple time steps, making the image gradually become blurred and finally approaching the distribution of Gaussian noise.
[0156] Step 1 is specifically implemented according to the following steps:
[0157] Step 1.1: Respectively give a stacked sequence of a normal-light event stream and the stacked sequence of the corresponding low-light event stream Concatenate them along the channel dimension to obtain a sequence
[0158] Step 1.2: At each step t from 0 to T, add noise ε to the data t , and this process is expressed as:
[0159]
[0160] where is the sequence after adding noise ε to the data from the previous step t - 1 at the current step t; β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix; t is the scale parameter of the noise, set to a small positive number; I is the identity matrix;
[0161] Step 1.3: Gradually add noise, and the distribution of the data gradually approaches the standard Gaussian distribution N(0,Ι). When t reaches the maximum time step T, the data is approximated as the standard Gaussian distribution, obtaining a series of noise-polluted data of the video frames of the normal-illumination event stream.
[0162] At each step of the forward process, a certain amount of noise is accumulated. This cumulative effect causes the original data to gradually lose its structure and information and eventually be completely covered by noise. The result generated by the forward process is expressed as:
[0163]
[0164] where, is the data at the previous step \(t - 1\) at the current step \(t\) with noise \(\epsilon\) added t to form the sequence; is a parameter related to \(\beta\) t used to represent the degree of noise accumulation, and \(I\) is the identity matrix.
[0165] Step 2: Design a noise estimation network for the forward diffusion process. At multiple time steps, Gaussian noise added to the \(Y\) estimate of the concatenated sequence along the channel dimension of the stacked sequence of event streams under normal light conditions and the corresponding stacked sequence of low-light event streams is gradually estimated. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure and can process temporal and spatial information in video data.
[0166] Step 3: In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at step \(t\), and the second part is the stacked sequence of low-light event streams. They will be used as conditional inputs for the reverse diffusion model. These two parts are concatenated along the channel dimension as the conditional input for the above-mentioned noise estimation network to fit the current posterior mean and variance. Based on the mean, log variance, and random noise predicted by this noise estimation network, the denoised image is finally calculated and output.
[0167] Step 3 is specifically implemented according to the following steps:
[0168] In the reverse process, the above-trained noise estimation network is adopted, and the \(Y\) of the stacked sequence of every 5 video frames of the low-light event stream is used as the conditional input to gradually remove the noise.
[0169] Step 3.1: Estimate and remove the noise added in the forward process;
[0170] Step 3.1 is specifically implemented according to the following steps:
[0171] At each time step \(t\), the diffusion model proposed in this paper will learn to estimate the distribution of the original data from the noisy data . This estimation process is expressed as:
[0172]
[0173] Among them, μ θ and ∑ θ are the mean and variance of model parameterization;
[0174] The diffusion model gradually recovers from time step T to time step 0 by gradually denoising, that is:
[0175]
[0176] Step 3.2, after all time steps of the reverse process are completed, the finally denoised image is obtained.
[0177] Step 3.2 is specifically implemented according to the following steps:
[0178] At each step, by sampling, using the mean and variance estimated in the previous step, from generate This process starts from the initial noise and gradually recovers to by reverse diffusion steps, that is, new data samples are generated. This process reflects the path of the data gradually recovering from the Gaussian noise distribution to the original data distribution. When all time steps of the reverse process are completed, the finally obtained is the finally denoised image;
[0179] Each step of the reverse process is described in the following form:
[0180]
[0181] Among them, α t and β t are the noise parameters in the forward process, is the cumulative noise parameter, ε t is the noise estimated by the model, σ t is the standard deviation of each step, used to control the denoising intensity, Z t is the additional noise term to ensure the randomness of the generation process. When predicting the noise ε t at the t-th step, the method is the same as that described in step 2.2.
Claims
1. A low-light event flow enhancement method based on a conditional diffusion model, characterized in that: Follow the steps below to implement it: First, the data set is preprocessed by stacking every five consecutive video frames of the event stream under normal light conditions in time sequence. The same operation is performed for the low-light event stream. In the forward process, noise is added to the stacked sequence of each event stream under normal light conditions in turn. In the reverse process, the denoising operation is performed on the stacked sequence of each low-light event stream in turn, and the denoised image is finally output.
2. The low-light event flow enhancement method based on the conditional diffusion model according to claim 1 is characterized in that: Follow the steps below to implement it: Step 1: All frames of event streams of different scenes under normal light conditions and all frames of low-light event streams of each corresponding scene are taken as data sets, and the data sets are preprocessed first. The event streams of each scene under normal light conditions are stacked in time sequence for every 5 consecutive video frames to obtain a stacked sequence; the same operation is performed for all frames of the low-light event stream of each scene. In the forward process, the stacked sequence of the event stream under normal light conditions and the stacked sequence of the corresponding low-light event stream are spliced along the channel dimension to enhance the texture and details of each video frame. The joint diffusion model gradually adds Gaussian noise to each spliced sequence at multiple time steps, so that the image gradually becomes blurred and eventually approaches the distribution of Gaussian noise; Step 2: Design a noise estimation network for the forward diffusion process, and gradually estimate the Gaussian noise added to the Y of the sequence after the stacked sequence of the event stream under normal light conditions and the corresponding stacked sequence of the low-light event stream are spliced along the channel dimension in multiple time steps. This noise estimation network adopts a 3D U-Net architecture with an encoder-decoder structure, which can process the temporal and spatial information in the video data. Step 3. In the reverse reconstruction process, the input consists of two parts. The first part is the sequence containing noise at the tth step, and the second part is the stacked sequence of low-light event streams. They will be used as conditional inputs of the reverse diffusion model. These two parts are spliced along the channel dimension as conditional inputs of the above noise estimation network, and the current posterior mean and variance are fitted. According to the mean, logarithmic variance and random noise predicted by this noise estimation network, the denoised image is finally calculated and output.
3. The low-light event flow enhancement method based on the conditional diffusion model according to claim 2 is characterized in that: The step 1 is specifically implemented according to the following steps: Step 1.1: Given a stacking sequence of normal lighting event streams Stacked sequence with corresponding low light event stream Concatenate them along the channel dimension to get a sequence Step 1.2: Add noise ε to the data at each step t from 0 to T t , this process is expressed as: in, is the data from the current step t to the next step t-1 Add noise ε t The sequence after β t is the scale parameter of the noise, set to a small positive number; I is the identity matrix; Step 1.3: Add noise gradually, data The distribution of gradually approaches the standard Gaussian distribution N(0,Ι). When t reaches the maximum time step T, the data Approximated as a standard Gaussian distribution, a series of noise-contaminated data of normal illumination event stream video frames are obtained The result generated by the forward process is expressed as: in, is the data from the current step t to the next step t-1 Add noise ε t The sequence after is a t Related parameters are used to indicate the degree of noise accumulation, and I is the identity matrix.
4. The low-light event flow enhancement method based on the conditional diffusion model according to claim 2, characterized in that: The step 2 is specifically implemented according to the following steps: Step 2.1, use 3D U-Net architecture as the noise estimation network; Step 2.2: Use the spliced sequence containing noise in the t-step process As the input of the noise estimation network, it outputs the predicted noise, expressed as Step 2.3: Based on the training objective of the noise estimation network, design the loss function during the training process of the noise estimation network.
5. The low-light event flow enhancement method based on the conditional diffusion model according to claim 4 is characterized in that: The loss function in the training process of the noise estimation network in step 2.3 is expressed as: Use ||.||2 as the loss function to measure the difference between the predicted noise and the true noise.
6. The low-light event flow enhancement method based on the conditional diffusion model according to claim 5, characterized in that: The step 3 is specifically implemented according to the following steps: Step 3.1, estimate and remove the noise added in the forward process; Step 3.2: After the reverse process completes all time steps, the denoised image is finally obtained.
7. The low-light event flow enhancement method based on the conditional diffusion model according to claim 6, characterized in that: The step 3.1 is specifically implemented according to the following steps: At each time step t, the diffusion model proposed in this paper will learn from the noisy data Estimated original data The distribution of , this estimation process is expressed as: Among them, μ θ and∑ θ are the mean and variance of the model parameterization; The diffusion model is gradually restored from time step T to time step 0 by stepwise denoising, that is:
8. The low-light event flow enhancement method based on the conditional diffusion model according to claim 7, characterized in that: The step 3.2 is specifically implemented according to the following steps: Each step uses sampling to obtain the mean and variance estimated in the previous step. generate This process starts with the initial noise Start, and gradually recover to That is, generate new data samples. When the reverse process completes all time steps, the final result is This is the final denoised image; Each step of the reverse process is described in the following form: Among them, α t and β t is the noise parameter in the forward process, is the cumulative noise parameter, ε t is the noise of the model estimate, σ t is the standard deviation of each step, used to control the denoising strength, Z t is an additional noise term to ensure the randomness of the generation process and predict the noise ε at step t t The method is the same as step 2.2.
Citation Information
Cited By
Video processing method based on multi-condition control diffusion model and related equipment
CN121418573A