A standardized stream based event camera data simulation method and apparatus

By using a standardized flow-based event camera data simulation method, event camera data that closely approximates the target domain is adaptively generated, solving the problem of high simulation costs for event cameras and achieving efficient and accurate simulation generation results.

CN116309356BActive Publication Date: 2026-05-29BEIHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2023-02-02
Publication Date
2026-05-29

Smart Images

  • Figure CN116309356B_ABST
    Figure CN116309356B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a standardized flow-based event camera data simulation method and device. A specific implementation of the method comprises: decoding video and event information in to-be-processed event camera data; inputting an event accumulation graph and corresponding front and rear video frames into a scene brightness coding network; inputting obtained scene brightness coding feature sets and the event accumulation graph into a standardized flow network; inputting the obtained scene brightness coding feature sets and event camera threshold sets and event camera noise sets of all pixels in target event camera data into the standardized flow network to reversely generate event data; performing maximum likelihood learning on the reversely generated event data and pre-acquired event camera data distribution; and generating simulated event camera data close to a target domain event camera. The implementation avoids too large distribution differences between simulation domain event camera data and target domain event camera data in different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the fields of computer vision and multimodal generation, specifically to a method and apparatus for simulating event camera data based on standardized streams. Background Technology

[0002] Event cameras, due to their imaging principle, offer advantages over traditional color cameras such as high dynamic range and low time delay. They have broad application prospects in fields such as national defense, film and television production, and public safety. However, event cameras are currently expensive and not yet widely available. Utilizing event data simulation algorithms to quickly and inexpensively generate large amounts of event camera data is of great significance in applications such as image restoration, video surveillance, and smart cities.

[0003] Normalized flow is a generative model used to describe the variations between different data distributions. It possesses advantages such as strong theoretical foundation and lossless, reversible information representation. It has great application potential in fields such as computer vision and theoretical machine learning. Conditional normalized flow models, due to their more generalized input conditions, have been developed and used in many fields in recent years. Utilizing normalized flow to describe the distribution of event data, event thresholds, and noise distribution is of great significance to the development of event camera applications. Summary of the Invention

[0004] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of this disclosure propose a method and apparatus for simulating event camera data based on standardized streams to address one or more of the technical problems mentioned in the background section above.

[0006] Based on the aforementioned practical needs and key issues, the purpose of this disclosure is to propose an event camera data simulation method based on standardized streams. The method takes as input the event camera data to be processed in the target domain, including video frame sequences and event cumulative graph sequences, and adaptively generates the target domain event camera threshold and noise set pixel by pixel. Combined with a pre-learned event camera data generation model, the method outputs simulated event camera data.

[0007] Some embodiments of this disclosure provide a method for simulating event camera data based on a normalized stream. The method includes: decoding video and event information in the event camera data to be processed to obtain a video frame sequence and an event accumulation map sequence; for each event accumulation map in the event accumulation map sequence, performing the following processing steps: inputting the event accumulation map and corresponding preceding and following video frames in the video frame sequence into a scene brightness coding network to obtain a scene brightness coding feature set for all pixels; inputting the obtained scene brightness coding feature set and the event accumulation map into a normalized stream network to obtain an event camera threshold set and event information set for all pixels in the target event camera data. Camera noise set; leveraging the reversibility of the normalized flow network, the obtained scene brightness encoding feature set and the event camera threshold set and event camera noise set of all pixels in the target event camera data are input into the normalized flow network to generate event data in reverse; maximum likelihood learning is performed using the above event data and the pre-acquired event camera data distribution to update the above event camera threshold set and event camera noise set, resulting in updated event camera threshold set and event camera noise set; based on the above updated event camera threshold set and event camera noise set and video frame sequence, simulated event camera data that approximates the target domain event camera is generated.

[0008] Some embodiments of this disclosure provide an event camera data simulation apparatus based on a normalized stream. The apparatus includes: a decoding unit configured to decode video and event information in the event camera data to be processed, obtaining a video frame sequence and an event accumulation map sequence; and a processing unit configured to perform the following processing steps for each event accumulation map in the event accumulation map sequence: inputting the event accumulation map and corresponding preceding and following video frames in the video frame sequence into a scene luminance coding network to obtain a scene luminance coding feature set for all pixels; and inputting the obtained scene luminance coding feature set and the event accumulation map into a normalized stream network to obtain the events of all pixels in the target event camera data. Camera threshold set and event camera noise set; leveraging the invertibility of the normalized flow network, the obtained scene brightness encoding feature set and the event camera threshold set and event camera noise set of all pixels in the target event camera data are input into the normalized flow network to generate event data in reverse; maximum likelihood learning is performed using the above event data and the pre-acquired event camera data distribution to update the above event camera threshold set and event camera noise set, resulting in updated event camera threshold set and event camera noise set; based on the above updated event camera threshold set and event camera noise set and video frame sequence, simulated event camera data that approximates the target domain event camera is generated.

[0009] The method disclosed herein utilizes color camera data for event camera data simulation generation. Compared to existing simulation methods that rely excessively on manually adjusted simulation parameters, it offers the following advantages: 1) It eliminates the need for manually set simulated event camera thresholds and noise; realistic simulated event camera thresholds and noise can be obtained using only a small amount of real event camera data. 2) It eliminates the need for separate modeling for various existing event cameras and different shooting scenarios; a unified model can accurately estimate the target domain event camera data. 3) It eliminates the need for complex parameter adjustments and algorithm simulations; no additional parameters or computational load are added during the event data generation process. Reliable simulation of event camera data is achieved through threshold and noise estimation, and adaptive learning is performed using information from the real event data itself, avoiding excessive differences in the distribution of simulated domain event camera data and target domain event camera data in different scenarios. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a flowchart of some embodiments of the event camera data simulation method based on standardized streams according to the present disclosure;

[0012] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the event camera data simulation apparatus based on standardized streams according to this disclosure:

[0013] Figure 3 This is a schematic diagram of an application scenario of an event camera data simulation method based on standardized streams according to some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 Flowchart 100 illustrates some embodiments of the event camera data simulation method based on normalized streams disclosed herein:

[0021] Step 101: Decode the video and event information in the camera data of the event to be processed to obtain the video frame sequence and the event cumulative map sequence.

[0022] In some embodiments, the entity executing the event camera data can decode the acquired event camera data to obtain a video frame sequence and an event sequence. The event accumulation map sequence can be derived from the event sequence. The event accumulation map in the above event accumulation map sequence corresponds to two video frames; that is, each event accumulation map is obtained by accumulating all events between the corresponding video frame and the next video frame. The width and height of the event accumulation map are the same as the video frame, and the depth of the event accumulation map is 16-dimensional, representing that the time difference between two adjacent video frames is divided into 16 equal parts.

[0023] An event camera is a novel type of visual sensor. Unlike traditional color cameras that capture images by recording the brightness received by the sensor within a certain exposure time, an event camera records changes in brightness. When the brightness change of a pixel exceeds the event camera's threshold, the event camera records the pixel coordinates, time, and polarity of the event (increased brightness is positive, decreased brightness is negative). The basic principle of simulating event camera data is to mimic the way a real event camera records data, comparing consecutive frames in a video sequence, and recording an event when the brightness change of a corresponding pixel exceeds the threshold. Some works have attempted to generate event camera data using color camera data. For example, in 2018, Rebecq of ETH Zurich first proposed a reliable method for simulating event camera data. This method first proposed that the threshold of the event camera is not fixed but contains noise, and used a Gaussian distribution to fit this noise. It also proposed that the time of the simulated event should be distributed between simulated video frame sequences, rather than the time of occurrence of video frames. In 2020, Hu of ETH Zurich proposed a more refined method for simulating event camera data. This method points out that in low-light conditions, some events may not be recorded by event cameras due to the limited optical sensing bandwidth. A low-pass filter is designed to simulate this bandwidth limitation. Furthermore, the method proposes a Poisson-based event camera temporal noise model and an event camera leakage noise model, further enhancing the reliability of the simulated event data.

[0024] Step 102: For each event accumulation graph in the above event accumulation graph sequence, perform the following processing steps:

[0025] Step 1021: Input the event accumulation map and the corresponding preceding and following video frames in the video frame sequence into the scene brightness coding network to obtain the scene brightness coding feature set of all pixels.

[0026] In some embodiments, the execution entity can input the event accumulation map and corresponding preceding and following video frames from the video frame sequence into a scene luminance coding network to obtain a scene luminance coding feature set for all pixels. The scene luminance coding network consists of a scene luminance encoder and a scene luminance decoder, which are connected by short connections. This scene luminance coding network performs feature encoding on each pixel in the video frame and the event accumulation map, ultimately obtaining a scene luminance coding feature set for all pixels.

[0027] Step 1022: Input the obtained scene brightness coding feature set and event accumulation map into the normalized flow network to obtain the event camera threshold set and event camera noise set of all pixels in the target event camera data.

[0028] In some embodiments, the execution entity can input the obtained scene brightness encoding feature set and event accumulation map into a normalized flow network to obtain the event camera threshold set and event camera noise set for all pixels in the target event camera data. The normalized flow network consists of four normalized flow modules of different depths, connected by a depth multiplication module. For each pixel p in the event accumulation map, the depth multiplication module samples a noise vector of the same dimension as the feature vector output by the normalized flow module from a preset standard normal distribution N = (0, 1), and concatenates the feature vector output by the normalized flow module with the sampled noise vector to achieve feature vector depth multiplication.

[0029] Each standard stream module consists of six identical conditional reversible modules connected by short connections. Each conditional reversible module includes: for each pixel p in the event accumulation graph, dividing the feature vector h into two equal vectors h''. A and h B Then, the scene brightness encoding feature u corresponding to pixel p is used as the condition for this reversible module, and a feature fusion network f and formula h are used. B '=h B +f(h A ;u) Update h B ', and then update h B 'and the original h A The feature vector h' is reassembled and updated. Then, for the feature vector h', each feature value of the feature vector h' is reordered and updated in a predefined order W to obtain the new feature vector h. 2 This process is reversible, meaning that h can be generated from h. 2 Conversely, according to h 2 h can be obtained through reverse reasoning.

[0030] For obtaining the event camera threshold set and event camera noise set of all pixels in the target event camera data, the event camera threshold set includes the mean of the threshold distribution of each pixel of the event camera, and the event camera noise set includes the probability value of noise occurrence of each pixel of the event camera.

[0031] Step 1023: Utilizing the reversibility of the normalized flow network, the obtained scene brightness encoding feature set, the event camera threshold set of all pixels in the target event camera data, and the event camera noise set are input into the normalized flow network to generate event data in reverse.

[0032] In some embodiments, the aforementioned executing entity can reverse-engineer event data. For example... Figure 3As shown, the scene brightness encoding feature set obtained above is also needed in the reverse generation process. Furthermore, by inputting the event camera threshold set and event camera noise set of all pixels in the target event camera data into a normalized flow network, event data can be generated in reverse.

[0033] Step 1024: Perform maximum likelihood learning using the event data and the pre-acquired event camera data distribution to update the event camera threshold set and the event camera noise set, resulting in the updated event camera threshold set and event camera noise set.

[0034] In some embodiments, the execution entity may perform maximum likelihood learning on the reverse-generated event data and the pre-acquired event camera data distribution to update the event camera threshold set and the event camera noise set, thereby obtaining updated event camera threshold set and event camera noise set. Here, the event data can be simulated event camera data.

[0035] First, the simulated event camera data and the pre-acquired event camera data are processed to obtain event representation information. This event representation information is then segmented at multiple scales to obtain multi-scale segmented event representation information. The segmentation process includes the following steps: dividing the event representation information into blocks to obtain multiple event representation blocks; performing feature analysis on each of these event representation blocks; and updating the event camera threshold set and noise set based on the different feature analysis results. The feature in the feature analysis is gradient strength.

[0036] Then, feature analysis is performed on each event representation block in the multiple event representation blocks. This includes: an event representation block consists of an event cumulative graph and corresponding video frames. The event representation blocks are divided into 4 blocks of 2×2 each using grid partitioning, and then further divided into 16 blocks of 4×4 each, resulting in a total of 21 event representation blocks of multiple scales (1+4+16). Then, feature gradient intensity analysis is performed on each of the obtained multi-scale event representation blocks. The specific analysis process includes: calculating the horizontal and vertical gradients of the video frame portion in each feature block to obtain the gradient matrix of each feature block; next, determining the eigenvalues ​​λ1 and λ2 of each feature block; and calculating the gradient intensity of each feature block using the following formula:

[0037] Finally, based on the different feature analysis results, different weights are assigned to different event representation blocks, and the above-mentioned event camera threshold set and noise set are updated according to the pre-acquired data distribution, including updating the above-mentioned event camera threshold set and noise set using the following formula:

[0038]

[0039] Where m represents a feature block. M represents the set of feature blocks. γ is the gradient strength. represents the weights of the two learning methods. This represents the mean absolute value error loss function. This indicates that the loss function is estimated using maximum likelihood. σ represents the ELU activation function. ∑ represents the summation of the loss function over all feature blocks. This represents the overall loss function.

[0040] Step 1025: Based on the updated event camera threshold set, event camera noise set, and video frame sequence, generate simulated event camera data that approximates the target domain event camera.

[0041] In some embodiments, the aforementioned execution entity may generate simulated event camera data approximating the target domain event camera based on the updated event camera threshold set, event camera noise set, and video frame sequence, including:

[0042] Based on the updated event camera threshold set, each video frame in the above video frame sequence is input into the pre-trained event camera simulation generation model to generate event camera data.

[0043] Based on the aforementioned event camera data and the updated event camera noise set, simulated event camera data is generated. This simulated event camera data is data whose distribution similarity to real event camera data exceeds a predetermined threshold.

[0044] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an event camera data simulation device based on a standardized stream. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this standardized stream-based event camera data simulation device can be specifically applied to various electronic devices.

[0045] like Figure 2As shown, an event camera data simulation device 200 based on a normalized stream in some embodiments includes a decoding unit 201 and a processing unit 202. The decoding unit 201 is configured to decode video and event information in the event camera data to be processed, obtaining a video frame sequence and an event accumulation map sequence. The processing unit 202 is configured to perform the following processing steps for each event accumulation map in the event accumulation map sequence: inputting the event accumulation map and corresponding preceding and following video frames in the video frame sequence into a scene brightness coding network to obtain a scene brightness coding feature set for all pixels; inputting the obtained scene brightness coding feature set and the event accumulation map into a normalized stream network to obtain an event camera threshold set and an event camera noise set for all pixels in the target event camera data. Leveraging the reversibility of the standardized flow network, the obtained scene brightness encoding feature set, the event camera threshold set, and the event camera noise set of all pixels in the target event camera data are input into the standardized flow network to generate event data in reverse. Maximum likelihood learning is then performed using the event data and the pre-acquired event camera data distribution to update the event camera threshold set and event camera noise set, resulting in updated event camera threshold sets and event camera noise sets. Based on the updated event camera threshold set and event camera noise set, and the video frame sequence, simulated event camera data that approximates the target domain event camera is generated.

[0046] It is understandable that the units and references described in the event camera data simulation device 200 based on standardized streams... Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the event camera data simulation device 200 based on standardized streams and the units contained therein, and will not be repeated here.

[0047] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for simulating event camera data based on normalized streams, comprising: The video and event information in the camera data of the event to be processed are decoded to obtain the video frame sequence and the event cumulative map sequence; For each event accumulation graph in the event accumulation graph sequence, the following processing steps are performed: The event accumulation map and the corresponding preceding and following video frames in the video frame sequence are input into the scene brightness coding network to obtain the scene brightness coding feature set of all pixels; The obtained scene brightness encoding feature set and event accumulation map are input into a normalized flow network to obtain the event camera threshold set and event camera noise set for all pixels in the target event camera data. The normalized flow network consists of four normalized flow modules of different depths, which are connected by a depth doubling module. This depth doubling module performs a depth multiplication on each pixel in the event accumulation map. From the preset standard normal distribution The sampling and normalized stream modules output a noise vector of the same dimension as the feature vector. The feature vector output by the normalized stream module is concatenated with the sampled noise vector. Each normalized stream module consists of six identical conditional invertible modules, which are connected by short connections. By leveraging the reversibility of the normalized flow network, the obtained scene brightness coding feature set, the event camera threshold set of all pixels in the target event camera data, and the event camera noise set are input into the normalized flow network to generate event data in reverse. Maximum likelihood learning is performed using the event data and the pre-acquired event camera data distribution to update the event camera threshold set and the event camera noise set, resulting in updated event camera threshold set and event camera noise set. Based on the updated event camera threshold set, event camera noise set, and video frame sequence, simulated event camera data that approximates the target domain event camera is generated.

2. The method according to claim 1, wherein, The event accumulation graph sequence corresponds to two video frames. Each event accumulation graph is obtained by accumulating all events between the corresponding video frame and the next video frame. The width and height of the event accumulation graph are the same as those of the video frames. The depth of the event accumulation graph is 16-dimensional, which means that the time difference between two adjacent video frames is divided into 16 parts.

3. The method according to claim 2, wherein, The scene luminance coding network consists of a scene luminance encoder and a scene luminance decoder, which are connected by short connections. The scene luminance coding network performs feature encoding on each pixel in the video frame and the event accumulation map, and finally obtains the scene luminance coding feature set of all pixels.

4. The method according to claim 3, wherein, The event camera threshold set includes the mean of the threshold distribution for each pixel of the event camera; the event camera noise set includes the probability value of noise occurring for each pixel of the event camera.

5. The method according to claim 4, wherein, Each standard stream module consists of six identical conditional reversible modules, wherein the conditional reversible modules include: For each pixel in the event accumulation map eigenvectors The feature vector is divided into two equal vectors. and Then the pixels Corresponding scene brightness encoding features As a condition for this reversible module, a feature fusion network is used. and formula renew Then the updated And the original Reassemble and update the feature vector Then, for the feature vector , to feature vector Each feature value is in a pre-defined order. The feature vector is obtained by reordering and updating. This process is reversible, that is, from Can generate Conversely, according to Reverse reasoning yields .

6. The method according to claim 4, wherein, The four normalized flow modules are connected by a depth multiplication module, which includes: For each pixel in the event accumulation graph From the preset standard normal distribution The noise vector with the same dimension as the feature vector output by the sampling and normalization stream modules is concatenated with the feature vector output by the normalization stream module to achieve the purpose of doubling the feature vector depth.

7. The method according to claim 4, wherein maximum likelihood learning is performed using the event data and the pre-acquired event camera data distribution, comprising: The event data and the pre-acquired event camera data are processed to obtain event characterization information; The event representation information is segmented at multiple scales to obtain multi-scale segmented event representation information. The segmentation process includes the following steps: dividing the event representation information into blocks to obtain multiple event representation blocks; performing feature analysis on each of the multiple event representation blocks; setting different weights for different event representation blocks based on different feature analysis results; and updating the event camera threshold set and noise set according to the pre-acquired data distribution to obtain updated event camera threshold set and noise set. The feature in the feature analysis is gradient strength.

8. The method according to claim 7, wherein performing feature analysis on each of the plurality of event representation blocks comprises: An event representation block consists of an event accumulation graph and the corresponding video frame; The event representation block is divided into 4 blocks of 2×2 by grid partitioning, and then further divided into 16 blocks of 4×4, resulting in a total of 21 event representation blocks of multiple scales (1+4+16). Then, feature gradient intensity analysis is performed on the obtained multi-scale event representation blocks. The specific analysis process includes: calculating the horizontal and vertical gradients of the video frame portion in each feature block to obtain the gradient matrix of each feature block. Next, the feature values ​​of each feature block are determined. , The gradient intensity of each feature block is calculated using the following formula: .

9. The method according to claim 7, wherein different weights are set for different event representation blocks based on different feature analysis results, and the event camera threshold set and noise set are updated according to the pre-acquired data distribution, comprising: The event camera threshold set and noise set are updated using the following formula: , in, Represents a feature block. Represents a set of feature blocks. The gradient strength represents the weights of the two learning methods. This represents the mean absolute value error loss function. This indicates that the loss function is estimated using maximum likelihood estimation. This represents the ELU activation function. This represents summing the loss function over all feature blocks. This represents the overall loss function.

10. The method according to claim 1, wherein, The step of generating simulated event camera data that approximates the target domain event camera based on the updated event camera threshold set, event camera noise set, and video frame sequence includes: Based on the updated event camera threshold set, each video frame in the video frame sequence is input into a pre-trained event camera simulation generation model to generate event camera data. Based on the event camera data and the updated event camera noise set, simulated event camera data is generated, wherein the simulated event camera data is data whose similarity to the distribution of real event camera data is greater than a predetermined threshold.

11. An event camera data simulation device based on standardized streams, comprising: The decoding unit is configured to decode the video and event information in the camera data of the event to be processed, and obtain the video frame sequence and the event cumulative map sequence. The processing unit is configured to perform the following processing steps for each event accumulation map in the event accumulation map sequence: inputting the event accumulation map and the corresponding preceding and following video frames in the video frame sequence into the scene brightness coding network to obtain a scene brightness coding feature set for all pixels; The obtained scene brightness encoding feature set and event accumulation map are input into a normalized flow network to obtain the event camera threshold set and event camera noise set for all pixels in the target event camera data. The normalized flow network consists of four normalized flow modules of different depths, which are connected by a depth doubling module. This depth doubling module performs a depth multiplication on each pixel in the event accumulation map. From the preset standard normal distribution The sampling and normalization flow modules output a noise vector with the same dimension as the feature vector. The feature vector output by the normalization flow module is concatenated with the sampled noise vector. Each normalization flow module consists of six identical conditional invertible modules connected by short connections. Utilizing the invertibility of the normalization flow network, the obtained scene brightness encoding feature set, the event camera threshold set of all pixels in the target event camera data, and the event camera noise set are input into the normalization flow network to generate event data in reverse. Maximum likelihood learning is performed using the event data and the pre-acquired event camera data distribution to update the event camera threshold set and the event camera noise set, resulting in updated event camera threshold set and event camera noise set. Based on the updated event camera threshold set, event camera noise set, and video frame sequence, simulated event camera data that approximates the target domain event camera is generated.