A high-dimensional light field event camera based on a microlens array and an extraction method
By installing a microlens array in front of the event camera and constructing a multi-view triangle voting measurement method, combining the light field geometric characteristics and asynchronous response characteristics, the efficient multi-view high-dimensional light field event extraction of a single-lens light field event camera is achieved, solving the problems of high hardware costs and complex algorithms in the existing technology, and improving robustness and efficiency.
Patent Information
- Application Number
- CN202211219900.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-08
AI Technical Summary
In the prior art, the combination of event cameras and light field cameras has high hardware costs and complex synchronization methods. The asynchronous response of event cameras leads to complex feature matching algorithms in light field application scenarios and large delays, making it difficult to effectively extract multi-view high-dimensional light field events.
A single-lens combined with a microlens array is used to adopt a light field event camera structure, and by placing a microlens array in front of a dynamic vision sensor, the light field geometric characteristics and the asynchronous response characteristics of the event camera are used to construct an efficient multi-view high-dimensional light field event extraction method, and a diffusion weight mechanism is introduced to reduce the impact of noise and calibration errors.
It realizes efficient multi-view high-dimensional light field event acquisition and extraction based on single lens, reduces algorithm complexity, improves robustness, reduces delay and power consumption, and adapts to light field applications in cases of large noise and inaccurate calibration.
Smart Images

Figure CN115598744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computational vision, and particularly to a single-lens high-dimensional light field event camera based on a microlens array and a corresponding high-dimensional light field event extraction algorithm. Background Art
[0002] Traditional cameras usually acquire and display images in the form of frames. For example, the standard of traditional movies is to shoot and play at a standard of 24 frames per second. Such an acquisition form usually results in a large amount of redundant information between consecutive frames. In some application scenarios of computer vision algorithms such as anomaly detection and motion tracking, such redundancy often brings relatively large latency and additional computational power consumption. In addition, this synchronous acquisition method also causes the amount of data collected and the acquisition cost to increase with the increase in the frame rate, thus making the costs associated with high-frame-rate video shooting and analysis extremely high.
[0003] Different from the synchronous acquisition method of traditional cameras, in various biological vision systems in nature, an asynchronous acquisition method in which each visual nerve can independently generate a response is a more common visual mechanism. The dynamic vision sensor is exactly inspired by this asynchronous vision mechanism.
[0004] Compared with the method of synchronously acquiring each pixel at a fixed frame rate on a traditional camera, each pixel of an event camera based on a dynamic vision sensor can detect the brightness change and independently generate a response output, describing the trend of brightness change in the scene in an asynchronous manner. The brightness change responses detected by each pixel will be encoded into an event stream for output, encoding the position, time, and polarity of the brightness change. For example, the event (x, y, p, t) is a typical event description form, where x and y represent the two-dimensional coordinates of the current event on the dynamic vision sensor, p represents the polarity of the current event, and t represents the timestamp of the event. The event camera can provide a time resolution of the order of microseconds, a larger dynamic range, lower power consumption, and a higher pixel readout bandwidth. These characteristics make the event camera more advantageous in low-latency, high-speed, and high-dynamic range scenarios. Therefore, the event camera has great potential in the field of computer vision. At present, certain accumulations have been formed in the application research of event cameras in fields such as feature extraction and tracking, high-dynamic range image reconstruction, and motion detection, but there is still great potential to be explored.
[0005] The emergence of light field cameras has broken through the limitations of traditional cameras in describing positions using two-dimensional planar coordinates. At the cost of sacrificing some spatial resolution, they can capture the complete light radiation distribution in space. Light field imaging can provide the angular and spatial information of light rays, thus offering a larger depth of field compared to traditional cameras. The data captured by light field cameras can provide high-dimensional information such as the depth of the image and the refocusing results at different depths after subsequent processing such as phase transformation and projection integration. Early light field data acquisition methods had limitations such as large array volumes and complex synchronization control. However, handheld light field cameras based on a simplified version of the four-dimensional light ray function use microlens arrays to reduce the acquisition cost of light field images, expand the acquisition scenarios, and thus provide more abundant light field applications. At the same time, the data acquisition of light field cameras also poses higher requirements for the algorithms of subsequent processing. Applications such as three-dimensional reconstruction, motion velocity measurement, particle microscopy, and target recognition based on light field images are all popular research directions.
[0006] The development of event cameras and light field cameras has fully demonstrated that computer vision tasks require more powerful information acquisition methods to tap potential. With the development of software and hardware technologies, the production and usage costs of handheld light field cameras and event cameras are gradually decreasing. More and more computer vision algorithm applications use the shooting data of light field cameras or event cameras as input to obtain richer information than traditional imaging data. Based on this trend, designing a collection device that combines the advantages of handheld light field cameras and event cameras will be able to fully integrate the high-dimensionality of light field data and the high temporal resolution and high dynamic characteristics of event data, which is of great significance for fully tapping the potential of these two imaging devices and improving the algorithm performance of related vision tasks.
[0007] However, current related research that can provide multi-view high-dimensional outputs similar to light fields based on dynamic vision sensors requires the use of multiple event cameras to build, with high hardware costs and complex synchronization methods. The structure of a single lens combined with a microlens array in handheld light field cameras can provide some inspiration, but the characteristics of dynamic vision sensors in light field application scenarios still need to be overcome. Currently, the noise description model of event cameras is still in a relatively preliminary research and analysis stage, and the asynchronous response of event cameras also makes it impossible for light field event cameras to use the intra-frame feature matching algorithm of traditional light field cameras to match multi-view feature points. For these problems, more complex algorithm processing will introduce a large delay and eliminate it. Therefore, being able to construct a light field event camera based on a single lens combined with a microlens array and its high-dimensional light field event extraction algorithm has become a very important topic. Summary of the Invention
[0008] In view of the changes and characteristics of the above-mentioned prior art, the object of the present invention is to propose a light field event camera based on a single lens combined with a microlens array and a high-dimensional light field event matching and extraction method thereof. The present invention can combine the advantages of a light field camera and an event camera, and fully reduce the algorithm complexity, avoid introducing large delays and power consumption, and still has excellent robustness in the case of large noise and inaccurate calibration.
[0009] The present invention uses a microlens array placed in front of a dynamic vision sensor, and constructs an efficient algorithm based on the geometric characteristics of the light field and the asynchronous response characteristics of the event camera to extract multi-view high-dimensional light field events from the continuously output event stream.
[0010] The specific technical solution adopted by the camera of the present invention is as follows:
[0011] A high-dimensional light field event camera based on a microlens array, comprising a main lens, a microlens array and a dynamic vision sensor, wherein the microlens array is placed between the main lens and the dynamic vision sensor, and among them, the microlens array is located behind the imaging space of the main lens, and each microlens in the array can perform secondary imaging on the imaging of the main lens; a certain brightness change area in the target scene is projected onto the dynamic vision sensor after two optical imagings by the main lens and the microlens array, and different positions of the dynamic vision sensor can sense the brightness change of the same area, forming multiple asynchronous outputs.
[0012] The present invention also provides a method for extracting high-dimensional light field events using the above camera, and the method includes the following steps:
[0013] Step 1, assuming that the microlens array is N rows and M columns, a total of N*M sub-images of viewpoints, initialize N*M corresponding event buffers for subsequent event management;
[0014] Step 2, after reading out the output event e i (x i , y i , p i , t i ) of the dynamic vision sensor, store it in the buffer of the corresponding viewpoint according to its coordinates, where i represents the i-th event read out, x i , y i represents the two-dimensional coordinates of the current event on the dynamic vision sensor, p i represents the polarity of the current event, and t i represents the timestamp of the event;
[0015] Step 3, check the events that have been stored in the buffer. If the events in the buffer exceed the maximum queue length, clear the earliest stored events in the buffer according to the exceeded length;
[0016] Step 4: After determining that the current event enters the buffer, check whether it can activate the number of perspectives in the buffer multi-perspective event combination coverage within the defined domain that reaches the preset threshold. If the threshold limit condition can be met, proceed to Step 5; otherwise, proceed to Step 2.
[0017] Step 5: Take out all the events involved in the multi-perspective combination activated by the event that newly enters the buffer. In the world coordinate system constructed according to the application scenario, based on the center coordinates of the microlenses corresponding to the perspectives, construct the geometric space of the multi-perspective triangular voting measurement method, that is, connect the coordinates of each response pixel of the dynamic vision sensor with the corresponding microlens center to form multiple straight lines in space, and each straight line corresponds to a ray of light collected.
[0018] Step 6: Introduce a diffusion weight mechanism on each ray of light, that is, take each point on the straight line as the center and k unit lengths as the defined range, and assign weights to the points around the straight line within the range.
[0019] Step 7: Screen the point weights in the multi-perspective triangular voting measurement space according to the threshold. If there are spatial points exceeding the preset threshold, determine that the events corresponding to the relevant rays involved in the spatial points constitute a high-dimensional light field event, which can describe the brightness change in the real space from multiple perspectives.
[0020] Step 8: If the dynamic vision sensor still has output, return to Step 2 to process the next event; otherwise, end the processing.
[0021] Through two key points of placing a microlens array in front of the dynamic vision sensor and constructing a multi-perspective light field event triangular voting measurement method by using the asynchronous output characteristics, the present invention can realize the acquisition and extraction of multi-perspective high-dimensional light field events based on a single lens. In addition, in order to enhance the robustness of the method, the present invention method also introduces a diffusion weight mechanism, which is proven in experiments to be able to effectively reduce the influence of noise and calibration errors on the algorithm performance. The present invention not only innovatively constructs a brand-new data acquisition device, the light field event camera, but also constructs a corresponding light field event extraction method based on the geometric structure characteristics of the light field and the characteristics of the dynamic vision sensor, realizing a light field event camera hardware structure and data processing algorithm with strong practicability, wide application scenarios, and high robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow schematic diagram of the present invention;
[0023] Figure 2 It is a schematic diagram of a certain world coordinate system;
[0024] Figure 3 It is a schematic diagram of multi-perspective buffer status management;
[0025] Figure 4 Schematic diagram of the geometric space and diffusion weight for the multi-view triangular voting measurement method. Detailed implementation mode
[0026] This embodiment provides a single-lens high-dimensional light field event camera based on a microlens array. As shown in the optical structure in the appendix Figure 1 , it includes a main lens, a microlens array, and a dynamic vision sensor. Among them, the microlens array is placed between the main lens and the dynamic vision sensor, so as to realize the acquisition of high-dimensional light field events based on a single lens. At the same time, the microlens array can project a certain brightness change in the target scene onto different positions of the dynamic vision sensor simultaneously based on a single main lens to form an output. Specifically, the microlens array in this system is placed according to the principle of focused light field imaging, and can be divided into a Kepler-type or Galileo-type focused light field imaging system according to the positional relationship between the microlens array and the imaging of the main lens. Each microlens under this configuration forms an imaging result with an independent view on the dynamic vision sensor.
[0027] Using the above camera, this embodiment also provides a high-dimensional light field event extraction method based on a microlens array, including the following steps:
[0028] Step 1, initialize the event buffer storage space. Here, it is assumed that the microlens array can form sub-images with N rows and M columns, a total of N*M views. Then, N*M corresponding event buffers need to be initialized for subsequent event management. As shown in the appendix Figure 2 , if the microlens array is of 3*3 structure, then 3*3 buffers need to be initialized correspondingly.
[0029] Step 2, after reading out the output event e i (x i , y i , p i , t i ) of the dynamic vision sensor, the event queue manager stores it in the buffer corresponding to the view generated by the microlens according to its coordinates. Here, i represents the i-th event read out, and x i , y i represent the two-dimensional coordinates of the current event on the dynamic vision sensor, p i represents the polarity of the current event, and t i represents the timestamp of the event.
[0030] Step 3, the event queue manager checks the events that have entered the buffer for storage. If the events in the buffer exceed the maximum queue length, the earliest stored events in the buffer are cleared according to the exceeded length.
[0031] Step 4, after the event queue manager determines that the current event enters the buffer, it checks whether it can form multiple perspectives that cover a preset threshold with the events in the adjacent perspective buffer within a preset domain range. If a multi-perspective structure can be formed, proceed to Step 5; otherwise, proceed to Step 2. For example, if the camera parameters are set such that each point in the real space can be captured by 3*3 microlenses and generate a response on the dynamic vision sensor, the domain radius range can be set to 3. When the content is stored in more than a certain threshold number of buffer areas within a 3*3 buffer area, the cached events involved in this area are taken out and Step 5 is performed. As Figure 3 shown in the example, here it is assumed that the extraction threshold is 4. The dark color in the figure indicates that there is already event storage, which is in the active state and may contain events with at least 5 perspectives. Since it exceeds the threshold of 4, all the events involved in this area will be taken out and processed in Step 5.
[0032] Step 5, take out all the events involved in the multi-perspective combination activated by the event that newly enters the buffer. As Figure 2 shown, in the world coordinate system of the application scenario, as Figure 4 shown, based on the center coordinates of the microlenses corresponding to the perspectives, construct the geometric space of the multi-perspective triangular voting measurement method, that is, connect the coordinates of each response pixel of the dynamic vision sensor with the corresponding microlens center to form multiple straight lines in space, and each straight line corresponds to a ray of light collected. Subsequently, record the coordinates of the space points passed by the straight lines in the memory.
[0033] Step 6, based on the space point coordinates recorded in Step 5, introduce a diffusion weight mechanism at each point passed by each ray of light, that is, with each passed point as the center and k unit lengths as the limited range, assign weights to the points around the straight line within the range according to certain distributions. The form of this distribution can be specified as needed. For example, a Gaussian distribution with distance as the independent variable, etc. As Figure 4 shown on the right, different depths of color mark different weighted points. The closer to the place where the light is intensive, the higher the weight, so that the possible real event points can be screened out.
[0034] Step 7, screen the point weights in the multi-perspective triangular voting measurement space according to the threshold. If there are space points exceeding the preset threshold, it can be determined that the events corresponding to the relevant rays involved in this space point constitute a high-dimensional light field event output, which can describe the brightness change in the real space from multiple perspectives.
[0035] Step 8, if the dynamic vision sensor still has output, return to Step 2 to process the next event; otherwise, end the processing.
Claims
1. An extraction method for a high-dimensional light field event camera based on a microlens array. The camera includes a main lens, a microlens array, and a dynamic vision sensor. The microlens array is placed between the main lens and the dynamic vision sensor. Among them, The microlens array is located behind the main lens, and each microlens in the array can perform secondary imaging on the imaging of the main lens; A certain brightness change area in the target scene is projected onto the dynamic vision sensor after two optical imaging processes by the main lens and the microlens array. Different positions of the dynamic vision sensor can sense the brightness change of the same area, forming multiple asynchronous outputs. The method is characterized in that it includes the following steps: Step 1, assuming that the microlens array is N rows and M columns, with a total of N*M sub-images of different viewing angles, initialize N*M corresponding event buffers for subsequent event management; Step 2: After reading out the output event e i (x i , y i , p i , t i ) of the dynamic vision sensor, store it in the buffer corresponding to the perspective according to its coordinates, where i represents the i-th event read out, x i, y i represents the two-dimensional coordinates of the current event on the dynamic vision sensor, p i represents the polarity of the current event, and t i represents the timestamp of the event; Step 3, check the events stored in the buffer. If the events in the buffer exceed the maximum queue length, clear the earliest stored events in the buffer according to the exceeded length; Step 4, determine whether, after the current event enters the buffer, it can activate the number of viewing angles covered by the multi-view event combination within the defined domain range to reach the preset threshold. If the threshold limit condition can be met, proceed to Step 5; otherwise, proceed to Step 2; Step 5, take out all the events involved in the multi-view combination activated by the latest event entering the buffer. In the world coordinate system constructed according to the application scenario, based on the center coordinates of the microlenses corresponding to the viewing angles, construct the geometric space of the multi-view triangulation voting measurement method, that is, connect the coordinates of each response pixel of the dynamic vision sensor with the corresponding microlens center to form multiple straight lines in space, and each straight line corresponds to a ray of light collected; Step 6, introduce a diffusion weight mechanism on each ray of light, that is, taking each point on the straight line as the center and k unit lengths as the defined range, assign weights to the points around the straight line within the range; Step 7, screen the point weights in the multi-view triangulation voting measurement space according to the threshold. If there are spatial points exceeding the preset threshold, determine that the events corresponding to the relevant rays involved in the spatial points constitute a high-dimensional light field event, which can describe the brightness change in the real space from multiple viewing angles; Step 8, if the dynamic vision sensor still has output, return to Step 2 to process the next event; otherwise, end the process.
Citation Information
Patent Citations
Event camera imaging device and method and event camera
CN114401358A