Method for Evaluating Denoising Effect of Dynamic Visual Event Stream Based on Event Spatio-Temporal Synchronization
Through the evaluation method of noise reduction effect of dynamic visual event stream based on event space-time synchronization, combined with the Event-based Multi View Stereo algorithm for three-dimensional reconstruction and confidence map generation, the problem of noise interference in the event stream of dynamic vision sensors is solved, and the objective evaluation of noise reduction effect and the credibility of event stream is achieved.
Patent Information
- Application Number
- CN202211076662.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-05
AI Technical Summary
The event stream output by the dynamic vision sensor contains a large amount of noise interference, which makes it difficult to effectively reduce noise and evaluate the noise reduction effect, limiting the further development and application of event stream noise reduction algorithm and dynamic vision sensor.
The dynamic visual event stream noise reduction effect evaluation method based on event space-time synchronization is adopted. By reading the event stream output from DVS and obtaining pose information, three-dimensional reconstruction and confidence map generation are carried out in combination with the Event-based Multi View Stereo algorithm to calculate the rationality of the event stream and evaluate the noise reduction effect.
In the case where the specific noise distribution and reference event stream are unknown, objective evaluation of the noise reduction effect of event stream is achieved, the overall credibility of event stream is improved, and the noise reduction accuracy of different algorithms is evaluated.
Smart Images

Figure CN115375581B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sensor signal processing, and particularly relates to a method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization. Background Art
[0002] A Dynamic Vision Sensor (DVS), also known as an event camera, is a bio-inspired sensor that has a completely different working mode from traditional cameras by simulating the visual system of organisms in nature. The dynamic vision sensor does not output images at a fixed rate. Based on asynchronous event-driven, it is sensitive to the illumination changes of each pixel. When the logarithmic change of brightness reaches a preset threshold, an "event" is triggered with ultra-low latency (less than 1 microsecond). Each event is represented by a four-dimensional vector e(x, y, t, p), including the pixel coordinates (x, y) of the event, the trigger time t, and the polarity p. Among them, the polarity p ∈ {1, -1} represents the increase or decrease of brightness on the pixel. The DVS only outputs the relevant information of local brightness changes, and has the characteristics of fast response speed, ultra-low latency, high dynamic range, only capturing dynamic changes, low power consumption, etc. It can overcome the deficiencies of traditional cameras such as high redundancy, low frame rate, large latency, and low dynamic range, and thus has been widely used in fields such as unmanned driving and robotics.
[0003] Due to its own structure, the dynamic vision sensor is very sensitive to environmental brightness changes and is limited by the hardware level. A large amount of noise interference is included in the output asynchronous event stream. The noise may come from pulse noise during digital signal transmission and Gaussian noise caused by photodiodes, etc., which has a great impact on the further application and visualization of the event stream. Therefore, noise reduction processing of the event stream information is a very important link, and the noise reduction of the event stream has also become an important research topic in the field of dynamic vision sensors.
[0004] However, due to the large data load of the DVS (millions of events are output per second), it is difficult to manually mark the validity of each event, and the specific distribution of the noise and the reference event stream cannot be obtained. Therefore, there is currently no effective method to measure the noise reduction effect of the event stream and compare the noise reduction effects of different algorithms, which further restricts the further development and application of the event stream noise reduction algorithm and the dynamic vision sensor. Summary of the Invention
[0005] Aiming at the above problems, the present invention provides a method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization. By using the high-frequency advantage of the dynamic vision sensor and combining the pose information, it is possible to objectively evaluate the noise reduction effect of the event stream when the specific distribution of the noise and the reference event stream are unknown.
[0006] The present invention provides a method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization, and the specific steps are as follows:
[0007] Step 1: Read the event stream output by the dynamic vision sensor DVS, and obtain the pose information of the DVS through a motion capture system, visual odometer, inertial navigation, or indoor positioning method;
[0008] Step 2: Based on the Event-based Multi View Stereo algorithm, use the event stream combined with the pose information to perform three-dimensional reconstruction of the actual scene, project the events triggered at different times to the reference time for spatio-temporal synchronization, obtain a confidence map, and realize the sharpening of real events and the blurring of noise events, as a reference benchmark for noise reduction;
[0009] Step 3: Since the local maxima and edge regions in the DSI often correspond to the intensity gradients in the scene and the probability of triggering events is greater, after obtaining the confidence map c(x, y) at the reference time t r , by calculating the probability of each pixel being a local maximum or an edge, convert it into an event probability map, which represents the probability of events triggered on each pixel of the DVS in the real scene at the ideal time t r ;
[0010] Step 4: Based on the consistency between the event stream at the reference time t r and the event probability map, calculate the rationality of the event stream;
[0011] Step 5: A high-precision event stream noise reduction method can remove noise events with low rationality and retain valid events with high rationality, thereby improving the overall credibility of the event stream. Therefore, by comparing the rationality of the event stream before and after noise reduction, calculate the improvement of the noise reduction algorithm on the rationality of the event stream, obtain the noise reduction accuracy index, and use it to evaluate and compare the noise reduction effects of different algorithms:
[0012]
[0013] where, e original and e denoised represent the event streams before and after noise reduction respectively. The higher the noise reduction accuracy index, the more obvious the improvement of the noise reduction algorithm on the rationality of the event stream, and the better the noise reduction effect.
[0014] As a further improvement of the present invention, the three-dimensional reconstruction using the event stream combined with the pose information in Step 2 includes the following process:
[0015] Before performing three-dimensional reconstruction, first detect the repeated events on each pixel and only retain the first triggered event among them:
[0016] IE = {e i (xi , y i , t i ) | (t i -t i-1 ) > τ IE ^(t i+1 -t i ) < τ IE}}
[0017] Among them, IE represents the first event in the repeated trigger event, representing t i The timestamp of the i-th event triggered on a certain pixel, and the time threshold parameter t IE is set to 20 ms;
[0018] After that, event-based three-dimensional reconstruction is performed to generate a confidence map. The specific steps include:
[0019] (2-1) Select the observation view at the reference time as the reference view, and discretize the observation camera system along its optical axis direction into a grid map to construct a disparity space image DSI. The disparity space image DSI discretizes the reference view into N depth planes Each depth plane is divided into w×h spatial units, which is consistent with the pixel resolution of the DVS. Therefore, the DSI is divided into w×h×N spatial voxel units, where N is set to 100;
[0020] (2-2) Then project all events from the pixel plane to the disparity space image DSI according to the pose at the corresponding time, and calculate the number of intersections between each unit voxel in the disparity space image DSI and the event back-projection ray. The more intersections, the more times the corresponding area is observed and responded to by the DVS, and the greater the probability that the unit voxel contains the scene edge. Correspondingly, the probability of triggering an event on the DVS at the reference view is also greater;
[0021] In the process of projecting and composing events, the effective events triggered by the actual scene continuously at a high frequency will be automatically synchronized to the corresponding spatial positions, while the noise in the event stream will not produce spatio-temporal persistent voting on the fixed areas in the DSI and will be diluted by the effective events;
[0022] Finally, by recording the maximum value of the DSI along the optical axis direction of each pixel at the reference view, the confidence map at the reference view is obtained.
[0023] As a further improvement of the present invention, the step of projecting events from the pixel plane to the DSI in step 2 is as follows:
[0024] Use homography to solve the intersection cells between the event projection ray and each depth plane of the DSI. Each depth plane has the following expressions respectively:
[0025] Zi = [n, d i T = [(0, 0, 1), z i T
[0026] where n and z i are the normal vector and depth of each plane respectively;
[0027] During the projection process, first, use the relative pose [R|t] between each event e i (x i , y i ) between the observation time and the reference time to calculate the homography matrix between the two relative to the Z0 plane
[0028]
[0029] After that, combined with the projection matrix P of the DVS, through homography transformation, obtain the projection coordinates on the Z0 plane from the pixel coordinates of the event:
[0030]
[0031]
[0032] In the formula, (x i , y i ) and (x(z0), y(z0)) are the pixel coordinates of the event and the projection coordinates on the Z0 plane respectively;
[0033] The projection coordinates of the event on the remaining depth planes of the DSI are calculated again through homography transformation from the coordinates on the Z0 plane:
[0034]
[0035] where After simplification, obtain the coordinates of the event on the Z i plane:
[0036]
[0037] In the formula, (c x , c y , c z ) T = -R T t, which is the coordinate of the DVS relative to the reference time.
[0038] As a further improvement of the present invention, converting the confidence map to an event probability map in step 3 includes the following process:
[0039] (3-1) For each pixel x, y in the confidence map, the spatial proximity and pixel similarity between it and each pixel (x i , y i ) ∈ Ω within the neighborhood window Ω are respectively used to construct the spatial domain Gaussian kernel G d and the range domain Gaussian kernel G r :
[0040]
[0041]
[0042] Among them,
[0043] (3-2) Then, using the normalized product of the spatial domain Gaussian kernel and the range domain Gaussian kernel as the weight W(x i , y i ), all pixels within the window Ω are weighted and fused to obtain the adaptive threshold T(x, y). By comparing it with the confidence value c(x, y) at the central pixel, the probability that the pixel (x, y) is a local maximum or an edge region of the DSI is calculated, representing the event trigger probability p(x, y):
[0044]
[0045]
[0046]
[0047] Among them, the window size is set to 7x7, and the above steps are repeated for all pixels on the confidence map to obtain the event probability map at the reference time.
[0048] As a further improvement of the present invention, the event stream rationality calculation in step 4 includes the following process:
[0049] (4-1) For the event stream e i (x i , y i , t i ), use I:Z 2 → {0, 1} to represent the event trigger situation on the pixel plane Z 2 of the dynamic vision sensor within a period of time before and after the reference time:
[0050]
[0051] Among them, τ represents the time range, and 1 and 0 respectively represent the presence and absence of events at this pixel within the time period [t r - τ, t r + τ];
[0052] (4-2) When I(x, y) = 1, use the time distance between the timestamp of the event at the corresponding pixel and the reference time to construct an exponential decay kernel, representing the temporal correlation between the event stream at this pixel and the event probability map:
[0053]
[0054] where the decay rate parameter δ t is set to 20 ms, and the rationality of the event stream triggering an event at the pixel is quantified by the product of the event probability p(x, y) and Γ(x, y) at the corresponding position. The greater the rationality, the greater the likelihood that the event is an effective signal triggered by the actual scene;
[0055] When I(x, y) = 0, use the inverse event probability on the event probability map: to represent the rationality of the absence of events at this pixel;
[0056] (4-3) Therefore, the logarithmic rationality of the event stream on the pixel (x, y) within the time period [t r - τ, t r + τ] is obtained:
[0057]
[0058] Calculate the logarithmic rationality at all pixels on the pixel plane Z 2 respectively, and obtain the rationality of the event stream e i :
[0059]
[0060] logP(e i ) The smaller it is, the better the consistency between the event stream e i and the event probability map, and the higher the rationality.
[0061] Beneficial effects:
[0062] The present invention makes full use of the high-frequency characteristics of the dynamic vision sensor, synchronizes the effective events continuously triggered by the actual scene on the pixel plane to the three-dimensional space at the reference time, realizes the highlighting and sharpening of the effective events and the blurring and filtering of the noise events, thereby quantifying the rationality of the event stream and the noise reduction accuracy of the algorithm, and can objectively evaluate the noise reduction effect of the event stream in the case where the specific distribution of the noise and the reference event stream are unknown. Description of the drawings
[0063] Figure 1 is the flowchart of the method for evaluating the noise reduction effect of the event stream provided by the present invention. Detailed implementation manners
[0064] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0065] The present invention discloses a method for evaluating the denoising effect of a dynamic visual event stream based on event spatiotemporal synchronization. The method of the present invention is as follows: Figure 1 As shown, the specific steps include:
[0066] Step 1: Read the event stream output by the dynamic vision sensor (DVS) and obtain the pose information of the DVS through methods such as motion capture system, visual odometer, inertial navigation or indoor positioning.
[0067] Step 2: Based on the Event-based MultiView Stereo (EMVS) algorithm, the actual scene is reconstructed in three dimensions using event streams combined with pose information. Events triggered at different times are projected to the reference time for spatiotemporal synchronization to obtain a confidence map, which sharpens real events and blurs noise events as a reference for noise reduction.
[0068] Since repeated events on the same pixel in a short period of time will be repeatedly projected to the reference time, thus affecting the construction of the confidence map and the subsequent event validity evaluation, before 3D reconstruction, repeated events on each pixel are detected and only the first triggering event is retained:
[0069] IE={e i (x i ,y i ,t i )(t i -t i-1 )>τ IE ^(t i+1 -t i )<τ IE}
[0070] Among them, IE represents the first event in the repeated triggering event, and t i The timestamp of the i-th event triggered on a certain pixel, the time threshold parameter t IE Set to 20ms.
[0071] After that, event-based 3D reconstruction is performed to generate a confidence map. The specific steps include:
[0072] (2-1) The observation angle at the reference time is selected as the reference angle, and the observation camera system along its optical axis is discretized into a grid map to construct a disparity space image (DSI). DSI discretizes the reference angle into N depth planes. Each depth plane is divided into w×h spatial units, which is consistent with the pixel resolution of the DVS. Therefore, the DSI is divided into w×h×N spatial voxel units. Here, N is set to 100.
[0073] (2-2) Then project all events from the pixel plane to the DSI according to the pose at the corresponding moment, and calculate the number of intersections (also called "voting") between each voxel in the DSI and the event back-projection ray. The more intersections, the more times the corresponding area is observed and responded to by the DVS, the greater the probability that the voxel contains the scene edge, and correspondingly, the greater the probability of triggering an event on the DVS at the reference view.
[0074] Use the homography to solve the intersection cells of the event projection ray and each depth plane of the DSI. Each depth plane has the following expressions:
[0075] Z i =[n,d i T =[(0,0,1),z i T
[0076] where n and z i are the normal vector and depth of each plane respectively.
[0077] During the projection process, first use the relative pose [R|t] between each event e i (x i ,y i ) between the observation moment and the reference moment to calculate the homography matrix H Z0 with respect to the Z0 plane:
[0078]
[0079] Then, combined with the projection matrix P of the DVS, through the homography transformation, obtain the projection coordinates on the Z0 plane from the pixel coordinates of the event:
[0080]
[0081]
[0082] In the formula, (x i ,y i ) and (x(z0),y(z0)) are the pixel coordinates of the event and the projection coordinates on the Z0 plane respectively.
[0083] The projection coordinates of the event on the remaining depth planes of the DSI are calculated again through the homography transformation from the coordinates on the Z0 plane:
[0084]
[0085] Among them After simplification, the coordinates of the event on the Z i plane are obtained:
[0086]
[0087] In the formula, (c x , c y , c z ) T =-R T t, which is the coordinate of the DVS relative to the reference time.
[0088] In the process of projecting and constructing the events, the effective events triggered by the actual scene at a high frequency and continuously will be automatically synchronized to the corresponding spatial positions, while the noise in the event stream will not generate spatio-temporal persistent votes in the fixed areas of the DSI and will be diluted by the effective events. Therefore, the algorithm has good robustness to noise, providing a reference for objectively evaluating the noise reduction effect of the event stream.
[0089] Finally, by recording the maximum value of the DSI along the optical axis direction for each pixel in the reference view, the confidence map in the reference view is obtained.
[0090] Step 3: Since the local maxima and edge regions in the DSI often correspond to the intensity gradients in the scene and have a greater probability of triggering events, after obtaining the confidence map c(x, y) at the reference time t r , we convert it into an event probability map by calculating the probability of each pixel being a local maximum or an edge, representing the event probability of the true scene triggering on each pixel of the DVS at the ideal time t r . The specific steps are as follows:
[0091] (3-1) For each pixel (x, y) in the confidence map, the spatial proximity and pixel similarity between it and each pixel (x i , y i ) ∈ Ω in the neighborhood window Ω are used to construct the spatial domain Gaussian kernel G d and the range domain Gaussian kernel G r :
[0092]
[0093]
[0094] Among them,
[0095] (3-2) After that, the normalized product of the spatial domain Gaussian kernel and the range domain Gaussian kernel is used as the weight W(x i , yi ), perform weighted fusion on all pixels within the window Ω to obtain the adaptive threshold T(x, y), compare it with the confidence value c(x, y) at the central pixel, and calculate the probability that the pixel (x, y) is a local maximum or an edge region of the DSI, representing the event trigger probability p(x, y):
[0096]
[0097]
[0098]
[0099] Among them, the window size is set to 7x7. And repeat the above steps for all pixels on the confidence map to obtain the event probability map at the reference time. Therefore, by introducing spatial proximity and pixel similarity, the confidence map can be converted into an event probability map according to the event trigger characteristics and combined with local scene information.
[0100] Step 4: Based on the reference time t r The consistency between the event stream and the event probability map, calculate the rationality of the event stream at the corresponding time, which specifically includes the following steps:
[0101] (4-1) For the event stream e i (x i , y i , t i ), use I:Z 2 →{0, 1} to represent the event trigger situation on the pixel plane Z of the dynamic vision sensor within a certain time range before and after the reference time: 2 On the event trigger situation on the pixel plane Z of the dynamic vision sensor within a certain time range before and after the reference time:
[0102]
[0103] Among them, τ represents the time range, and 1 and 0 represent the presence and absence of events at this pixel within the time period [t r -τ, t r +τ], respectively.
[0104] (4-2) When I(x, y) = 1, use the time distance between the timestamp of the event at the corresponding pixel and the reference time to construct an exponentially decaying kernel to represent the temporal correlation between the event stream and the event probability map at this pixel:
[0105]
[0106] Among them, the decay rate parameter δt is set to 20 ms. The rationality of the event stream triggering an event at a pixel is quantified by the product of the event probability p(x, y) and Γ(x, y) at the corresponding position. The greater the rationality, the greater the likelihood that the event is an effective signal triggered by the actual scene.
[0107] When I(x, y) = 0, the inverse event probability on the event probability map is used: It represents the rationality of the absence of events at this pixel.
[0108] (4 - 3) Therefore, the logarithmic rationality of the event stream at the pixel (x, y) within the time period [t r - τ, t r + τ] is obtained:
[0109]
[0110] Calculate the logarithmic rationality at all pixels on the pixel plane Z 2 respectively, and the rationality of the event stream e i is obtained:
[0111]
[0112] The smaller logP(e i ) is, the better the consistency between the event stream e i and the event probability map, and the higher the rationality.
[0113] Step 5: A high-precision event stream noise reduction method can remove noise events with low rationality and retain effective events with high rationality, thereby improving the overall credibility of the event stream. Therefore, by comparing the rationality of the event stream before and after noise reduction, the improvement of the noise reduction algorithm on the rationality of the event stream is calculated, and the noise reduction accuracy index is obtained, which is used to evaluate and compare the noise reduction effects of different algorithms:
[0114]
[0115] Among them, e original and e denoised represent the event streams before and after noise reduction respectively. The higher the noise reduction accuracy index, the more obvious the improvement of the noise reduction algorithm on the rationality of the event stream, and the better the noise reduction effect.
[0116] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.
Claims
1. A method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization, the specific steps are as follows, characterized in that: Step 1: Read the event stream output by the dynamic vision sensor DVS, and obtain the pose information of the DVS through an action capture system, visual odometer, inertial navigation or indoor positioning method; Step 2: Based on the Event-based Multi View Stereo algorithm, use the event stream combined with the pose information to perform three-dimensional reconstruction of the actual scene, project the events triggered at different times to the reference time for spatio-temporal synchronization, obtain a confidence map, and realize the sharpening of real events and the blurring of noise events, as a reference benchmark for noise reduction; Step 3: Since the local maxima and edge regions in DSI often correspond to the intensity gradients in the scene and have a greater probability of triggering events, after obtaining the confidence map c(x, y) at the reference time t r , it is converted into an event probability map by calculating the probability of each pixel being a local maximum or an edge, representing the probability of events triggered by the true scene at each pixel of the DVS at time t r in an ideal situation; The process of converting the confidence map into an event probability map in Step 3 includes the following: (3-1) For each pixel x, y in the confidence map, the spatial proximity and pixel similarity between it and each pixel (x i , y i ) ∈ Ω within the neighborhood window Ω are respectively used to construct the spatial domain Gaussian kernel G d and the range domain Gaussian kernel G r : Among them, After (3-2), using the normalized product of the spatial Gaussian kernel G d and the range Gaussian kernel Gr as the weight W(x i , y i ), all pixels within the window Ω are weighted and fused to obtain the adaptive threshold T(x, y), which is compared with the confidence value c(x, y) at the central pixel to calculate the probability that the pixel (x, y) is a local maximum or an edge region of the DSI, representing the event trigger probability p(x, y): Among them, the window size is set to 7x7, and the above steps are repeated for all pixels on the confidence map to obtain the event probability map at the reference time; Step 4: Based on the reference time t r Calculate the rationality of the event stream based on the consistency between the event stream and the event probability map; Step 5: A high-precision event stream noise reduction method can remove noise events with low rationality and retain effective events with high rationality, thereby improving the overall credibility of the event stream. Therefore, by comparing the rationality of the event stream before and after noise reduction, calculate the improvement of the noise reduction algorithm on the rationality of the event stream, and obtain the noise reduction accuracy index, which is used to evaluate and compare the noise reduction effects of different algorithms: Among them, e original and e denoised respectively represent the event streams before and after noise reduction. The higher the noise reduction accuracy index, the more obvious the improvement of the noise reduction algorithm on the rationality of the event stream, and the better the noise reduction effect.
2. The method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization according to claim 1, characterized in that: The process of using the event stream combined with the pose information for three-dimensional reconstruction in Step 2 includes the following: Before performing three-dimensional reconstruction, first detect the repeated events on each pixel and only retain the first triggered event among them: IE = {e i (x i , y i , t i ) | (t i - t i-1 ) > τ IE ∧(t i+1 - t i ) < τ IE} Among them, IE represents the first event in the repeated trigger event, indicating t i The timestamp of the i-th event triggered on a certain pixel, and the time threshold parameter τ IE is set to 20 ms; After that, perform event-based three-dimensional reconstruction to generate a confidence map. The specific steps include: (2-1) Select the observation perspective at the reference moment as the reference perspective, discretize the observation camera system along its optical axis direction into a grid map, and construct a disparity space image (DSI). The DSI discretizes the reference perspective into N depth planes. Each depth plane is divided into w×h spatial units, which is consistent with the pixel resolution of the DVS. Therefore, the DSI is divided into w×h×N spatial voxel units, where N is set to 100. (2-2) After that, project all events from the pixel plane to the disparity space image DSI according to the pose at the corresponding time, and calculate the number of intersections between each voxel in the disparity space image DSI and the event back-projection ray. The more intersections, the more times the corresponding area is observed and responded to by the DVS, and the greater the probability that the voxel contains the scene edge. Correspondingly, the probability of triggering an event on the DVS at the reference view is also greater; During the process of projecting and composing the events, the effective events triggered by the actual scene at a high frequency will be automatically synchronized to the corresponding spatial positions, while the noise in the event stream will not produce spatio-temporal persistent votes for the fixed areas in the DSI and will be diluted by the effective events; Finally, by recording the maximum value of the DSI along the optical axis direction of each pixel at the reference view, the confidence map at the reference view is obtained.
3. The method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization according to claim 2, wherein: The steps of projecting the event from the pixel plane to the DSI in Step 2 are as follows: Using homography to solve the intersection cells of the event projection rays and each depth plane of the DSI. Each depth plane has the following expressions respectively: Z i = [n, d i T = [(0, 0, 1), z i T where n and z i are the normal vector and depth of each plane, respectively; During the projection process, first, for each event e i (x i ,y i ) the relative pose [R|t] between the observation time and the reference time is used to calculate the homography matrix of the two relative to the Z0 plane After that, combined with the projection matrix F of the DVS, through the homography transformation, obtain the projection coordinates on the Z0 plane from the pixel coordinates of the event: where (x i , y i ) and (x(z0), y(z0)) are the pixel coordinates of the event and the projected coordinates on the Z0 plane, respectively; The projection coordinates of the event on the remaining depth planes of the DSI are calculated again through the homography transformation from the coordinates on the Z0 plane: Among them After simplification Obtain the coordinates of the event in the Z i plane: where (c x , c y , c z ) T = -R T t, which is the coordinate of the DVS relative to the reference time.
4. The method for evaluating the noise reduction effect of a dynamic vision event stream based on event spatio-temporal synchronization according to claim 1, wherein: The calculation of the rationality of the event stream in Step 4 includes the following process: (4-1) For the event stream e i (x i , y i , t i ), use I:Z 2 →{0, 1} to represent the event triggering situation on the dynamic vision sensor pixel plane Z within a period of time before and after the reference time: 2 where τ represents the time range, 1 and 0 represent the presence and absence of events at this pixel within the time period r [t r - τ, t + τ], respectively; (4-2) When I(x,y) = 1, use the time distance between the timestamp of the event at the corresponding pixel and the reference time to construct an exponential decay kernel, which represents the temporal correlation between the event stream at this pixel and the event probability map: Among them, the decay rate parameter δ t is set to 20 ms. The rationality of the event stream triggering an event at a pixel is quantified by the product of the event probability p(x, y) and Γ(x, y) at the corresponding position. The greater the rationality, the greater the possibility that the event is an effective signal triggered by the actual scene; When I(x, y) = 0, the inverse event probability on the event probability map is used: indicating the rationality of the pixel missing event; (4-3) Therefore, the log-likelihood of the event stream on pixel (x, y) in the time period r -τ,t r +τ] is obtained: Calculate the log-likelihood at all pixels on the pixel plane Z 2 to obtain the event stream e i likelihood: logP(e i ) The smaller it is, the better the consistency between the event stream e i and the event probability map, and the higher the rationality.
Citation Information
Patent Citations
Image vision processing method, device and equipment
US20180137639A1