Event stream denoising method and system based on point voxel attention mechanism
Through the event flow denoising method based on point voxel attention mechanism, time window sampling and feature alignment technology are used to solve the background noise problem in the event camera, accurately distinguishing events and noise, and improving imaging quality.
Patent Information
- Application Number
- CN202510474381.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
There is background activity noise in the data of existing event cameras, which affects imaging quality and occupies communication bandwidth and computing resources. The existing denoising methods destroy the spatiotemporal correlation of events and lead to loss of geometric information.
The event flow denoising method based on the point voxel attention mechanism is adopted, through time window sampling, T-Net network feature alignment, point voxel attention module feature extraction and fusion, the sum pooling operation is used to aggregate spatiotemporal information to distinguish real events from noise.
It effectively solves the background activity noise problem during the event camera imaging process, retains the asynchronousness and sparseness of event data, realizes the accurate distinction between events and noise, and improves imaging quality.
Smart Images

Figure CN120387947A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an event stream denoising method and system based on a point voxel attention mechanism. Background Art
[0002] The event camera is a novel visual sensor inspired by biomimetic technology, originating from research on biomimetic retinas in neuromorphic engineering. Unlike traditional cameras that synchronously generate image frames, event cameras asynchronously output event streams by sensing changes in scene brightness. Specifically, they only output event information when the brightness change at a pixel location exceeds a threshold. The output event includes the pixel location, timestamp, and event polarity. Compared to traditional visual sensors, event cameras offer advantages such as high temporal resolution, low latency, high dynamic range, low energy consumption, and low data volume. They hold broad application prospects in areas such as drone inspection, autonomous driving, and public security video surveillance. However, due to the sensor's photosensitive hardware characteristics and digital signal transmission contamination, event stream noise can be generated, compromising the data representation performance of event cameras. The most serious of these noises is background activity noise, which not only consumes additional communication bandwidth and computational resources but also introduces useless information into the image reconstruction process, negatively impacting image quality.
[0003] Although some event denoising methods have been proposed, most of them are based on image or voxel grid representations, which destroy the spatiotemporal correlation of events and suffer from problems such as loss of geometric information and insufficient temporal consistency. In order to improve the quality of event-based data, it is necessary to develop appropriate denoising methods. Summary of the Invention
[0004] Purpose of the invention: In order to overcome the problems existing in the prior art, the present invention discloses an event stream denoising method based on a point voxel attention mechanism, which can directly denoise events while retaining the asynchrony and sparsity of event data.
[0005] To achieve the above objectives, this paper proposes an event stream denoising method based on a point voxel attention mechanism, which can directly process the current event. By focusing on the current event, we can perform the following operations without losing global information:
[0006] Collect data of several event points within a predetermined time range to generate an event point set;
[0007] performing spatial alignment on all event point data and input features in the event point set;
[0008] Construct a point voxel attention module to extract the local features of each event point data in the voxel domain through its voxel branch, and extract the global features of each event point data through its point branch;
[0009] The local features and all features are integrated to obtain output spatial features; the geometric information of the output spatial features is aggregated through a sum pooling operation to obtain a potential feature vector containing spatiotemporal information;
[0010] Real events and noise are distinguished from the latent feature vector containing spatiotemporal information.
[0011] In a further embodiment, collecting data of several event points within a predetermined time range to generate an event point set specifically includes:
[0012] Calculate the time window range based on the event stream;
[0013] Within a certain time window, collect the N events that are closest in time to the current event in the spatial neighborhood of the current event;
[0014] N events are filtered based on time correlation, and event point data that meet the predetermined time correlation requirements are retained to form an event point set.
[0015] In a further embodiment, spatial alignment of all event point data and input features in the event point set specifically includes:
[0016] The N×3 event point set collected in the time window is passed into the T-Net network;
[0017] The T-Net network passes the input event point set through three convolutional layers in sequence, gradually increasing it to the first channel dimension, the second channel latitude, and the third channel dimension. The first channel dimension can be 64, the second channel latitude can be 128, and the third channel dimension can be 1024.
[0018] Aggregate the extracted features into global features through sum pooling;
[0019] Convert the global features into vector form and input them into the fully connected layer to obtain a 3×3 transformation matrix;
[0020] The transformation matrix is multiplied by the input event point set to spatially align the event points.
[0021] In a further embodiment, the sampled event point set is passed through a point voxel attention module for feature extraction, in which:
[0022] The input event point set is first voxelized; then the voxelized events are subjected to a three-dimensional convolution operation and then passed into the convolutional attention module for feature extraction.
[0023] In the convolutional attention module, the input feature V is first passed to the channel attention module, which performs maximum pooling and average pooling operations on the input feature V respectively, calculates the average and maximum values of the input feature V in the last three dimensions, and thus generates two feature vectors with the size of the channel number. Then, the maximum pooling feature V obtained after pooling is max and average pooling feature V avg The shared multi-layer perceptron is sent to learn the attention weight of each channel, and the output channel attention feature is obtained by adding and passing the sigmoid function. Finally, the input feature V is multiplied by the obtained channel attention feature to obtain the channel feature V C .
[0024] In a further embodiment, further comprising:
[0025] The channel is characterized by V C A spatial attention module is passed in to extract key information at different locations in the space;
[0026] The spatial attention module takes as input V C Perform two-dimensional convolution and use Sigmoid activation function to generate channel attention weights to generate a two-dimensional spatial attention feature;
[0027] The attention weights are combined with the input features V C Multiply each channel of , aggregate the spatiotemporal features through the multi-layer perceptron layer, and obtain the voxel features to be extracted.
[0028] In a further embodiment, a trilinear interpolation method is used to map voxel features back to point features, specifically including:
[0029] The input event point set is subjected to a one-dimensional convolution operation to extract the point features, which are then passed into a BN layer for batch normalization processing, and the sigmoid function is used for nonlinear activation to extract the point features.
[0030] In a further embodiment, the local features and the overall features are fused to obtain output spatial features, specifically including:
[0031] The voxel features and the point features are added together and passed into three PVA modules and one MLP module to obtain an N×1280 feature matrix;
[0032] The geometric information is aggregated through the sum pooling operation in the N-dimensional direction to obtain the potential feature vector containing spatiotemporal information.
[0033] In addition, the present invention also discloses an event stream denoising system, which is used to execute the above-mentioned event stream denoising method based on the point voxel attention mechanism. The event stream denoising system consists of an event point sampling module, a feature alignment module, a feature extraction module, a feature fusion module, and an event denoising module.
[0034] The event point sampling module is used to collect data of several event points within a predetermined time range and generate an event point set;
[0035] The feature alignment module embeds a T-Net module, and uses the T-Net module to spatially align all event point data and input features in the event point set;
[0036] The feature extraction module includes three point voxel attention modules and an MLP module. The voxel branch of the point voxel attention module extracts the local features of each event point data in the voxel domain, and the point branch of the point voxel attention module extracts the global features of each event point data.
[0037] The feature fusion module obtains the output spatial features after passing through three point voxel attention modules and an MLP module, aggregates the geometric information of the output spatial features through the sum pooling operation, and obtains the latent feature vector containing spatiotemporal information;
[0038] The event denoising module is used to distinguish real events from noise from the potential feature vector containing spatiotemporal information.
[0039] In addition, the present invention also discloses an electronic device, which includes a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the above-mentioned event stream denoising method based on the point voxel attention mechanism.
[0040] Compared with the prior art, the present invention has at least the following beneficial effects:
[0041] The present invention introduces a time window to perform local sampling of event points, avoiding the large amount of noise and irrelevant events mixed in the output data of the event camera, which will bring erroneous information prompts to the network; the T-Net network is used to ensure the consistency of feature distribution; the point voxel attention module is used to extract features, which can solve the problem of information loss in the voxelization process. By fusing the point branch and the voxel branch, the point branch is used to extract the global information of the event point to complement the voxel branch and make up for the information lost in the voxelization process; after three PVA modules and one MLP module, spatial features containing spatiotemporal information are obtained; finally, the fully connected layer is used as a binary classifier to achieve accurate distinction between events and noise. The present invention can solve the problems of event geometric information loss and insufficient temporal consistency, achieve accurate distinction between events and noise, and effectively solve the problem of background activity noise generated by the event camera during the imaging process. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Flowchart of the method of the present invention.
[0043] Figure 2 This is a network structure diagram of the present invention.
[0044] Figure 3 Schematic diagram of channel attention.
[0045] Figure 4 Schematic diagram of spatial attention.
[0046] Figure 5 Visualization of the event flow after denoising the simulated data.
[0047] Figure 6 This is a visualization of the event flow after denoising the real data.
[0048] Figure 7 Visualization of the event stream after denoising the DVSNOISE20 dataset. DETAILED DESCRIPTION
[0049] In the following description, numerous specific details are provided to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, certain technical features well known in the art have not been described to avoid confusion with the present invention.
[0050] The present invention discloses an event stream denoising system based on a point voxel attention mechanism, and discloses an event stream denoising method based on a point voxel attention mechanism based on the system. The flow chart is as follows: Figure 1 shown.
[0051] The denoising neural network proposed in this paper consists of a temporal window (TW) module, a T-Net module, a point-voxel attention (PVA) module, a multi-layer perceptron (MLP) module and a fully connected layer. Figure 2 , where the time window module performs the following operations:
[0052] Step S1, event point extraction: Since the data output by the event camera often contains a lot of noise and irrelevant events, in order to avoid affecting the network training, a time window is used to locally sample the event stream, filter out the time-related events in the spatial domain of the current event, and define this closed area as the spatiotemporal domain. After sampling the time window, a size of The set of event points, where N represents the number of event points and D represents its feature dimension. Sampling the spatial neighborhood through a time window and only retaining the events that are closely related to the current event in time for local feature extraction, as shown in the following formula:
[0053]
[0054] Time limit t lim The calculation formula is:
[0055]
[0056] where t max and t min are the maximum and minimum timestamps of the event stream respectively, and M is the total number of events in the event stream.
[0057] Based on the time window, the screening rules for the time-related events of the current event are as follows:
[0058]
[0059] After completing the event sampling in the spatio-temporal domain, the N events closest to the current event are screened out. If the number of some event points is insufficient in the spatio-temporal domain, these events are very likely to belong to background activity noise. However, the movement trajectories of some small objects or the extreme edges of objects may also result in insufficient number of event points in the spatio-temporal domain, thus being misjudged as noise. To prevent this situation from being missed, the missing parts of these event points are set to the same value as the current event. To judge the relationship between the current point and its adjacent events, the spatial information of the neighboring events is centered relative to the current event by subtracting the x and y coordinate values of the current event. Finally, N×3 spatial information is extracted from the N neighboring events that are closest in time.
[0060] Step S2, Feature alignment: Spatially align the input event points and their features through the T-Net module.
[0061] T-Net can be regarded as a small feature extraction network, which consists of a multi-layer perceptron, a pooling layer and a fully connected layer. The N×3 set of event points collected by the time window is passed into the T-Net network. T-Net performs one-dimensional convolution operations on the input data to obtain features with 64, 128, and 1024 channels respectively. Then, a sum pooling operation is performed on the extracted features to obtain a 1×1024 feature vector. Through the fully connected layer, the global feature is mapped to a 2×2 rotation matrix, and the generated rotation matrix is multiplied by the input feature to achieve feature alignment.
[0062] Step S3, Feature Extraction: To fully exploit the spatio-temporal characteristics in event data, the present invention employs a point-voxel attention module, which consists of a voxel branch and a point branch. Event points can well preserve position information in the point-based representation, but are prone to losing the geometric information of irregular event points; while in the voxel-based representation, they have regularity and good local memory ability. When extracting 3D voxel features, it is mainly divided into voxelization, feature aggregation, and de-voxelization.
[0063] In the voxelization stage, the input event points are mapped to a new set of voxel features V, and the features f of all points k are normalized according to the voxel grid cells (u, v, w) corresponding to their coordinates e k =(x k , y k , t k ), and the eigenvalue of this voxel unit is obtained:
[0064]
[0065] where r represents the voxel resolution, which determines the size of each voxel; I[·] is a binary indicator function used to determine whether the point e k belongs to the voxel unit (u, v, w); f k,c is the feature of the c-th channel corresponding to the event point e k ; N u,v,w is the normalization factor, representing the number of points falling into this voxel unit.
[0066] After completing the voxelization process of event points, the features of regular 3D voxels are extracted by introducing a convolutional attention module, and the structure of the convolutional attention module is as Figure 2 shown. In the convolutional attention module, the input is first passed into the channel attention module. The channel attention module will perform max-pooling and average-pooling operations on the input respectively to calculate the average and maximum values of the input features in the last three dimensions, thereby generating two feature vectors of the size of the number of channels. Then, the max-pooling feature and the average-pooling feature obtained after pooling are sent into a shared multi-layer perceptron to learn the attention weights of each channel, and its formula is as follows:
[0067]
[0068] where, represents the channel attention function, is the weight matrix of the shared MLP layer, and r is the dimensionality reduction factor used to control the computational cost.
[0069] These weights are multiplied by each channel of the input feature V to obtain the attention-weighted channel feature map, as Figure 3 , and its formula is as follows:
[0070] V c = Sigmoid(M(V)) · V
[0071] Subsequently, the channel-wise features are fed into a spatial attention module to extract the key information at different positions in the space. The structure of the spatial attention module is as shown in Figure 4 shown. The spatial attention module will take the input V C and use a two-dimensional convolution to calculate the spatial attention weights, generating a spatial matrix The specific calculation formula is as follows:
[0072]
[0073] where is the spatial attention function. W i,j are the learnable parameters of the convolution kernel, and K is the size of the two-dimensional convolution kernel.
[0074] Similar to the channel attention module, the sigmoid function is also used to limit the spatial attention weights between 0 and 1, and then the features at each spatial position of the input feature V C are weighted and multiplied by the input feature V C . The formula is as follows:
[0075]
[0076] Then, the spatio-temporal features are further aggregated through a multi-layer perceptron layer. The specific formula is as follows:
[0077] V ′ = MLP(CAM(Conv3D(V)))
[0078] The voxel features to be extracted are obtained. To fuse the voxel features with the branch based on point feature transformation, the voxel features need to be remapped back to the event point features. Considering that using the simplest nearest neighbor interpolation may cause the points within the same voxel unit to share exactly the same features, thus reducing the distinctiveness of the features, the present invention adopts the trilinear interpolation method to map the voxel features back to the point features, ensuring that the feature values of each point are different.
[0079] When extracting the features of points, first, a one-dimensional convolution operation is performed on each point of the input N×3 event point set to extract the features of the points, then it is fed into a BN layer for batch normalization processing, and finally, the sigmoid function is used for non-linear activation. Finally, the features of the points are extracted. Although this method is simple, it can generate independent and differentiated feature expressions for each point.
[0080] Step S4, feature fusion: The two branches are fused together using local features and aggregated global context. The voxel-based feature transformation and the point-based feature transformation are efficiently fused together through addition operations. The fused feature is expressed as:
[0081] E ′ =E loaal +E global
[0082] Among them, E local is the voxel-based feature transformation, E global is a point-based feature transformation.
[0083] After three PVA modules and one MLP module, the spatial features output by all modules are concatenated into a matrix of size N × 1280. The geometric information is aggregated through the sum pooling operation to obtain a latent feature vector containing spatiotemporal information.
[0084] Step S5, event denoising: The fully connected layer is used as a binary classifier to process the spatiotemporal features and distinguish real events from noise.
[0085] The following is a comparative experiment. On the simulated dataset DVSCLEAN, the background activity filter (BAF), probabilistic undirected graph model (PUGM), event denoising convolutional neural network (EDnCNN), and multilayer perceptron denoising filter (MLPF) methods are used for comparative experiments with the present invention. The signal-to-noise ratio (SNR) is calculated as the denoising index to measure the denoising performance. A higher SNR represents better denoising performance. The formula is:
[0086]
[0087] Where M represents the number of real events and N represents the number of noises.
[0088] The experimental results are shown in Table 1, where the bold fonts represent the best results in each column, and the underlined fonts represent the suboptimal results in each column.
[0089] Table 1: Comparison of experimental results of various methods
[0090]
[0091] As can be seen from Table 1, the present invention is optimal in terms of 50% noise ratio, 100% noise ratio, and the average score. The EDnCNN method is comparable to the present invention at a 50% noise ratio, but the EDnCNN performs poorly at a 100% noise ratio. Therefore, the present invention is superior. It can be known from Table 1 that BAF and MLPF are comparable to the present invention, but Figure 5 it can be seen that these methods have the phenomenon of removing real events as noise. Therefore, the present invention has the best effect on DVSCLEAN.
[0092] Comparative experiments were conducted on a real dataset, Figure 6 showing the visualization results. Figure 6 In (a), it shows a simple indoor scene, (b) shows a complex indoor scene, and (c) shows a complex outdoor scene. Since real events have no labels, the polarity of the event stream is used for coloring. It can be seen from the results that in indoor scenes, except for PUGM, the other methods can effectively remove noise. However, in outdoor scenes, the differences are obvious. The other several methods cannot completely remove noise. Therefore, the present invention has a better denoising effect on real data.
[0093] Five frames of event streams in the exposure scene were selected from the DVSNOISE20 dataset for denoising effect testing. The results are as Figure 7 shown. Figure 7 In (a), the original image is a chessboard, and in (b), the original image is a bicycle. It can be clearly seen that the denoising effects of EDnCNN and PUGM are very poor. Although BAF and MLPF can effectively remove noise, there are still problems in distinguishing noise from real events. Therefore, the present invention has the best effect on the DVSNOISE20 dataset.
[0094] The technical process of the event stream denoising method based on the point-voxel attention mechanism disclosed in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination.
[0095] When implemented using hardware, the above embodiments can, in whole or in part, run the working logic and calculation process on an electronic device after being compiled by software. The electronic device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the technical process of the event stream denoising method based on the point-voxel attention mechanism disclosed in the above embodiments.
[0096] When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. If the above method is implemented in the form of software functional modules and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the related art can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination among hardware, software, and firmware.
[0097] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes may be made in its form and details without departing from the spirit and scope of the present invention as defined by the appended claims.
Claims
1. An event stream denoising method based on a point voxel attention mechanism, characterized in that Including: Collecting data of a number of event points within a predetermined time range to generate an event point set; Performing spatial alignment on all the event point data and their input features in the event point set; Constructing a point-voxel attention module, extracting local features of each event point data in the voxel domain through its voxel branch, and extracting global features of each event point data through its point branch; Fusing the local features and all features to obtain output spatial features; Aggregating geometric information of the output spatial features through a sum pooling operation to obtain a latent feature vector containing spatio-temporal information; Distinguishing real events from noise in the latent feature vector containing spatio-temporal information.
2. The event stream denoising method based on the point voxel attention mechanism according to claim 1, characterized in that, Collecting data of a number of event points within a predetermined time range to generate an event point set, specifically including: Calculating the time window range according to the event stream; Within the determined time window range, collecting the N events that are closest in time among the spatial neighboring events of the current event; Filtering the N events based on time correlation, retaining the event point data that meets the predetermined time correlation requirements, and forming an event point set.
3. The event stream denoising method based on a point voxel attention mechanism according to claim 1, wherein Performing spatial alignment on all the event point data and their input features in the event point set, specifically including: Feeding the event point set collected by the time window into the T-Net network; The T-Net network sequentially passes the input event point set through three convolutional layers, gradually elevating it to the first channel dimension, the second channel dimension, and the third channel dimension; where the third channel dimension > the second channel dimension > the first channel dimension; Aggregating the extracted features into global features through sum pooling; Converting the global features into vector form and inputting them into a fully connected layer to obtain a transformation matrix of a predetermined size; Multiplying the transformation matrix by the input event point set to align the event points spatially.
4. A method for denoising event streams based on a point voxel attention mechanism according to claim 1, characterized in that, Passing the sampled event point set through a point-voxel attention module for feature extraction. In the point-voxel attention module: First, voxelize the input event point set; then perform three-dimensional convolution operations on the voxelized events, and then feed them into the convolutional attention module for feature extraction.
5. A method for denoising event streams based on a point voxel attention mechanism according to claim 4, characterized in that, In the convolutional attention module, first, the input feature V is passed into the channel attention module. The channel attention module performs max-pooling and average-pooling operations on the input feature V respectively, calculates the average value and the maximum value of the input feature V in the last three dimensions, thereby generating two feature vectors of the size of the number of channels. Then, the max-pooled feature V max and the average-pooled feature V avg are sent to a shared multi-layer perceptron to learn the attention weights of each channel. The output channel attention feature is obtained by adding them and passing through the sigmoid function. Finally, the input feature V is multiplied by the obtained channel attention feature to get the channel-oriented feature V C .
6. The event stream denoising method based on the point voxel attention mechanism according to claim 5, characterized in that, Also including: Input the feature V in the channel direction C into a spatial attention module to extract key information at different positions in space; The spatial attention module takes the input V C performs two-dimensional convolution and uses the Sigmoid activation function to generate channel attention weights, generating a two-dimensional spatial attention feature; Multiply the attention weights with each channel of the input feature V C to aggregate the spatio-temporal features through a multi-layer perceptron layer, obtaining the voxel features to be extracted.
7. The event stream denoising method based on the point voxel attention mechanism according to claim 6, wherein, Using trilinear interpolation method to map voxel features back to point features, specifically including: Performing one-dimensional convolution operation on the input event point set to extract point features, feeding them into a BN layer for batch normalization processing, and using a sigmoid function for non-linear activation to extract point features.
8. A denoising method for event streams based on a point voxel attention mechanism according to claim 1 or 7, characterized in that Fusing the local features and all features to obtain output spatial features, specifically including: Adding the voxel features and the point features, and feeding them into three PVA modules and one MLP module to obtain a feature matrix of N×1280; Aggregating geometric information through sum pooling operation in the N-dimensional direction to obtain a latent feature vector containing spatio-temporal information.
9. An event stream denoising system that can automatically execute the event stream denoising method based on the point voxel attention mechanism according to any one of claims 1 to 8, characterized in that, The event stream denoising system includes: An event point sampling module for collecting data of a number of event points within a predetermined time range to generate an event point set; A feature alignment module, which embeds a T-Net module and uses the T-Net module to perform spatial alignment on all the event point data and their input features in the event point set; A feature extraction module, which includes three point-voxel attention modules and an MLP module; the local features of each event point data in the voxel domain are extracted through the voxel branch of the point-voxel attention module, and the global features of each event point data are extracted through the point branch of the point-voxel attention module; A feature fusion module, which obtains the output spatial features after passing through three point-voxel attention modules and an MLP module, and aggregates the geometric information of the output spatial features through a sum pooling operation to obtain a latent feature vector containing spatio-temporal information; An event denoising module, which is used to distinguish real events from noise in the latent feature vector containing spatio-temporal information.
10. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the event stream denoising method based on the point-voxel attention mechanism as described in any one of claims 1 to 8.