Semantic segmentation method, system and equipment and storage medium
The event data stream obtained by the event camera is converted into voxel grid data and combined with the event attention map and cross attention mechanism, solving the accuracy of semantic segmentation under extreme lighting and high-speed motion, achieving higher segmentation accuracy and dynamic object capture.
Patent Information
- Application Number
- CN202510513530.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
AI Technical Summary
The existing semantic segmentation methods perform poorly under extreme lighting conditions and when capturing high-speed moving objects, resulting in loss of details of some scenes and blurred motion, affecting accuracy and real-time, and lack of discovery of the rich temporal dynamics of event flow.
The event camera is used to obtain the event data flow, and by converting it into voxel grid data and setting different time ranges, combining the event attention map and cross attention mechanism for feature fusion, enhancing the model's attention in significantly changing areas, and using event sparseness and temporal dynamics to improve segmentation accuracy.
It improves the accuracy of semantic segmentation, enhances the adaptability to changes in different scenarios and captures dynamic objects, and ensures the effective utilization of event information.
Smart Images

Figure CN120431327A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of semantic segmentation, and specifically relates to a semantic segmentation method, system, device and storage medium. Background Art
[0002] Semantic segmentation is a core task in computer vision. Its goal is to classify each pixel in an image into a predefined semantic category, thereby achieving pixel-level scene understanding. It has widespread applications in areas such as autonomous driving and medical image analysis. While deep learning methods have made significant progress in semantic segmentation using traditional frame-based sensors, the success of these methods relies on excellent image quality. However, frame-based cameras often perform poorly in extreme lighting conditions and for capturing high-speed moving objects. This can lead to loss of scene details and motion blur, compromising semantic segmentation accuracy and real-time performance. To address the limitations of frame-based images, event cameras have been introduced for semantic segmentation. Event cameras are a novel and unique biomimetic vision sensor. Unlike traditional RGB cameras, they detect changes in light intensity at the pixel level and asynchronously output an event stream. Event sensors offer the advantage of a high dynamic range (>120 dB), far exceeding the 60 dB of frame-based cameras. This allows them to effectively perceive in extreme lighting conditions, such as low illumination and overexposure.
[0003] In related technologies, speech segmentation mainly uses events represented by a single time window, which lacks the exploration of the rich temporal dynamics of the event stream and does not take into account that the sparsity of event data can reflect the dynamically changing areas in the scene, thus making the semantic segmentation results inaccurate. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a semantic segmentation method, system, device and storage medium, which utilize the temporal dynamics and sparsity of event data to improve the accuracy of semantic segmentation.
[0005] A semantic segmentation method, comprising:
[0006] Get event data stream and image data;
[0007] The event data flow is represented as:
[0008]
[0009] in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, ei =(x i ,y i ,t i ,p i ), represents the event data at a certain moment, i represents a specific event, and N is the total number of events;
[0010] Parsing the event data stream to obtain event data at all times;
[0011] Converting the event data into voxel grid data using a conversion formula;
[0012] The conversion formula is:
[0013]
[0014] Where σ is the Diracdelta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, x and y are variables used to represent all possible pixel positions on the entire image plane, and t represents the event time;
[0015] Setting a first time range and a second time range, wherein the first time range and the second time range do not overlap;
[0016] Obtaining first voxel grid data according to the first time range and the voxel grid data;
[0017] obtaining second voxel grid data according to the second time range and the voxel grid data;
[0018] performing pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fusing the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data;
[0019] Extract features from the fused data and the image data through a semantic segmentation network to obtain fused feature data and image feature data;
[0020] Get event attention map;
[0021] Calibrate the fused feature data according to the event attention map to obtain first event calibration data;
[0022] Performing channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and performing channel calibration and spatial calibration on the image feature data to obtain image calibration data;
[0023] Fusing the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data;
[0024] The cross-attention data is fed into the decoder to obtain the segmentation result.
[0025] Optionally, obtaining the event attention map includes:
[0026] Get the number of times an event occurs at each pixel;
[0027] According to the number of events and the statistical formula, the cumulative number of times for each pixel is obtained;
[0028] The statistical formula is:
[0029]
[0030] Where σ is the Diracdelta function, x and y are variables used to represent all possible pixel positions on the entire image plane, i is each specific event, N is the total number of events, C is the event count map, (x i ,y i ) are the horizontal and vertical coordinate values of the event data;
[0031] The accumulated times are normalized to generate an event attention map.
[0032] Optionally, calibrating the fused feature data according to the event attention map to obtain first event calibration data includes:
[0033] According to the event attention map, obtaining the allocation weight;
[0034] Obtaining first event calibration data according to the assigned weight, the fused feature data, and a calibration formula;
[0035] The calibration formula is:
[0036] F′ in =F in ⊙C A +F in ;
[0037] Among them, F' in is the first event calibration data, F in To fuse feature data, C A is the event attention map.
[0038] Optionally, performing pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fusing the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data includes:
[0039] Using a pooling layer of a first preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively, to obtain first-scale pooled data and second-scale pooled data;
[0040] Using a pooling layer of a second preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively to obtain third-scale pooled data and fourth-scale pooled data;
[0041] splicing the first-scale pooled data and the second-scale pooled data to obtain first pooled voxel grid data;
[0042] splicing the third-scale pooled data and the fourth-scale pooled data to obtain second-pooled voxel grid data;
[0043] The first pooled voxel grid data and the second pooled voxel grid data are upsampled to obtain fused data.
[0044] Optionally, performing channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and performing channel calibration and spatial calibration on the image feature data to obtain image calibration data includes:
[0045] Perform channel calibration and spatial calibration on the event calibration data through a convolution calibration formula to obtain second event calibration data;
[0046] Perform channel calibration and spatial calibration on the image feature data using a convolution calibration formula to obtain image calibration data;
[0047] The convolution calibration formula is:
[0048] F out =SA(CA(F′ in ))+F′ in
[0049] I out =SA(CA(I in ))+I in
[0050] Among them, F out is the second event calibration data, I out is the image calibration data, CA is the channel calibration, SA is the spatial calibration, F'in is the first event calibration data, I in is the image feature data.
[0051] Optionally, fusing the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data includes:
[0052] Activate the second event calibration data and the image calibration data using a 1×1 convolutional layer to obtain activation event data and activation image data;
[0053] multiplying the activation event data and the activation image data element-wise to obtain event modal data and image modal data;
[0054] The event modal data and the image modal data are fused using a cross-attention mechanism to obtain cross-attention fused data.
[0055] Optionally, fusing the event modal data and the image modal data using a cross-attention mechanism to obtain cross-attention fused data includes:
[0056] Projecting the event modal data to obtain event values, event keys, and event queries;
[0057] Projecting the image modality data to obtain an image value, an image key, and an image query;
[0058] Performing a dot product on the event query and the event key to obtain an event dot product value;
[0059] Performing a dot product between the image query and the image key to obtain an image dot product value;
[0060] Normalizing the event dot product and the image dot product value respectively to obtain an event normalized value and an image normalized value;
[0061] Performing weighted fusion of the event normalization value and the image value to obtain an event weighted value, and performing weighted fusion of the image dot product value and the event value to obtain an image weighted value;
[0062] The event weighted value and the image weighted value are spliced using a cross-attention mechanism to obtain cross-attention fusion data.
[0063] A semantic segmentation system, comprising:
[0064] A first acquisition module is used to acquire event data stream and image data;
[0065] The event data flow is represented as:
[0066]
[0067] in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, e i =(x i ,y i ,t i ,p i ), represents the event data at a certain moment, i represents a specific event, and N is the total number of events;
[0068] A conversion module, configured to convert the event data into voxel grid data through a conversion formula;
[0069] The conversion formula is:
[0070]
[0071] Where σ is the Diracdelta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, and x and y are variables used to represent all possible pixel positions on the entire image plane;
[0072] a setting module, configured to set a first time range and a second time range, wherein the time range of the first time range is smaller than the time range of the second time range;
[0073] a data selection module, configured to obtain first voxel grid data and second voxel grid data according to the first time range, the second time range, and the voxel grid data;
[0074] a scale processing module, configured to perform pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fuse the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data;
[0075] An extraction module, configured to extract features from the fused data and the image data through a semantic segmentation network to obtain fused feature data and image feature data;
[0076] The second acquisition module is used to obtain the event attention map;
[0077] A first calibration module is used to calibrate the fused feature data according to the event attention map to obtain first event calibration data;
[0078] A second calibration module is used to perform channel calibration and spatial calibration on the event calibration data and the image feature data respectively to obtain second event calibration data and image calibration data;
[0079] a fusion module, configured to fuse the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data;
[0080] The segmentation module is used to input the cross-attention data into the decoder to obtain a segmentation result.
[0081] A terminal device includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a semantic segmentation method is adopted.
[0082] A computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, a semantic segmentation method is adopted.
[0083] The beneficial effects of the present invention are:
[0084] An event camera is used to obtain an event data stream, and the event data in the event data stream is converted into first voxel grid data and second voxel grid data of different time ranges, and the data are fused to obtain fused data. The fused data is calibrated according to the event attention map to obtain first event calibration data, and the event calibration data and image feature data are respectively subjected to channel calibration and spatial calibration to obtain second event calibration data and image calibration data; the second event calibration data and the image calibration data are fused using a cross-attention mechanism to obtain cross-attention fusion data, and the cross-attention data is sent to the decoder to obtain a segmentation result. Compared with traditional semantic segmentation methods, the present application first uses rich temporal dynamics to extract event representations that can adapt to changes in different scenes from the original event stream data, ensuring comprehensive dynamic object capture and effective utilization of event information, and secondly uses event sparsity to enhance the model's attention on significantly changed areas, i.e., the event attention map. Finally, the event features and image features are fused so that the obtained cross-attention fusion data contains richer semantic information, thereby improving the accuracy of semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 Schematic diagram of the overall structural framework of the semantic segmentation method of the present invention;
[0086] Figure 2 This is a schematic diagram of the results of quantitative evaluation of the present application and other methods;
[0087] Figure 3 A visual diagram comparing the present invention with other methods;
[0088] Figure 4 Schematic diagram of ablation experiment results of the present invention;
[0089] Figure 5 This is a schematic diagram of the qualitative analysis results of the dual time window values of the event representation module DT-EI in the present invention;
[0090] Figure 6 This is a schematic diagram of the visual qualitative analysis results of the present invention for different time ranges;
[0091] Figure 7 This is a visualization diagram of the event branch before and after correction of the present invention;
[0092] Figure 8 Schematic diagram of the comparison results of using different MiT backbones in the present invention. DETAILED DESCRIPTION
[0093] A semantic segmentation method, the present invention comprises:
[0094] S1. Obtain event data stream and image data.
[0095] The event data flow is represented as:
[0096]
[0097] in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, e i =(x i ,y i ,t i ,p i ), which represents the event data at a certain moment, i represents a specific event, and N is the total number of events.
[0098] Specifically, the event data stream is obtained through an event camera. In the event camera, the event e is the basic representation unit of the brightness change in the scene, defined as a four-tuple e i =(x i ,y i ,t i ,p iThe polarity of the event, i.e., the direction of the brightness change. An event camera generates an event stream ε by asynchronously capturing brightness changes in the scene, which is defined as the set of all events in a period of time.
[0099] S2. Parse the event data stream to obtain event data at all times.
[0100] S3. Convert the event data into voxel grid data using a conversion formula.
[0101] The conversion formula is:
[0102]
[0103] Where σ is the Diracdelta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, and x and y are variables used to represent all possible pixel positions on the entire image plane.
[0104] Specifically, the event stream is usually converted into a frame-based voxel grid representation F, that is, within the time window [0, T], ε→F∈R B×H×W , where B represents the number of channels after the time dimension is discretized, (H, W) represent the height and width of the event camera imaging, and R is a real set.
[0105] S4. Set a first time range and a second time range, wherein the first time range and the second time range do not overlap.
[0106] Specifically, two time windows [0, T] are selected: a smaller time window T s and a larger time window T l The choice of T is based on the time range that best fits the motion of different objects in the dataset. Small-scale events capture fast-moving objects, while large-scale events capture slower-moving objects.
[0107] S5. Obtain first voxel grid data according to the first time range and the voxel grid data.
[0108] S6. Obtain second voxel grid data according to the second time range and the voxel grid data.
[0109] Specifically, the first voxel grid data is recorded as F in this embodiment. s , the second voxel grid data is recorded as F in this embodiment l .
[0110] S7. Perform pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fuse the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data.
[0111] The first voxel grid data and the second voxel grid data are pooled at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and the first pooled voxel grid data and the second pooled voxel grid data are fused to obtain fused data including:
[0112] S71 , using a pooling layer of a first preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively to obtain first-scale pooled data and second-scale pooled data.
[0113] Specifically, the first preset scale in this embodiment is a 3*3 average pooling layer.
[0114] S72 , using a pooling layer of a second preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively to obtain third-scale pooled data and fourth-scale pooled data.
[0115] Specifically, the second preset scale in this example is a 5*5 average pooling layer.
[0116] The specific processing process is expressed as:
[0117]
[0118] Among them, when i=1, it indicates that a pooling layer processing of 3*3 scale is performed, when i=2, it indicates that a pooling layer processing of 5*5 scale is performed, and pool indicates pooling processing.
[0119] S73 : Concatenate the first-scale pooled data and the second-scale pooled data to obtain first-scale pooled voxel grid data.
[0120] S74 , concatenating the third-scale pooled data and the fourth-scale pooled data to obtain second-pooled voxel grid data.
[0121] S75 . Upsample the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data.
[0122] Specifically, after performing pooling layer processing at different scales, four data are obtained. The data of the same scale are first spliced together to obtain the first pooled voxel grid data and the second pooled voxel grid data. At this time, the scales of the first pooled voxel grid data and the second pooled voxel grid data are different. Then, upsampling and matching the projection dimension are performed for fusion. Splicing and sampling are expressed as:
[0123]
[0124] in, Indicates fused data, Concat means concatenation, and UP means upsampling. represents the first scale pooled data, represents the second scale pooling data, represents the third scale pooled data, Represents the fourth scale pooled data.
[0125] The spatial resolution of the original vector {F s , F l}. In order to strengthen the key information, we use spatial attention to Filter and finally obtain the refined event representation F e :
[0126]
[0127] Among them, up is upsampling and SA is spatial calibration.
[0128] S8. Extract features from the fused data and image data through a semantic segmentation network to obtain fused feature data and image feature data.
[0129] S9. Obtain event attention map.
[0130] Obtaining the event attention map includes:
[0131] S91. Obtain the number of times an event occurs at each pixel.
[0132] S92. Obtain the cumulative number of times for each pixel according to the number of events and the statistical formula.
[0133] The statistical formula is:
[0134]
[0135] Where σ is the Diracdelta function, x and y are variables used to represent all possible pixel positions on the entire image plane, i is each specific event, N is the total number of events, C is the event count map, (x i ,y i ) are the horizontal and vertical coordinate values of the event data;
[0136] S93. Normalize the accumulated times to generate an event attention map.
[0137] Specifically, such as Figure 1As shown in the figure, SegFormer is used as the semantic segmentation backbone network in this embodiment. It is a hierarchical structure based on Transformer. There are four stages, which extract high-resolution shallow features to low-resolution fine features. The overall architecture is a two-stream network. Both the image branch and the event branch use SegFormer as the encoder to extract features, and the two branches calibrate and fuse after obtaining the intermediate features of each layer.
[0138] In the dual-branch semantic segmentation SegFormer backbone network encoder, the input of the EGRF module is the event and image feature tensor extracted at each stage of the backbone network {F in ,I in} and the corresponding event count attention map C A . C A The event count graph initially counts the number of times an event occurs on a pixel. The Map accumulates the number of events that occur at each pixel.
[0139] It is worth noting that there is not only one event attention map, which depends on how many times the event and image features are fused. In this embodiment, the event features and image features are set to be processed 4 times, so each fusion stage requires an event attention map area to adjust the weight.
[0140] The original event count map is processed by sigmoid normalization and maximum pooling with kernel sizes of 7×7, 5×5, 3×3 and 3×3 (corresponding to the four fusion stages respectively) to generate an attention map C that matches the output feature size and dimension of each stage of the encoder. A .
[0141] S10. Calibrate the fused feature data according to the event attention map to obtain first event calibration data.
[0142] The fused feature data is calibrated according to the event attention map, and the first event calibration data obtained includes:
[0143] S101. Obtain allocation weights based on the event attention graph.
[0144] Specifically, the representation of events within a certain time range will produce some noisy data that does not conform to the actual situation. To address this problem, we give full play to the characteristics and advantages of event cameras, use the sparsity of events to calculate the frequency of events at each pixel, generate attention maps, and enhance the model's attention to areas with key event features. Event Count Attention Map C ABy assigning weights to the event branch, it filters out noisy data and focuses on spatial regions with a higher number of event occurrences. These regions represent more reliable events, significantly reducing noise data and highlighting object edge features, which facilitates more accurate and detailed semantic segmentation of object boundaries and contours.
[0145] S102: Obtain first event calibration data according to the assigned weights, the fused feature data, and the calibration formula.
[0146] The calibration formula is:
[0147] F′ in =F in ⊙C A +F in .
[0148] Among them, F' in is the first event calibration data, F in To fuse feature data, C A is the event attention map.
[0149] S11 . Perform channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and perform channel calibration and spatial calibration on the image feature data to obtain image calibration data.
[0150] Performing channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and performing channel calibration and spatial calibration on the image feature data to obtain image calibration data includes:
[0151] S111, performing channel calibration and spatial calibration on the event calibration data using a convolution calibration formula to obtain second event calibration data;
[0152] S112 , performing channel calibration and spatial calibration on the image feature data using a convolution calibration formula to obtain image calibration data.
[0153] The convolution calibration formula is:
[0154] F out =SA(CA(F′ in ))+F′ in
[0155] I out =SA(CA(I in) )+I in
[0156] Among them, F out is the second event calibration data, I out is the image calibration data, CA is the channel calibration, SA is the spatial calibration, F' in is the first event calibration data, Iin is the image feature data.
[0157] Specifically, for event feature F' in and image features I in At the same time, we calibrate the channel dimension to enhance feature complementarity and reduce the impact of noise data in the two modalities. We use the channel attention block in CBAM to generate a fusion vector
[0158] Where C and D represent the dimensions of event features and image features, respectively. In the subsequent process, they can be split into vectors that conform to the event and image modalities to guide the calibration of the input feature map.
[0159] Next, in order to solve the spatial offset problem between event and image data due to differences in shooting angles and equipment, we conducted a cr ,I cr} calibrate the channel dimension to enhance the consistency of spatial features. We use global average pooling and global maximum pooling to fully capture feature information and accurately identify important spatial areas. Similarly, the fusion vector The two vectors are split to redistribute the weights for the two modalities.
[0160] Use a structure similar to the residual to add the calibrated vector to the vector elements before the input channel calibration module to obtain the output {F out ,I out}. The calibrated vector {F out ,I out} is fed into the next encoder layer to extract finer low-resolution features.
[0161] S12. The second event calibration data and the image calibration data are fused using a cross-attention mechanism to obtain cross-attention fused data.
[0162] The second event calibration data and the image calibration data are fused using the cross attention mechanism to obtain the cross attention fusion data including:
[0163] S121 : Activate the second event calibration data and the image calibration data using a 1×1 convolutional layer to obtain activation event data and activation image data.
[0164] S122. Multiply the activation event data and the activation image data element-wise to obtain event modal data and image modal data.
[0165] S123. Fuse the event modal data and the image modal data using a cross-attention mechanism to obtain cross-attention fused data.
[0166] Specifically, in the coarse-to-fine EGRF module, we use a cross-attention mechanism to fuse event and image features after modality calibration. First, a 1×1 convolutional layer is used to activate both modalities. To enhance feature representation, we multiply the activated event and image features element-wise to calculate global attention and add them to the respective modality features to obtain F. enh and I enh The activated event and image features are element-wise multiplied to compute the global attention, which is then added to the respective modality features.
[0167] Specifically expressed as:
[0168] F conv =Con 1×1 (F out ), I conv =Con 1×1 (I out )
[0169]
[0170] Among them, F conv To activate event data, I conv To activate image data, F enh is event modal data, I enh is image modality data.
[0171] The event modal data and image modal data are fused using the cross-attention mechanism, and the cross-attention fusion data obtained includes:
[0172] S1230. Project the event modal data to obtain event value, event key, and event query.
[0173] S1231. Project the image modality data to obtain image value, image key, and image query.
[0174] S1232. Perform a dot product on the event query and the event key to obtain an event dot product value.
[0175] S1233. Perform a dot product on the image query and the image key to obtain an image dot product value.
[0176] S1234 . Normalize the event dot product and the image dot product values respectively to obtain an event normalized value and an image normalized value.
[0177] S1235 . Perform weighted fusion of the event normalization value and the image value to obtain an event weighted value, and perform weighted fusion of the image dot product value and the event value to obtain an image weighted value.
[0178] S1236. Concatenate the event weighted value and the image weighted value using a cross-attention mechanism to obtain cross-attention fusion data.
[0179] Specifically, following the convention of multimodal feature fusion, we use the cross-attention mechanism to connect information from different sources to obtain a fused feature M, which helps to enhance the complementarity between the two modalities by dynamically adjusting the feature weights. The enhanced features are projected as values, keys, and queries respectively. In the calculation of cross-attention, we first obtain the similarity matrix within the two modalities by calculating the dot product of Q and K from the same source. Then, the normalized attention weight is applied to V of the other modality, dynamically transferring the feature information of one modality to the other modality, which is specifically expressed as:
[0180] Q=XW q , K=YW k , V=XW v
[0181] CA(X,Y)=ρ q (Q)(ρ k (K) T V)
[0182] In the cross-attention mechanism, Q is the query, K is the key, V is the value, W represents the learnable linear projection, CA represents the multi-head operation, and ρ q , ρ k is a normalized function of the query and key features, X represents one modality of the event or image, Y represents the other modality, and W q 、W k 、W v are the learnable parameters of the three linear layers, and T is the transposed matrix.
[0183] Depending on the number of fusions, multiple different fusion features can be obtained. That is, each time the image feature and the event feature are fused, a fusion feature M is obtained. In this embodiment, four fusions are performed, so the fusion feature is expressed as: M i (i∈1,2,3,4).
[0184] S13. Send the cross-attention data into the decoder to obtain the segmentation result.
[0185] Specifically, the lightweight MLP decoder of the semantic segmentation backbone network SegFormer is used. It first unifies the four feature vectors to the same dimension, upsamples them to the same size, and concatenates them. The MLP layer then transforms them into an H / 4xW / 4xC shape. Finally, the MLP layer classifies the pixels to achieve semantic segmentation. This lightweight decoder significantly reduces computational complexity while maintaining accuracy.
[0186] The validation was performed on two datasets, DDD17 and DSEC-Semantic. DDD17 (DAVISDrivingDataset 2017) is the first publicly released real-world dataset recorded by DAVIS driving. It uses a DAVIS sensor with a resolution of 346×260 pixels, which can simultaneously record standard APS (Active Pixel Sensor) images and DVS (Dynamic Vision Sensor) events. The DDD17 dataset uses a trained model to generate semantic pseudo labels for DAVIS grayscale frames. A subset of DDD17 was selected to provide them with semantic labels for semantic segmentation tasks. The dataset contains more than 12 hours of driving records, including 6 categories: ground, background, objects, vegetation, pedestrians and vehicles. 15,950 event image pairs were used for training and 3,890 pairs were used for testing.
[0187] Specifically, the EGRF module (fusion module) mentioned below is related to the use of event attention maps for calibration and the use of attention mechanisms for cross-fusion, and the DT-EI module is related to the use of the first time range, the second time range, and the voxel grid data to obtain the first voxel grid data and the second voxel grid data.
[0188] The DSEC-Semantic dataset is a semantic segmentation dataset generated based on the DSEC dataset. It contains semantic labels for 11 sequences (a total of 10,891 frames). The training set consists of eight sequences totaling 8,082 frames, and the test set consists of three sequences totaling 2,809 frames. It is acquired from the real world and has a resolution of 440×640. This dataset first warps images from the frame-based left camera to the left event camera's perspective. State-of-the-art semantic segmentation methods are then applied to the warped images to generate the final semantic labels. DSEC-Semantic provides two types of semantic labels; in this article, we use labels from 11 categories.
[0189] A mature semantic segmentation network was selected as the backbone of the dual-branch model. The pre-trained MiT-B0 was used for the event branch, and MiT-B2 was used for the image branch. The encoder consists of four stages, and the resulting feature vectors of input sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 are fed into a simple MLP decoder to produce segmentation results. Post-processing of both datasets was performed by cropping the bottom of the images to prevent noise. To capture rich event information across large and small time ranges, the dual time window of DDD17 in the DT-EI module was set to {50ms, 250ms}, while that of DSEC-Semantic was set to {30ms, 50ms}. The number of channels B in both models is 3. Event count graphs are generated from event data within the large event range.
[0190] Experiments were conducted using PyTorch on an NVIDIA RTX-4090D GPU, using the AdamW optimizer for cross-entropy loss. Following the EISNet setup, we used a batch size of 16 for DDD17, an initial learning rate of 0.0002, and trained for 80 epochs. We also used a batch size of 8 for DSEC, an initial learning rate of 0.00006, and also trained for 80 epochs.
[0191] The mean intersection over union (mIoU) and pixel accuracy are used as the main metrics for evaluating model performance. mIoU measures the overlap between the predicted segmented region and the ground-truth labeled region, while pixel accuracy measures the proportion of pixels correctly classified by the model to the total number of pixels.
[0192] Quantitative evaluation: This application is compared with the state-of-the-art methods for 2D semantic segmentation with single image modality and H containing event modality, including event-only models EV-SegNet, ESS, and event image multimodal methods HALSIE, CMX, CMNeXt, EISNet, EDCNET-S2D, etc. The experimental results are shown in the figure. Figure 2 shown.
[0193] On the DDD17 and DSEC-Semantic datasets, our method achieved mIoU scores that were 2.33% and 0.2% higher than the second-place EISNet, respectively. Furthermore, its pixel accuracy reached 96.3% and 95.24%, respectively, achieving the best performance among all methods. This demonstrates the effectiveness of our method. First, the DT-EI module's dual-timescale event feature extraction across large and small timescales compensates for the underutilization of event information. Second, the concept of extracting event count maps effectively guides the recalibration of event features. Third, the EGRF module's multimodal calibration and cross-attention mechanism successfully combine the complementary strengths of the two modalities, enhancing the fusion of features from different sources.
[0194] Qualitative assessment: Figure 3 This figure shows a visual comparison of our method with the baseline and ground truth. Compared to EISNet, our proposed method more accurately segments the "object" (yellow) and "pedestrian" (red) categories, which occupy a smaller pixel ratio. The boundaries between different categories and the edge contours of objects are more consistent with the ground truth. These visualization examples intuitively demonstrate that our method successfully combines the advantages of event and image modalities to more accurately distinguish pixels.
[0195] Ablation experiment: Figure 5 As shown in Figure 3, to verify the effectiveness of our method, we conducted ablation experiments on different input modalities and components.
[0196] The mIoU of single modality, i.e. image and event segmentation, is only 72.86% and 55.69%, respectively. This shows that semantic segmentation combined with multimodality effectively makes up for the shortcomings of single modality and emphasizes the necessity of adding event modality.
[0197] In terms of event representation, regardless of the way the two modalities are calibrated and fused, the dual time scale (i.e., long time range and short time range) using the DT-EI module outperforms the voxel grid generated by the fixed time interval by more than 1.14% in segmentation results, which proves the effectiveness of the DT-EI module. Figure 4 We can find that for variants 3, 4, 6, and 7, regardless of whether the event is represented as a voxel grid or the DT-EI module, the event count map guided event modality calibration improves mIoU by 0.89% and 1.33% respectively. Figure 4 Variants 5 and 8 in
[15] show that the cross-attention feature fusion design can better leverage the complementary effects between the two modalities than directly concatenating the features, achieving mIoU of 75.31% and 77.36% respectively. This demonstrates the importance and effectiveness of our EGRF module (calibration of the event attention map).
[0198] We conducted a qualitative analysis on the dual time window values of the event representation module DT-EI. Figure 5In the DSEC-Semantic dataset, when the dual time window values are shorter (15ms, 30ms), the mIoU and accuracy are 72.70% and 95.05%, respectively. This may be due to the limited number of events accumulated in the short time window or the insufficient span between T1 and T2. When the time window values are 30ms and 50ms, the number of events captured is more consistent with the movement of objects in the dataset, improving the mIoU by 0.57%. When the T2 value is further expanded to 70ms, the segmentation results of the model deteriorate. This shows that the tailing caused by the longer time interval misleads the learning performance of the model and affects the performance of the model.
[0199] We present a qualitative visualization of event voxel grids at both large and small time scales from the DDD17 dataset. Figure 6 The following are three different scenes, with a small time window of 50ms and a large time window of 250ms. In the first scene, the small time window generates fewer events, and while the object's shape can be seen, it's not distinct. However, the large time window triggers more events, clearly capturing the edge features of the vehicle and building. The second scene contains two objects with significantly different moving speeds. The vehicle in the lower left corner is moving faster, generating a large number of events within the small time window. Meanwhile, the building is moving slower than the vehicle, and fewer events are generated, forming a clear outline. In the large time window, the building is more clearly visible, but the vehicle becomes blurred. This requires our DT-EI module to adaptively learn advantageous features at two time scales for objects with different speeds. Scene three shows that the scene in front of the event camera is undergoing rapid lighting changes. The events in the small time window are already dense enough. If the time window is enlarged, severe smearing will occur. The DT-EI module also effectively overcomes smearing when fusing multi-scale features and learns event features that are consistent with real-world conditions.
[0200] Use the event count graph feature to calibrate event branches. Figure 7 The importance of event count guided recalibration in the EGRF module is illustrated. The event camera generates an event stream through light changes, which sometimes produces some noisy data. Figure 7 In the first scene, the vehicles and houses in front of the camera are stationary, with only two cars moving to the right, so the event voxel grid represents both cars. However, the event data contains some noise outside the cars. After event count-guided recalibration, the noise is nearly eliminated, revealing clear car outlines. The event features in scene two are more cluttered, but after weighting the event count attention map, unimportant areas are reduced and edge features are more prominent.
[0201] We use different MiT backbones for the encoders of event and image modality branches to evaluate the contribution of different backbones in segmentation results and computational efficiency. The results using the DDD17 dataset are reported in Figure 8 Variants 1 and 3 both use symmetric structures. While the MiTB0 encoder reduces computational complexity, the shallower structure does not fully extract features from both modalities. Conversely, while Variant 3 offers comparable accuracy, it is computationally complex. Notably, segmentation accuracy is highest with a shallower encoder for the event branch and a deeper encoder for the image branch, while computational complexity is also moderate.
[0202] A semantic segmentation system, comprising:
[0203] A first acquisition module is used to acquire event data stream and image data;
[0204] The event data flow is represented as:
[0205]
[0206] in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, e i =(x i ,y i ,t i ,p i ), represents the event data at a certain moment, i represents a specific event, and N is the total number of events;
[0207] A conversion module, used to convert event data into voxel grid data through a conversion formula;
[0208] The conversion formula is:
[0209]
[0210] Where σ is the Dirac delta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, and x and y are variables used to represent all possible pixel positions on the entire image plane.
[0211] a setting module for setting a first time range and a second time range, wherein the time range of the first time range is smaller than the time range of the second time range;
[0212] a data selection module, configured to obtain first voxel grid data and second voxel grid data according to the first time range, the second time range, and the voxel grid data;
[0213] a scale processing module, configured to perform pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fuse the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data;
[0214] An extraction module is used to extract features from the fused data and image data through a semantic segmentation network to obtain fused feature data and image feature data;
[0215] The second acquisition module is used to obtain the event attention map;
[0216] A first calibration module is used to calibrate the fused feature data according to the event attention map to obtain first event calibration data;
[0217] A second calibration module is used to perform channel calibration and spatial calibration on the event calibration data and the image feature data respectively to obtain second event calibration data and image calibration data;
[0218] a fusion module, configured to fuse the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data;
[0219] The segmentation module is used to feed the cross-attention data into the decoder to obtain the segmentation result.
[0220] An embodiment of the present application also discloses a terminal device, including a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a semantic segmentation method is adopted.
[0221] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer or a cloud server, and the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input and output devices, network access devices and buses, etc.
[0222] Among them, the processor can adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.
[0223] Among them, the memory can be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device, or it can be an external storage device of the terminal device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD) or flash memory card (FC) equipped on the terminal device, etc., and the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store computer programs and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or is to be output. This application does not impose any restrictions on this.
[0224] Among them, through this terminal device, a semantic segmentation method in the above embodiment is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device for easy use.
[0225] An embodiment of the present application further discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, a semantic segmentation method in the above embodiment is adopted.
[0226] Among them, the computer program can be stored in a computer-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The computer-readable medium includes any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that computer-readable medium includes but is not limited to the above-mentioned components.
[0227] Among them, through this computer-readable storage medium, a semantic segmentation method in the above embodiment is stored in a computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.
[0228] Those skilled in the art will appreciate that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of this application is limited to these examples. Within the context of this application, the technical features of the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more of the above embodiments of this application, which are not provided in detail for the sake of clarity.
[0229] The one or more embodiments of this application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application should be included in the scope of protection of this application.
Claims
1. A semantic segmentation method, characterized in that: include: Get event data stream and image data; The event data flow is represented as: in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, e i =(x i ,y i ,t i ,p i ), represents the event data at a certain moment, i represents a specific event, and N is the total number of events; Parsing the event data stream to obtain event data at all times; Converting the event data into voxel grid data using a conversion formula; The conversion formula is: Where σ is the Diracdelta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, x and y are variables used to represent all possible pixel positions on the entire image plane, and t represents the event time; Setting a first time range and a second time range, wherein the first time range and the second time range do not overlap; Obtaining first voxel grid data according to the first time range and the voxel grid data; obtaining second voxel grid data according to the second time range and the voxel grid data; performing pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fusing the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data; Extract features from the fused data and the image data through a semantic segmentation network to obtain fused feature data and image feature data; Get event attention map; Calibrate the fused feature data according to the event attention map to obtain first event calibration data; Performing channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and performing channel calibration and spatial calibration on the image feature data to obtain image calibration data; Fusing the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data; The cross-attention data is fed into the decoder to obtain the segmentation result.
2. A semantic segmentation method according to claim 1, characterized in that: The acquiring of the event attention map comprises: Get the number of times an event occurs at each pixel; According to the number of events and the statistical formula, the cumulative number of times for each pixel is obtained; The statistical formula is: Where σ is the Diracdelta function, x and y are variables used to represent all possible pixel positions on the entire image plane, i is each specific event, N is the total number of events, C is the event count map, (x i ,y i ) are the horizontal and vertical coordinate values of the event data; The accumulated times are normalized to generate an event attention map.
3. A semantic segmentation method according to claim 1, characterized in that: The calibrating the fused feature data according to the event attention map to obtain first event calibration data includes: According to the event attention map, obtaining the allocation weight; Obtaining first event calibration data according to the assigned weight, the fused feature data, and a calibration formula; The calibration formula is: F′ in =F in ☉C A +F in ; Among them, F' in is the first event calibration data, F in To fuse feature data, C A is the event attention map.
4. A semantic segmentation method according to claim 1, characterized in that: The performing pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fusing the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data includes: Using a pooling layer of a first preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively, to obtain first-scale pooled data and second-scale pooled data; Using a pooling layer of a second preset scale to process the first pooled voxel grid data and the second pooled voxel grid data respectively to obtain third-scale pooled data and fourth-scale pooled data; splicing the first-scale pooled data and the second-scale pooled data to obtain first pooled voxel grid data; splicing the third-scale pooled data and the fourth-scale pooled data to obtain second-pooled voxel grid data; The first pooled voxel grid data and the second pooled voxel grid data are upsampled to obtain fused data.
5. A semantic segmentation method according to claim 1, characterized in that: The performing channel calibration and spatial calibration on the first event calibration data to obtain second event calibration data, and performing channel calibration and spatial calibration on the image feature data to obtain image calibration data comprises: Perform channel calibration and spatial calibration on the event calibration data through a convolution calibration formula to obtain second event calibration data; Perform channel calibration and spatial calibration on the image feature data using a convolution calibration formula to obtain image calibration data; The convolution calibration formula is: F out =SA(CA(F′ in ))+F′ in F out =SA(CA(F in ))+F in Among them, F out is the second event calibration data, I out is the image calibration data, CA is the channel calibration, SA is the spatial calibration, F' in is the first event calibration data, I in is the image feature data.
6. A semantic segmentation method according to claim 1, characterized in that: The fusing the second event calibration data and the image calibration data using a cross attention mechanism to obtain cross attention fused data includes: Activate the second event calibration data and the image calibration data using a 1×1 convolutional layer to obtain activation event data and activation image data; multiplying the activation event data and the activation image data element-wise to obtain event modal data and image modal data; The event modal data and the image modal data are fused using a cross-attention mechanism to obtain cross-attention fused data.
7. A semantic segmentation method according to claim 6, characterized in that: The cross-attention fusion data obtained by fusing the event modality data and the image modality data using a cross-attention mechanism includes: Projecting the event modal data to obtain event values, event keys, and event queries; Projecting the image modality data to obtain an image value, an image key, and an image query; Performing a dot product on the event query and the event key to obtain an event dot product value; Performing a dot product between the image query and the image key to obtain an image dot product value; Normalizing the event dot product and the image dot product value respectively to obtain an event normalized value and an image normalized value; Performing weighted fusion of the event normalization value and the image value to obtain an event weighted value, and performing weighted fusion of the image dot product value and the event value to obtain an image weighted value; The event weighted value and the image weighted value are spliced using a cross-attention mechanism to obtain cross-attention fusion data.
8. A semantic segmentation system, characterized by: include: A first acquisition module is used to acquire event data stream and image data; The event data flow is represented as: in, and Indicates the pixel coordinates where the event occurs, t i ∈R + Indicates the timestamp of the event, p i ∈{-1,+1} represents the polarity of the event, e i =(x i ,y i ,t i ,p i ), represents the event data at a certain moment, i represents a specific event, and N is the total number of events; A conversion module, configured to convert the event data into voxel grid data through a conversion formula; The conversion formula is: Where σ is the Diracdelta function, is the scaled event timestamp with time buckets, t1 is the timestamp of the first event, B is the time bucket, t N is the timestamp of the Nth event, (x i ,y i ) are the horizontal and vertical coordinate values of the event data, and x and y are variables used to represent all possible pixel positions on the entire image plane; a setting module, configured to set a first time range and a second time range, wherein the time range of the first time range is smaller than the time range of the second time range; a data selection module, configured to obtain first voxel grid data and second voxel grid data according to the first time range, the second time range, and the voxel grid data; a scale processing module, configured to perform pooling processing on the first voxel grid data and the second voxel grid data at different scales to obtain first pooled voxel grid data and second pooled voxel grid data, and fuse the first pooled voxel grid data and the second pooled voxel grid data to obtain fused data; An extraction module, configured to extract features from the fused data and the image data through a semantic segmentation network to obtain fused feature data and image feature data; The second acquisition module is used to obtain the event attention map; A first calibration module is used to calibrate the fused feature data according to the event attention map to obtain first event calibration data; A second calibration module is configured to perform channel calibration and spatial calibration on the event calibration data and the image feature data, respectively, to obtain second event calibration data and image calibration data; a fusion module, configured to fuse the second event calibration data and the image calibration data using a cross-attention mechanism to obtain cross-attention fused data; The segmentation module is used to input the cross-attention data into the decoder to obtain a segmentation result.
9. A terminal device comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the segmentation method according to any one of claims 1 to 7 is adopted.
10. A computer-readable storage medium storing a computer program, wherein: When the computer program is loaded and executed by a processor, the segmentation method according to any one of claims 1 to 7 is adopted.