An image semantic segmentation method based on event camera
By using SwinTransformer, event image attention module and graph inference module in the image semantic segmentation method based on event camera, the problem of uneven distribution of event frame image data is solved, and more efficient feature extraction and contextual relationship capture is achieved, and segmentation accuracy is improved.
Patent Information
- Application Number
- CN202310606274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-05-26
AI Technical Summary
The semantic segmentation method based on the event camera has low segmentation accuracy in complex scenarios, mainly due to the uneven distribution of the image data of the event frame, resulting in the lack of sufficient detailed information in local features, making it difficult to capture long-distance contextual relationships.
SwinTransformer is used as a feature extraction network, combining the event frame image attention module and graph reasoning module, event features are extracted from a global perspective, and the attention mechanism is used to enhance the global connection of each point in the image, and a position-independent attention diagonal matrix is introduced through the improved Laplace formula to capture long-distance dependencies.
It effectively improves the performance of using event data for image semantic segmentation, enhances the network's expression ability, provides more accurate semantic clues, and improves the accuracy of segmentation prediction.
Smart Images

Figure CN116597144B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular to a method for deep learning and semantic segmentation using event data output by an event camera, specifically an image semantic segmentation method based on an event camera. Background Art
[0002] Image semantic segmentation is an important topic in computer vision. It is to classify all pixels in an input image and assign different colors to different categories so that the distribution and contour information of different categories in the image can be intuitively seen. Methods based on convolutional neural networks (CNNs) have shown excellent performance in this field. Most methods rely on traditional frame images (RGB images or grayscale images) to complete semantic segmentation. However, the segmentation accuracy of semantic segmentation algorithms based on traditional frame images will drop sharply in complex scenes (e.g., low light, fast motion, etc.).
[0003] The event camera is a novel bionic visual sensor that asynchronously captures changes in light intensity in a scene and then outputs events. The event camera provides very high temporal resolution (up to 1MHz) and consumes very little power. Since changes in light intensity are calculated in a logarithmic scale, it can operate at a very high dynamic range (140dB). Therefore, the event camera can still maintain image clarity in complex scenes such as extremely dim light, overexposure, and sudden changes in light. Compared with traditional cameras, event cameras do not produce motion blur, which has great advantages in capturing high-speed moving objects. At present, most semantic segmentation methods based on event data still use CNN to extract local features for segmentation prediction. However, the data distribution of event frame images is not uniform, and some areas are relatively sparse, which also causes a large gap in detail information in different areas. If local features are extracted, problems such as the lack of sufficient detail information in the extracted features may be faced, which in turn poses challenges to subsequent segmentation predictions. Therefore, the present invention further extracts event features from a global perspective, strengthens the global connectivity of each point in the image, and also plays a role in filtering event frame image noise; and through the improved graph Laplace formula, introduces a position-independent attention diagonal matrix, which can better capture long-distance dependencies, thereby better extracting high-level event features and effectively improving the performance of image semantic segmentation using event data. Next, the relevant background technology in this field is introduced in detail.
[0004] (1) Image semantic segmentation based on traditional RGB camera
[0005] With the emergence of Fully Convolutional Networks (FCN), semantic segmentation algorithms based on deep learning have shown superior feature extraction capabilities than traditional segmentation methods and have become the mainstream method in the field of semantic segmentation. Compared with traditional image segmentation methods, fully convolutional neural networks can extract high-level semantic information of images and improve the segmentation accuracy of images. After that, the U-Net network proposed the concept of "encoding-decoding", which obtains low-resolution feature maps through convolution operations, and then upsamples the feature maps obtained by the convolution layer to restore them to the input size. Moreover, the feature maps obtained by each convolution layer will be cascaded to the corresponding upsampling layer, thereby providing more sufficient semantic information. DeepLab and DeepLab v2 proposed the concept of dilated convolution, which expands the receptive field through the dilated convolution algorithm without reducing the resolution of the feature map. At the same time, a dilated spatial pyramid pooling module was proposed to improve the performance of semantic segmentation. DeepLab v3 improved the dilated spatial pyramid pooling module, combining feature maps of different receptive fields generated by different dilated convolutions, thereby obtaining richer contextual information.
[0006] Since convolutional neural networks can only obtain the relationship between pixels in a local area of the size of the convolution kernel through convolution operations, and cannot directly model the relationship between pixels over long distances, this inherent defect has also become a bottleneck affecting the performance of semantic segmentation. As ViT introduced the Transformer model in natural language processing into the field of computer vision, in recent years, there have also been studies that have attempted to apply the Transformer model to the field of semantic segmentation, using the Transformer's self-attention mechanism to establish long-distance dependencies. SETR (SETR: Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspe) replaced the encoder's CNN with a Transformer model for the first time to solve the image semantic segmentation task, that is, directly using the Transformer model to learn information features, so that the network can obtain global contextual relationships from the beginning, further improving the segmentation accuracy. However, due to the large resolution of the image, the model will require extremely high computing power, resulting in a large overhead.
[0007] (2) Image semantic segmentation based on event camera
[0008] Ev-SegNet is the first algorithm to use event cameras for image semantic segmentation tasks. It improves the Xception network to solve image semantic segmentation tasks and proves the feasibility and effectiveness of using event cameras to solve image semantic segmentation tasks. VID2E improves on Ev-SegNet by using synthetic events converted from videos to enhance training data, supplementing event data to a certain extent, and using jump connections to help optimize the original network depth architecture. However, the lack of large-scale and accurate semantic segmentation datasets with pixel-level annotations is an important factor limiting the development of this task. Therefore, most of the subsequent research on image semantic segmentation algorithms based on event cameras is based on unsupervised and self-supervised training. Rebecq et al. (Video to Events: Recycling Video Datasets for Event Cameras) attempted to complete the semantic segmentation task by migrating event data to RGB images. The use of unsupervised domain adaptation from the labeled image source domain to the unlabeled target event domain can solve the problem of lack of labeled data to a certain extent. Hu et al. (EvDistill: Asynchronous Events to End-task Learning via BidirectionalReconstruction-guided Crossmodal Knowledge Distillation) proposed a network grafting algorithm, which replaced the front-end network of the pre-trained deep network that processes event frame images with a new grafted network that can process event data, and maximized the feature similarity between the pre-trained network and the grafted network by synchronously recording intensity frames and event data through self-supervised training. However, this type of algorithm is greatly affected by data distribution and its robustness needs to be improved. Summary of the invention
[0009] To solve the above problems, the present invention provides an image semantic segmentation method based on event camera. This method aims at the sparse local data of event frame images, strengthens the global correlation of each point in the image through the attention mechanism, and captures the long-distance contextual relationship through the graph reasoning strategy, so as to better achieve the semantic segmentation task.
[0010] The technical solution of the present invention:
[0011] An image semantic segmentation method based on event camera, the steps are as follows:
[0012] (1) Feature extraction module
[0013] The event image feature extraction module is used to extract features from event images stacked with event data. The event data is expressed using the following formula:
[0014] Among them, (x k ,y k ) is the pixel coordinate of the event, t k is the timestamp of the event, p k = ±1 is the polarity of the event; in order to input the captured asynchronous event data into the feature extraction module, firstly, the event data within the same time interval are superimposed at a fixed time interval to form an event frame image, and then the event frame image is sent to the feature extraction module; SwinTransformer is used as the feature extraction network, and the feature extraction module outputs four stages of features, the low-level features of the first stage carry rich detail information, and the high-level features of the remaining stages carry rich semantic information;
[0015] (2) Event Frame Image Attention Module
[0016] In order to better utilize the information carried by event frame images, the present invention designs an event frame image attention module to extract event features from a global perspective and enhance the relevance of context. Each point in the image can better perceive the surrounding semantic information and filter noise at the same time.
[0017] For the first-stage event features output by the feature extraction module First, pass them through a 3×3 convolution layer to obtain two features of the same dimension, and then merge the last two dimensions of the two features to reshape them into Where n = h × w; Next, E 1 After transposition, 2 Multiply, and the result is passed through the Softmax function to obtain the feature At this point, we get an attention feature map that establishes the global spatial connection of each pixel. The specific process is as follows:
[0018] E 1 =Reshape(Conv3(E))(2)
[0019] E 2 =Reshape)Conv3(E))(3)
[0020]
[0021] Among them, Conv3 represents 3×3 convolution, Reshape represents reshaping operation, and Softmax represents Softmax function;
[0022] Next, the first-stage event feature E is passed through a 3×3 convolutional layer and reshaped into E 1 、E 2 Features of the same dimension Then with E ′Multiply them together, reshape the result into the dimension of the first-stage event feature E, perform element-wise multiplication with a learnable parameter α, and then add it to the first-stage event feature E. At this time, the feature vector calculated by self-attention is obtained. The specific calculation process is as follows:
[0023] E 3 =Reshape(Conv3(E))(5)
[0024] E″=α·(Reshape(E 3 E′))+E(6)
[0025] Among them, α is a learnable parameter with a value of 0 to 1 and is initialized to 0; α will gradually increase from 0, aiming to gradually add the attention mechanism to the model.
[0026] Finally, E″ is subjected to an adaptive average pooling operation in the spatial dimension, and then passes through a 1×1 convolution layer, and then activated by a Sigmoid activation function, and is again element-wise multiplied with the input first-stage event feature E to correct the features containing contextual information and retain valuable features to obtain the final output result. The process is expressed as:
[0027] M=E·σ(Conv1(AvgPool(E″)))(7)
[0028] Among them, Conv1 represents 1×1 convolution, AvgPool represents adaptive average pooling operation, and σ represents Sigmoid activation function;
[0029] (3) Graph Reasoning Module
[0030] The receptive field of the convolution operation is limited. Only by stacking multiple layers into a deep model can the convolutional network have the ability to aggregate rich global context information. However, the superposition of local clues cannot accurately handle long-distance contextual relationships, especially for semantic segmentation tasks based on event data. Performing long-distance interactions is an important factor in reasoning in complex scenes. Since graph-based propagation has the advantage of storing clear semantic reasoning in a graph structure, and graph propagation has the ability to capture global information, the present invention introduces graph reasoning into the semantic segmentation task and proposes a graph reasoning module. Through an improved Laplace formula, a diagonal matrix of a position-independent attention mechanism is introduced into the inner product, which can achieve a better distance metric. And by performing graph reasoning in the original feature space, long-distance contextual relationships can be further captured, thereby better extracting high-level event features.
[0031] Graph convolution is a simulation of convolution operation on graph structured data. Given a graph G = (V, E) and its adjacency matrix A and degree matrix D, the normalized graph Laplacian matrix can be expressed as:
[0032]
[0033] Where I is the identity matrix, and the layer-by-layer propagation rule in the multi-layer graph convolutional network (GCN) is expressed as:
[0034]
[0035] Among them, H (l) is the vertex feature of the lth layer, Θ (l) is the trainable weight matrix of layer l, and σ is the nonlinear activation function;
[0036] Apply the propagation rule in the above formula to the convolutional neural network (CNN) features, that is, The only difference between the GCN layer and the convolutional layer is that the Laplacian matrix is multiplied on the left side of the feature vector. In order to better capture the spatial structure in the event frame image, the present invention designs an improved Laplacian matrix Ensure that the long-range contextual relationships to be learned depend on the input features rather than being restricted to specific features. It can be expressed as:
[0037]
[0038] in, diag represents diagonal operation, Represents the similarity matrix, dimension n = H × W;
[0039] For the similarity matrix, the similarity between position i and position j is expressed as:
[0040]
[0041] Among them, for the input features First, a 1×1 convolutional layer is used to reduce the dimension and obtain a feature vector with a channel number of C. Then, a reshaping operation is performed to obtain feature vectors and C is set to 64;
[0042] Next, the feature vector X is globally averaged and then reduced in dimension through a 1×1 convolutional layer to obtain a feature vector with C channels. Next, the feature vector is obtained through a diagonal operation. Finally, the three eigenvectors are multiplied to obtain the similarity matrix
[0043] The entire graph reasoning module is expressed as:
[0044]
[0045] Among them, Y represents the feature vector output by the module, τ represents the ReLu activation function, and Θ represents the trainable weight matrix;
[0046] The high-level features of the three stages obtained by the feature extraction module are sent to the graph reasoning module respectively. The feature vectors passing through the graph reasoning module carry richer long-distance context dependencies, which can provide more accurate semantic clues for the final segmentation prediction.
[0047] (4) Full sensor module
[0048] The full perceptron module only contains a multi-layer perceptron (MLP), which can fuse features at different levels through linear operations. It is simple and efficient in design, and can retain most feature information while reducing the complexity of the decoder and the amount of computation.
[0049] Beneficial effects of the present invention:
[0050] Due to the uneven distribution of event data, the detail information provided by different areas of the event frame image is quite different. The local features extracted in areas with sparse event data may not have sufficient semantic and detail information. In order to effectively extract event frame image features, the present invention extracts event features from a global perspective, and uses the attention mechanism to establish rich global contextual relationships for each point in the image, thereby enhancing the expressive power of the network; then, the contextual dependencies of high-level features are strengthened through the graph reasoning module, providing more accurate semantic clues for subsequent segmentation predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is the network structure diagram of the image semantic segmentation algorithm based on event camera. DETAILED DESCRIPTION
[0052] The present invention will be further described in detail below in conjunction with specific implementation modes, but the present invention is not limited to the specific implementation modes.
[0053] The present invention uses the public DDD17-Semantic dataset for experiments, with a training set and a test set ratio of 7:3. The dataset is collected using a DAVIS dynamic visual sensor (with a resolution of 346×260 pixels), and the results of grayscale image prediction by the model trained in the public dataset Cityscapes dataset are used as semantic labels. A total of 9440 images are collected, including 6 semantic categories. In order to train this network, the learning rate is set to 0.00006, the weight decay is set to 0.01, and the learning strategy uses a poly strategy with a power value of 1.0 to update the learning rate in real time. The number of training images per batch is 4, the number of iterations is set to 50, and the loss function uses the cross entropy loss function (Cross Entropy Loss). In the test phase, the input event frame image is sequentially passed through the feature extraction module, the event frame image attention module, and the graph reasoning module, and finally the segmentation prediction result of the entire image is output through the full perceptron module.
Claims
1. An image semantic segmentation method based on event camera, It is characterized in that Here are the steps: (1) Feature extraction module The event image feature extraction module is used to extract features from event images stacked with event data. The event data is expressed using the following formula: Among them, (x k ,y k ) is the pixel coordinate of the event, t k is the timestamp of the event, p k = ±1 is the polarity of the event; in order to input the captured asynchronous event data into the feature extraction module, firstly, the event data within the same time interval are superimposed at a fixed time interval to form an event frame image, and then the event frame image is sent to the feature extraction module; SwinTransformer is used as the feature extraction network, and the feature extraction module outputs four stages of features, the low-level features of the first stage carry rich detail information, and the high-level features of the remaining stages carry rich semantic information; (2) Event Frame Image Attention Module Design an event frame image attention module to extract event features from a global perspective and enhance the relevance of context; For the first-stage event features output by the feature extraction module First, pass them through a 3×3 convolution layer to obtain two features of the same dimension, and then merge the last two dimensions of the two features to reshape them into Where n = h × w; Next, E 1 After transposition, 2 Multiply, and the result is passed through the Softmax function to obtain the feature At this point, we get an attention feature map that establishes the global spatial connection of each pixel. The specific process is as follows: AND 1 =Reshape(Conv3(E)) (2) AND 2 =Reshape(Conv3(E)) (3) Among them, Conv3 represents 3×3 convolution, Reshape represents reshaping operation, and Softmax represents Softmax function; Next, the first-stage event feature E is passed through a 3×3 convolutional layer and reshaped into E 1 、E 2 Features of the same dimension Then multiply it with E′, reshape the result into the dimension of the first-stage event feature E, perform element-wise multiplication with a learnable parameter α, and then add it to the first-stage event feature E. At this time, the feature vector calculated by self-attention is obtained. The specific calculation process is as follows: AND 3 =Reshape(Conv3(E)) (5) E″=α·(Reshape(E 3 E′))+E (6) Among them, α is a learnable parameter with a value of 0 to 1 and is initialized to 0; α will gradually increase from 0, aiming to gradually add the attention mechanism to the model; Finally, E″ is subjected to an adaptive average pooling operation in the spatial dimension, and then passes through a 1×1 convolution layer, and then activated by a Sigmoid activation function, and is again element-wise multiplied with the input first-stage event feature E to correct the features containing contextual information and retain valuable features to obtain the final output result. The process is expressed as: M=E·σ(Conv1(AvgPo0l(E″))) (7) Among them, Conv1 represents 1×1 convolution, AvgPool represents adaptive average pooling operation, and σ represents Sigmoid activation function; (3) Graph Reasoning Module Graph reasoning is introduced into the semantic segmentation task, and a graph reasoning module is proposed. Through an improved Laplace formula, a diagonal matrix of the position-independent attention mechanism is introduced into the inner product to achieve better distance measurement. And by performing graph reasoning in the original feature space, long-distance contextual relationships are further captured, thereby better extracting high-level event features. Graph convolution is a simulation of convolution operation on graph structured data; given a graph G = (V, E) and its adjacency matrix A and degree matrix D, the normalized graph Laplacian matrix is expressed as: Among them, I is the identity matrix, and the layer-by-layer propagation rule in the multi-layer graph convolutional network (GCN) is expressed as: Among them, H (l) is the vertex feature of the lth layer, Θ (l) is the trainable weight matrix of layer l, and σ is the nonlinear activation function; Apply the propagation rule in the above formula to the convolutional neural network (CNN) features, that is, Design an improved Laplace matrix Ensure that the long-range contextual relationships to be learned depend on the input features and are not restricted to specific features; It is expressed as: in, diag represents diagonal operation, Represents the similarity matrix, dimension n = H × W; For the similarity matrix, the similarity between position i and position j is expressed as: Among them, for the input features First, a 1×1 convolutional layer is used to reduce the dimension and obtain a feature vector with a channel number of C. Then, a reshaping operation is performed to obtain feature vectors and C is set to 64; Next, the feature vector X is globally averaged and then reduced in dimension through a 1×1 convolutional layer to obtain a feature vector with C channels. Next, the feature vector is obtained through a diagonal operation. Finally, the three eigenvectors are multiplied to obtain the similarity matrix The entire graph reasoning module is expressed as: Among them, Y represents the feature vector output by the module, τ represents the ReLu activation function, and Θ represents the trainable weight matrix; The high-level features of the three stages obtained by the feature extraction module are respectively sent to the graph reasoning module. The feature vectors after the graph reasoning module carry richer long-distance context dependencies, providing more accurate semantic clues for the final segmentation prediction; (4) Full sensor module The full perceptron module only contains multi-layer perceptrons, which fuse features at different levels through linear operations. It is simple and efficient in design, and reduces the complexity of the decoder and the amount of computation while retaining most feature information.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Moving target visual tracking method based on multi-source information fusion
CN112686928A