A multi-modal road scene object detection method based on images and events
By combining the multimodal data of traditional cameras and event cameras in object detection, and using efficient attention mechanisms and long-term memory mechanisms to fusion information, the problem of traditional cameras' detection difficulties in harsh scenarios and the poor effect of event single-modal detection when stationary objects is achieved, achieving stronger detection performance and robustness.
Patent Information
- Application Number
- CN202310296018.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-03-24
AI Technical Summary
Traditional cameras find it difficult to identify the target of interest in scenarios such as high-speed motion, overexposure and low-light. At the same time, the single-modal object detection algorithm based on event mode has poor detection effect when the object is stationary.
A multimodal road scene object detection method based on images and events is adopted, and data is obtained through traditional cameras and event cameras, image feature extraction and event data representation are performed. Combined with efficient attention mechanism and long-term memory mechanism, multimodal fusion is performed, and the final input detection head is used for object detection.
It improves the accuracy and robustness of target detection, can achieve good results in challenging scenarios such as high-speed motion, low light and overexposure, and enhances the practical value of the detection algorithm.
Smart Images

Figure CN116453014B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and event cameras, and in particular, to a multi-modal road scene object detection method based on images and events. Background Art
[0002] Object detection technology is one of the important components in the field of computer vision. It involves identifying objects of interest in an image or video, determining their positions, and classifying the objects. Object detection can be achieved through various technologies, including deep learning-based algorithms, traditional computer vision algorithms, and machine learning algorithms. Currently, deep learning-based detection algorithms have become the most popular methods because of their high recognition rate and the ability to easily handle relatively complex scenarios.
[0003] In deep learning-based object detection models, there are mainly two types: two-stage detectors and one-stage detectors. Existing two-stage detectors, such as R-CNN, Fast R-CNN, Faster R-CNN, and Mask R-CNN, etc., first generate a set of region proposal boxes, and then classify and refine them to identify the objects. One-stage detectors such as the YOLO series (You Only Look Once) and SSD (Single Shot Detector) directly generate the bounding boxes and class probabilities of all objects in the image in one step. Object detection has many practical applications, such as autonomous vehicles, security systems, robotics, and medical imaging and pathology analysis, etc.
[0004] The above-mentioned object detection algorithms are all designed based on traditional image sensors. These sensors are inexpensive, and the captured images can provide rich scene information. However, most traditional cameras require a certain exposure time, have a relatively low frame rate, and a small dynamic range. Therefore, it is difficult to capture effective information in scenarios such as extreme lighting conditions and high-speed movements, and they are no longer suitable for object detection tasks in harsh scenarios.
[0005] The emergence of event cameras perfectly solves the problems existing in the above traditional cameras. Different from the continuous image frames output by traditional cameras, event cameras output a series of continuous asynchronous event streams, which reflect the change of light brightness in a dynamic scene, equivalent to a traditional camera operating at thousands of frames per second. However, its data volume is much smaller, so the overall power of event cameras is smaller, the requirements for computer hardware are lower, and it has advantages such as high dynamic range, low power consumption, and microsecond-level latency, and is almost not affected by motion blur. These advantages enable event cameras to still capture rich information in some scenarios with poor lighting conditions and high-speed movements. Combined with advanced object detection algorithms, it can complete fast and accurate object recognition and positioning.
[0006] Since an event camera reflects the change in pixel brightness of a certain scene, when all the objects in the captured scene are stationary, the camera cannot generate enough events, that is, it cannot obtain enough scene information. Therefore, in order to achieve the complementary advantages of traditional cameras and event cameras, the data of these two modalities, images and events, are fused to achieve information complementarity, which can further improve the accuracy and robustness of target detection. Summary of the Invention
[0007] The present invention aims to provide a multi-modal road scene target detection method based on images and events, which solves the problem that it is difficult for traditional cameras to identify targets of interest in scenes such as high-speed movement, overexposure, and low light, and to a certain extent solves the problem that the detection effect of event-based single-modal target detection algorithms is poor when objects are stationary. This method includes technologies such as target detection, event representation, and multi-modal fusion, and can complete the target perception and recognition of road scenes for any images and events.
[0008] The technical solution of the present invention is as follows: A multi-modal road scene target detection method based on images and events, including the following steps:
[0009] Step 1: Continuously acquire color image data and event data through a traditional camera and an event camera respectively;
[0010] Step 2: Preprocess the color image data, and encode the information of the color image data through an image feature extraction module; the image feature extraction module includes three parts: a convolution module, a residual module, and a pooling module; the convolution module includes a convolutional layer with a stride of 2 and a convolutional kernel size of 3×3, a batch normalization layer, and a Leaky ReLU function; the residual module includes a convolutional layer and a skip connection structure; the pooling module includes 1 convolutional layer and 3 max-pooling layers with different sizes of convolutional kernels; each time the input color image data passes through a residual module, the color image data features are downsampled by a factor of 2; the image feature extraction module outputs color image data features downsampled by 8 times, 16 times, and 32 times respectively;
[0011] Input the color image data features downsampled by 32 times into an efficient attention mechanism module, and output the color image data features after temporal enhancement; the color image data features after temporal enhancement are respectively input into an image-event fusion module and a convolutional module. After being input into the convolutional module, bilinear interpolation is used for upsampling by a factor of 2, and then concatenated with the 16-times downsampled features output by the image feature extraction module along the channel direction, and then respectively input into the image-event fusion module and the convolutional module. After passing through the convolutional module, the concatenated features are upsampled by a factor of 2 and then concatenated with the 8-times downsampled features output by the image feature extraction module, and input into the image-event fusion module;
[0012] Step 3: Divide the event data along the time dimension into multiple groups of the same size. Perform rasterization on all events within each group and map them to a two-dimensional plane, representing the input event data as multi-channel event data features. Extract information from the event data features through a feature extraction module with the number of channels halved. The overall structure of the feature extraction module with the number of channels halved is the same as that of the feature extraction module, and the input channel numbers of its convolutional module, residual module, and pooling module are all half of the corresponding module channel numbers in the feature extraction module. Each time the input event data features pass through a residual module, the event data features are downsampled by a factor of 2, and the feature extraction module with the number of channels halved finally outputs event data features downsampled by a factor of 32.
[0013] Input the event data features downsampled by a factor of 32 into the long short-term memory mechanism module to output enhanced event data features. For the convenience of information fusion with image data features, perform upsampling by a factor of 4 on the enhanced event data features.
[0014] Step 4: Input the temporally enhanced color image data features and the enhanced event data features into the image-event fusion module. The image-event fusion module includes feature fusion modules with three different scales of input, and each feature fusion module is connected by a convolutional layer with a size of 1×1. In the feature fusion module, the temporally enhanced color image data features and the enhanced event data features are concatenated into one kind of feature, and the feature dimensions are shuffled along the channel direction. The shuffled features are evenly divided into two parts, which are respectively input into the spatial attention module and the channel attention module, and using a convolutional layer with a size of 3×3 and a sigmoid activation function, score maps of h×w×1 and 1×1×c are calculated respectively. Here, h and w represent the length and width of the score map respectively, and c represents the number of channels of the score map. These two score maps are respectively weighted to the features with shuffled channel dimensions to complete the information interaction between the color image data and the event data. The image-event fusion module finally outputs the features after fusing the two modalities.
[0015] Step 5: Input the fused features into the detection head. The detection head includes three branches: a classification branch, a regression branch, and an IoU branch, which are used for object classification, regression, and IoU calculation respectively. Each branch is mainly composed of multiple convolutional modules and activation function layers. The classification branch outputs features of size h×w×cls, where cls represents the number of object categories. The regression branch outputs features of size h×w×4, and the IoU branch outputs features of size h×w×1. The features output by the three branches are concatenated through feature concatenation to output features of h×w×(5 + cls). Finally, the method of non-maximum suppression is used to filter redundant bounding boxes, and the detection results are output.
[0016] The input of the efficient attention mechanism module includes the color image data features at time t and the historical color image data features at time t-1; the color image data features at time t-1 are respectively input into the convolutional layer The convolutional layer ψ serves as the key and value for feature transformation. The feature vectors output by the two are subjected to dot product and softmax calculations to obtain a feature map with a dimension of d×d; the color image data features at time t are used as the query volume, input into the convolutional layer θ for feature transformation, and the output feature vector is multiplied by the feature map to complete feature modulation, and finally added to the original color image data features at time t to output the temporally enhanced color image data features.
[0017] The input of the long short-term memory mechanism module consists of three parts, namely the event data feature F at time t t , the hidden state h at time t-1 t-1 and the cell state c t-1 ; F t and h t-1 are feature concatenated along the channel direction, passed through a separable convolutional layer of size 3×3, and respectively fed to the forget gate, input gate and output gate. Each gate is mainly composed of a sigmoid activation function; the forget gate controls the amount of information lost by the hidden state h t-1 , the input gate controls the amount of information added to the cell state c t , and the output gate controls the output amount of information of the hidden state h at time t t and the cell state c t ; the long short-term memory mechanism module completes the temporal information interaction of the event features at adjacent times and outputs the enhanced event data features.
[0018] In step 3, the event data is divided into 5 groups of the same size along the time dimension, and the input event data is represented as event data features with 5 channels.
[0019] Advantages of the present invention: It can be more widely applied to actual scenarios, has strong detection performance, and can also achieve good results for some challenging scenarios such as high-speed movement, low light, and overexposure. The present invention can improve the robustness of existing object detection algorithms and has certain practical value. Description of the Drawings
[0020] Figure 1 is the system block diagram of the multi-modal road scene object detection method based on images and events.
[0021] Figure 2 is the network structure diagram of the multi-modal road scene object detection method based on images and events.
[0022] Figure 3 is the corresponding diagram of event data and image data.
[0023] Figure 4 It is the network structure diagram of the efficient attention mechanism module.
[0024] Figure 5 It is the network structure diagram of the long short-term memory mechanism module.
[0025] Figure 6 It is the schematic diagram of the structure of the image-event feature fusion module. Specific implementation manners
[0026] The following further describes the specific implementation manners of the present invention in combination with the accompanying drawings and technical solutions.
[0027] The system block diagram of the multi-modal road scene object detection method based on images and events of the present invention is as Figure 1 shown, and the specific content is as follows:
[0028] (1) Data preparation:
[0029] First, with the help of a traditional image camera and an event camera, continuous color image data and event data are respectively obtained. Before data acquisition, the two cameras are time-synchronized, placed horizontally on the same reference plane, and a checkerboard is used as a calibration board to complete automatic calibration. As Figure 3 shown is the correspondence between the image data and the event data. The traditional camera outputs an image at time t-3, and after a certain exposure time, outputs another image at time t-2. During this exposure time, the event camera outputs a series of event data.
[0030] (2) Image preprocessing and feature extraction:
[0031] In the object detection task, the quality of the input image will directly affect the detection performance of the entire model. Usually, before the image is input into the network, image preprocessing is required. Common methods include grayscale transformation, geometric correction, image filtering, and image enhancement, etc. Among them, grayscale transformation is a point transformation, and the output value is related to the specified pixel point, while image filtering is a local processing method, and the pixel value of this point is determined by all the pixels in the neighborhood. The processed series of color image data is input into the feature extraction module to encode the key information in the picture. Since the image quality obtained by the camera shooting varies in different scenarios, and at the same time, too many vehicles in the road scene will cause occlusion of some objects. To solve the problems of image quality degradation and object occlusion, it is necessary to perform temporal feature fusion on adjacent image frames taken. The present invention introduces an efficient attention mechanism module to calculate the attention map of two adjacent image frames and weight it on the original image to enhance the expression ability of the features.
[0032] (3) Event representation and feature extraction:
[0033] Different from the image frames output by traditional cameras, the event output is a series of continuous and spatially discrete event streams, which have characteristics such as high temporal resolution, low latency, and low power consumption. Event data is asynchronous and sparse. As the spatial resolution of the camera increases, the number of events will increase exponentially. Directly inputting events into existing detection models will greatly increase the computational load. In addition, before processing event data, in order to more conveniently input the event stream into a deep convolutional neural network for operation, the event stream is usually transformed into different representations. The voxel grid representation method is a spatio-temporal histogram representation based on events. By rasterizing event data along the time dimension, it can effectively avoid compressing the timestamp information of the event stream, thus better retaining the temporal and spatial information of the event data. After completing the representation of the event, it is input into a feature extraction module with the number of channels halved to extract the key information in the event. According to the characteristic of the event camera that only reflects the brightness change of pixels, it is difficult to generate enough events in a static scene. In order to obtain more accurate detection results, it is necessary to perform temporal fusion on events in adjacent time periods. A multi-layer long short-term memory mechanism module is introduced to use the hidden state and cell state to transfer the features of events at different historical moments to enhance the representation ability of the features of the current moment event.
[0034] (4) Image-event feature fusion:
[0035] After the image data is preprocessed, it is input into a feature extraction module to complete the encoding of key information, and the efficient attention mechanism is used to perform temporal information fusion on images at adjacent moments, and output feature data of color images at multiple different scales. After the event data is represented, it is input into a feature extraction module with the number of channels halved for information extraction, and the long short-term memory mechanism module is used to aggregate the features of events at different historical moments, and output feature data of event data at multiple different scales. In order to complete the information interaction and complementarity between the image data and the event data, the present invention inputs the above two multi-scale features into an image-event fusion module. After shuffling the features along the channel dimension, the temporal attention mechanism and the spatial attention mechanism are used for deep feature fusion, and high-level features are output.
[0036] (5) Post-processing:
[0037] Whether it is a two-stage object detection algorithm or a one-stage object detection algorithm, a series of overlapping boxes will ultimately be generated near the same object. Usually, the non-maximum suppression method is used to eliminate redundant prediction boxes and screen out high-quality detection results. After a series of operations in the present invention, deep features are obtained, which are input into a detection head to complete feature decoding, and non-maximum suppression is used for redundant box filtering, and the detection results are output.
[0038] Figure 2 It is the network structure diagram of a multi-modal road scene object detection method based on images and events. The overall network input part is color image data and event data. First, the method of bilinear interpolation is used to scale the color image data proportionally. It is stipulated that the unified scaling scale of all images is 640×640, and at the same time, random horizontal and vertical flips are performed. Then, the processed picture is input into the network. First, it passes through a convolutional module, which is composed of a common convolutional layer, a batch normalization layer, and an activation function Leaky ReLU. Then, it passes through a series of residual modules, and the stride of each residual module is 2, and a 32-fold downsampled feature map is output. These residual modules are based on the convolutional module as the basic unit and combined with a simple residual structure. The obtained 32-fold downsampled feature map is input into the pooling module to complete the information aggregation between the channels of the features.
[0039] To complete the fusion of image information at adjacent times, a multi-layer efficient attention mechanism module is introduced. As Figure 4 shown, this module uses the image features at time t-1 as the key and value, and the image features at time t as the query volume. They respectively pass through convolutional layers convolutional layer ψ and convolutional layer θ for feature transformation. The feature vectors output by the key and value are multiplied point by point and softmax calculation is performed to obtain a feature map with a dimension of d×d. Then, it is multiplied point by point with the feature vector output by the query volume to complete feature modulation. Finally, it is added to the original feature at time t to output the enhanced color image data feature The entire image fusion process can be expressed by the following formula, where i represents the features at different scales, r is a constant, and c i represents the number of channels of the input feature.
[0040]
[0041]
[0042] For event data, first rasterize it. The specific operation is to evenly divide the events along their timestamps into a series of groups. Each group has the same spatial size, and there is a plane interval between adjacent groups. For all events within each group, map them to the corresponding plane respectively. Each position on the mapped plane represents the number of events at that place. After the event mapping is completed, a multi-channel event representation is obtained. Scale it proportionally to a size of 640×640, and after performing feature standardization operations, input it into a feature extraction module with the number of channels halved. The overall structure of this part of the network is generally the same as the network structure of the feature extraction module of the above image branch. However, the number of input and output feature channels of each module is 1 / 2 of that of the image branch. This is because events have a high time resolution, and a lightweight network structure is used for operation, which has a faster processing speed and is more in line with the characteristics of event data itself. In addition, different from the information fusion method of image data, event data undergoes 32-fold downsampling through multiple residual modules and is connected to a long short-term memory mechanism module after the pooling module. As Figure 5 shown, the input of the long short-term memory mechanism module consists of three parts, namely the event feature F t at time t, the hidden state h t-1 at time t-1, and the cell state c t-1 . Concatenate F t and h t-1 along the channel direction for feature splicing. After passing through a separable convolution of size 3×3, send them to the forget gate, input gate, and output gate respectively. Each gate is mainly composed of a sigmoid activation function. The forget gate controls the amount of information lost in the hidden state h t-1 , the input gate controls the amount of information added to the cell state c t , and the output gate controls the output amount of information of the hidden state h t and the cell state c t at time t. The long short-term memory mechanism module completes the temporal information interaction of the event features at adjacent times and outputs the enhanced event data features.
[0043] After the color image data and event data respectively complete feature extraction and fusion of temporal information, input the features of the above two modalities into the image-event fusion module. As Figure 6As shown in the figure, in order to make the channel dimensions of the color image data features and the event data features equal, the event data features are input into a module composed of a convolutional layer and a batch normalization layer to complete the dimensionality increase of the event data features. Then, the dimensionality-increased event data features and the color image data features are concatenated along the channel dimension, and then the feature dimensions of the concatenated features are shuffled along the channel direction. The shuffled features are evenly divided into two parts, which are respectively input into the spatial attention module and the channel attention module, and using a 3×3 convolution and a sigmoid activation function, score maps of h×w×1 and 1×1×c are respectively calculated, where h and w represent the size of the score map, and c represents the number of channels of the score map. These two score maps are respectively weighted to the input features to complete the information interaction between the color image data and the event data. The image-event fusion module outputs the features after fusing the two modalities.
[0044] The fused features are input into the detection head to complete feature decoding. The detection head contains 3 branches, which are respectively used for object classification, regression, and IoU calculation. Each branch is mainly composed of multiple convolutional modules and the sigmoid activation function. The classification branch outputs features of size h×w×cls, where cls represents the number of object categories. The regression branch outputs features of size h×w×4, and the IoU branch outputs features of size h×w×1. After feature concatenation, features of size h×w×(5 + cls) are output. Before decoding, a relatively low threshold is set to filter out the prediction boxes with low quality, and the prediction boxes are sorted by the classification scores, and the top 100 prediction boxes with the highest scores are selected. Then, the prediction boxes are decoded with the assigned anchor boxes and restored to the original image through scale scaling. The non-maximum suppression method is used to further filter redundant boxes. Based on the prediction box with the highest score, all the remaining boxes are traversed. If the IoU with the current highest-scoring box is greater than the set threshold, the prediction box is deleted. The above process is repeated continuously, and finally the detection results are output.
[0045] For the image-event multi-modal object detection network proposed by the present invention, the initial learning rate is set to 0.001, and it is trained for 300 epochs, and the learning rate gradually decays during the training process. During the training stage and the inference stage, the size of the input pictures of the network is 640×640. During inference, the confidence score threshold is 0.1, and the non-maximum suppression threshold is set to 0.5. During the experiment, the mAP metric is used to evaluate the network performance, and the experimental results are as follows:
[0046]
[0047] Experiment 1 and Experiment 2 respectively represent that the input data is only image data and event data, and there is no efficient attention mechanism, long short-term memory mechanism, and image-event fusion module in the network. The mAP values are only 0.291 and 0.328. Experiment 3 and Experiment 4 respectively introduce an efficient attention mechanism and a long short-term memory mechanism, and the mAP values reach 0.309 and 0.351. Experiment 5 shows that the input data contains image data and event data, and an image-event fusion module is introduced, and the mAP is further improved, reaching the highest 0.382.
Claims
1. A multi-modal road scene object detection method based on images and events, characterized in that, it includes the following steps: Step 1: Obtain continuous color image data and event data through a traditional camera and an event camera respectively; Step 2: Preprocess the color image data, and encode the information of the color image data through an image feature extraction module; The image feature extraction module includes three parts: a convolution module, a residual module, and a pooling module; The convolution module includes a convolutional layer with a stride of 2 and a convolutional kernel size of 3×3, a batch normalization layer, and a Leaky ReLU function; The residual module includes a convolutional layer and a skip connection structure; The pooling module includes 1 convolutional layer and 3 max pooling layers with different sizes of convolutional kernels; Each time the input color image data passes through a residual module, the color image data features are downsampled by 2 times; The image feature extraction module outputs color image data features downsampled by 8 times, 16 times, and 32 times respectively; Input the color image data features downsampled by 32 times into an efficient attention mechanism module, and output the color image data features after temporal enhancement; The color image data features after temporal enhancement are respectively input into an image-event fusion module and a convolutional module. After being input into the convolutional module, bilinear interpolation is used for upsampling by 2 times, and then it is concatenated with the 16-times downsampled features output by the image feature extraction module along the channel direction, and then respectively input into the image-event fusion module and the convolutional module. After passing through the convolutional module, the concatenated features are upsampled by 2 times and then concatenated with the 8-times downsampled features output by the image feature extraction module, and input into the image-event fusion module; Step 3: Divide the event data along the time dimension into multiple groups of the same size, perform rasterization operations on all events within each group and map them to a two-dimensional plane, and represent the input event data as multi-channel event data features; Extract information from the event data features through a feature extraction module with the number of channels halved; The overall structure of the feature extraction module with the number of channels halved is the same as the overall structure of the feature extraction module, and the input channel numbers of its convolution module, residual module, and pooling module are all half of the channel numbers of the corresponding modules in the feature extraction module; Each time the input event data features pass through a residual module, the event data features are downsampled by 2 times, and the feature extraction module with the number of channels halved finally outputs event data features downsampled by 32 times; Input the event data features downsampled by 32 times into a long short-term memory mechanism module, and output the enhanced event data features; The enhanced event data features are upsampled by 4 times; Step 4: Input the temporally enhanced color image data features and the enhanced event data features into the image-event fusion module; the image-event fusion module includes feature fusion modules with three different scales of input, and each feature fusion module is connected by a convolutional layer with a size of 1×1; in the feature fusion module, the temporally enhanced color image data features and the enhanced event data features are concatenated into one kind of feature, and the feature dimensions are shuffled along the channel direction; the shuffled features are evenly divided into two parts, which are respectively input into the spatial attention module and the channel attention module, and using a convolutional layer with a size of 3×3 and a sigmoid activation function, score maps of h×w×1 and 1×1×c are respectively calculated; where h and w respectively represent the length and width of the score map, and c represents the number of channels of the score map, and these two score maps are respectively weighted to the shuffled features to complete the information interaction between the color image data and the event data; the image-event fusion module finally outputs the features after fusing the two modalities. Step 5: Input the fused features into the detection head. The detection head includes three branches: a classification branch, a regression branch, and an IoU branch, which are respectively used for the classification, regression of the target, and the calculation of IoU. Each branch is mainly composed of multiple convolutional modules and activation function layers; the classification branch outputs features with a size of h×w×cls, where cls represents the number of target categories; the regression branch outputs features with a size of h×w×4, and the IoU branch outputs features with a size of h×w×1; the features output by the three branches are concatenated through feature concatenation, and features with a size of h×w×(5 + cls) are output; finally, the method of non-maximum suppression is used to filter redundant bounding boxes, and the detection results are output.
2. The multi-modal road scene object detection method based on images and events according to claim 1, characterized in that, The input of the efficient attention mechanism module includes the color image data features at time t and the historical color image data features at time t - 1; the color image data features at time t - 1 are respectively input into the convolutional layer φ and the convolutional layer ψ as the key and value for feature transformation, and the output feature vectors of the two are multiplied and softmax calculated to obtain a feature map with a dimension of d×d; the color image data features at time t are used as the query amount, input into the convolutional layer θ for feature transformation, and the output feature vector is multiplied by the feature map to complete feature modulation, and finally added to the original color image data features at time t to output the temporally enhanced color image data features.
3. The multi-modal road scene object detection method based on images and events according to claim 1 or 2, characterized in that, The input of the long short-term memory mechanism module consists of three parts, namely, the event data feature F at time t t , the hidden state h at time t-1 t-1 , and the cell state c t-1 ; Feature concatenation of F t and h t-1 is performed along the channel direction, passed through a separable convolutional layer of size 3×3, and respectively fed to the forget gate, input gate, and output gate. Each gate is mainly composed of a sigmoid activation function; The forget gate controls the amount of information lost in the hidden state h t-1 , the input gate controls the amount of information added to the cell state c t , and the output gate controls the output amount of information of the hidden state h t and the cell state c t at time t; The long short-term memory mechanism module completes the temporal information interaction of the event features at adjacent times and outputs the enhanced event data features.
4. The multi-modal road scene object detection method based on images and events according to claim 3, characterized in that, In step 3, the event data is divided into 5 groups with the same size along the time dimension, and the input event data is represented as event data features with 5 channels.
Citation Information
Patent Citations
Neuromorphic visual target classification method based on improved spiking neural network
CN112699956A
Motion blurred image line segment detection method and system fusing event and image
CN114913342A