A target detection method, device, medium and product in a driving scenario
Data is collected through event cameras and processed using pulsed neural networks, which solves the problem of insufficient object detection capabilities of the existing technology in challenging scenarios, achieving higher detection accuracy and lower energy consumption.
Patent Information
- Application Number
- CN202410784760.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-06-18
AI Technical Summary
Existing object detection methods perform poorly in scenarios such as overexposure, low light conditions, and high-speed motion, and the asynchronousness of event cameras poses a challenge to the deep learning framework, resulting in a comprehensive loss of information.
The event camera is used to collect data and process it through pulse neural networks (SNNs). The feature extraction module includes area merging blocks, pulse self-attention blocks, multi-layer perceptron blocks and timing memory pulse neuron blocks to operate directly on the pulse input sequence to avoid complicated preprocessing.
Improve the accuracy of target detection, especially in challenging scenarios, enhance the detection capability of fast moving targets and reduce energy consumption.
Smart Images

Figure CN118609090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and particularly to an object detection method, device, medium and product in a driving scenario. Background Art
[0002] Object detection plays a crucial role in many fields, including autonomous driving, intelligent drone patrol and navigation, and is of extremely important significance in modern society. Existing literature mainly focuses on using image data generated by RGB cameras as the input of the detection model. However, such methods fail to better solve the many limitations faced by object detection tasks. For example: Traditional cameras generate low-quality image data in challenging scenarios involving overexposure, low light conditions, and high-speed movement, greatly reducing the ability and performance of the detection model.
[0003] Compared with RGB cameras, event cameras (such as DVS and ATIS) have significant advantages. These cameras capture brightness changes as distinct, timestamped events, providing precise information about the pixel locations and polarities of these changes. Event cameras operate at sub-millisecond intervals, generating a large number of asynchronous events, ensuring the capture of absolute intensity information of moving targets based on rapid changes in light intensity. Therefore, event cameras are inherently suitable for detecting fast-moving objects.
[0004] The asynchrony of event representation poses a challenge to deep learning frameworks and requires a large amount of preprocessing. However, some studies solve this problem by converting discrete asynchronous events into an image-like format, which sacrifices the inter-frame time information in multi-channel tensors, thus compromising the comprehensiveness of information. Compared with traditional deep learning methods, spiking neural networks (SNNs), inspired by the structure and operating principle of the human brain, offer higher computational efficiency and lower energy consumption. As a network that can essentially process time-related asynchronous event data, SNNs operate directly on spike input sequences without the need for complex preprocessing of the data generated by event cameras. Therefore, in object detection scenarios involving the use of event data as input, spiking neural networks have shown great potential.
[0005] Currently, there is a lack of sufficient exploration of object detection based on event cameras and SNNs. Summary of the Invention
[0006] The object of the present invention is to provide an object detection method, device, medium and product in a driving scenario to improve the accuracy of object detection.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A target detection method in a driving scenario, comprising:
[0009] Obtain event data in a driving scenario; the event data is collected by an event camera;
[0010] Perform format conversion on the event data to obtain processed event data;
[0011] According to the processed event data, use a target detection model to determine the target bounding box, target category, and confidence of the processed event data; wherein, the target detection model is obtained by training an initial model using a training dataset; the training dataset includes event data in a training driving scenario and corresponding target bounding box labels, target categories, and confidences; the initial model includes a feature extraction module and a YOLOX detection head module; the feature extraction module includes a region merging block, a pulse self-attention block, a multi-layer perceptron block, and a temporal memory pulse neuron block connected in sequence.
[0012] Optionally, training the initial model using a training dataset specifically includes:
[0013] Construct an event dataset in a training driving scenario;
[0014] Preprocess the event dataset to obtain a processed event dataset;
[0015] Divide the processed event dataset into a training dataset and a test dataset;
[0016] Based on the training dataset and the test dataset, use the cross-validation method to train the initial model to obtain a target detection model.
[0017] Optionally, preprocessing the event dataset to obtain a processed event dataset specifically includes:
[0018] Delete the event dataset with the side length and diagonal length of the bounding box less than the corresponding set threshold to obtain a filtered event dataset;
[0019] Scale the event data with a resolution greater than the set resolution in the filtered event dataset to obtain a scaled event dataset;
[0020] Delete the scaled event data corresponding to the category with a frequency less than the set frequency to obtain a deleted event dataset;
[0021] Perform format conversion on the deleted event dataset to obtain a processed event dataset.
[0022] Optionally, the pulse self-attention block includes a first linear projection layer, three first batch normalization layers, three first pulse neuron layers, a second pulse neuron layer, a second linear projection layer, a second batch normalization layer, and a third pulse neuron layer;
[0023] The first linear projection layer is respectively connected to the three first batch normalization layers; one of the first batch normalization layers is connected to one of the first pulse neuron layers; the three first pulse neuron layers are all connected to the second pulse neuron layer; the second pulse neuron layer is connected to the second linear projection layer; the second linear projection layer is connected to the second batch normalization layer; the second batch normalization layer is connected to the third pulse neuron layer.
[0024] Optionally, the temporal memory pulse neuron block includes a first convolutional layer, a third batch normalization layer, a fourth pulse neuron layer, a fourth batch normalization layer, a fifth pulse neuron layer, a second convolutional layer, a fifth batch normalization layer, and a sixth pulse neuron layer;
[0025] The first convolutional layer, the third batch normalization layer, and the fourth pulse neuron layer are connected in sequence; the output of the fourth pulse neuron layer is connected to the fourth batch normalization layer through two residual links; the fourth batch normalization layer is connected to the fifth pulse neuron layer; the fifth pulse neuron layer is connected to the second convolutional layer through a residual link; the second convolutional layer, the fifth batch normalization layer, and the sixth pulse neuron layer are connected in sequence.
[0026] A computer device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned object detection method in a driving scenario.
[0027] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the above-mentioned object detection method in a driving scenario.
[0028] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned object detection method in a driving scenario.
[0029] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0030] The present invention provides a target detection method, device, medium and product in a driving scenario. The method includes obtaining event data in the driving scenario, where the event data is collected by an event camera; converting the format of the event data to obtain processed event data; using a target detection model to determine the target bounding box, target category and confidence of the processed event data according to the processed event data. The target detection model is obtained by training an initial model using a training data set, where the training data set includes event data in the training driving scenario and corresponding target bounding box labels, target categories and confidences. The initial model includes a feature extraction module and a YOLOX detection head module. The feature extraction module includes a region merging block, a pulsed self-attention block, a multi-layer perceptron block and a temporal memory pulsed neuron block connected in sequence. By using an event camera to collect data and using a pulsed neural network to identify the target bounding box and target category, the present invention improves the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic flowchart of the target detection method in the driving scenario provided in Embodiment 1 of the present invention;
[0033] Figure 2 It is a flowchart of the target detection model training;
[0034] Figure 3 It is a schematic diagram of the target detection model structure;
[0035] Figure 4 It is a schematic diagram of the feature extraction module structure;
[0036] Figure 5 It is a schematic diagram of the pulsed self-attention block structure;
[0037] Figure 6 It is a schematic diagram of the temporal memory pulsed neuron structure;
[0038] Figure 7 It is a curve graph of the comparison results;
[0039] Figure 8 It is an internal structure diagram of a computer device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] The object of the present invention is to provide a target detection method, device, medium and product in a driving scenario to improve the accuracy of target detection.
[0042] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0043] Embodiment 1
[0044] As Figure 1 shown, the target detection method in the driving scenario in this embodiment includes:
[0045] Step 101: Obtain event data in the driving scenario; the event data is collected by an event camera.
[0046] Step 102: Perform format conversion on the event data to obtain processed event data.
[0047] Step 103: According to the processed event data, use a target detection model to determine the target bounding box, target category and confidence of the processed event data; wherein, the target detection model is obtained by training an initial model using a training data set; the training data set includes event data in the training driving scenario and corresponding target bounding box labels, target categories and confidences; the initial model includes a feature extraction module and a YOLOX detection head module; the feature extraction module includes a region merging block, a pulse self-attention block, a multi-layer perceptron block and a temporal memory pulse neuron block connected in sequence. In practical applications, the target category can be a car, a pedestrian, a person riding a bicycle, a motorcycle or other riding tools, a bicycle, a motorcycle, a train, a truck, a bus, a traffic sign or a traffic signal, etc.
[0048] The role of confidence mainly has two aspects:
[0049] 1. Predict the probability of the existence of the target: The confidence value represents the credibility of the model detecting the target within the prediction box. Specifically, the confidence is the product of two prediction values: the confidence probability of the existence of the object and the confidence probability of correct classification. The confidence probability of the existence of the object reflects the possibility of the actual existence of the target within the given prediction box, while the confidence probability of correct classification reflects the matching degree between the predicted target category and the true category.
[0050] 2. Filter low-quality predictions: In object detection tasks, the model may generate many low-quality predictions that may not match the actual objects or have very low confidence. By setting a confidence threshold, these low-quality predictions can be filtered out, and only the prediction boxes with high confidence and more likely to be correct detections are retained. This helps reduce false detections and improve the accuracy of detection.
[0051] As an alternative implementation, the pulse self-attention block includes a first linear projection layer, three first batch normalization layers, three first pulse neuron layers, a second pulse neuron layer, a second linear projection layer, a second batch normalization layer, and a third pulse neuron layer.
[0052] The first linear projection layer is respectively connected to the three first batch normalization layers; one of the first batch normalization layers is connected to one of the first pulse neuron layers; the three first pulse neuron layers are all connected to the second pulse neuron layer; the second pulse neuron layer is connected to the second linear projection layer; the second linear projection layer is connected to the second batch normalization layer; the second batch normalization layer is connected to the third pulse neuron layer.
[0053] As an alternative implementation, the temporal memory pulse neuron block includes a first convolutional layer, a third batch normalization layer, a fourth pulse neuron layer, a fourth batch normalization layer, a fifth pulse neuron layer, a second convolutional layer, a fifth batch normalization layer, and a sixth pulse neuron layer.
[0054] The first convolutional layer, the third batch normalization layer, and the fourth pulse neuron layer are connected in sequence; the output of the fourth pulse neuron layer is connected to the fourth batch normalization layer through two residual links; the fourth batch normalization layer is connected to the fifth pulse neuron layer; the fifth pulse neuron layer is connected to the second convolutional layer through a residual link; the second convolutional layer, the fifth batch normalization layer, and the sixth pulse neuron layer are connected in sequence.
[0055] As an alternative implementation, as Figure 2 and Figure 3 shown, the initial model is trained using the training dataset, specifically including:
[0056] (1) Construct an event dataset for the training driving scenario.
[0057] (2) Preprocess the event dataset to obtain the processed event dataset.
[0058] In practical applications, the event dataset in the real driving scenario collected by the event camera will be preprocessed. Among them, abnormal data is filtered, including incorrect target bounding boxes and inappropriate target bounding boxes; datasets with too large resolution are scaled; for target bounding boxes in small batches compared to those in large batches, the small-batch target bounding boxes are deleted. A suitable time window is selected to convert discrete events into an easy-to-train format, and the time window is selected as 10 ms.
[0059] Specifically, the preprocessing process is as follows:
[0060] The event camera captures a large number of discrete asynchronous events, which can basically be represented as a tuple (x k , y k , t k , p k ). When the logarithmic intensity value lnL(x k , y k ) at the pixel point (x k , y k , t k ) exceeds the threshold K within any time interval Δt k starting from the time point t k , an event is triggered. This event triggering process can be expressed as follows:
[0061] lnL(x k , y k , t k + Δt k ) - lnL(x k , y k , t k ) ≥ p k K.
[0062] In the above formula, p k represents the polarity of the event e k , which can take values of 1 or -1, indicating the change in light intensity, that is, whether it increases or decreases.
[0063] The entire dataset consists of a large number of events continuous in the time domain. At the same time, an RGB camera is still used to collect frame-based 2D images. The target bounding boxes of the dataset are generated based on the optimal target detection boxes in the frames.
[0064] The event dataset with the side length and diagonal length of the bounding box less than the corresponding set threshold is deleted to obtain the filtered event dataset. In practical applications, for different datasets, those with the side length and diagonal length of the bounding box less than a certain number of pixels will be regarded as unreasonable bounding boxes, which affect the performance of the model and will be deleted.
[0065] Scale the event data with a resolution greater than the set resolution in the filtered event dataset to obtain a scaled event dataset. In practical applications, for datasets with too high a resolution, such as: 1Mpx, with a resolution of 720×1280, it will be scaled down proportionally to half of the original.
[0066] Delete the scaled event data corresponding to the category with a frequency less than the set frequency to obtain a deleted event dataset. In practical applications, for categories with very few label categories, that is, targets with very low frequencies, such as: traffic lights, trains, motorcycles, etc., due to the inconsistent number of categories, it will cause bias in the learning process of the model, so they will be regarded as inappropriate labels and deleted.
[0067] Convert the format of the deleted event dataset to obtain a processed event dataset.
[0068] In addition, event data that is continuous in the time domain cannot be directly sent to the model for training due to too large a time window. Therefore, it should be divided according to an appropriate time window. The expression for this process is as follows:
[0069]
[0070] The above formula describes the processing of a set of events ε within the time period [t a , t b ), where t k is between t a and t b . δ(·) is the Dirac function. The value of T depends on the selected number of discrete time steps, that is, the above time window, and is fixed at 10ms.
[0071] After the above processing, a four-dimensional tensor E∈[2,T,H,W] is obtained, where T, H, and W represent the aggregation time, the height and width obtained after data conversion, respectively. Merge the first dimension representing polarity to obtain a new tensor E∈[T,H,W]. This merge operation aims to ensure the sequentiality of the frames in the time dimension, which is consistent with the computational nature of SNNs.
[0072] Since the dimensions representing polarity are merged, the two polarity information is mixed together. Therefore, a spatial mapping kernel is used to mitigate the impact of dimension merging. The expression for this process is as follows:
[0073]
[0074] In the above formula, k(x,y,t) represents the kernel function. However, it is too complex to use a manually designed kernel function, so a convolutional layer is used to learn this process and complete the spatial mapping.
[0075] (3) Divide the processed event data set into a training data set and a test data set.
[0076] (4) Based on the training data set and the test data set, use the cross-validation method to train the initial model to obtain a target detection model.
[0077] After preprocessing, the resulting tensor will be used as the input to the feature extraction module. Before that, relative position encoding is added to the input. The feature extraction module is as Figure 4 shown.
[0078] The feature extraction module consists of four identical learning stages based on spiking neural networks. Each stage includes a set of stacked blocks, namely: Patch Merging (PM) block, Spiking Self-Attention (SSA) block, Multi-Layer Perceptron (MLP) block, and a Temporal Memory Spiking Neuron (TMSN) block. The processing process specifically includes the following steps:
[0079] Step 1: Before the data obtained after data preprocessing and data conversion is input into the feature extraction module, an additional relative position encoding block is implemented to generate relative position embeddings. This relative position embedding is added to the tensor after data preprocessing and data conversion.
[0080] Step 2: Taking the first learning stage as an example, the input is fed into the PM block, which plays a key role in capturing spatial information and merging adjacent patches into larger patches. This merging process doubles the size of each patch, increasing the receptive field and facilitating multi-scale feature extraction. Therefore, the PM block realizes downsampling of the input tensor. In the first learning stage, its PM block connects the features from each group of adjacent blocks, resulting in a 2-fold reduction in the resolution of the tensor in the height (H) and width (W) dimensions, and a 4-fold reduction overall.
[0081] Similarly, in the second learning stage, the PM block performs the same block merging operation, resulting in a tensor resolution of H / 8 × W / 8. This process continues in the third and fourth learning stages, resulting in tensors with resolutions of H / 16 × W / 16 and H / 32 × W / 32 respectively.
[0082] Step 3: The SSA processes the output feature map from the PM block, denoted as X. As Figure 5As shown, this operation generates three matrices, denoted as Q, K, and V respectively. Generating Q, K, and V involves passing the input through a linear projection layer, a Batch Normalization (BN) layer, and a Leaky Integrate-and-Fire (LIF) spiking neuron layer. The spiking neurons selectively enhance significant features while weakening weaker features. The expressions for this process are as follows:
[0083] Q′ = XW Q
[0084] K′ = XW K 。
[0085] V′ = XW V
[0086] Q = LIF Q (BN(Q′))
[0087] K = LIF K (BN(K′)).
[0088] V = LIF V (BN(V′))
[0089] The parameters W Q 、W K and W V represent the linear mapping matrices obtained through training. BN represents the batch normalization calculation, while LIF Q 、LIF K and LIF V represent the feature enhancement operations implemented using LIF spiking neurons. These operations respectively produce three updated feature matrices: Q, K, and V. The calculation process expression of SSA is described as follows:
[0090]
[0091] The updated matrices Q, K, and V are scaled by a factor s and input into a LIF neuron layer with a voltage threshold of 0.5, different from other LIF neurons set to 1.0. Subsequently, the output passes through a Linear layer, a BN layer, and another LIF neuron layer, finally forming the SSA output. The time complexity of calculating the attention weight matrix can be controlled at the minimum value Min(O(N 2 *d), O(N * d 2 ))), where d is the dimension of the head and N is the number of patches.
[0092] The features calculated by the SSA module will be added to the output of the PM module as the input to the next module.
[0093] Step 4: The output obtained in Step 3 is used as the input and passed to the MLP module. The role of the MLP is to capture the hidden correlation information from the temporal and spatial dimensions and introduce more additional learnable parameters.
[0094] Step 5: The output obtained in Step 4 is used as the input, denoted by X input and passed to the TMSN module. As shown in Figure 6 , this module also receives the residual information in the previous time domain, denoted by , where s represents the current time domain. However, although the features in the previous time domain are very similar to those in the current time domain, there are certain deviations between them. Therefore, a linear projection matrix is used for correction, and the linear projection matrix is denoted by W seq . After correction, a memory that can act on the time domain for feature enhancement is obtained, denoted by V memory . Finally, X input and V memory are added together for further feature extraction. The expression of the above process is as follows:
[0095]
[0096] X' in the above formula is the feature map to be processed next.
[0097] The TMSN module extracts spatio-temporal information from X'. It first applies the BN layer, LIF activation layer, and residual connection to obtain the intermediate feature map I. Then, the intermediate feature map is further processed through the convolutional layer, BN layer, and LIF activation layer. The output of the module consists of two components of the LIF layer. The first component is the activated feature denoted by S, which is combined with I through the residual connection to generate the output. The second component is the activated feature V, which serves as the interaction feature for the four time domains of the subsequent sequence. The expression of this process is as follows:
[0098]
[0099] In the above formula, S t and V t are generated by the final LIF layer. This process can be represented by the following equation, assuming that the feature map obtained through the LIF layer is denoted by I'. The operation Concat refers to the stacking of 2D feature maps along the time axis.
[0100]
[0101] In the above equation, H t and V t represent the membrane voltages before and after the spike, respectively. The parameter τ represents the membrane time constant and is set to 2. V treset Represents the reset voltage of the neuron layer and is set to 0. S t Represents the output pulse, Θ(·) represents the Heaviside step function. The variable T corresponds to the total number of time steps in the asynchronous step function.
[0102] Step 6: Finally, transfer the feature maps of different scales in the last three stages in MFE to the YOLOX detection head module. This detection head fuses the features of different scales and finally outputs bounding box information, class information, and confidence information.
[0103] In the present invention, training is carried out on existing public datasets. Currently, the datasets collected by event cameras in the driving scenario include: Gen1, 1Mpx, DSEC-Detection. Training has been completed on the Gen1 and 1Mpx datasets and results have been obtained. Validation has been carried out on the DSEC-Detection dataset using the weights obtained from training on the 1Mpx dataset, and the final validation results are 0.386 mAP and 0.669 AP 50 .
[0104] In the gen1 dataset: The target classes include car and pedestrian.
[0105] In the 1Mpx dataset: The target classes include Car, Pedestrian, Two-wheeler, Truck, Bus, Traffic Sign, and Traffic Light.
[0106] In the DSEC-Detection dataset: The target classes include pedestrian, rider, car, bus, truck, bicycle, motorcycle, and train.
[0107] Note: In the experiments on the 1Mpx and DSEC-Detection datasets, only the first three classes are retained to avoid the problem of class imbalance during model training (the same setting as previous work).
[0108] The results are compared using mAP and AP 50As an indicator. As shown in Table 1, the method of the present invention achieves the highest accuracy among the methods that also use spiking neural networks for related research. On the Gen1 dataset, the mAP indicator is 8.4 percentage points higher than the best previous result, and the AP 50 indicator is increased by 2.6 percentage points. In Table 1, C represents the variable time step. In the models with three different parameter sizes, the initial values of C (i.e., stage 1) are: 32, 48, 64; as the stage increases, C doubles, that is, in the models with three different parameter sizes in stage 2: 64, 96, 128; in stage 3: 128, 192, 256; in stage 4: 256, 384, 512. The data marked in bold represents the optimal result, and the data marked with an underline represents the sub-optimal result.
[0109] Table 1 Comparison results of indicators between different methods
[0110]
[0111] However, other methods have not been verified on the 1Mpx dataset for the following reasons:
[0112] 1. In terms of the design of the network structure, the methods studied in previous work are mainly based on convolutional neural networks and do not perform information aggregation operations in the time domain. In this work, a hierarchical spiking Transformer structure is adopted. At present, the feature extraction ability of the Transformer method is stronger than that of convolutional neural networks of the same magnitude. At the same time, the designed TMSN module exchanges information in the time domain, further improving the feature extraction ability of the model.
[0113] 2. Since the time window is selected as 10 ms, it means that the dimension of the tensor after data conversion is E∈[2,T,H,W]. Compared with the tensor dimension E∈[3,H,W] of ordinary frame-based images, the computing video memory occupancy of event data will be T times that of ordinary images (when the channel dimension is stretched to the same scale). At the same time, due to the sparsity of event data itself, if no relevant processing is done, the gradient will fluctuate greatly during the training process, resulting in unstable training. By fusing the polarity dimension, although some information is lost, the change of the gradient during the training process is relatively stable, and the detection accuracy is improved instead.
[0114] Such as Figure 7As shown, where (A) reflects the changes in the losses (i.e., the differences between the predicted values and the true values) of two methods, retaining the polarity dimension and performing polarity dimension fusion, during the training process, as well as the differences in gradient stability (i.e., the fluctuation range of the loss values), and (B) reflects the changes in the detection index mAP of the aforementioned two methods. The results show that the polarity dimension fusion method (the blue curve in the figure) has more stable gradients, smaller losses, and higher accuracy during the training process. On the contrary, the method of retaining the polarity dimension (the orange curve in the figure) has a larger fluctuation range of gradients, larger losses, and lower accuracy during the training process.
[0115] Finally, an artificial neural network (ANNs) of the same scale was built on the basis of the method of the present invention as a benchmark model, and the energy consumption of both was calculated respectively. As shown in Table 2, taking the energy consumption of the artificial neural network benchmark model as the unit, it is obtained that the model of the method of the present invention is 2.28 times more energy-efficient than the artificial neural network of the same order of magnitude. The calculation method is shown in the following formula.
[0116]
[0117] Where the subscript AC represents the addition operation, and MAC represents the multiplication and addition operation. In SNNs, pulse accumulation only involves the addition operation, while ANNs involve the multiplication and addition operation. fr represents the firing rate, T is the time step, FLOPs represents the number of floating-point operations, so SOPs represents the number of floating-point operations in SNNs, and the superscript l represents the layer in SNNs. On a 45-nm, 32-bit neuromorphic chip, E AC is 0.9 pJ, E MAC is 4.6 pJ, and the subscripts Conv, SSA, and det_head represent the convolution operation, the pulse-only attention operation, and the operation in the detection head, respectively.
[0118] Table 2 Energy efficiency comparison with the benchmark model
[0119]
[0120] Example 2
[0121] A computer device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the object detection method in the driving scenario of Example 1.
[0122] Example 3
[0123] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the object detection method in the driving scenario of Example 1.
[0124] Example 4
[0125] A computer program product includes a computer program which, when executed by a processor, implements the object detection method in the driving scenario in Embodiment 1.
[0126] Embodiment 5
[0127] A computer device, which can be a database, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store transactions to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements the object detection method in the driving scenario in Embodiment 1.
[0128] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the object or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present invention can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present invention can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0131] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A target detection method in a driving scenario, characterized in that: include: Acquiring event data in a driving scene; the event data is collected by an event camera; Converting the event data into a format to obtain processed event data; According to the processed event data, a target bounding box, a target category and a confidence of the processed event data are determined using a target detection model; wherein the target detection model is obtained by training an initial model using a training data set; the training data set includes event data in a training driving scenario and corresponding target bounding box labels, target categories and confidences; the initial model includes a feature extraction module and a YOLOX detection head module; the feature extraction module includes a region merging block, a pulse self-attention block, a multi-layer perceptron block and a temporal memory pulse neuron block connected in sequence; The spiking self-attention block includes a first linear projection layer, three first batch normalization layers, three first spiking neuron layers, a second spiking neuron layer, a second linear projection layer, a second batch normalization layer, and a third spiking neuron layer; The first linear projection layer is connected to three first batch normalization layers respectively; one first batch normalization layer is connected to one first pulse neuron layer; and three first pulse neuron layers are all connected to the second pulse neuron layer; The second spiking neuron layer is connected to the second linear projection layer; the second linear projection layer is connected to the second batch normalization layer; the second batch normalization layer is connected to the third spiking neuron layer.
2. The target detection method in a driving scenario according to claim 1, characterized in that: Use the training data set to train the initial model, including: Construct an event dataset for driving scenarios for training; Preprocessing the event data set to obtain a processed event data set; Dividing the processed event data set into a training data set and a test data set; Based on the training data set and the test data set, the initial model is trained using a cross-validation method to obtain a target detection model.
3. The target detection method in a driving scenario according to claim 2, characterized in that: Preprocessing the event data set to obtain a processed event data set specifically includes: The event data sets whose bounding box side length and diagonal length are less than the corresponding set thresholds are deleted to obtain a filtered event data set; Scaling the event data with a resolution greater than a set resolution in the filtered event data set to obtain a scaled event data set; The scaled event data corresponding to the categories whose occurrence frequency is less than the set frequency are deleted to obtain a deleted event data set; The deleted event data set is format converted to obtain a processed event data set.
4. The target detection method in a driving scenario according to claim 1, characterized in that: The temporal memory spike neuron block includes a first convolutional layer, a third batch normalization layer, a fourth spike neuron layer, a fourth batch normalization layer, a fifth spike neuron layer, a second convolutional layer, a fifth batch normalization layer, and a sixth spike neuron layer; The first convolutional layer, the third batch normalization layer and the fourth spiking neuron layer are connected in sequence; the output of the fourth spiking neuron layer is connected to the fourth batch normalization layer through two residual links; the fourth batch normalization layer is connected to the fifth spiking neuron layer; The fifth pulse neuron layer is connected to the second convolutional layer via a residual link; the second convolutional layer, the fifth batch normalization layer and the sixth pulse neuron layer are connected in sequence.
5. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target detection method in a driving scenario according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method in the driving scenario described in any one of claims 1 to 4 is implemented.
7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the target detection method in the driving scenario described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Event data processing method, device and system
CN115546248A
Target tracking method and device and storage medium
CN117934540A