Low-delay moving target detection method and system for navigation
Through the on-board event camera and pulse neural network, the event flow is processed, combined with feature pyramid network and nonlinear model prediction control, the detection delay and misjudgment problems of autonomous driving systems in emergencies and extreme environments are solved, and the low-latency target detection and trajectory extrapolation are achieved.
Patent Information
- Application Number
- CN202510734039.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing autonomous driving system is difficult to effectively deal with high-speed and mutational moving targets in emergencies and extreme environments, resulting in high detection delays and misjudgment. The sensor fusion solution is complex and cannot respond to emergency scenarios in a timely manner.
The on-board event camera is used to collect event streams, and the event images are processed by adjusting the time window length and pulse neural network, combining feature pyramid network and nonlinear model prediction control to achieve low-latency object detection and trajectory extrapolation.
It significantly improves the real-time response ability to high-speed and mutational moving targets and the detection robustness in complex lighting environments, reducing detection delay and misjudgment rate.
Smart Images

Figure CN120260016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a low-latency moving target detection method and system for navigation. Background Art
[0002] Existing autonomous driving systems mainly use natural image cameras and millimeter-wave radars to achieve environmental target perception and displacement prediction, so as to plan navigation paths, perform obstacle avoidance and emergency handling. For emergencies, such as pedestrians and motorcycles suddenly crossing, vehicles cutting in, and objects suddenly appearing after being blocked, or in extreme environments, such as strong backlight, sudden changes in tunnel brightness and darkness, due to the limitations of the device perception principles of natural image cameras and millimeter-wave radars, it is often difficult to effectively handle them.
[0003] Due to the fixed frame rate and dependent exposure mechanism of in-vehicle natural image cameras, it is easy to cause high detection latency due to motion blur in high-speed scenarios, and the insufficient dynamic range makes it difficult to capture the target contour in environments such as strong backlight and sudden changes in tunnel brightness and darkness. Although millimeter-wave radars can measure speed, their sparse point cloud data cannot identify the morphological features of targets, which is prone to misjudgment and further affects the trajectory prediction of targets. The multi-sensor fusion scheme has latency superposition due to data alignment and processing complexity, and it is difficult to respond to emergency scenarios in a timely manner.
[0004] Currently, the main event camera target detection methods mainly focus on the accuracy and recognition rate of target detection. There are still few studies on the low-latency detection methods for event cameras. In addition, there is currently a lack of corresponding solutions for the event sparsity difference and temporal position offset problems in event target detection. Summary of the Invention
[0005] In view of the above problems, the object of the present invention is to provide a low-latency moving target detection method and system for navigation, which can achieve ultra-low-latency target detection and trajectory extrapolation, and significantly improve the real-time response ability to high-speed and mutant moving targets and the detection robustness in complex lighting environments.
[0006] The present invention provides a low-latency moving target detection method for navigation, including: Collecting an event stream through an in-vehicle event camera; Real-time adjusting the time window length of the event stream according to the vehicle speed, vehicle acceleration and event density gradient to obtain an event image; Inputting the event image into a pulsed neural network to obtain an output feature map; Detecting the feature map through a feature pyramid network and a detection head to obtain a moving trajectory stream of the target; Performing non-linear model predictive control according to the moving trajectory stream to obtain an optimized moving trajectory stream of the target.
[0007] In a possible implementation, the real-time adjustment of the time window length of the event stream according to the vehicle speed, vehicle acceleration, and event density gradient includes: Adjust the time window length according to the following formula: ; Where, is the time window length, is the upper limit length of the time window, is the lower limit length of the time window, is the upper threshold of the speed, is the lower threshold of the speed, is the upper threshold of the acceleration, is the lower threshold of the acceleration, is the two-norm of the spatial gradient of the event, is the time gradient of the event, is the scaling coefficient, is the parameter for balancing space and time.
[0008] In a possible implementation, the real-time adjustment of the time window length of the event stream according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the spatial gradient of the event according to the following formula : ; ; Where, is the time gradient of the event, is the event direction, is the polar coordinate encoding of each event point.
[0009] In a possible implementation, the real-time adjustment of the time window length of the event stream according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the time gradient of the event according to the following formula : ; Where, is the time encoding of the event, is the second derivative of the event with respect to time and is the differential component of time.
[0010] In a possible implementation, the spiking neural network includes 4 cascaded convolutional modules; the inputting the event image into the spiking neural network to obtain the output feature map includes: Input the event image at the current moment into a convolutional module to obtain the features at the current moment; Decompose the features at the current moment into low-frequency features and high-frequency features through a wavelet attention mechanism, and concatenate the low-frequency features and high-frequency features to obtain the output features at the current moment; Perform convolution on the features at the previous moment through a spatio-temporal displacement attention mechanism to obtain the output features at the previous moment; Fuse the output features at the current moment and the output features at the previous moment through concatenation and a convolutional module to obtain a feature map.
[0011] In a possible implementation, the detection head consists of three convolutional layers, which respectively output the detected category, confidence, coordinates, and the offset of the bounding box; detecting the feature map through the feature pyramid network and the detection head to obtain the motion trajectory flow of the target includes: Perform upsampling on the feature map step by step to obtain an image interpolation; Fuse the image interpolation and the bilinear interpolation aligned with the channels to obtain a fused feature map; Optimize the fused feature map through the detection head to obtain the motion trajectory flow of the target.
[0012] In a possible implementation, optimizing the fused feature map through the detection head to obtain the motion trajectory flow of the target includes: Optimize the fused feature map according to the following formula: ; where is the motion trajectory flow of the target, is the weight of the coordinates and the offset of the bounding box, is the weight of the coordinates and the confidence of the bounding box, is the weight of the category loss of the coordinates and the bounding box, is the offset of the coordinates and the bounding box, is the confidence of the coordinates and the bounding box, is the category loss of the coordinates and the bounding box.
[0013] In a possible implementation, optimizing the fused feature map through the detection head to obtain the motion trajectory flow of the target includes: Calculate the offset of the coordinates and the bounding box according to the following formula : ; where is the intersection over union of the predicted target bounding box and the true target bounding box, The Euclidean distance between the center point of the predicted target bounding box and the center point of the ground-truth target bounding box, is the diagonal length of the smallest predicted target bounding box and the ground-truth target bounding box, where the predicted target bounding box is and the ground-truth target bounding box is; Calculate the coordinates and the confidence of the bounding box according to the following formula : ; wherein, is the predicted grid resolution, is the number of predicted targets per grid, is the ground-truth target confidence, is the predicted target confidence, and are the coordinate indices of the grid; Calculate the coordinates and the class loss of the bounding box according to the following formula : ; wherein, is the number of classes, is the class index, is the predicted class, is the ground-truth class, is the focusing parameter.
[0014] In a possible implementation manner, the obtaining of the optimized motion trajectory stream of the target by performing non-linear model predictive control according to the motion trajectory stream includes: Dynamically adjust the time window length of the motion trajectory stream according to the jerk and the sliding overlap mechanism; Construct a non-linear motion model based on a cubic polynomial according to the time window length; the model parameters of the non-linear motion model are jointly determined according to the fusion of physical dynamics constraints and data-driven optimization; Predict the motion trajectory stream according to the non-linear motion model to obtain a predicted trajectory; Correct the predicted trajectory according to a Kalman filter to obtain the optimized motion trajectory stream of the target.
[0015] The present invention also provides a low-latency moving target detection system for navigation, which is characterized in that it is applied to any one of the above-mentioned low-latency moving target detection methods, and includes: An acquisition module, configured to acquire an event stream through an in-vehicle event camera; A time window adjustment module, configured to adjust the time window length of the event stream in real time according to the vehicle speed, the vehicle acceleration, and the event density gradient to obtain an event image; A feature map extraction module for inputting the event image into a spiking neural network to obtain an output feature map; A motion trajectory flow detection module for detecting the feature map through a feature pyramid network and a detection head to obtain the motion trajectory flow of the target; A motion trajectory flow optimization module for performing nonlinear model predictive control according to the motion trajectory flow to obtain an optimized motion trajectory flow of the target.
[0016] The low-latency moving target detection method and system for navigation provided by the present invention are designed for the high-density and directionally continuous spatio-temporal features (such as event clusters aggregating along the motion trajectory and density gradient changes caused by acceleration) exhibited by fast-moving targets (such as suddenly lane-changing vehicles, crossing pedestrians, high-speed motorcycles, and obstacles that suddenly appear after being occluded) in short-time event slices. An extremely lightweight spiking neural network (SNN) is designed to directly process the asynchronous spatio-temporal signals of the event stream, avoiding the computational redundancy of traditional image format conversion and complex models. Combining the physical laws of short-time trajectories (such as uniform / constant acceleration modes and steering angle continuity) to construct a dynamic prediction model, realizing ultra-low-latency target detection and trajectory extrapolation, and significantly improving the real-time response ability to high-speed and mutant moving targets and the detection robustness in complex lighting environments. Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of the low-latency moving target detection method provided by the embodiment of the present invention; Figure 2 It is a schematic working diagram of the spiking neural network provided by the embodiment of the present invention; Figure 3 It is a schematic simulation diagram of the model training process provided by the embodiment of the present invention; Figure 4 It is a schematic diagram of the experimental results provided by the embodiment of the present invention. Detailed Embodiments
[0018] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and embodiments. The detailed description and drawings of the following embodiments are used to exemplarily illustrate the principles of the present invention, but cannot be used to limit the scope of the present invention, that is, the present invention is not limited to the described preferred embodiments, and the scope of the present invention is defined by the claims.
[0019] In the description of the present invention, it should be noted that unless otherwise specified, the meaning of "a plurality" is two or more; the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance; for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0020] Figure 1The flowchart of the low-latency moving target detection method provided by the embodiment of the present invention is shown as follows. Figure 1 As shown, the present invention provides a low-latency moving target detection method for navigation, including: Step S1, collecting an event stream through an in-vehicle event camera; In a possible implementation, traditional target detection often uses a fixed time window, such as 30 ms, for event accumulation, resulting in blurred features of high-speed targets or insufficient signal-to-noise ratio of low-speed targets. For example, a too-long window causes motion blur, and a too-short window leads to too few valid events. The asynchronous event stream collected by the in-vehicle event camera of the present invention has a time resolution at the microsecond level.
[0021] Step S2, adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient to obtain an event image; In a possible implementation, the motion speed of the target is calculated based on the event density gradient of the in-vehicle event camera. A speed-window non-linear mapping function is constructed to achieve adaptive adjustment of the time window length.
[0022] In a possible implementation, a dual-threshold method is used to truncate the window time, and it can also be implemented in the form of a soft threshold / no threshold. For example, the time window is normalized to between 0 and 1 using a sigmoid function and then mapped back to the time upper and lower limits.
[0023] The speed-driven time window adaptive mechanism of the present invention dynamically adjusts the window length by real-time sensing of the target motion speed, solving the dual contradictions of motion blur in high-speed scenarios and insufficient signal-to-noise ratio in low-speed scenarios of the traditional fixed window method. It effectively reduces the window response time, trajectory positioning errors caused by motion blur, and slow-moving scenarios.
[0024] In a possible implementation, adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient includes: Adjusting the time window length according to the following formula: ; This formula means that when the speed is too fast or the acceleration is too large, the minimum time window is taken; when the speed is too slow and the acceleration is too slow, the maximum time window is taken; and in other cases, it is jointly calculated by the acceleration and speed; Wherein, is the time window length, is the upper limit length of the time window, 10 ms, is the lower limit length of the time window, 50 ms, is the upper threshold of the speed, is the lower threshold of the speed, is the upper threshold of acceleration, is the lower threshold of acceleration, is the two-norm of the spatial gradient of the event, is the temporal gradient of the event, is the scaling coefficient, is a parameter that balances space and time.
[0025] In a possible implementation, the real-time adjustment of the time window length of the event stream according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the spatial gradient of the event according to the following formula : ; ; where, is the temporal gradient of the event, is the event direction, is the polar coordinate encoding of each event point.
[0026] In a possible implementation, the real-time adjustment of the time window length of the event stream according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the temporal gradient of the event according to the following formula : ; where, is the temporal encoding of the event, is the second derivative of the event with respect to time and is the differential component of time.
[0027] Step S3, input the event image into the spiking neural network to obtain the output feature map; In a possible implementation, the spiking neural network includes 4 cascaded SNN Blocks, namely convolutional modules, Figure 2 is the working schematic diagram of the spiking neural network provided by the embodiment of the present invention.
[0028] Each convolutional module simultaneously fuses the output of the previous moment to construct the features in the target feature time.
[0029] Specifically, input the event image of the current moment into a convolutional module and obtain the features of the current moment; In order to align the feature frequency distributions in different environments, the present invention proposes a wavelet domain cross-attention mechanism. Decompose the features of the current moment into low-frequency features and high-frequency features through the wavelet attention mechanism, and splice the low-frequency features and high-frequency features to obtain the output features of the current moment; To align the features in time series, the features of the previous moment are convolved through a spatio-temporal displacement attention mechanism to obtain the output features of the previous moment; The output features of the current moment and the output features of the previous moment are fused through a concatenation and convolution module to obtain a feature map.
[0030] Specifically, for the wavelet-domain cross-attention mechanism, the input feature map F is decomposed into low-frequency features Flf and high-frequency features Fhf through the discrete wavelet transform (DWT). Since the low-frequency features contain more background noise and the high-frequency features contain more discrete noise, in this study, the self-attention mechanism is used to extract the important information in the low-frequency and high-frequency respectively, and their outputs are concatenated, and the final output is obtained through the inverse wavelet transform (IDWT).
[0031] The self-attention mechanism consists of three convolutions, which map the input features to Q, K, and V respectively, representing different distributions of the features. Subsequently, Q is transposed and multiplied by K and then softmax is applied to obtain the attention map, and the attention map is further multiplied by V to obtain the attention result. The attention result is added to the input to obtain the output.
[0032] Specifically, for the spatio-temporal displacement attention mechanism, each channel of the input feature map will pass through a convolution to obtain the spatial offset, which will be encoded as the spatial offset coordinates and applied to the convolution to obtain an offset feature. Subsequently, an attention mechanism is applied to the offset feature to obtain the final model output.
[0033] The wavelet-domain cross-attention mechanism proposed by the present invention calculates the channel-spatial cross-attention matrix in the wavelet domain, enhances the response weight of the sudden motion features, and aligns the sparsity of the event images. Through the collaborative calculation of multi-scale wavelet decomposition and spatio-temporal attention weights, it breaks through the sensitivity bottleneck of traditional neural networks to sudden motion features, and effectively improves the signal-to-noise ratio of motion edge features in the sharp acceleration / emergency braking scenarios. In actual use, the discrete Fourier transform can also be used to optimize the density difference problem of the event stream.
[0034] The present invention adopts the spatio-temporal displacement attention mechanism to significantly improve the robustness of trajectory prediction and the accuracy of temporal prediction effect in complex scenarios through spatio-temporal position encoding and dynamic association modeling.
[0035] Step S4, detecting the feature map through a feature pyramid network and a detection head to obtain the motion trajectory stream of the target; In a possible implementation, the detection head is composed of three convolutional layers, which respectively output the detected category, confidence, coordinates, and the offset of the bounding box; the feature map is upsampled step by step to obtain the image interpolation; the image interpolation is fused with the bilinear interpolation aligned with the channels to obtain the fused feature map; the detection head optimizes the fused feature map to obtain the motion trajectory stream of the target.
[0036] Among them, the feature pyramid performs bilinear interpolation fusion of upsampling and channel alignment of multi-scale feature maps output by multiple SNN Blocks through a top-down path and lateral cross-layer connections. Among them, high-level semantic features perform element-wise addition with shallow high-resolution features through 2× nearest neighbor interpolation, and at the same time introduce a 1×1 convolutional kernel to achieve cross-channel information interaction. The finally output fused feature map aggregates global context while retaining spatial details, forming a unified representation space with multi-scale perception ability, effectively solving the size sensitivity problem in the object detection task.
[0037] In a possible implementation, the detection head optimizes the fused feature map to obtain the motion trajectory flow of the target, including: Optimize the fused feature map according to the following formula: ; Among them, is the motion trajectory flow of the target, is the weight of the coordinate and the offset of the bounding box, is the weight of the coordinate and the confidence of the bounding box, is the weight of the coordinate and the class loss of the bounding box, is the coordinate and the offset of the bounding box, is the coordinate and the confidence of the bounding box, is the coordinate and the class loss of the bounding box.
[0038] In a possible implementation, the detection head optimizes the fused feature map to obtain the motion trajectory flow of the target, including: Calculate the offset of the coordinate and the bounding box according to the following formula : ; Among them, is the intersection over union of the predicted target bounding box and the ground truth target bounding box, is the Euclidean distance between the center point of the predicted target bounding box and the center point of the ground truth target bounding box, is the diagonal length of the smallest predicted target bounding box and the ground truth target bounding box, is the predicted target bounding box, is the ground truth target bounding box; Calculate the confidence of the coordinate and the bounding box according to the following formula : ; Among them, is the predicted grid resolution, is the number of predicted targets per grid, is the true target confidence level, is the predicted target confidence level, and is the coordinate index of the grid; Calculate the coordinate and the class loss of the bounding box according to the following formula : ; wherein, is the number of classes, is the class index, is the predicted class, is the true class, is the focusing parameter.
[0039] Step S5, perform non - linear model predictive control according to the motion trajectory flow to obtain the optimized motion trajectory flow of the target.
[0040] In a complex dynamic scenario, the non - linear modeling of short - time motion trajectories needs to balance real - time performance, accuracy, and anti - interference ability. Traditional linear models are difficult to depict sudden motion patterns such as rapid acceleration and emergency steering, while global non - linear models are difficult to meet real - time requirements due to their high computational complexity. Therefore, this paper proposes a solution that integrates hierarchical modeling and dynamic correction, and realizes high - precision and adaptive trajectory modeling through four core steps.
[0041] In a possible implementation, for the local non - linear characteristics of the motion trajectory, dynamically adjust the time - window length of the motion trajectory flow according to the jerk, that is, the acceleration change rate, and the sliding overlap mechanism; Specifically, the traditional fixed - time window performs stably in uniform - speed or uniform - acceleration scenarios, but cannot quickly respond to instantaneous behaviors such as emergency braking and sudden direction change. The present invention dynamically adjusts the time - window length by calculating the jerk in real time: when the jerk increases significantly (such as a vehicle making an emergency evasive maneuver), the window length automatically shortens to capture instantaneous motion details; when the motion tends to be stable (such as a drone cruising), the window extends to reduce the computational load. At the same time, a sliding overlap mechanism is introduced to ensure the continuity of the trajectory when the window switches, and avoid the data break problem caused by segmented modeling. This dynamic partitioning strategy can not only ensure a sensitive response to sudden motions but also effectively balance the computational efficiency under limited hardware resources.
[0042] Construct a non - linear motion model based on a cubic polynomial according to the time - window length; Specifically, based on dynamic window partitioning, a non-linear motion model based on a cubic polynomial is constructed. The cubic polynomial model can mathematically characterize the characteristic of acceleration changing with time (such as non-uniform acceleration motion when a vehicle makes a sharp turn), and its degrees of freedom provide sufficient flexibility for trajectory fitting. The model parameters of the non-linear motion model are jointly determined by integrating physical dynamics constraints and data-driven optimization: on the one hand, the theoretical acceleration derived from Newton's laws of motion is introduced as a prior constraint to avoid physical unreasonableness caused by pure data-driven approaches (such as exceeding the speed limit of a vehicle instantaneously); on the other hand, the least squares method with regularization is used to fit the observed data to suppress the overfitting risk caused by sensor noise. This physical-data hybrid modeling method not only retains the complex non-linear characteristics in actual motion but also ensures the feasibility of the model in terms of dynamics.
[0043] The non-linear modeling method of the present invention based on cubic polynomials and physical dynamics constraints overcomes the problem of inaccurate prediction of traditional linear models in short-term mutant trajectories. The coefficient matrix is optimized and solved by the Newton-Raphson iteration method. A dual-gain adjustment mechanism for model prediction and sensor observation is designed to automatically switch to the observation-dominated mode under abnormal disturbances to achieve spatial coordinate alignment in time series.
[0044] Predict the motion trajectory stream according to the non-linear motion model to obtain the predicted trajectory; Specifically, a multi-level optimization strategy is adopted in the trajectory fitting stage to improve the modeling robustness. First, the original sensor data (such as GPS position, IMU acceleration) are aligned in space-time and outlier filtered to eliminate hardware transmission delays and wild point interferences. Subsequently, the model parameter estimation is carried out by the iteratively reweighted least squares method, assigning greater weights to high-confidence data points (such as high-precision position measurements of lidar), and at the same time suppressing high-frequency jitter noise through L1 regularization constraints. For the problem of kinematic parameter jumps (such as a robot suddenly stopping), a local parameter smoother based on a sliding window is designed to optimize the consistency of the model parameters in adjacent windows on the premise of retaining the motion mutation characteristics, avoiding the trajectory step phenomenon caused by independent window modeling.
[0045] In addition to the non-linear motion model, the present invention can also use a multi-layer neural network for modeling as a trajectory fitting solution.
[0046] To overcome the contradiction between the cumulative error of long-term model prediction and the instantaneous noise of sensors, a dual-mode Kalman filter correction framework is designed. The predicted trajectory is corrected according to the Kalman filter to obtain the optimized motion trajectory stream of the target.
[0047] In the normal motion mode, the predicted output of the aforementioned non-linear model is used as the prior estimate of the Kalman filter. Through the covariance adaptive adjustment mechanism, multi-source sensor observation data (such as the relative pose of the visual odometer and the speed information of the millimeter-wave radar) is dynamically fused, and the Kalman gain matrix is used to balance the weights of the model prediction and the real-time data. When a severe motion disturbance is detected (such as a sudden increase in the model residual caused by an emergency brake), it automatically switches to the observation-dominated mode, temporarily weakening the model prediction weight, and preferentially updating the state based on the high-frequency sensor data to prevent the trajectory divergence caused by model mismatch. In addition, a kinematic rationality verification module is introduced to review physical rules such as speed continuity and acceleration boundaries for the corrected trajectory, further ensuring the reliability of the output trajectory.
[0048] Embodiment 1 The operating system of the experimental environment is Windows 10 (64-bit), the CPU is Intel Core i7-8750H, the memory is 32GB, the hard disk is a 932GB mechanical hard disk, and the graphics card is Nvidia RTX 3090. The development language is Python 3.6.0, and the machine learning environment is numpy 1.19.5, tensorflow 1.2.0, pandas 1.1.5, torch 1.7.1, and Scikit-learn 1.0.
[0049] The simulation uses the Gen 1 event camera autonomous driving dataset. The Gen 1 dataset is recorded using the PROPHESEE GEN1 sensor installed on the car dashboard, with a resolution of 304×240 pixels. The labels are obtained through manual annotation using the gray-scale estimation function of the ATIS camera. The Gen 1 dataset contains 39 hours of open roads and various driving scenarios, including urban, highway, suburban, and rural scenarios. The parameter settings are shown in Table 1 below.
[0050] Table 1
[0051] Figure 3 It is a schematic diagram of the simulation of the model training process provided by the embodiment of the present invention. Figure 3 Among them, Total Loss is the convergence of the loss function, Cls Loss is the loss of target classification, IoU Loss is the loss of bounding box regression, and Obj Loss is the loss of target discrimination.
[0052] Figure 4This is a schematic diagram of the experimental results provided by the embodiments of the present invention. The PR curve showing the prediction results is designed to evaluate the changes in precision and recall at different thresholds. The vertical axis represents precision, the horizontal axis represents recall, and the area under the PR curve is the AP metric. The comparison results between the model of the present invention and other models are shown in Table 2 below.
[0053] Table 2
[0054] In the experiment, we used average precision (Average Precision), precision, recall, and running time (Time) as the criteria to evaluate the performance of the model.
[0055] Among them, YOLOX: Open-sourced by Megvii Technology, etc., integrates the progress in object detection fields such as decoupled heads, data augmentation, anchor-free, and label classification with YOLO. Experiments on various datasets have proven that YOLOX has excellent generalization and efficiency.
[0056] SNN-YOLO: Proposed in 2024, it is the SNN version of YOLO for object detection in event cameras. It has excellent computing speed but lower performance than the original YOLO.
[0057] The model of the present invention: Integrates the speed advantage of SNN and significantly improves detectability through multiple improvements.
[0058] The present invention also provides a low-latency moving target detection system for navigation, which is characterized in that it is applied to any of the above-mentioned low-latency moving target detection methods, including: An acquisition module for acquiring an event stream through an in-vehicle event camera; A time window adjustment module for adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient to obtain event images; A feature map extraction module for inputting the event images into a spiking neural network to obtain the output feature map; A moving trajectory stream detection module for detecting the feature map through a feature pyramid network and a detection head to obtain the moving trajectory stream of the target; A moving trajectory stream optimization module for performing nonlinear model predictive control based on the moving trajectory stream to obtain the optimized moving trajectory stream of the target.
[0059] The low-latency moving target detection method and system for navigation provided by the present invention address the challenges of sparsity and variability in the event stream. The event stream collected by the in-vehicle event camera passes through a speed-driven time window adaptive module, which extracts an event window based on the vehicle's moving speed and constructs it into an event image. To construct a feature map, the present invention proposes that the event image will be gradually input into 4 SNN Blocks, and each SNN Block simultaneously fuses the output of the previous moment to construct the temporal correlation of the target features. To align the frequency-domain and temporal features, the present invention proposes to use a wavelet-domain cross-attention mechanism and a spatio-displacement temporal attention mechanism in the SNN Block. The trajectory of the target is predicted by a feature pyramid and a detection head. To optimize the predicted trajectory of the target, the present invention further introduces a non-linear modeling of the short-term motion trajectory, ultimately achieving low-latency and high-precision temporal tracking of fast-moving targets.
[0060] The present invention designs an extremely lightweight spiking neural network (SNN) for fast-moving targets (such as suddenly lane-changing vehicles, crossing pedestrians, high-speed motorcycles, and obstacles that suddenly appear after being occluded), which exhibit high-density and directionally continuous spatio-temporal features (such as event clusters aggregating along the motion trajectory and density gradient changes caused by acceleration) in short-time event slices. It directly processes the asynchronous spatio-temporal signals of the event stream, avoiding the computational redundancy of traditional image format conversion and complex models, and constructs a dynamic prediction model in combination with the physical laws of short-term trajectories (such as uniform / constant acceleration modes and steering angle continuity) to achieve ultra-low-latency target detection and trajectory extrapolation, significantly enhancing the real-time response ability to high-speed and abruptly moving targets and the detection robustness in complex lighting environments.
[0061] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A low-latency moving target detection method for navigation, characterized in that Including: Collecting an event stream through an in-vehicle event camera; Adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient to obtain an event image; Inputting the event image into a spiking neural network to obtain an output feature map; Detecting the feature map through a feature pyramid network and a detection head to obtain a motion trajectory stream of the target; Performing nonlinear model predictive control according to the motion trajectory stream to obtain an optimized motion trajectory stream of the target.
2. The low-latency moving object detection method according to claim 1, wherein The adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient includes: Adjusting the time window length according to the following formula: ; Among them, is the time window length, is the upper limit length of the time window, is the lower limit length of the time window, is the upper threshold of the speed, is the lower threshold of the speed, is the upper threshold of the acceleration, is the lower threshold of the acceleration, is the two-norm of the spatial gradient of the event, is the temporal gradient of the event, is the scaling factor, is the parameter for balancing space and time.
3. The low-latency moving object detection method according to claim 2, wherein The adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the spatial gradient of an event according to the following formula : ; ; Among them, is the time gradient of the event, is the event direction, is the polar coordinate encoding of each event point.
4. The low-latency moving object detection method according to claim 2, wherein The adjusting the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration, and event density gradient includes: Calculate the time gradient of the event according to the following formula :[[]]END]] ; Among them, is the time encoding of the event, is the second derivative of the event with respect to time and is the differential of time.
5. The low-latency moving target detection method according to claim 1, characterized in that The spiking neural network includes 4 cascaded convolutional modules; the inputting the event image into the spiking neural network to obtain an output feature map includes: Inputting the event image at the current moment into a convolutional module and obtaining the feature at the current moment; Decomposing the feature at the current moment into a low-frequency feature and a high-frequency feature through a wavelet attention mechanism, and splicing the low-frequency feature and the high-frequency feature to obtain the output feature at the current moment; Performing convolution on the feature at the previous moment through a spatio-temporal displacement attention mechanism to obtain the output feature at the previous moment; Fusing the output feature at the current moment and the output feature at the previous moment through splicing and a convolutional module to obtain a feature map.
6. The low-latency moving object detection method according to claim 1, characterized in that, The detection head is three convolutional layers, which respectively output the detected category, confidence, coordinates, and the offset of the bounding box; The detecting the feature map through a feature pyramid network and a detection head to obtain a motion trajectory stream of the target includes: Performing progressive upsampling on the feature map to obtain an image interpolation; Fusing the image interpolation and the bilinear interpolation aligned with the channels to obtain a fused feature map; Optimizing the fused feature map through the detection head to obtain a motion trajectory stream of the target.
7. The low-latency moving target detection method according to claim 6, wherein The optimizing the fused feature map through the detection head to obtain a motion trajectory stream of the target includes: Optimizing the fused feature map according to the following formula: ; Among them, is the motion trajectory stream of the target, is the weight of the offset of the coordinates and the bounding box, is the weight of the confidence of the coordinates and the bounding box, is the weight of the class loss of the coordinates and the bounding box, is the offset of the coordinates and the bounding box, is the confidence of the coordinates and the bounding box, is the class loss of the coordinates and the bounding box.
8. The low-latency moving object detection method according to claim 7, wherein The optimizing the fused feature map through the detection head to obtain a motion trajectory stream of the target includes: Calculate the coordinates and the offset of the bounding box according to the following formula : ; Among them, is the intersection over union of the predicted target bounding box and the ground truth target bounding box, is the Euclidean distance between the center point of the predicted target bounding box and the center point of the ground truth target bounding box, is the diagonal length of the smallest predicted target bounding box and the ground truth target bounding box, is the predicted target bounding box, is the ground truth target bounding box; Calculate the coordinates and the confidence of the bounding box according to the following formula : ; Among them, is the predicted grid resolution, is the number of predicted targets for each grid, is the true target confidence, is the predicted target confidence, and is the coordinate index of the grid; Calculate the coordinate and bounding box class losses according to the following formula : ; Among them, is the number of categories, is the category index, is the predicted category, is the true category, is the focusing parameter.
9. The low-latency moving object detection method according to claim 1, characterized in that The performing nonlinear model predictive control according to the motion trajectory stream to obtain an optimized motion trajectory stream of the target includes: Dynamically adjusting the time window length of the motion trajectory stream according to the jerk and the sliding overlap mechanism; Constructing a nonlinear motion model based on a cubic polynomial according to the time window length; the model parameters of the nonlinear motion model are jointly determined according to the fusion of physical dynamics constraints and data-driven optimization; Predicting the motion trajectory stream according to the nonlinear motion model to obtain a predicted trajectory; Correcting the predicted trajectory according to a Kalman filter to obtain an optimized motion trajectory stream of the target.
10. A low-latency moving target detection system for navigation, characterized in that, Applied to the low-latency moving target detection method according to any one of claims 1-9, including: An acquisition module for collecting an event stream through an in-vehicle event camera; A time window adjustment module, which is used to adjust the time window length of the event stream in real time according to the vehicle speed, vehicle acceleration and event density gradient, so as to obtain an event image; A feature map extraction module, which is used to input the event image into a spiking neural network to obtain an output feature map; A motion trajectory stream detection module, which is used to detect the feature map through a feature pyramid network and a detection head to obtain the motion trajectory stream of the target; A motion trajectory stream optimization module, which is used to perform nonlinear model predictive control according to the motion trajectory stream to obtain an optimized motion trajectory stream of the target.
Citation Information
Patent Citations
High-dynamic target detection method based on event camera
CN111582300A
Multi-modal vehicle trajectory prediction and training method and device based on visual perception
CN118736520A
Robot alarm processing method and system based on target detection algorithm and cloud platform
CN119418174A
Neural Network Inference Acceleration Method, Target Detection Method, Device, and Storage Medium
US20240161474A1
Automatic trajectory prediction method based on graph spatial-temporal pyramid
WO2024193334A1
Cited By
Event camera target tracking method and device, storage medium and electronic equipment
CN121458754A
Event camera target tracking methods, devices, storage media and electronic equipment
CN121458754B