Event camera target tracking methods, devices, storage media and electronic equipment
By employing adaptive temporal aggregation, noise reduction, and motion compensation methods, the problem of event cameras losing spatiotemporal information during target tracking was solved, achieving high-precision target trajectory extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BANKER FUTURE TECH (BEIJING) CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing event cameras are prone to losing spatiotemporal information during target tracking, making it difficult to accurately and stably extract target trajectories.
An adaptive time aggregation module is used to dynamically adjust the event aggregation time window, and noise reduction is performed by combining a spatiotemporal neighborhood consistency strategy and a polarity time difference constraint strategy. Edge ghosting is eliminated by a motion compensation and feature representation module, and trajectory prediction is performed by a hybrid attention network.
Maintaining the temporal continuity and spatial integrity of event information under different motion states improves the accuracy and stability of target tracking and enables precise extraction of target trajectories.
Smart Images

Figure CN121458754B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular to an event camera target tracking method, apparatus, storage medium, and electronic device. Background Technology
[0002] An event camera (Dynamic Vision Sensor, DVS) is a visual sensor that records brightness changes in a scene with microsecond-level temporal resolution, based on asynchronous output triggered by changes in light intensity. Compared to traditional frame-based cameras, event cameras offer significant advantages such as high temporal resolution, low latency, wide dynamic range, and low power consumption, making them particularly suitable for visual perception tasks in extreme environments such as fast-moving scenes, strong light, or low light. Because its output data is generated only when pixel brightness changes, event cameras effectively reduce redundant information, enabling efficient information transmission and processing. This provides new possibilities for achieving real-time dynamic scene understanding and high-speed target perception.
[0003] Currently, when using event cameras to track targets, the event stream is typically converted into pseudo-frames with fixed time intervals. However, the target's movement speed varies, and the image frames generated by existing methods are inconsistent under different conditions, easily losing spatiotemporal information, thus making it difficult to accurately and stably extract the target trajectory. Summary of the Invention
[0004] In view of this, this application provides an event camera target tracking method, device, storage medium and electronic device, the main purpose of which is to avoid losing spatiotemporal information and accurately and stably extract the target trajectory.
[0005] According to a first aspect of this application, an event camera target tracking method is provided, the method comprising:
[0006] Acquire the raw event stream of moving targets captured by the event camera;
[0007] Based on the spatiotemporal neighborhood consistency strategy and / or polarity time difference constraint strategy, the original event stream is denoised to obtain the denoised event stream.
[0008] Based on multi-dimensional statistical features, the event aggregation time window is dynamically adjusted, and the denoised event stream is aggregated based on the adjusted event aggregation time window to obtain the aggregated event stream;
[0009] Motion compensation is applied to the aggregated event stream to obtain a motion-compensated event stream;
[0010] Based on the motion-compensated event stream and the preset motion trajectory prediction model, the trajectory information of the moving target is predicted.
[0011] According to a second aspect of this application, an event camera target tracking device is provided, the device comprising:
[0012] The acquisition unit is used to acquire the raw event stream of moving targets captured by the event camera;
[0013] Based on the spatiotemporal neighborhood consistency strategy and / or polarity time difference constraint strategy, the original event stream is denoised to obtain the denoised event stream.
[0014] The adjustment unit is used to dynamically adjust the event aggregation time window based on multi-dimensional statistical features, and to aggregate the denoised event stream based on the adjusted event aggregation time window to obtain the aggregated event stream.
[0015] The compensation unit is used to perform motion compensation on the aggregated event stream to obtain a motion-compensated event stream;
[0016] The prediction unit is used to predict the trajectory information of the moving target based on the motion-compensated event stream and the preset motion trajectory prediction model.
[0017] According to a third aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described event camera target tracking method.
[0018] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described event camera target tracking method.
[0019] By employing the above technical solutions, this application provides an event camera target tracking method, apparatus, storage medium, and electronic device. Compared with existing technologies, it can dynamically adjust the event aggregation time window according to the event rate. When the target or camera moves quickly and the event density is high, the time window is automatically shortened to prevent motion blur. When the scene is static or events are sparse, the aggregation time window is appropriately extended to enhance spatial information. Through adaptive time aggregation, this application can maintain the temporal continuity and spatial integrity of event information under different motion states of the target, thereby avoiding the loss of spatiotemporal information and enabling accurate and stable extraction of target trajectories. In addition, this application uses a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to denoise the original event stream, which can remove false events and interference information in the event stream and retain real brightness change events, i.e., effective signals. This can more accurately reflect the movement and changes of the target in the scene, providing a more reliable data foundation for subsequent trajectory prediction tasks and further improving the accuracy of target tracking. At the same time, by performing motion compensation on the aggregated event stream, this application can align events of the same target at the reference time, thereby eliminating edge blur caused by motion.
[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0022] Figure 1 This illustration shows a schematic diagram of a system architecture for target tracking using an event camera, provided in an embodiment of this application.
[0023] Figure 2 A flowchart illustrating an event camera target tracking method provided in an embodiment of this application is shown.
[0024] Figure 3 A schematic diagram illustrating adaptive time aggregation under different event densities provided in an embodiment of this application is shown;
[0025] Figure 4 A schematic diagram illustrating the use of motion compensation to eliminate artifacts according to an embodiment of this application is shown;
[0026] Figure 5 A schematic diagram of the structure of the hybrid attention network provided in an embodiment of this application is shown;
[0027] Figure 6 A schematic diagram of the training method for the preset motion trajectory prediction model provided in an embodiment of this application is shown.
[0028] Figure 7 A schematic diagram of the structure of an event camera target tracking device provided in an embodiment of this application is shown. Detailed Implementation
[0029] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0030] The image frames generated by existing methods vary in quality under different conditions and are prone to losing spatiotemporal information, making it difficult to accurately and stably extract the target trajectory.
[0031] To address the aforementioned problems, embodiments of the present invention provide a system architecture for target tracking using an event camera, such as... Figure 1 As shown, the system architecture mainly includes an adaptive temporal aggregation module, a motion compensation and feature representation module, an event motion transformer module, a differentiable geometric tracking and association module, and a self-supervised learning and optimization module. The adaptive temporal aggregation module dynamically adjusts the time window size according to the rate of change of the event stream. It automatically shortens the aggregation duration when the event density is high and expands the time window when the events are sparse, thus maintaining information integrity and temporal accuracy under different motion speeds. The motion compensation and feature representation module uses short-term motion prediction of the target or camera to perform spatiotemporal back-projection (motion) of event points. The event motion transformer module eliminates edge blurring caused by target motion. Furthermore, it constructs the compensated events into a learnable voxel structure and extracts multi-scale spatial and temporal dynamic features using a deep neural network. The event motion transformer module introduces a feature modeling structure based on a self-attention mechanism (preset motion trajectory prediction model) to globally correlate the temporal and spatial distribution of the event stream, automatically distinguishing between noise events and target motion events. It also adds a motion prediction branch to the network to generate instantaneous velocity and direction information of the target. The differentiable geometric tracking and association module optimizes the trajectory based on the predicted position, scale, and orientation parameters of the moving target using a differentiable reprojection model, achieving joint updates of the event domain and target geometric parameters. It then uses the Hungarian algorithm or Bayesian filtering to achieve multi-target data association, ensuring the continuity and stability of the trajectory. The self-supervised learning and optimization module achieves self-supervised training under unlabeled data conditions by reconstructing a differentiable loss function for the event distribution. It also introduces multimodal contrastive loss to improve the model's generalization ability under different lighting and motion scenarios.
[0032] based on Figure 1The system architecture shown in this invention provides an event camera target tracking method, which mainly uses a trained motion trajectory prediction model for target tracking, such as... Figure 2 As shown, the method includes:
[0033] Step 10: Obtain the raw event stream of moving targets captured by the event camera.
[0034] The moving targets include rapidly moving hands, robotic arms, rapidly rotating fan blades, and speeding vehicles. It should be noted that the moving targets in this embodiment of the invention are not limited to the above list; they include all targets that fit the high temporal resolution characteristics of the event camera and are easy to track. Furthermore, the original event stream is an asynchronous data sequence, containing multiple events.
[0035] In this embodiment of the invention, an event camera is used to acquire brightness change events in the scene at a microsecond-level temporal resolution. During the acquisition process, the pixel coordinates, timestamp, and polarity of each event are recorded. The original event stream can be specifically represented as follows: ,in, Indicates an event pixel coordinates, Indicates an event timestamp, Indicates an event polarity, polarity Used to indicate the direction of brightness change. N1 This indicates the number of original events. In this embodiment of the invention, the original event stream of the moving target is an abnormal data sequence, which serves as the original input.
[0036] Step 20: Based on the spatiotemporal neighborhood consistency strategy and / or polarity time difference constraint strategy, the original event stream is denoised to obtain the denoised event stream.
[0037] In this embodiment of the invention, since the original event stream may contain issues such as photosensitive noise, flickering false triggers, and background drift, it is necessary to perform noise reduction processing before aggregating the event stream. This embodiment of the invention provides several implementation methods for this noise reduction process. The first implementation method uses a spatiotemporal neighborhood consistency strategy to reduce noise in the original event stream; the second implementation method uses a polarity time difference constraint strategy to reduce noise in the original event stream; and the third implementation method uses both a spatiotemporal neighborhood consistency strategy and a polarity time difference constraint strategy to reduce noise in the original event stream.
[0038] For the first implementation method, namely, the process of denoising the original event stream using a spatiotemporal neighborhood consistency strategy, the process includes: for any event within the current event aggregation time window of the original event stream, identifying other events within the current event aggregation time window that have the same polarity as the event; calculating the time difference based on the timestamp of the event and the timestamps of the other events; calculating the same-polarity event density corresponding to the event within the current event aggregation time window based on the time difference and the width of the current event aggregation time window; and determining whether the event is a noise event based on the same-polarity event density.
[0039] Specifically, for each event in the original event stream Calculate the density of events of the same polarity within the current event aggregation time window. When the density of events of the same polarity Less than the preset event density ,Right now < At that time, the event Events identified as noise are removed. The preset event density can be set according to actual business needs, and this embodiment of the invention does not impose specific limitations on it.
[0040] Same polarity event density The specific calculation formula is as follows:
[0041]
[0042] Here, II(·) is an indicator function, the purpose of which is to filter events with the same polarity, ensuring that only events with the same polarity are included in the density calculation. If events polarity With the event polarity Same (i.e.) If the condition is met, the function returns 1; otherwise, it returns 0. It is a Gaussian kernel function used to weight events based on time differences. and These are events and Timestamp. This is a time scale parameter used to control the rate of time decay, i.e., the width of the current event aggregation time window. This exponential term assigns time to events that are closer in time. Events with higher weight, i.e., time difference The smaller the value, the closer the weight is to 1; time difference The larger the value, the lower the weighting index gradually becomes (to 0). Indicates The spatiotemporal neighborhood centered on the center has a spatial radius of . The time half-window is 。
[0043] Therefore, based on the above calculation formula and judgment method, it can be determined whether each event within the current event aggregation time window is a noise event.
[0044] The second implementation method, which employs a polarity time difference constraint strategy to denoise the original event stream, includes: determining the time interval between consecutive events at the same pixel position based on the pixel coordinates and timestamps corresponding to each event in the original event stream; determining whether sensor jitter occurs or whether the consecutive events are false triggering events based on the time interval; and filtering the consecutive events to obtain the denoised event stream if sensor jitter occurs or the consecutive events are false triggering events.
[0045] Specifically, the time interval of consecutive events at the same pixel location is constrained; if the time interval... Less than the preset time interval ,Right now If such an event occurs, it is considered a sensor jitter or a false trigger event. The preset time interval can be set according to actual business needs; this embodiment of the invention does not impose specific limitations on it.
[0046] For the third implementation method, which simultaneously employs a spatiotemporal neighborhood consistency strategy and a polarity time difference constraint strategy to denoise the original event stream, the spatiotemporal neighborhood consistency strategy can be used first for denoising, followed by a polarity time difference constraint strategy for further denoising, or vice versa. This embodiment of the invention simultaneously employs both spatiotemporal neighborhood consistency and polarity time difference constraint strategies for denoising, which further ensures the denoising effect and improves data quality.
[0047] The denoised event stream obtained from the three implementation methods described above can be represented as follows: ,in, N2 This indicates the number of events after noise reduction.
[0048] This invention employs a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to denoise the original event stream. This removes false events and interference information from the event stream, retaining the real brightness change events, i.e., the effective signals. As a result, it can more accurately reflect the motion and changes of targets in the scene, providing a more reliable data foundation for subsequent trajectory prediction tasks and further improving the accuracy of target tracking.
[0049] Step 30: Based on multi-dimensional statistical features, dynamically adjust the event aggregation time window, and aggregate the denoised event stream based on the adjusted event aggregation time window to obtain the aggregated event stream.
[0050] The multi-dimensional statistical features include the event rate, noise level, and event density map within the current event aggregation time window, as well as the motion speed of moving targets within the historical event aggregation time window.
[0051] To overcome the shortcomings of existing methods that easily lose spatiotemporal information, embodiments of the present invention dynamically adjust the event aggregation time window based on the aforementioned multi-dimensional statistical characteristics, such as... Figure 3 As shown. Based on this, step 30 specifically includes: statistically analyzing the event rate and noise level within the current event aggregation time window, and the motion speed of the moving target within the historical event aggregation time window; calculating the event density map within the current event aggregation time window; determining the event rate, noise level, motion speed, and event density map as the multi-dimensional statistical features; calculating a global dynamic coefficient based on the event rate, noise level, motion speed, and event density map; calculating the optimal aggregation window width for the current time based on the global dynamic coefficient; and dynamically adjusting the event aggregation time window based on the optimal aggregation window width for the current time.
[0052] In the specific calculation of the global dynamic coefficient, the average event density is calculated based on the event density corresponding to each event in the event density map, and the global dynamic coefficient is calculated based on the average event density, the event rate, the noise level, and the motion speed.
[0053] The specific formula for calculating the global dynamic coefficient is as follows.
[0054]
[0055] in, Represents the global dynamic coefficient. Represents the event rate within the current event aggregation time window. The speed of the moving target within the time window of the historical event aggregation. This represents the average event density within the current event aggregation time window. This represents the noise level within the current event aggregation time window. Represents the reference event rate. Represents reference motion speed. Represents the reference noise level. Represents the density of reference events. , , , These represent the weights of each item and can be determined through self-supervised training or experimental tuning. The reference event rate, reference motion speed, reference noise level, and reference event density can be set according to actual business needs.
[0056] Furthermore, when calculating the event density map within the current event aggregation time window, the pixel coordinate difference between each event and other events within the current event aggregation time window is calculated respectively; based on the width of the current event aggregation time window and the pixel coordinate difference, the event density corresponding to each event within the current event aggregation time window is determined; and based on the event density corresponding to each event within the current event aggregation time window, the event density map is determined.
[0057] The specific formula for calculating the event density for each event is as follows:
[0058]
[0059] in, This represents the event density corresponding to a certain event. Represents the width of the current event aggregation time window, such as 5 pixels or 9 pixels. The pixel coordinates representing a specific event. Pixel coordinates representing other events, This is an indicator function. Therefore, the event density corresponding to each event within the current event aggregation time window can be calculated using the above formula, thus determining the event density map. When calculating the optimal aggregation window width for the current time based on the global dynamic coefficient, the optimal aggregation window width for the current time can be calculated based on the standard time window and the global dynamic coefficient, as shown in the following formula.
[0060]
[0061] in, The optimal aggregation window length for the current time. Represents a standard time window. Represents the global dynamic coefficient.
[0062] In this embodiment of the invention, when the moving target or camera is moving fast and the event density is high, the time window is automatically shortened to prevent motion blur; when the scene is static or events are sparse, the aggregation time window is appropriately extended to enhance spatial information. Through the above-mentioned adaptive temporal aggregation, the temporal continuity and spatial integrity of event information can be maintained under different motion states.
[0063] This invention achieves a dynamic balance between spatiotemporal resolution and noise suppression by dynamically updating the time window, which automatically shortens the window when the target is moving violently or the noise is high, and automatically extends the window when the scene is stable.
[0064] Step 40: Perform motion compensation on the aggregated event stream to obtain a motion-compensated event stream.
[0065] In this embodiment of the invention, after aggregating the event stream, spatiotemporal back-projection compensation is performed on the aggregated event stream based on the pixel coordinates and timestamps corresponding to each event in the aggregated event stream, the relative motion velocity between the moving target and the event camera, and a reference time point, to obtain the motion-compensated event stream. For events within the current aggregation time window, the specific spatiotemporal back-projection formula for event motion compensation is as follows:
[0066]
[0067] in, These are the original pixel coordinates and timestamp of the event. The relative velocity between the moving target and the event camera. For reference time points. For example... Figure 4 As shown, the compensated event flow is obtained. .
[0068] This invention eliminates edge blurring caused by motion by performing spatiotemporal back-projection compensation on events within the current time window, so that events of the same target are aligned at the reference time.
[0069] Step 50: Based on the motion-compensated event stream and the preset motion trajectory prediction model, predict the trajectory information of the moving target.
[0070] The preset motion trajectory prediction model is a hybrid attention network, which includes a local convolutional encoder, a temporal attention network, a spatial attention network, and a motion attention network.
[0071] In this embodiment of the invention, after motion compensation of the event stream, the motion-compensated event stream is mapped to a multi-channel voxel grid according to its temporal distribution to obtain voxel grid data corresponding to the motion-compensated event stream; the voxel grid data is input to the local convolutional encoder for feature encoding to obtain the encoded features corresponding to the voxel grid data; the encoded features are input to the spatial attention network to extract the edge texture features of the moving target; the encoded features are input to the temporal attention network to capture the dynamic change pattern of the moving target; the encoded features are input to the motion attention network to estimate the local motion direction of the moving target; based on the edge texture features, the dynamic change pattern, and the local motion direction, the position information, size information, and motion direction of the moving target are predicted; and the trajectory information of the moving target is determined according to the position information, size information, and motion direction of the moving target.
[0072] In this embodiment of the invention, the motion-compensated event stream is mapped to a multi-channel voxel grid according to its temporal distribution to maintain temporal continuity and polarity characteristics. The temporal kernel weights of the voxel channels are adaptively learned by a neural network, thereby achieving optimal encoding for different motion models. The formula for constructing the voxel grid is as follows:
[0073]
[0074] in, It is the polarity of the event. Indicates the target time slot τ With the time of the event The distance between them, K(·), is a learnable temporal kernel function that smoothly distributes the contribution of an event to its neighboring time bins. The output of this formula is the constructed voxel grid data. .
[0075] Furthermore, the compensated, speed-up mesh data is input into a hybrid attention network, which includes a local convolutional encoder, a temporal attention network, a spatial attention network, and a motion attention network. Specifically, such as... Figure 5 As shown, a lightweight convolutional neural network is used for feature encoding. Then, a spatial attention module is used to extract edge texture features in the spatial dimension, while a temporal attention module is used to capture the dynamic change pattern of the target in the temporal dimension. A motion attention module is used to estimate the local motion direction and velocity information, thereby achieving implicit modeling of the target's motion trend and suppressing the interference of noise events.
[0076] By using voxel grid data The input is fed into a hybrid attention network, which outputs a multi-scale spatial feature map F of the moving target and motion prediction parameters, including the position information, size information and direction of motion of the moving target.
[0077] Through the above process, the system of this embodiment can output the position information, size information, direction of movement, and confidence score of a moving target in real time. By performing short-term prediction and state smoothing of the moving target's trajectory and dynamically updating the tracking structure based on subsequent event data, continuous, low-latency target tracking can be achieved. Furthermore, by combining temporal attention networks, spatial attention networks, and motion attention networks, this embodiment can jointly analyze event distribution, enabling the model to understand the direction of movement and maintain the target structure, thereby achieving spatiotemporally consistent target modeling.
[0078] Furthermore, embodiments of the present invention also provide a training method for a preset motion trajectory prediction model, for which, as follows: Figure 6 As shown, it includes:
[0079] Step 60: Obtain the historical event stream of the moving target and construct an initial motion trajectory prediction model.
[0080] The initial motion trajectory prediction model is an initial hybrid attention network, which includes an initial local convolutional encoder, an initial temporal attention network, an initial spatial attention network, and an initial motion attention network.
[0081] In this embodiment of the invention, the historical event stream of moving targets captured by the event camera is collected and used as training samples. At the same time, a neural network algorithm is used to build an initial hybrid attention network.
[0082] Step 70: Based on the historical event stream and the initial motion trajectory prediction model, predict the historical trajectory information of the moving target.
[0083] In this embodiment of the invention, after acquiring the historical event stream, the historical event stream is denoised based on a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to obtain a denoised historical event stream. Then, the denoised historical event stream is aggregated to obtain an aggregated historical event stream. Next, motion compensation is performed on the aggregated historical event stream to obtain a motion-compensated historical event stream. Finally, based on the motion-compensated historical event stream and a preset motion trajectory prediction model, the position information, size information, and motion direction of the moving target are predicted, thereby determining the historical trajectory information of the moving target.
[0084] Step 80: Based on the historical trajectory information, construct a differentiable reprojection error function;
[0085] In this embodiment of the invention, a differentiable projection model is used to project historical trajectory information into the event domain, generating an event probability distribution. Based on the event probability distribution and the observed event probability distribution, a differentiable reprojection error function is constructed. The specific calculation formula for the differentiable reprojection error function is as follows:
[0086]
[0087] in, This indicates that the target parameter (rotation angle) is used to represent the target parameter (rotation angle). ,scale ,center The generated event probability distribution, The pixel coordinates of the historical event. Represents the probability distribution of observed events. This represents the differential reprojection error.
[0088] Step 90: Based on the differentiable reprojection error function, iteratively train the initial motion trajectory prediction model to construct the preset motion trajectory prediction model.
[0089] For embodiments of the present invention, in a real environment, the differentiable reprojection error function is used. Differential calculations and gradient direction propagation are performed to optimize model parameters. This embodiment of the invention achieves joint optimization of the event domain and target geometric parameters through a differentiable reprojection model, dynamically suppressing noise events and maintaining stable tracking in scenarios such as high-speed motion, strong light, and occlusion.
[0090] In multi-target scenarios, a matching cost matrix based on spatiotemporal distance and appearance features is used, and the Hungarian algorithm or Bayesian filtering is employed to achieve trajectory association and update, thereby ensuring the continuity of moving targets and tracking stability.
[0091] Under conditions of unlabeled or weakly labeled data, embodiments of the present invention reconstruct a differentiable loss function for the event distribution. Self-supervised training is performed to enable the model to automatically learn the motion patterns of the event flow and the target contour structure.
[0092] The embodiments of the present invention, through the design of self-supervised loss, can reduce the dependence on manually labeled data, support unsupervised or weakly supervised training, and thus significantly reduce training costs.
[0093] To further improve the training accuracy of the model, this embodiment of the invention also introduces a multimodal contrastive learning mechanism. Specifically, when frame-based images are present, cross-modal feature consistency loss is used to further enhance the robustness of the model in complex lighting, occlusion, and high-speed motion scenarios. Based on this, the method includes: acquiring image data and inertial sensing data of the moving target; and constructing a multimodal contrastive consistency loss function and a trajectory smoothing regularization loss function based on the image data and the inertial sensing data, respectively.
[0094] The differentiable reprojection error function, the multimodal contrast consistency loss function, and the trajectory smoothing regularization loss function are weighted and summed to obtain the total loss function. Based on the total loss function, the initial motion trajectory prediction model is trained to construct the preset motion trajectory prediction model. The specific formula for the total loss function of model learning is:
[0095]
[0096] in, For the total loss function, Let be a differentiable reprojection error function. The multimodal contrast consistency loss function is... Let the trajectory smoothing regularization loss function be used. and It is the weighting coefficient.
[0097] The above-described technical solutions of this invention can run directly on the asynchronous data stream output by the event camera, or they can be combined with traditional image and inertial sensor data to achieve end-to-end robust target tracking.
[0098] This invention can dynamically adjust the event aggregation time window based on the event rate. When the target or camera moves quickly and the event density is high, the time window is automatically shortened to prevent motion blur. When the scene is static or events are sparse, the aggregation time window is appropriately extended to enhance spatial information. Through adaptive temporal aggregation, this invention can maintain the temporal continuity and spatial integrity of event information under different target motion states, thereby avoiding the loss of spatiotemporal information and enabling accurate and stable extraction of target trajectories. Furthermore, this invention employs a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to denoise the original event stream, eliminating false events and interference information while retaining genuine brightness change events, i.e., effective signals. This more accurately reflects the movement and changes of targets in the scene, providing a more reliable data foundation for subsequent trajectory prediction tasks and further improving target tracking accuracy. Simultaneously, by performing motion compensation on the aggregated event stream, this invention can align events of the same target at the reference time, thereby eliminating edge blur caused by motion.
[0099] Furthermore, as Figure 2 and Figure 6 The specific implementation of the method shown in this embodiment provides an event camera target tracking device, such as... Figure 7 As shown, the device includes: an acquisition unit 101, a noise reduction unit 102, an adjustment unit 103, a compensation unit 104, and a prediction unit 105.
[0100] The acquisition unit 101 can be used to acquire the original event stream of moving targets captured by the event camera.
[0101] The noise reduction unit 102 can be used to reduce the noise of the original event stream based on the spatiotemporal neighborhood consistency strategy and / or the polarity time difference constraint strategy to obtain the noise-reduced event stream.
[0102] The adjustment unit 103 can be used to dynamically adjust the event aggregation time window based on multi-dimensional statistical features, and aggregate the denoised event stream based on the adjusted event aggregation time window to obtain the aggregated event stream.
[0103] The compensation unit 104 can be used to perform motion compensation on the aggregated event stream to obtain a motion-compensated event stream.
[0104] The prediction unit 105 can be used to predict the trajectory information of the moving target based on the motion-compensated event stream and the preset motion trajectory prediction model.
[0105] In some embodiments, the noise reduction unit 102 may be specifically configured to: determine other events with the same polarity as any event within the current event aggregation time window of the original event stream; calculate the time difference based on the timestamp of the event and the timestamps of the other events; calculate the density of events of the same polarity corresponding to the event within the current event aggregation time window based on the time difference and the width of the current event aggregation time window; determine whether the event is a noise event based on the density of events of the same polarity; if the event is a noise event, remove the event to obtain the noise-reduced event stream; and / or determine the time interval of consecutive events at the same pixel position based on the pixel coordinates and timestamps of each event in the original event stream; determine whether sensor jitter occurs or whether the consecutive events are false triggering events based on the time interval; if sensor jitter occurs or the consecutive events are false triggering events, filter the consecutive events to obtain the noise-reduced event stream.
[0106] In some embodiments, the adjustment unit 103 includes: a statistics module, a calculation module, a determination module, and an adjustment module.
[0107] The statistics module can be used to count the event rate and noise level within the current event aggregation time window, as well as the movement speed of the moving target within the historical event aggregation time window.
[0108] The calculation module can be used to calculate the event density map within the current event aggregation time window.
[0109] The determining module can be used to determine the event rate, the noise level, the motion speed, and the event density map as the multi-dimensional statistical features.
[0110] The calculation module can also be used to calculate global dynamic coefficients based on the event rate, the noise level, the motion speed, and the event density map.
[0111] The calculation module can also be used to calculate the optimal aggregation window width for the current time based on the global dynamic coefficients.
[0112] The adjustment module can also be used to dynamically adjust the event aggregation time window based on the optimal aggregation window width for the current time.
[0113] In some embodiments, the calculation module may be specifically used to calculate the pixel coordinate difference between each event and other events within the current event aggregation time window; determine the event density corresponding to each event within the current event aggregation time window based on the width of the current event aggregation time window and the pixel coordinate difference; and determine the event density map based on the event density corresponding to each event within the current event aggregation time window.
[0114] In some embodiments, the calculation module may also be specifically used to calculate the average event density based on the event density corresponding to each event in the event density map; and to calculate the global dynamic coefficient based on the average event density, the event rate, the noise level, and the motion speed.
[0115] In some embodiments, the calculation module may also be specifically used to calculate the optimal aggregation window width for the current time based on the standard time window and the global dynamic coefficient.
[0116] In some embodiments, the compensation unit 104 may be specifically used to perform spatiotemporal back-projection compensation on the aggregated event stream based on the pixel coordinates and timestamps corresponding to each event in the aggregated event stream, the relative motion speed between the moving target and the event camera, and a reference time point, to obtain the motion-compensated event stream.
[0117] In some embodiments, the preset motion trajectory prediction model is a hybrid attention network, which includes a local convolutional encoder, a temporal attention network, a spatial attention network, and a motion attention network. The prediction unit 105 can be specifically used to map the motion-compensated event stream to a multi-channel voxel grid according to its temporal distribution, obtaining voxel grid data corresponding to the motion-compensated event stream; input the voxel grid data into the local convolutional encoder for feature encoding, obtaining encoded features corresponding to the voxel grid data; input the encoded features into the spatial attention network to extract the edge texture features of the moving target; input the encoded features into the temporal attention network to capture the dynamic change pattern of the moving target; input the encoded features into the motion attention network to estimate the local motion direction of the moving target; based on the edge texture features, the dynamic change pattern, and the local motion direction, predict the position information, size information, and motion direction of the moving target; and determine the trajectory information of the moving target based on the position information, size information, and motion direction of the moving target.
[0118] In some embodiments, the apparatus further includes a training unit.
[0119] The training unit can be used to acquire the historical event stream of the moving target and construct an initial motion trajectory prediction model; predict the historical trajectory information of the moving target based on the historical event stream and the initial motion trajectory prediction model; construct a differentiable reprojection error function based on the historical trajectory information; and iteratively train the initial motion trajectory prediction model according to the differentiable reprojection error function to construct the preset motion trajectory prediction model.
[0120] The training unit can be specifically used to project the historical trajectory information into the event domain using a differentiable projection model to generate an event probability distribution; and to construct a differentiable reprojection error function based on the event probability distribution and the observed event probability distribution.
[0121] It should be noted that other corresponding descriptions of the functional units involved in the event camera target tracking device provided in this embodiment can be found in [reference needed]. Figure 2 and Figure 6 The corresponding descriptions in [the document] will not be repeated here.
[0122] Based on the above, Figure 2 and Figure 6 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 2 and Figure 6The event camera target tracking method shown.
[0123] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) and includes several instructions to cause an electronic device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0124] Based on the above, Figure 2 and Figure 6 The method shown, and Figure 7 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 2 and Figure 5 The event camera target tracking method shown.
[0125] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0126] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0127] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0129] This invention can dynamically adjust the event aggregation time window based on the event rate. When the target or camera moves quickly and the event density is high, the time window is automatically shortened to prevent motion blur. When the scene is static or events are sparse, the aggregation time window is appropriately extended to enhance spatial information. Through adaptive temporal aggregation, this invention can maintain the temporal continuity and spatial integrity of event information under different target motion states, thereby avoiding the loss of spatiotemporal information and enabling accurate and stable extraction of target trajectories. Furthermore, this invention employs a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to denoise the original event stream, eliminating false events and interference information while retaining genuine brightness change events, i.e., effective signals. This more accurately reflects the movement and changes of targets in the scene, providing a more reliable data foundation for subsequent trajectory prediction tasks and further improving target tracking accuracy. Simultaneously, by performing motion compensation on the aggregated event stream, this invention can align events of the same target at the reference time, thereby eliminating edge blur caused by motion.
[0130] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0131] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. An event camera target tracking method, characterized by, include: Acquire the raw event stream of moving targets captured by the event camera; Based on the spatiotemporal neighborhood consistency strategy and / or polarity time difference constraint strategy, the original event stream is denoised to obtain the denoised event stream. Based on multi-dimensional statistical features, the event aggregation time window is dynamically adjusted, and the denoised event stream is aggregated based on the adjusted event aggregation time window to obtain the aggregated event stream; Motion compensation is applied to the aggregated event stream to obtain a motion-compensated event stream; Based on the motion-compensated event stream and the preset motion trajectory prediction model, predict the trajectory information of the moving target; The method of dynamically adjusting the event aggregation time window based on multi-dimensional statistical features includes: The event rate and noise level within the current event aggregation time window, as well as the motion speed of the moving target within the historical event aggregation time window, are statistically analyzed. Calculate the event density map within the current event aggregation time window; The event rate, the noise level, the motion speed, and the event density map are defined as the multi-dimensional statistical features; Calculate the average event density based on the event density corresponding to each event in the event density map; The global dynamic coefficients are calculated based on the average event density, the event rate, the noise level, and the motion speed, as well as the reference event density, reference event rate, reference noise level, and reference motion speed. Based on the standard time window and the global dynamic coefficient, calculate the optimal aggregation window width for the current time; The event aggregation time window is dynamically adjusted based on the optimal aggregation window width at the current time. The method further includes: Obtain the historical event stream of the moving target and construct an initial motion trajectory prediction model; Based on the historical event stream and the initial motion trajectory prediction model, predict the historical trajectory information of the moving target; Based on the historical trajectory information, a differentiable reprojection error function is constructed, wherein the historical trajectory information is projected onto the event domain using a differentiable projection model to generate an event probability distribution; and a differentiable reprojection error function is constructed based on the event probability distribution and the observed event probability distribution. Image data and inertial sensing data of the moving target are collected; based on the image data and the inertial sensing data, a multimodal contrast consistency loss function and a trajectory smoothing regularization loss function are constructed respectively; The total loss function is obtained by weighted summing of the differentiable reprojection error function, the multimodal contrast consistency loss function, and the trajectory smoothing regularization loss function. Based on the total loss function, the initial motion trajectory prediction model is trained to construct the preset motion trajectory prediction model.
2. The method according to claim 1, characterized in that, Based on a spatiotemporal neighborhood consistency strategy, the original event stream is denoised to obtain a denoised event stream, including: For any event within the current event aggregation time window of the original event stream, determine other events within the current event aggregation time window that have the same polarity as the given event; Calculate the time difference based on the timestamp of any one of the events and the timestamps of the other events; Based on the time difference and the width of the current event aggregation time window, calculate the same-polarity event density corresponding to any event within the current event aggregation time window; Based on the density of events of the same polarity, determine whether any one of the events is a noise event; If any of the events is a noisy event, then the event is removed to obtain the denoised event stream.
3. The method according to claim 1, characterized in that, Based on a polarity time difference constraint strategy, the original event stream is denoised to obtain a denoised event stream, including: Based on the pixel coordinates and timestamps corresponding to each event in the original event stream, the time interval between consecutive events at the same pixel position is determined; Based on the time interval, it is determined whether sensor jitter has occurred or whether the continuous events are false triggering events; If sensor jitter occurs or the continuous events are false triggering events, the continuous events are filtered to obtain the noise-reduced event stream.
4. The method according to claim 1, characterized in that, The calculation of the event density map within the current event aggregation time window includes: Calculate the pixel coordinate difference between each event and other events within the current event aggregation time window; Based on the width of the current event aggregation time window and the pixel coordinate difference, determine the event density corresponding to each event within the current event aggregation time window; The event density map is determined based on the event density corresponding to each event within the current event aggregation time window.
5. The method according to claim 1, characterized in that, The step of performing motion compensation on the aggregated event stream to obtain a motion-compensated event stream includes: Based on the pixel coordinates and timestamps corresponding to each event in the aggregated event stream, the relative motion velocity between the moving target and the event camera, and a reference time point, spatiotemporal back-projection compensation is performed on the aggregated event stream to obtain the motion-compensated event stream; and / or The preset motion trajectory prediction model is a hybrid attention network, which includes a local convolutional encoder, a temporal attention network, a spatial attention network, and a motion attention network. The step of predicting the trajectory information of the moving target based on the motion-compensated event stream and the preset motion trajectory prediction model includes: The motion-compensated event stream is mapped to a multi-channel voxel grid according to the time distribution to obtain the voxel grid data corresponding to the motion-compensated event stream. The voxel grid data is input into the local convolutional encoder for feature encoding to obtain the encoded features corresponding to the voxel grid data; The encoded features are input into the spatial attention network to extract the edge texture features of the moving target; The encoded features are input into the temporal attention network to capture the dynamic change patterns of the moving target; The encoded features are input into the motion attention network to estimate the local motion direction of the moving target; Based on the edge texture features, the dynamic change patterns, and the local motion direction, the position information, size information, and motion direction of the moving target are predicted. Based on the position information, size information, and direction of movement of the moving target, the trajectory information of the moving target is determined.
6. An event camera target tracking device, characterized in that, include: The acquisition unit is used to acquire the raw event stream of moving targets captured by the event camera; A noise reduction unit is used to reduce the noise of the original event stream based on a spatiotemporal neighborhood consistency strategy and / or a polarity time difference constraint strategy to obtain a noise-reduced event stream. The adjustment unit is used to dynamically adjust the event aggregation time window based on multi-dimensional statistical features, and to aggregate the denoised event stream based on the adjusted event aggregation time window to obtain the aggregated event stream. The compensation unit is used to perform motion compensation on the aggregated event stream to obtain a motion-compensated event stream; The prediction unit is used to predict the trajectory information of the moving target based on the motion-compensated event stream and the preset motion trajectory prediction model. The adjustment unit is specifically used to statistically analyze the event rate and noise level within the current event aggregation time window, as well as the movement speed of the moving target within the historical event aggregation time window. Calculate the event density map within the current event aggregation time window; determine the event rate, the noise level, the motion speed, and the event density map as the multi-dimensional statistical features; Calculate the average event density based on the event density corresponding to each event in the event density map; calculate the global dynamic coefficient based on the average event density, the event rate, the noise level, and the motion speed, as well as the reference event density, reference event rate, reference noise level, and reference motion speed. Based on the standard time window and the global dynamic coefficient, calculate the optimal aggregation window width for the current time; and dynamically adjust the event aggregation time window according to the optimal aggregation window width for the current time. The training unit is used to acquire the historical event stream of the moving target and construct an initial motion trajectory prediction model; based on the historical event stream and the initial motion trajectory prediction model, predict the historical trajectory information of the moving target; and based on the historical trajectory information, construct a differentiable reprojection error function, wherein the differentiable projection model is used to project the historical trajectory information onto the event domain to generate an event probability distribution. Based on the event probability distribution and the observed event probability distribution, a differentiable reprojection error function is constructed; image data and inertial sensing data of the moving target are collected; based on the image data and the inertial sensing data, a multimodal contrast consistency loss function and a trajectory smoothing regularization loss function are constructed respectively; the differentiable reprojection error function, the multimodal contrast consistency loss function, and the trajectory smoothing regularization loss function are weighted and summed to obtain a total loss function; based on the total loss function, the initial motion trajectory prediction model is trained to construct the preset motion trajectory prediction model.
7. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
8. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Low-delay moving target detection method and system for navigation
CN120260016A
Autonomous positioning method and system of underwater robot, electronic equipment, storage medium and program product
CN120997287A