European-hyperbolic motion perception-based event camera target detection method
By employing graph convolutional networks based on B-spline kernel functions and hyperbolic graph convolutional networks in event cameras, combined with dynamic sampling mechanisms and Markov vector field modeling, the problems of noise suppression and hierarchical structure representation in target detection of event cameras are solved, achieving high-precision target detection in complex scenes, which is applicable to fields such as autonomous driving and robot navigation.
Patent Information
- Application Number
- CN202511004016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing event camera target detection methods have shortcomings in noise interference suppression, graph modeling, and hierarchical structure representation. They are difficult to maintain high accuracy and robustness in complex scenes, especially under high dynamic or low illumination conditions. Noise interference seriously affects detection accuracy, local mapping methods are difficult to capture global semantic consistency and behavioral chain integrity, and traditional Euclidean space modeling is difficult to represent multi-scale structures.
We employ a graph convolutional network based on B-spline kernel function to extract local perceptual features and project them onto hyperbolic space. We then combine a learnable hyperbolic graph convolutional network with curvature to extract global event-dependent features. We reduce noise through a dynamic event sampling mechanism, construct an event graph, and use Markov vector fields to describe the consistency of event motion. Finally, we combine a graph neural network architecture based on Euclidean and hyperbolic space for object detection.
It significantly reduces background noise, enhances system robustness, improves target detection accuracy and generalization ability in complex scenes, accurately characterizes cross-temporal and spatiotemporal semantic dependencies and multi-scale structures between events, and improves the real-time target perception performance of the system in complex environments such as autonomous driving.
Smart Images

Figure CN120931941A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer vision technology, and in particular to an event camera target detection method based on Euclidean hyperbolic motion perception. Background Technology
[0002] Event-based camera-based object detection is a method for real-time identification and localization of dynamic objects using asynchronous event stream data. Event cameras trigger events through pixel-level brightness changes, outputting sparse spatiotemporal data (x, y, t, p), offering advantages such as microsecond-level latency, high dynamic range, and low power consumption. Compared to traditional frame cameras, event cameras perform better in high-speed motion and low-light scenarios, and this technology is widely used in autonomous driving, robot navigation, and virtual reality. The detection algorithm needs to process unstructured event streams; existing mainstream methods can be summarized as object detection strategies based on convolutional neural networks (CNN), spiking neural networks (SNN), and graph neural networks (GNN).
[0003] Existing technologies either fail to effectively suppress noise interference in event data or suffer from significant deficiencies in graph modeling and hierarchical structure representation, making it difficult to achieve a balance between the two, thus limiting the performance of perception systems in complex scenarios. Therefore, this invention mainly focuses on solving the following three key problems:
[0004] (1) Noise interference in event data is difficult to suppress effectively: In event-based camera-based perception tasks, the asynchronous sampling mechanism brings microsecond-level temporal resolution and extremely high data sensitivity, but it also inevitably introduces a large amount of background noise and redundant events. These noisy events are often caused by changes in background illumination, random triggering of static areas, or meaningless small movements. Their characteristics lack significant differences from the real targets, making them easily misjudged as abnormal areas or key targets, seriously affecting detection accuracy and system robustness. Especially in high-dynamic or low-light scenarios, this noise interference phenomenon is more obvious, leading to an increase in the false alarm rate of the perception system and an increased risk of missing key behaviors.
[0005] (2) Insufficient global topological perception capability based on traditional graph modeling: In current event perception tasks, graph structure construction methods based on nearest neighbor constraints and spatial adjacency relationships are commonly used. Although this type of graph construction method is suitable for modeling local spatial relationships, it shows obvious insufficient topological expressiveness when facing multi-target collaboration, cross-regional movement, or long-term dependent behavior. If events do not co-occur in a short time or in a nearby region, even if there is potential global semantic consistency and behavioral association, it is difficult to effectively capture them through existing graph modeling methods, thus affecting the model's ability to understand high-level movement patterns. Especially in open scenarios with frequent interactions, event graphs that rely solely on local edge construction strategies are prone to "fragmented" perspectives, making it difficult to restore the integrity of the behavior chain.
[0006] (3) Hierarchical modeling based on pure Euclidean space suffers from representational defects: Data collected by event cameras naturally possesses tree-like, hierarchical structural characteristics. Especially during the continuous motion of object generation, events exhibit significant multi-scale evolutionary relationships in an asynchronous and non-uniform manner. However, most current event visual modeling methods are still limited to feature learning and relationship modeling in Euclidean space. This linear metric system suffers from severe representational defects when facing event flows with complex hierarchical structures and non-linear spatial dependencies. Specifically, this manifests as distance degradation, class indistinguishability, and feature collapse, making it difficult for the system to identify structural differences between multi-level targets, thereby affecting the completeness of the expression of abnormal events and the clarity of classification boundaries. In open environmental perception tasks, such as autonomous driving and Mars exploration, the system must process various levels of target and behavioral variations in real time, including macroscopic phenomena such as trajectory evolution and group anomalies, extending from single behavioral fragments to trajectory evolution. Systems lacking hierarchical modeling capabilities cannot support the requirements of such tasks, exhibiting problems such as failure to handle low-frequency, long-term dependent events and false detection of ambiguous boundary anomaly areas, reducing the overall safety and reliability of the system. Summary of the Invention
[0007] To address at least one problem in the existing technology, this invention proposes an event camera target detection method based on Euclidean-hyperbolic motion perception. The method involves downsampling the original event data, constructing an event graph using the downsampled data, and inputting the event graph into a dual-space network for target detection. The detection process includes:
[0008] A graph convolutional network based on the B-spline kernel function is used to extract local perceptual features from the event graph, and the extracted features are projected onto hyperbolic space.
[0009] Global event-dependent features are extracted from local perceptual features mapped to hyperbolic space using a learnable hyperbolic graph convolutional network.
[0010] After decoding the global event-dependent features, the data is input into the perceptron, which then outputs the target detection results.
[0011] Furthermore, a dynamic event sampling mechanism based on event density and motion intensity estimation downsamples the original event data, specifically including:
[0012] Density estimation is performed based on local constraints, calculating the local spatial density of each event, and introducing the time dimension variance as an indicator to measure the intensity of motion, thereby quantifying the dynamics in the event flow.
[0013] By constructing a sampling probability function using a sigmoid function, spatial density and temporal variability are fused to obtain the sampling weight for each event.
[0014] Furthermore, when constructing the event graph, the sampled time nodes are used as event nodes. In the case of high consistency between events, there is an edge relationship between two event nodes. The transition probability between events is described based on the Markov vector field, and this transition probability is used as the edge weight between two event nodes.
[0015] Furthermore, the graph convolutional network based on the B-spline kernel function consists of multiple graph convolutional modules based on the B-spline kernel function. Each graph convolutional module based on the B-spline kernel function is composed of spline convolutional units, normalization units, and activation units cascaded together.
[0016] Furthermore, the hyperbolic graph convolutional network comprises a first hyperbolic graph convolutional module, a first pooling module, a second hyperbolic graph convolutional module, a third hyperbolic graph convolutional module, a fourth hyperbolic graph convolutional module, and a second pooling module cascaded together. The output of the first pooling module is connected to the output of the third hyperbolic graph convolutional module via a skip connection and then used as the input of the fourth hyperbolic graph convolutional module.
[0017] Furthermore, the curvature hyperbolic graph convolution module is composed of cascaded hyperbolic units, hyperbolic aggregation units, hyperbolic activation units, and normalization units.
[0018] Compared with existing technologies, this invention addresses the performance bottleneck of event-aware systems in complex scenarios and offers the following advantages:
[0019] First, the present invention adopts a dynamic event adaptive sampling mechanism to adjust the sampling strategy based on event density and motion intensity. While significantly reducing background noise and redundant events, it retains effective information in key dynamic areas, improves the signal-to-noise ratio and utilization efficiency of event data, and effectively enhances the robustness of the system in complex scenarios.
[0020] Secondly, the geometric hypergraph modeling method proposed in this invention is no longer limited to local adjacency relationships, but establishes a multi-point collaborative structure based on the consistency of motion states, thereby accurately depicting the semantic dependencies between events across time and space, and improving the system's ability to model complex interaction relationships and behavioral chains between targets.
[0021] Next, by introducing a graph neural network architecture that integrates Euclidean and hyperbolic spaces, this invention breaks through the limitations of traditional Euclidean space in expressing hierarchical structures and global topology, and achieves effective modeling of multi-scale structures and long-distance dependencies, significantly improving the accuracy and generalization ability of the system in target detection tasks.
[0022] In summary, this invention solves the problems of noise suppression, topology mapping, and hierarchical representation in event data, and has good application prospects and promotion value. It is especially suitable for real-time target perception tasks in complex dynamic environments such as autonomous driving and robot navigation. Attached Figure Description
[0023] Figure 1 A flowchart for constructing an event graph for this invention;
[0024] Figure 2 This is a flowchart illustrating the process of inputting an event graph into a dual-space network for target detection according to the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] This invention proposes an event camera target detection method based on Euclidean-hyperbolic motion perception. The method downsamples the original event data, constructs an event graph using the downsampled data, and inputs the event graph into a dual-space network for target detection. The detection process includes:
[0027] A graph convolutional network based on the B-spline kernel function is used to extract local perceptual features from the event graph, and the extracted features are projected onto hyperbolic space.
[0028] Global event-dependent features are extracted from local perceptual features mapped to hyperbolic space using a learnable hyperbolic graph convolutional network.
[0029] After decoding the global event-dependent features, the data is input into the perceptron, which then outputs the target detection results.
[0030] To address the noise interference and redundancy issues in event camera data, while ensuring that critical motion information in high dynamic range regions is not missed, this embodiment employs a dynamic event sampling mechanism based on event density and motion intensity estimation. The overall process is represented as follows:
[0031] Density estimation is performed based on local constraints, calculating the local spatial density for each event. This process is represented as follows:
[0032]
[0033] Where, d i Indicates event e i Local spatial density; k represents event e i The number of nearest neighbors; ij Indicates event e i To its nearest neighbor event e jThe Euclidean distance; ε is used to prevent the minimum value where the denominator is 0.
[0034] Simultaneously, time dimension variance is introduced as an indicator to measure exercise intensity, which is expressed as:
[0035]
[0036] Among them, MD i Indicates event e i Exercise intensity index; Var(t) is the function used to calculate variance; N is the total number of events; t i Let i be the timestamp of the i-th event; It is the average of the timestamps of N events.
[0037] The dynamics of the event flow are quantified, and the index for measuring the intensity of motion is expressed as:
[0038]
[0039] Among them, P i Indicates event e in the event stream i The dynamics; α is a constant that adjusts the sampling sensitivity, and β is an offset value used to adjust the effect of motion on the sampling rate. When the motion intensity index of an event is greater than β, the sampling rate increases, and vice versa.
[0040] A sampling probability function is constructed using a sigmoid function, fusing spatial density and temporal variability to obtain the sampling weight for each event. This achieves an adaptive sampling strategy of "dynamic preservation and efficient compression." The calculation of the event sampling weight includes:
[0041]
[0042] in, For event e i The sampling weights; d min d represents the minimum local spatial density in the event flow. max This represents the maximum value of the local spatial density in the event stream.
[0043] This invention employs the aforementioned sampling strategy to retain more events in high-motion regions to enhance detail representation, and to reduce redundant events in low-dynamic regions to suppress background noise. Furthermore, by introducing this mechanism, the perception system can significantly reduce its dependence on redundant data, improve the model's responsiveness to key motion information, and enhance stability and accuracy in resource-constrained or real-time-critical scenarios (such as autonomous driving and robot navigation). At the same time, this mechanism is highly transferable and can be embedded into existing event vision frameworks as a preprocessing module, laying a solid data foundation for subsequent event mapping and graph neural network modeling.
[0044] In dynamic scenes, the motion of objects exhibits certain regularities, especially within a short period. The speed of events triggered by objects shows consistency, and the event state strongly depends on the previous state of its nearest neighbor, demonstrating continuity and density in space and time. Event cameras, as sensors with extremely high spatiotemporal resolution, capture each event with precise spatiotemporal coordinates (x, y, t), making them suitable for capturing the boundaries and details of high-speed moving objects. Markov Vector Fields (MVFs) are probabilistic models used to describe the state transitions of a system. In dynamic system analysis, they represent the possible transitions of states over time. This is highly similar to the adjacent-state dependency of events, as the state transitions of events strongly depend on the motion patterns between adjacent events. Therefore, this project introduces Markov Vector Field theory, treating the event flow as a stochastic process in the state space, modeling the motion of events, capturing the motion characteristics of objects, and further quantifying the consistency of the motion states of events triggered by objects in dynamic scenes.
[0045] like Figure 1 As shown, this invention introduces Markov vector fields to model the motion state of events, defines the velocity vector and motion intensity of each event, and constructs the transition probability matrix in the event state space:
[0046]
[0047] Among them, Γ(F) j |F i ) represents the Markov transition matrix, where the element in the i-th row and j-th column describes the transition from state F. j To F i The probability of motion consistency between events, where the state of an event includes its velocity direction and velocity amplitude; ∝ represents the relevant symbol; ||·|| represents the Euclidean norm. Indicates event e j The direction of velocity, Indicates event e i The direction of velocity; σ v represents the tolerance for controlling the direction of velocity; |·| represents calculating the absolute value; s j Indicates event e j velocity amplitude, s i Indicates event e i velocity amplitude; σ s This indicates the tolerance level for controlling the amplitude of motion.
[0048] In this embodiment, state similarity is used as the criterion, and hyperedges are only established when there is high consistency among events, thereby generating a high-order event hypergraph structure with physical motion meaning. Compared with traditional graph structures, this method significantly enhances the event graph's ability to model target trajectories and behavioral consistency by connecting multiple event nodes with spatiotemporal coherence at once. By constructing such a motion consistency-driven hypergraph structure, the system can effectively capture the group behavior and potential motion patterns among multiple events, improve global topology awareness, and enhance the accuracy of identifying abnormal behaviors in complex scenes.
[0049] Specifically, in this embodiment, a set of nodes with high motion consistency is selected based on transition probability, only when the transition probability Γ(F) between nodes is high. j |F i When γ > 0, an edge relationship is established between the two nodes, and the transition probability is used as the edge weight. Combining all hyperedges, a hypergraph G can be obtained. H = (V,H), where V = {e1,e2,...e} M} is a set of event points. It is a set of hyperedges. In this hypergraph, the hyperedge contains a set of event points with consistent motion within a time window ΔT. This ensures that the motion patterns of the nodes inside the hyperedge satisfy both local smoothness and global Markov property, that is, the current state depends only on the previous state and is independent of earlier history.
[0050] While existing graph neural networks (GNNs) have achieved remarkable success in tasks such as event camera data detection and recognition, traditional Euclidean geometric space (with zero curvature) is difficult to effectively represent the hierarchical structure of event data. To address the challenges of representing the structural hierarchy and the weak ability to model global information in event data modeling, this invention designs a dual-structure event-aware network architecture that integrates Euclidean and hyperbolic geometric spaces.
[0051] The network structure employed in this invention combines the local precision of Euclidean space with the global extensibility of hyperbolic space, enabling it to capture both fine local motion and reconstruct macroscopic structural levels. This dual-space fusion network exhibits excellent detection robustness and interpretability. Taking autonomous driving scenarios as an example, when the system detects an irregular object's trajectory deviation, the Euclidean branch can accurately locate the anomaly, while the hyperbolic branch can identify its deviation from the global path topology, thus achieving semantic-level interpretation of the anomaly. Overall, this method breaks through the limitations of traditional geometric modeling, providing a novel path for constructing event perception systems with multi-scale cognitive capabilities.
[0052] like Figure 2This invention first extracts features from the constructed graph data in Euclidean space. In this embodiment, feature extraction is performed by a graph convolutional network based on the B-spline kernel function, consisting of multiple graph convolutional modules based on the B-spline kernel function. Each graph convolutional module based on the B-spline kernel function is composed of a cascaded spline convolutional unit, a normalization unit, and an activation unit. Then, the processed data is mapped to hyperbolic space, and the features are further processed by a hyperbolic graph convolutional network consisting of a first hyperbolic graph convolutional module, a first pooling module, a second hyperbolic graph convolutional module, a third hyperbolic graph convolutional module, a fourth hyperbolic graph convolutional module, and a second pooling module. Each hyperbolic graph convolutional module is composed of a cascaded hyperbolic unit, a hyperbolic aggregation unit, a hyperbolic activation unit, and a normalization unit. Finally, the processed data is decoded and processed by a perceptron to obtain the detection result.
[0053] Hyperbolic space offers superior hierarchical embedding capabilities, exponentially expanding to accommodate tree-like or fractal data within a finite dimension. Furthermore, the negative curvature of hyperbolic space maintains geometrical relationships between distant nodes within the embedding space, effectively capturing cross-regional and timestamp-based event dependencies and modeling long-range dependencies. Specifically, this method proposes a phased dual-space network architecture, which includes:
[0054] like Figure 2 As shown, based on the aforementioned downsampling and motion-aware geometry construction, a graph convolution module based on the B-spline kernel function (SplineCNN: Fast Geometric Deep Learning with Continuous B-Spline Kernels) is first constructed in Euclidean space to extract local perceptual features from the event graph, preserving the fine details of the neighborhood structure.
[0055] Subsequently, the Euclidean space feature Z was used to perform exponential mapping. e Projecting onto hyperbolic space yields hyperbolic features. This process can be represented by the following formula:
[0056]
[0057] in, This represents an exponential mapping with the origin as the base point, which maps a vector in the tangent space to a hyperbolic space with curvature c.
[0058] Subsequently, in this hyperbolic space, a learnable hyperbolic graph convolution mechanism is used to nonlinearly model the global event dependencies. The entire process can be summarized as hyperbolic transformation, hyperbolic convergence, and activation, and can be expressed by the following formula:
[0059]
[0060] in, Represents node e i The features obtained after the l-th layer undergoes a hyperbolic convolution; Indicates the activation function; This represents projecting the data onto the hyperbolic sphere; N(i) represents the set of neighboring nodes of the i-th event sampling point in the event graph; A i,j This represents the adjacency relationship between the i-th event sampling point and the j-th event sampling point in the adjacency matrix. If there is an edge relationship between the two events, this value is 1, otherwise it is 0. For exponential mapping The inverse process; W represents the learnable linear transformation matrix of the model; represents matrix multiplication with curvature c; b represents the bias term associated with W, used to perform additive offset in hyperbolic space. The Möbius stripe representing curvature c ( Addition is an important operation in hyperbolic transformations, and its specific formula is as follows:
[0061]
[0062] If let equal So It can be represented by the following:
[0063]
[0064] in, This indicates the calculation of projection onto hyperbolic space. The inner product of the last two data points.
[0065] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An event camera target detection method based on Euclidean-hyperbolic motion perception, characterized in that, The original event data is downsampled, and an event graph is constructed using the downsampled data. The event graph is then input into a dual-space network for target detection. The detection process includes: A graph convolutional network based on the B-spline kernel function is used to extract local perceptual features from the event graph, and the extracted features are projected onto hyperbolic space. Global event-dependent features are extracted from local perceptual features mapped to hyperbolic space using a learnable hyperbolic graph convolutional network. After decoding the global event-dependent features, the data is input into the perceptron, which then outputs the target detection results.
2. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 1, characterized in that, A dynamic event sampling mechanism based on event density and motion intensity estimation downsamples the original event data, specifically including: Density estimation is performed based on local constraints, calculating the local spatial density of each event, and introducing the time dimension variance as an indicator to measure the intensity of motion, thereby quantifying the dynamics in the event flow. By constructing a sampling probability function using a sigmoid function, spatial density and temporal variability are fused to obtain the sampling weight for each event.
3. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 1 or 2, characterized in that, When constructing the event graph, the sampled time nodes are used as event nodes. If there is high consistency between events, there is an edge relationship between two event nodes. The transition probability between events is described based on the Markov vector field, and this transition probability is used as the edge weight between two event nodes.
4. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 1, characterized in that, The graph convolutional network based on the B-spline kernel function consists of multiple graph convolutional modules based on the B-spline kernel function. Each graph convolutional module based on the B-spline kernel function is composed of spline convolutional units, normalization units, and activation units cascaded together.
5. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 1, characterized in that, The hyperbolic graph convolutional network comprises a first hyperbolic graph convolutional module, a first pooling module, a second hyperbolic graph convolutional module, a third hyperbolic graph convolutional module, a fourth hyperbolic graph convolutional module, and a second pooling module cascaded together. The output of the first pooling module is connected to the output of the third hyperbolic graph convolutional module via a skip connection and then used as the input of the fourth hyperbolic graph convolutional module.
6. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 5, characterized in that, The hyperbolic graph convolution module consists of cascaded hyperbolic units, hyperbolic aggregation units, hyperbolic activation units, and normalization units.
7. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 6, characterized in that, The data processing procedure of cascaded hyperbolic units, hyperbolic aggregation units, and hyperbolic activation units can be represented as follows: in, Represents node e i The features obtained after the l-th layer undergoes a hyperbolic convolution; Indicates the activation function; This indicates that the data is projected onto the hyperbola. This represents an exponential mapping with the origin as the base point, that is, mapping a vector in the tangent space to a hyperbolic space with curvature c; N(i) represents the set of neighbor nodes of the i-th event sampling point in the event graph; A i,j This represents the adjacency relationship between the i-th event sampling point and the j-th event sampling point in the adjacency matrix; For exponential mapping The inverse process; W represents the learnable linear transformation matrix of the model; Represents matrix multiplication with curvature c; represents the Möbius addition with curvature c; b represents the offset term that accompanies W, used to perform additive offset in hyperbolic space.
8. The event camera target detection method based on Euclidean-hyperbolic motion perception according to claim 7, characterized in that, The execution process of matrix multiplication with curvature c is represented as follows:
9. A method for event camera target detection based on Euclidean-hyperbolic motion perception according to claim 7 or 8, characterized in that, make The execution process of the Möbius method with curvature c is expressed as follows: in, This indicates the calculation of projection onto hyperbolic space. The inner product of the last two data; ||·|| denotes the Euclidean norm.
Citation Information
Cited By
Space debris positioning method and device
CN121708086A
A method and apparatus for locating space debris
CN121708086B
Event camera motion small target detection method and system based on topology constraint
CN121962585A