Target detection method and device, equipment and storage medium

By constructing a target spatiotemporal graph and utilizing a spatiotemporal graph neural network and a detector based on the Transformer architecture, the noise robustness and computational burden problems in event camera data processing are solved, achieving real-time, high-precision target detection.

CN120832472APending Publication Date: 2025-10-24HAINAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510993925.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively utilize the temporal evolution information in the event stream generated by event cameras, and are easily interfered by background noise, resulting in false alarms or missed detections. The heavy computational burden hinders their deployment on resource-constrained devices, and there is a lack of efficient and high-performance end-to-end target detection models.

Method used

The initial event stream is obtained through the preset neural vision sensor, data preprocessing is performed, the target spatiotemporal graph is constructed, features are extracted using the spatiotemporal graph neural network, and target detection is performed in combination with the detector of the Transformer architecture to generate target categories and bounding boxes.

Benefits of technology

It achieves real-time, robust, and high-precision target detection, improves detection performance in complex scenarios, conforms to the data characteristics of neural vision sensors, and generates more accurate target information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832472A_ABST
    Figure CN120832472A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and device, equipment and a storage medium, and relates to the field of target detection, and the method comprises the steps: obtaining an initial event stream through a preset neural vision sensor, and carrying out the data preprocessing of the initial event stream, and obtaining a processed event stream; determining effective events in the processed event stream, converting the effective events into target graph nodes, obtaining target edges between the target graph nodes according to a preset space-time proximity threshold, and constructing a target space-time graph based on the target graph nodes and the target edges; performing node information aggregation and graph pooling operation on the target space-time diagram by using a feature extraction module of the space-time diagram neural network to obtain a target high-dimensional feature matrix, and generating to-be-input data based on the target high-dimensional feature matrix and a preset sine position code; the method comprises the steps of determining a feature region and category information according to to-be-input data, and performing standard processing on the feature region and category information to generate a target category and a target bounding box; and real-time, robust and high-precision target detection can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of target detection, and in particular to a target detection method, device, equipment and storage medium. BACKGROUND

[0002] As a core task of computer vision, target detection has long relied on traditional frame-based image sensors and their supporting deep learning algorithms, especially convolutional neural networks. To break through these limitations, neuromorphic vision sensors, commonly known as event cameras, have emerged. Their biomimetic design discards the concept of global frames and adopts an asynchronous, event-driven perception mechanism. Each pixel works independently and only outputs an "event" containing the pixel position, microsecond-level accurate timestamp and brightness change polarity when the perceived light intensity change exceeds the set threshold.

[0003] Existing solutions attempt to convert the raw event stream into an intermediate representation suitable for CNN processing, but all have significant limitations. The deeper problem is that existing methods generally fail to effectively mine and utilize the precise, continuous temporal evolution information contained in the event stream, which is crucial for understanding the dynamic behavior of the target. At the same time, event data is susceptible to background noise, and existing methods lack noise robustness, easily producing false alarms or missed detections. Complex conversion or processing algorithms often result in heavy computational burden, negating the inherent low latency and low power consumption advantages of event cameras, hindering their deployment on resource-constrained edge devices. Ultimately, there is a lack of efficient, high-performance and versatile end-to-end event target detection models, whose performance or generalization ability often lags behind traditional frame detectors under ideal conditions in complex scenarios.

[0004] In summary, how to achieve real-time, robust and high-precision target detection is a problem to be solved at present. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a target detection method, device, equipment and storage medium, which can realize real-time, robust and high-precision target detection. The specific scheme is as follows:

[0006] In a first aspect, the present application provides a target detection method, comprising:

[0007] An initial event stream is obtained by a preset neural vision sensor, and a data preprocessing operation is performed on the initial event stream to obtain a processed event stream;

[0008] Valid events in the processed event stream are determined, the valid events are converted into target graph nodes, target edges between the target graph nodes are obtained according to a preset spatiotemporal proximity threshold, and a target spatiotemporal graph is constructed based on the target graph nodes and the target edges;

[0009] The feature extraction module of the spatio-temporal graph neural network is used for node information aggregation and graph pooling operation on the target spatio-temporal graph to obtain a target high-dimensional feature matrix; and the target high-dimensional feature matrix and preset sinusoidal position encoding are used to generate input data;

[0010] The feature region and category information of the target object are determined according to the input data, and the feature region and category information of the target object are standardized to generate a target category and a target bounding box of the target object.

[0011] Optionally, the initial event stream is obtained by using a preset neural vision sensor, and data preprocessing is performed on the initial event stream to obtain a processed event stream, including:

[0012] A preset neural vision sensor with a target resolution is determined according to a target user demand;

[0013] An initial event stream is obtained by using the preset neural vision sensor with the target resolution;

[0014] The initial event stream is divided into sub-event segments according to a preset time step window;

[0015] Noise filtering is performed on the sub-event segments to obtain filtered sub-event segments, and the filtered sub-event segments are integrated to generate a processed event stream.

[0016] Optionally, the noise filtering performed on the sub-event segments to obtain the filtered sub-event segments includes:

[0017] A spatio-temporal filter is used to determine a response difference between the sub-event segment and a front or rear time segment corresponding to the sub-event segment;

[0018] It is determined whether the sub-event segment is noise according to the determined response difference;

[0019] If the determination result indicates that the sub-event segment is not noise, a first filtering result is obtained;

[0020] If the determination result indicates that the sub-event segment is noise, the corresponding sub-event segment is removed to obtain a second filtering result;

[0021] The filtered sub-event segments are determined based on the first filtering result and the second filtering result.

[0022] Optionally, the effective event is converted into a target graph node, including:

[0023] A pixel space position coordinate of the effective event is obtained, and a spatial feature of the effective event is determined according to the pixel space position coordinate of the effective event;

[0024] acquire a timestamp of the effective event, determine a time feature of the effective event according to the timestamp of the effective event;

[0025] acquire a polarity attribute corresponding to the effective event, and determine a polarity feature of the effective event according to the polarity attribute of the effective event;

[0026] convert the effective event into a target graph node according to the spatial feature, the time feature and the polarity feature of the effective event.

[0027] Optionally, the acquiring of the target edge between the target graph nodes according to the preset spatio-temporal proximity threshold comprises:

[0028] determining a connection correlation degree between the target graph nodes satisfying the proximity condition;

[0029] judging whether the connection correlation degree between the target graph nodes satisfying the proximity condition is less than or equal to the preset spatio-temporal proximity threshold;

[0030] if the connection correlation degree between the target graph nodes is less than or equal to the preset spatio-temporal proximity threshold, generating a target edge between the target graph nodes.

[0031] Optionally, the node information aggregation and graph pooling operation of the target spatio-temporal graph by the feature extraction module of the spatio-temporal graph neural network comprises:

[0032] acquiring first node features of the target graph nodes under different spatial scales and time windows through multi-layer parallel graph convolution operation;

[0033] weighting and fusing each first node feature by using a preset attention mechanism, and aggregating second node features of nodes satisfying a proximity condition with the target graph node;

[0034] generating a target high-dimensional feature matrix according to the first node features and the second node features.

[0035] Optionally, the determining of the feature region and the category information of the target object according to the to-be-input data, and the standard processing of the feature region and the category information of the target object to generate a target category and a target bounding box of the target object comprise:

[0036] encoding the to-be-input data by using a preset Transformer encoder to obtain encoded data;

[0037] decoding the encoded data by using a preset real-time object detection model to determine the feature region and the category information of the target object;

[0038] The feature region of the target object and the category information are input into a classification Softmax function to output a target category of the target object and a target bounding box.

[0039] In a second aspect, the present application provides a target detection device, comprising:

[0040] An event stream acquisition module is configured to acquire an initial event stream by using a preset neural vision sensor, and perform a data preprocessing operation on the initial event stream to obtain a processed event stream.

[0041] A space-time graph construction module is configured to determine effective events in the processed event stream, convert the effective events into target graph nodes, acquire target edges between the target graph nodes according to a preset space-time proximity threshold, and construct a target space-time graph based on the target graph nodes and the target edges.

[0042] A data generation module is configured to perform node information aggregation and graph pooling operations on the target space-time graph by using a feature extraction module of a space-time graph neural network to obtain a target high-dimensional feature matrix, and generate input data to be input based on the target high-dimensional feature matrix and a preset sinusoidal position encoding.

[0043] A category and bounding box generation module is configured to determine a feature region and category information of a target object according to the input data to be input, and perform standard processing on the feature region and category information of the target object to generate a target category and a target bounding box of the target object.

[0044] In a third aspect, the present application provides an electronic device, comprising:

[0045] A memory is configured to save a computer program.

[0046] A processor is configured to execute the computer program to implement the target detection method as described above.

[0047] In a fourth aspect, the present application provides a computer readable storage medium configured to save a computer program; wherein the computer program is executed by a processor to implement the target detection method as described above.

[0048] In summary, the application obtains an initial event stream through a preset neural vision sensor, performs data preprocessing operations on the initial event stream to obtain a processed event stream, determines valid events in the processed event stream, converts the valid events into target graph nodes, acquires target edges between the target graph nodes according to a preset spatiotemporal proximity threshold, and constructs a target spatiotemporal graph based on the target graph nodes and the target edges. The feature extraction module of the spatiotemporal graph neural network is used to perform node information aggregation and graph pooling operations on the target spatiotemporal graph to obtain a target high-dimensional feature matrix, and the target high-dimensional feature matrix and a preset sinusoidal position encoding are used to generate input data. The feature region and category information of a target object are determined according to the input data, and the feature region and category information of the target object are processed to generate a target category and a target bounding box of the target object. As can be seen from the above, the application first obtains an initial event stream through a preset neural vision sensor and performs data preprocessing to obtain a processed event stream, then determines valid events and converts them into target graph nodes, and then acquires target edges between the target graph nodes according to a preset spatiotemporal proximity threshold, and then constructs a target spatiotemporal graph. Then, the feature extraction module of the spatiotemporal graph neural network is used to perform node information aggregation and graph pooling operations on the target spatiotemporal graph to obtain a target high-dimensional feature matrix, and the target high-dimensional feature matrix and a preset sinusoidal position encoding are used to generate input data. Finally, the feature region and category information of a target object are determined according to the input data, and the feature region and category information of the target object are processed to generate a target category and a target bounding box of the target object. In this way, the event sequence features are extracted using a graph neural network, and the relevant relationship between each event point is modeled. The detector based on the Transformer architecture can decode the target information contained in the event sequence faster, which is consistent with the data characteristics generated by the neural vision sensor, and can generate more accurate target information. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0050] Figure 1 A target detection method flowchart disclosed by the present application;

[0051] Figure 2 A specific event stream preprocessing flowchart disclosed by the present application;

[0052] Figure 3 A specific target detection method flowchart disclosed by the present application;

[0053] Figure 4 A target detection device structure diagram disclosed by the present application is shown in the figure;

[0054] Figure 5 An electronic device structure diagram disclosed by the present application is shown in the figure. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0056] At present, existing solutions attempt to transform the original event stream into an intermediate representation form suitable for CNN processing, but all have significant limitations. The deeper problem is that existing methods generally fail to effectively mine and utilize the precise and continuous time evolution information contained in the event stream, and understanding the dynamic behavior of the target is crucial for detection. At the same time, event data is susceptible to background noise interference, and the noise robustness of existing methods is insufficient, which is prone to false alarms or missed detection. Complex conversion or processing algorithms often result in heavy computational burden, which offsets the inherent low latency and low power consumption advantages of event cameras, hindering their deployment in resource-constrained edge devices. Finally, there is a lack of efficient, high-performance and highly versatile end-to-end event target detection model, and its performance or generalization ability often lags behind traditional frame detectors under ideal conditions in complex scenarios. In order to solve the above technical problems, the present application discloses a target detection method, device, equipment and storage medium, which can realize real-time, robust and high-precision target detection.

[0057] Referring to Figure 1 The embodiments of the present application disclose a target detection method, which comprises:

[0058] Step S11, obtaining an initial event stream through a preset neural vision sensor, and performing data preprocessing operation on the initial event stream to obtain a processed event stream.

[0059] In the embodiment, first, a preset neural vision sensor with a target resolution is determined according to the target user demand; an initial event stream is obtained through the preset neural vision sensor with the target resolution; the initial event stream is divided into sub-event segments according to a preset time step window; noise filtering processing is performed on the sub-event segments to obtain filtered sub-event segments, and the filtered sub-event segments are integrated to generate a processed event stream. Specifically, a neural vision sensor with a resolution of 346x260 or a resolution of 1280x720 is selected to output an initial event stream , sensors are selected according to scene requirements, for example, a neural vision sensor with a resolution of 346*260 is used in a high-speed low-power scene, a neural vision sensor with a resolution of 1280*720 is used in a high-resolution complex scene, an initial event stream is recorded in the form of a microsecond timestamp and a pixel coordinate, a time step window can be dynamically adjusted according to the event density, a sparse event uses a longer window, a dense event uses a shorter window, noise filtering is realized through a space-time consistency test, isolated events are removed and events meeting local space-time correlation are retained, and finally, a continuous and denoised processed event stream is generated by reordering the time stamp of the integrated sub-event segment , which can be directly used for target detection or high-level vision tasks such as SLAM.

[0060] It can be understood that, in order to determine the filtered sub-event segment, a space-time filter is used to determine the response difference between the sub-event segment and the front and rear time segments corresponding to the sub-event segment; whether the sub-event segment is noise is judged according to the determined response difference; if the judgment result represents that the sub-event segment is not noise, a first filtering result is obtained; if the judgment result represents that the sub-event segment is noise, the corresponding sub-event segment is removed to obtain a second filtering result; and the filtered sub-event segment is determined based on the first filtering result and the second filtering result. Specifically, first, based on the space-time filter, a space-time neighborhood threshold is set , the response difference between the sub-event segment and the adjacent events in the front and rear time segments is calculated, noise is distinguished by estimating the correlation degree of the event density and the polarity of the sub-event segment in the space-time neighborhood , isolated events and jitter noise with conflicting polarity are removed, that is, the second filtering result, and the sub-event segment that is not noise is retained, that is, the first filtering result. Wherein, i, j are natural random numbers, then the density threshold is used to filter hot spot noise with low activity, and the event cluster meeting the space-time continuity is retained. At the same time, the time stamp in the sub-event segment is standardized by a time scaling factor, the starting time of the event segment is reset to zero and the time interval is linearly scaled, the time reference difference caused by window division is eliminated, and finally the filtered sub-event segment with unified time reference and denoising is output.

[0061] In step S12, the valid events in the processed event stream are determined, the valid events are converted into target graph nodes, the target edges between the target graph nodes are obtained according to a preset space-time neighborhood threshold, and a target space-time graph is constructed based on the target graph nodes and the target edges.

[0062] In this embodiment, the pixel space position coordinates of the effective event are obtained, the spatial feature of the effective event is determined according to the pixel space position coordinates of the effective event, the timestamp of the effective event is obtained, the time feature of the effective event is determined according to the timestamp of the effective event, the polarity attribute corresponding to the effective event is obtained, and the polarity feature of the effective event is determined according to the polarity attribute of the effective event. The effective event is converted into a target graph node according to the spatial feature, the time feature and the polarity feature of the effective event. Specifically, the spatial feature The two-dimensional pixel coordinates (x, y) are directly mapped to the position attribute of the graph node, the time feature The relative difference Δt between the event timestamp t and the reference time is taken as the time weight of the node, and the polarity feature is encoded as a node attribute, and finally each effective event is converted into a target graph node containing a spatial feature, a time feature and a polarity feature .

[0063] Further, as Figure 2 indicated, the connection correlation degree between the target graph nodes satisfying the adjacent condition is determined, it is judged whether the connection correlation degree between the target graph nodes satisfying the adjacent condition is less than or equal to a preset spatio-temporal adjacent threshold, and if the connection correlation degree between the target graph nodes is less than or equal to the preset spatio-temporal adjacent threshold, a target edge between the target graph nodes is generated. Specifically, the connection correlation degree between nodes is calculated by using a composite distance metric , the preset spatio-temporal adjacent threshold d is determined, and when the target edge is established. Finally, after obtaining the target graph node and the target edge, a target spatio-temporal graph is constructed .

[0064] In step S13, a feature extraction module of a spatio-temporal graph neural network is used to perform node information aggregation and graph pooling operation on the target spatio-temporal graph to obtain a target high-dimensional feature matrix, and based on the target high-dimensional feature matrix and a preset sine position encoding, input data is generated.

[0065] In the embodiment, the first node features of the target graph node under different spatial scales and time windows are obtained through a multi-layer parallel graph convolution operation; each first node feature is weighted and fused by using a preset attention mechanism, and second node features of nodes meeting a neighboring condition of the target graph node are aggregated; and a target high-dimensional feature matrix is generated according to the first node features and the second node features. Specifically, the first node features of the target graph node under different spatial scales and time windows are calculated by using a 4-layer multi-scale fusion calculation module, the information attributes of neighboring nodes are aggregated, the message features of the neighboring nodes are updated, the event pixel space is divided into a clustering space by using a voxel grid, the nodes with high correlation are aggregated, the second node features are obtained, the accumulated node features are hierarchically converged in the last layer of the neural network, and all event node feature vectors are output, which are represented as the target high-dimensional feature matrix. .

[0066] Further, the obtained target high-dimensional feature matrix is superimposed with a sinusoidal position encoding, for each event node in the target high-dimensional feature matrix, a two-dimensional sinusoidal wave position encoding is superimposed according to a spatial coordinate, and a one-dimensional sinusoidal wave time encoding is superimposed according to a timestamp. The spatial encoding enables the network to perceive the specific pixel position of the event occurrence, and the time encoding preserves the accurate timing relationship between events. The two encodings are adaptively fused through learnable weight parameters, which not only maintains the integrity of the original feature semantics, but also explicitly injects the spatial and temporal position information to obtain the required input data.

[0067] In step S14, the feature region and the category information of the target object are determined according to the input data, and the feature region and the category information of the target object are standardized to generate the target category and the target bounding box of the target object.

[0068] In this embodiment, the preset Transformer encoder is used to encode the features of the to-be-input data to obtain encoded data; the preset real-time target detection model is used to decode the encoded data to determine the feature region and category information of the target object; and the feature region and category information of the target object are input into a classification Softmax function to output the target category and target bounding box of the target object. Specifically, a single-layer Transformer encoder is used to encode the features of the to-be-input data, a multi-head self-attention mechanism is used to capture the long-range spatiotemporal dependency relationship between event nodes, the obtained encoded data is input into a real-time detection model based on a graph structure, such as an RT-DETR (Real-Time DEtection Transformer) detection architecture, a learnable query initialization target information is introduced, and a 3-layer decoder is set to extract the feature region and category information of the target object at multiple scales. Finally, the feature region and category information of the target object are input into a classification Softmax, and the target category probability of the target object, i.e., the target category, can be output, and an independent prediction target box, i.e., the target bounding box, is also generated.

[0069] As can be seen from the above, in the embodiment of the application, an initial event stream is obtained by a preset neural vision sensor, and data preprocessing is performed to obtain a processed event stream. Then, valid events are determined from the processed event stream, and the valid events are converted into target graph nodes. Then, target edges between the target graph nodes are obtained according to a preset spatiotemporal proximity threshold, and a target spatiotemporal graph is constructed. Then, a feature extraction module of a spatiotemporal graph neural network is used to perform node information aggregation and graph pooling operations on the target spatiotemporal graph to obtain a target high-dimensional feature matrix, and a to-be-input data is generated in combination with a preset sinusoidal position encoding. Finally, the feature region and category information of the target object are determined according to the to-be-input data, and the target category and target bounding box of the target object are generated after standard processing. In this way, the graph neural network is used to extract event sequence features and model the correlation between each event point. The detector based on the Transformer architecture can decode the target information contained in the event sequence faster, which is consistent with the data characteristics generated by the neural vision sensor, and more accurate target information can be generated.

[0070] Based on the above embodiment, the application discloses a target detection method, which can realize real-time, robustness and high precision. Next, taking the detection of a vehicle as an example, the target detection method as shown in Figure 3 will be described in detail.

[0071] Firstly, an adaptive filter based on local spatio-temporal correlation is applied to the input raw asynchronous event stream for preprocessing. The adaptive filter dynamically analyzes the density, polarity consistency and time continuity of events in a micro spatio-temporal neighborhood, effectively filters out isolated noise events generated by background light flicker, thermal noise and the like, and generates a denoised event stream, laying a high-quality data foundation for subsequent processing.

[0072] Subsequently, the denoised event stream is modeled as a dynamic spatio-temporal graph structure: each valid event is taken as a graph node, and the node attributes include its accurate spatial features, time features and polarity features; the edges between nodes are based on a preset spatio-temporal proximity threshold, thereby accurately representing the continuity of the target motion trajectory and the potential interaction between events. Using ST-GNN (Spatial-Temporal Graph Neural Network), noise-aware spatio-temporal message passing and node state updating are iteratively performed: at each layer, the node aggregates the feature information transmitted by its spatio-temporal neighbors, and fuses the node state and the aggregated information to generate new node information embedding. In addition, an attention weight mechanism is embedded in the message passing to dynamically evaluate the importance of neighbor events and further suppress the interference of residual noise. Then, the graph pooling operation is used to orderly gather and convert the node-level features into structured feature representations suitable for target detection.

[0073] In the final detection stage, the traditional anchor-based or proposal-based detection head is abandoned, and the extracted graph features are directly input into the improved RT-DETR architecture. The architecture utilizes the powerful global relationship modeling capability of the Transformer: first, a lightweight encoder is used to enhance the input features; then, a set of learnable detection queries interacts with the encoder output in the decoder, and through a multi-scale deformable attention mechanism, it adaptively focuses on the key spatio-temporal feature regions related to the target; the decoder and the information classification module output the class of the target and its accurate bounding box in the spatial dimension, for example, the distinction between vehicles and pedestrians, and their respective bounding boxes.

[0074] Referring to Figure 4 The embodiment of the application discloses a target detection device, comprising:

[0075] An event stream acquisition module 11 is configured to acquire an initial event stream through a preset neural vision sensor, and perform a data preprocessing operation on the initial event stream to obtain a processed event stream.

[0076] A spatio-temporal graph construction module 12 is configured to determine valid events in the processed event stream, convert the valid events into target graph nodes, acquire target edges between the target graph nodes according to a preset spatio-temporal proximity threshold, and construct a target spatio-temporal graph based on the target graph nodes and the target edges.

[0077] The data generation module 13 is configured to perform node information aggregation and graph pooling operation on the target spatio-temporal graph by using a feature extraction module of a spatio-temporal graph neural network to obtain a target high-dimensional feature matrix, and generate input data to be input based on the target high-dimensional feature matrix and preset sinusoidal position encoding.

[0078] The category and bounding box generation module 14 is configured to determine feature regions and category information of the target object according to the input data to be input, and perform standard processing on the feature regions and category information of the target object to generate a target category and a target bounding box of the target object.

[0079] As can be seen from the above, the present application first obtains an initial event stream by using a preset neural vision sensor and performs data preprocessing to obtain a processed event stream, then determines valid events in the processed event stream and converts the valid events into target graph nodes, and then acquires target edges between the target graph nodes according to a preset spatio-temporal proximity threshold to construct a target spatio-temporal graph. Then, the feature extraction module of the spatio-temporal graph neural network is used to perform node information aggregation and graph pooling operation on the target spatio-temporal graph to obtain a target high-dimensional feature matrix, and the input data to be input is generated in combination with the preset sinusoidal position encoding. Finally, the feature regions and category information of the target object are determined according to the input data to be input, and the target category and the target bounding box of the target object are generated after standard processing. In this way, the event sequence features are extracted by using the graph neural network, and the related relationship between each event point is modeled. The detector based on the Transformer architecture decodes the target information contained in the event sequence faster, which is consistent with the data characteristics generated by the neural vision sensor, and more accurate target information is generated.

[0080] In some specific embodiments, the event stream acquisition module 11 can specifically include:

[0081] A preset neural vision sensor determination unit is configured to determine a preset neural vision sensor with a target resolution according to a target user demand.

[0082] An initial event stream acquisition unit is configured to acquire an initial event stream by using the preset neural vision sensor with the target resolution.

[0083] A sub-event segment segmentation unit is configured to segment the initial event stream into sub-event segments according to a preset time step window.

[0084] A processed event stream generation unit is configured to perform noise filtering processing on the sub-event segments to obtain filtered sub-event segments, integrate the filtered sub-event segments, and generate a processed event stream.

[0085] In some specific embodiments, the processed event stream generation unit can specifically include:

[0086] a response difference determination subunit, configured to determine a response difference between the sub-event segment and a time slice before and after the sub-event segment by using a spatio-temporal filter;

[0087] a noise judgment subunit, configured to judge whether the sub-event segment is noise according to the determined response difference;

[0088] a first filtering result acquisition subunit, configured to acquire a first filtering result if the judgment result indicates that the sub-event segment is not noise;

[0089] a second filtering result acquisition subunit, configured to eliminate the corresponding sub-event segment to acquire a second filtering result if the judgment result indicates that the sub-event segment is noise;

[0090] a filtered sub-event segment determination subunit, configured to determine filtered sub-event segments based on the first filtering result and the second filtering result.

[0091] In some specific embodiments, the spatio-temporal graph construction module 12 can specifically include:

[0092] a spatial feature determination unit, configured to acquire pixel spatial position coordinates of the effective event, and determine spatial features of the effective event according to the pixel spatial position coordinates of the effective event;

[0093] a time feature determination unit, configured to acquire time stamps of the effective event, and determine time features of the effective event according to the time stamps of the effective event;

[0094] a polarity feature determination unit, configured to acquire polarity attributes corresponding to the effective event, and determine polarity features of the effective event according to the polarity attributes of the effective event;

[0095] a target graph node transformation unit, configured to transform the effective event into a target graph node according to the spatial features, the time features and the polarity features of the effective event.

[0096] In some specific embodiments, the spatio-temporal graph construction module 12 can specifically include:

[0097] a connection correlation degree determination unit, configured to determine connection correlation degrees between the target graph nodes satisfying the proximity condition;

[0098] a connection correlation degree judgment unit, configured to judge whether the connection correlation degrees between the target graph nodes satisfying the proximity condition are less than or equal to a preset spatio-temporal proximity threshold;

[0099] The target edge generation unit is configured to generate a target edge between the target graph nodes if a connection relevance degree between the target graph nodes is less than or equal to the preset spatio-temporal proximity threshold.

[0100] In some specific embodiments, the data generation module 13 can specifically include:

[0101] The feature acquisition unit is configured to acquire first node features of the target graph nodes at different spatial scales and time windows through a multi-layer parallel graph convolution operation.

[0102] The feature aggregation unit is configured to weight and fuse each first node feature by using a preset attention mechanism, and aggregate second node features of nodes meeting a neighboring condition of the target graph nodes.

[0103] The target high-dimensional feature matrix acquisition unit is configured to generate a target high-dimensional feature matrix according to the first node features and the second node features.

[0104] In some specific embodiments, the category and bounding box generation module 14 can specifically include:

[0105] The encoded data acquisition unit is configured to perform feature encoding on the to-be-input data by using a preset Transformer encoder to obtain encoded data.

[0106] The feature region and category information determination unit is configured to decode the encoded data by using a preset real-time target detection model to determine feature regions and category information of the target object.

[0107] The target category and target bounding box output unit is configured to input the feature regions and category information of the target object into a classification Softmax function to output a target category and a target bounding box of the target object.

[0108] Further, the embodiments of the present application also disclose an electronic device, Figure 5 is a structural diagram of an electronic device 20 according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the use range of the present application.

[0109] Figure 5A structural schematic diagram of an electronic device 20 is provided in the embodiments of the present application. The electronic device 20 can specifically include at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25 and a communication bus 26. The memory 22 is configured to store a computer program, and the processor 21 is configured to load and execute the computer program to implement the related steps in the target detection method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in the embodiments of the present application can be specifically an electronic computer.

[0110] In the embodiments of the present application, the power supply 23 is configured to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 is capable of creating a data transmission channel between the electronic device 20 and external devices, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which is not specifically limited herein; the input / output interface 25 is configured to obtain external input data or output data to the outside, and the specific interface type can be selected according to the specific application needs, which is not specifically limited herein.

[0111] In addition, the memory 22 as a carrier for resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage mode can be temporary storage or permanent storage.

[0112] The operating system 221 is configured to manage and control each hardware device on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the target detection method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include a computer program capable of completing other specific work.

[0113] Further, the present application further discloses a computer readable storage medium for storing a computer program; wherein the computer program is executed by a processor to implement the target detection method disclosed in the foregoing embodiments. The specific steps of the method can refer to the corresponding contents disclosed in the foregoing embodiments, which will not be repeated here.

[0114] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can refer to the method part.

[0115] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples were described above generally in terms of their functionality, without referring to the corresponding

[0116] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is tangible.

[0117] Finally, it should be noted that the terms "comprises", "comprising", or other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0118] The above detailed description has set forth various examples of the technology disclosed herein. The description is purposefully rendered in this form for the purpose of providing clear and comprehensive disclosure of the technology disclosed herein, and, thus, no additional limitations or scope should be inferred therefrom for purposes of appropriate appreciation of the technology disclosed herein.

Claims

1. A target detection method characterized by, The method comprises the following steps: An initial event stream is acquired by a preset neural vision sensor, data preprocessing is performed on the initial event stream to obtain a processed event stream; Valid events in the processed event stream are determined, the valid events are converted into target graph nodes, target edges between the target graph nodes are acquired according to a preset spatiotemporal proximity threshold, and a target spatiotemporal graph is constructed based on the target graph nodes and the target edges; A node information aggregation and graph pooling operation is performed on the target spatiotemporal graph by using a feature extraction module of a spatiotemporal graph neural network to obtain a target high-dimensional feature matrix, and a to-be-input data is generated based on the target high-dimensional feature matrix and a preset sinusoidal position encoding; Feature regions and category information of a target object are determined according to the to-be-input data, and the feature regions and the category information of the target object are normalized to generate a target category and a target bounding box of the target object.

2. The object detection method of claim 1, wherein, The method of acquiring the initial event stream by the preset neural vision sensor and performing data preprocessing on the initial event stream to obtain the processed event stream comprises the following steps: A preset neural vision sensor with a target resolution is determined according to a target user demand; An initial event stream is acquired by the preset neural vision sensor with the target resolution; The initial event stream is divided into sub-event segments according to a preset time step window; Noise filtering is performed on the sub-event segments to obtain filtered sub-event segments, and the filtered sub-event segments are integrated to generate a processed event stream.

3. The object detection method of claim 2, wherein, The method of performing noise filtering on the sub-event segments to obtain the filtered sub-event segments comprises the following steps: A spatiotemporal filter is used to determine a response difference between the sub-event segments and time segments before and after the sub-event segments; It is determined whether the sub-event segments are noise according to the determined response difference; If the determination result indicates that the sub-event segments are not noise, a first filtering result is obtained; If the determination result indicates that the sub-event segments are noise, the corresponding sub-event segments are removed to obtain a second filtering result; The filtered sub-event segments are determined based on the first filtering result and the second filtering result.

4. The object detection method of claim 1, wherein, The method of converting the valid events into target graph nodes comprises the following steps: Pixel space position coordinates of the valid events are acquired, and spatial features of the valid events are determined according to the pixel space position coordinates of the valid events; Timestamps of the valid events are acquired, and temporal features of the valid events are determined according to the timestamps of the valid events; Polarity attributes corresponding to the valid events are acquired, and polarity features of the valid events are determined according to the polarity attributes of the valid events; The valid events are converted into target graph nodes according to the spatial features, the temporal features and the polarity features of the valid events.

5. The object detection method of claim 1, wherein, The method of acquiring target edges between the target graph nodes according to a preset spatiotemporal proximity threshold comprises the following steps: A connection correlation degree between the target graph nodes that meet a proximity condition is determined; It is determined whether the connection correlation degree between the target graph nodes that meet the proximity condition is less than or equal to a preset spatiotemporal proximity threshold; If the connection correlation degree between the target graph nodes is less than or equal to the preset spatiotemporal proximity threshold, target edges between the target graph nodes are generated.

6. The object detection method of claim 1, wherein, The feature extraction module using the spatio-temporal graph neural network performs node information aggregation and graph pooling operation on the target spatio-temporal graph to obtain a target high-dimensional feature matrix, including: A first node feature under different spatial scales and time windows of the target graph node is obtained through multi-layer parallel graph convolution operation; Each first node feature is weighted and fused by using a preset attention mechanism, and a second node feature of a node meeting a neighboring condition of the target graph node is aggregated; A target high-dimensional feature matrix is generated according to the first node feature and the second node feature.

7. The object detection method according to any one of claims 1 to 6, characterized in that, The feature region and category information of the target object are determined according to the to-be-input data, and the feature region and category information of the target object are standardized to generate a target category and a target bounding box of the target object, including: The to-be-input data is feature-encoded by using a preset Transformer encoder to obtain encoded data; The encoded data is decoded by using a preset real-time object detection model to determine the feature region and category information of the target object; The feature region and category information of the target object are input into a classification Softmax function to output the target category and the target bounding box of the target object.

8. A target detection apparatus characterized by comprising: It includes: An event stream acquisition module is configured to acquire an initial event stream by using a preset neural vision sensor, and perform data preprocessing operation on the initial event stream to obtain a processed event stream; A spatio-temporal graph construction module is configured to determine effective events in the processed event stream, convert the effective events into target graph nodes, acquire target edges between the target graph nodes according to a preset spatio-temporal proximity threshold, and construct a target spatio-temporal graph based on the target graph nodes and the target edges; A data generation module is configured to perform node information aggregation and graph pooling operation on the target spatio-temporal graph by using a feature extraction module of a spatio-temporal graph neural network to obtain a target high-dimensional feature matrix, and generate to-be-input data based on the target high-dimensional feature matrix and a preset sinusoidal position encoding; A category and bounding box generation module is configured to determine the feature region and category information of a target object according to the to-be-input data, and standardize the feature region and category information of the target object to generate a target category and a target bounding box of the target object.

9. An electronic device, comprising: It includes: A memory is configured to save a computer program; A processor is configured to execute the computer program to implement the object detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A memory is configured to save a computer program; wherein the computer program is executed by a processor to implement the object detection method according to any one of claims 1 to 7.