High-speed moving target recognition method, device, storage medium and electronic device
Through the combination of low-power dynamic vision sensors and embedded GPU edge computing platform, the spatiotemporal filtering core noise reduction and object detection segmentation model is adopted to solve the problems of high power consumption and large data volume of high-speed cameras, real-time, low-complexity high-speed abnormally dynamic target recognition on the edge computing platform.
Patent Information
- Application Number
- CN202510089797.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the prior art, high power consumption and large data volume of high-speed cameras have made the detection system complex structure and difficult to effectively deploy on edge computing platforms.
The low-power dynamic vision sensor and embedded GPU edge computing platform are used to process the original event stream data through spatiotemporal filtering core noise reduction, convert it into polar channel image frame data, and identify it using preset object detection and segmentation models.
It realizes high-speed abnormal target recognition with low power consumption and low data rate, simplifies the system structure, and supports real-time detection and video visualization on edge computing platforms.
Smart Images

Figure CN119540838B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular, to a method, device, storage medium, and electronic device for identifying high-speed moving targets. Background Art
[0002] The real-time identification of high-speed object movement has important application values in scenarios such as traffic management, autonomous driving, aircraft navigation and obstacle avoidance, production line quality inspection, and sports event analysis.
[0003] Currently, high-speed cameras are usually used for image capture. High-speed cameras can shoot at extremely high frame rates, capturing high-speed motion details that are invisible to the naked eye, thus providing clear image sequences for subsequent analysis. However, due to the large instantaneous data volume of high-speed cameras, usually above 10MB / s, and high power consumption, typical power consumption is 30 - 50W or even higher. Due to this characteristic, they can only work for a short time, and the requirements for the backend data transmission bandwidth and data processing ability are also relatively high, and they can only be processed on servers or desktop-level hosts, resulting in an overall overly complex detection system. Summary of the Invention
[0004] In view of this, this application provides a method, device, storage medium, and electronic device for identifying high-speed moving targets, mainly aiming to solve the problems of high power consumption, large data volume, and complex detection system structure.
[0005] According to the first aspect of this application, a method for identifying high-speed moving targets is provided, which is applied to an edge computing platform. The method includes:
[0006] Obtain the original event stream data collected by a dynamic vision sensor within a continuous time period;
[0007] Based on a preset spatio-temporal filtering kernel, detect whether there is a target event within the spatio-temporal neighborhood range of any event in the original event stream data, and perform noise reduction processing on the original event stream data according to the detection result to obtain the event stream data after noise reduction processing;
[0008] Determine the polarity corresponding to each denoised event in the event stream data after noise reduction processing, and map each denoised event to the corresponding image channel according to the polarity to obtain the polarity channel image frame data;
[0009] Input the polarity channel image frame data into a preset target detection model and a preset target segmentation model respectively for processing to obtain a target detection result and a target segmentation result;
[0010] Determine the identification result of the moving target according to the target detection result and the target segmentation result.
[0011] According to a second aspect of the present application, a high-speed moving target recognition device is provided, and the device includes:
[0012] An acquisition unit, configured to acquire raw event stream data collected by a dynamic vision sensor within a continuous time period;
[0013] A noise reduction unit, configured to detect whether there is a target event within the spatio-temporal neighborhood range of any event in the raw event stream data based on a preset spatio-temporal filtering kernel, and perform noise reduction processing on the raw event stream data according to the detection result to obtain denoised event stream data;
[0014] A mapping unit, configured to determine the polarity corresponding to each denoised event in the denoised event stream data, and map each denoised event to a corresponding image channel according to the polarity to obtain polarity channel image frame data;
[0015] A detection and segmentation unit, configured to input the polarity channel image frame data into a preset target detection model and a preset target segmentation model respectively for processing to obtain a target detection result and a target segmentation result;
[0016] A determination unit, configured to determine a moving target recognition result according to the target detection result and the target segmentation result.
[0017] According to a third aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned high-speed moving target recognition method is implemented.
[0018] According to a fourth aspect of the present application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the program, the above-mentioned high-speed moving target recognition method is implemented.
[0019] With the above technical solution, a high-speed moving target recognition method, device, storage medium and electronic device provided by the present application can, compared with the prior art, acquire the original event stream data collected by a dynamic vision sensor within a continuous time period, then perform noise reduction processing on the original event stream data, and convert the event stream data after noise reduction processing into polar channel image frame data. Then, target recognition and target segmentation are respectively performed on the polar channel image frame data. Finally, according to the target detection result and the target segmentation result, the moving target recognition result is determined. It can be seen that the present application uses a low-power dynamic vision sensor to capture high-speed object movement and picture change information, can achieve extremely low-power and extremely low-data-rate motion data acquisition, and can be input into an embedded edge computing platform in the form of an event stream. The embedded edge computing platform of the present application has AI computing power and good portability, so that a high-speed moving target detection system with low power consumption, miniaturization and lightweight deployment at the edge can be built. In addition, the present application can continuously track high-speed moving targets and can achieve real-time, video visualization and continuous detection and segmentation of high-speed moving targets.
[0020] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. Brief Description of the Drawings
[0021] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0022] Figure 1 Shows a schematic flow chart of a high-speed moving target recognition method provided by an embodiment of the present application;
[0023] Figure 2 Shows a schematic structural diagram of a high-speed moving target detection system provided by an embodiment of the present application;
[0024] Figure 3 Shows a schematic flow chart of noise reduction processing provided by an embodiment of the present application;
[0025] Figure 4 Shows a schematic overall flow chart of high-speed moving target recognition provided by an embodiment of the present application;
[0026] Figure 5 Shows a schematic structural diagram of a high-speed moving target recognition device provided by an embodiment of the present application. Detailed Description of the Invention
[0027] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0028] Existing solutions have high power consumption, large amounts of data, and the resulting problems of weak data processing capabilities and complex systems.
[0029] To solve the above problems, embodiments of the present invention provide a high-speed abnormal target recognition method, which is applied to an edge computing platform, such as Figure 1 As shown, the method includes:
[0030] Step 10: Obtain the original event stream data collected by a dynamic vision sensor within a continuous time period.
[0031] Among them, the dynamic vision sensor includes an event camera, and the original event stream data includes the spatial coordinates, label information, timestamps of data points, and the number of data points.
[0032] To solve the problems of high power consumption, large amounts of data, and complex detection system structure in the prior art, embodiments of the present invention provide a high-speed abnormal target detection system, such as Figure 2 As shown, it includes a low-power dynamic vision sensor, an embedded GPU edge computing platform, and a visualization platform. Among them, the dynamic vision sensor is a low-power event camera. The main function of the event camera is to capture high-speed object motion and picture change information. It has good sensitivity to dynamic changes, can achieve extremely low power consumption and extremely low data rate motion data acquisition, and input the acquired data into the embedded GPU edge computing platform for processing in the form of an event stream; the embedded GPU edge computing platform has AI computing power and good portability. Its main function is to perform real-time target detection and target segmentation based on the dynamic vision event stream data; the visualization platform is a set of software algorithms that can display the detection results and segmentation results in the form of continuous frame images.
[0033] When collecting data based on the dynamic vision sensor of the above system, obtain the spatial coordinates, timestamps, and label information of each event collected by the dynamic vision sensor within a continuous time period, and determine the original event stream data according to the spatial coordinates, timestamps, and label information of each event.
[0034] Among them, each event corresponds to a data point, and the label information is used to represent the direction of light intensity change. For example, the label information is represented by positive and negative. When the label information is positive, it means that the light intensity at the corresponding spatial position increases; when the label information is negative, it means that the light intensity at the corresponding spatial position decreases.
[0035] Specifically, based on a preset integration time, the acquired data for a period of time is accumulated to form the original event stream data (x, y, t, l), N0, where x and y represent the spatial coordinates of the data points, l represents the label information of the data points, t represents the timestamp of the data points, and N0 represents the number of data points. Each data point in the original event stream data represents an event.
[0036] Step 20: Based on a preset spatio-temporal filter kernel, detect whether there is a target event within the spatio-temporal neighborhood range of any event in the original event stream data, and perform noise reduction processing on the original event stream data according to the detection result to obtain the event stream data after noise reduction processing.
[0037] Wherein, the target event is an event among other events except any event itself in the original event stream data.
[0038] For the embodiments of the present invention, since the original event stream data collected by the dynamic vision sensor also contains sensor thermal noise and environmental noise, in order to improve the quality of the collected data and ensure the accuracy of subsequent target detection results and target segmentation results, it is necessary to perform noise reduction processing on the events in the original event stream data collected according to the spatio-temporal neighborhood information. The noise reduction processing process for the original event stream data is as Figure 3 shown and includes:
[0039] Step 21: Based on the size corresponding to the preset spatio-temporal filter kernel, determine whether there is a target event within the spatio-temporal neighborhood range of any event.
[0040] Wherein, the size corresponding to the preset spatio-temporal filter kernel includes a spatial size and a time size, and the spatial size and the time size can be set according to actual service requirements, and the embodiments of the present invention do not make specific limitations on this.
[0041] For the embodiments of the present invention, when determining whether there is a target event within the spatio-temporal neighborhood range of any event, according to the spatial coordinates and timestamp of any event, and the spatial coordinates and timestamp of other events except any event in the original event stream data, calculate the spatial position distance and time difference between any event and other events; if the spatial position distance is less than or equal to the spatial size, and the time difference is less than or equal to the time size, it is determined that there is a target event within the spatio-temporal neighborhood range of any event, and the other event is the target event.
[0042] Specifically, assume that the original event stream data is E = {e1, e2, …, ei, …, eN}, where ei represents any event in the original event stream data, and any event ei = (xi, yi, ti) has spatial coordinates (xi, yi) and a timestamp ti. For any event ei, detect the neighborhood within its spatial and temporal range, that is, under the influence of the preset spatio-temporal filtering kernel K(x, y, t), whether there exists a target event ej = (xj, yj, tj) such that:
[0043] |xi - xj| ≤ Δx, |yi - yj| ≤ Δy, |ti - tj| ≤ Δt
[0044] Among them, Δx and Δy represent the spatial dimensions of the preset spatio-temporal filtering kernel, and Δt represents the temporal dimension of the preset spatio-temporal filtering kernel. The spatio-temporal neighborhood range of any event can be defined by the spatial dimensions (Δx, Δy) and the temporal dimension (Δt).
[0045] If there exists a target event ej that satisfies the above relationship, it is determined that any event ei is not noise and does not need to be removed from the original event stream data.
[0046] Step 22: If there is no target event within the spatio-temporal neighborhood range, it is determined that the any event is noise, and the any event is removed from the original event stream data to obtain the denoised event stream data.
[0047] For the embodiments of the present invention, if there does not exist a target event ej that satisfies the above relationship, it is determined that any event ei is noise and is removed from the original event stream data. The embodiments of the present invention perform the above detection process for each event in the original event stream data, can determine whether it is a noise point, and in the case of determining that it is a noise point, remove it from the original event stream data, thereby obtaining the denoised event stream data (x, y, t, l) N1, where N1 represents the number of data points after denoising.
[0048] Step 30: Determine the polarity corresponding to each denoised event in the denoised event stream data, and map each denoised event to the corresponding image channel according to the polarity to obtain the polarity channel image frame data.
[0049] Among them, the polarity corresponding to each denoised event includes positive polarity and negative polarity.
[0050] For the embodiments of the present invention, after denoising the original event stream data, it is necessary to convert the denoised event stream data into polar channel image frame data. For this conversion process, the method includes: determining the previous several events corresponding to each denoised event according to the timestamp corresponding to each denoised event; determining the polarity corresponding to each denoised event according to the label information of the previous several events and the label information of each denoised event.
[0051] When determining the polarity corresponding to each denoised event according to the label information of the previous several events and the label information of each denoised event, determine the number of adjacent event label information that is the same among any two of the previous several events corresponding to each denoised event and the label information of each denoised event; if the number of the same label information is less than the preset number, determine the polarity corresponding to each denoised event according to the label information of the previous event corresponding to each denoised event and the label information of each denoised event; if the number of the same label information is greater than or equal to the preset number, determine the number of events corresponding to different label information according to the label information of the previous several events, and determine the target label information corresponding to the largest number of events according to the number of events corresponding to different label information; determine the polarity corresponding to each denoised event according to the target label information and the label information of each denoised event.
[0052] Among them, the preset number can be set according to actual business requirements, and the embodiments of the present invention do not make specific limitations on this.
[0053] Specifically, each denoised event in the denoised event stream data has a timestamp and label information. For the timestamp of any denoised event, a number of events before that event can be determined. Then, based on the label information of the previous several events and the label information of any denoised event, the polarity corresponding to any denoised event is determined. For example, for any denoised event ei, according to the label information of the 10 denoised events before it and the label information of event ei, the number of pairs of adjacent events with the same label is determined. Suppose there are 2 pairs of events with the same label among 10 pairs of any adjacent events. At this time, the number of consistent label information is determined to be 2. Since the number of consistent label information 2 is less than the preset number 5, the label information of the denoised event ei is compared with that of its previous event ej, where the previous event ej is adjacent to event ei. If the label information of event ei is the same as that of event ej, the polarity of the denoised event ei is determined to be positive; on the contrary, if the label information of event ei is opposite to that of event ej, the polarity of the denoised event ei is determined to be negative. Suppose there are 7 pairs of events with the same label among 10 pairs of any adjacent events. At this time, the number of consistent label information is determined to be 7. Since the number of consistent label information 7 is greater than the preset number 5, the number of events with positive and negative label information is determined respectively, and the label information with the largest number of events, that is, the target label information, is determined. Then, the target label information is compared with the label information of the denoised event ei. If the two are the same, the polarity of the denoised event ei is determined to be positive; on the contrary, if the two are not the same, the polarity of the denoised event ei is determined to be negative.
[0054] Thus, in the above manner, the polarity corresponding to each denoised event in the denoised event stream data can be accurately determined, thereby ensuring the conversion effect of subsequent image frames.
[0055] Further, after determining the polarity corresponding to each denoised event, according to this polarity, each denoised event is mapped to the corresponding image channel, thereby obtaining the polar channel image frame data. For this process, the method includes: creating a blank image frame with a size consistent with the field of view of the dynamic vision sensor for each image channel, where different image channels correspond to different polarities; mapping each denoised event to the blank image frame of the corresponding image channel according to the polarity and spatial coordinates corresponding to each denoised event, to obtain the polar channel image frame data;
[0056] Specifically, the size of the polar channel image is (H, W, C), where W and H represent the width and height of the image, and C represents the number of image channels. The number of image channels is 2, corresponding to positive and negative polarities respectively. Then, a blank image frame is created for each image channel.
[0057] For each denoised event \(e_i=(x_i,y_i,t_i,p_i)\), where \((x_i,y_i)\) represents the spatial coordinates, \(t_i\) represents the timestamp, and \(p_i\) represents the polarity, according to the polarity corresponding to each denoised event, and in accordance with its corresponding spatial coordinates, it is mapped into the blank image frame corresponding to the corresponding image channel. For example, if the polarity corresponding to the denoised event \(e_i\) is positive polarity, the pixel value at the corresponding position \((x_i,y_i)\) is updated in channel 1; if the polarity corresponding to the denoised event \(e_i\) is negative polarity, the pixel value at the corresponding position \((x_i,y_i)\) is updated in channel 2, and each update is performed in an accumulative manner.
[0058] In this way, each denoised event can be accurately mapped into the corresponding image channel according to the above method, thus completing the conversion of the image frame data of the polarity channel.
[0059] Step 40: Input the image frame data of the polarity channel into a preset target detection model and a preset target segmentation model respectively for processing to obtain a target detection result and a target segmentation result.
[0060] Among them, the targets to be detected and segmented specifically include vehicles, animals, etc., and the targets in the embodiments of the present invention are not limited to the above examples. The preset target detection model includes the R-CNN model, Single Shot MultiBox Detector (SSD), YOLO model, RetinaNet model, etc., and the preset target segmentation model includes the FCN model, U-Net model, Mask R-CNN model, PSPNet model, etc. It should be noted that the preset target detection model and the preset target segmentation model in the embodiments of the present invention are not limited to the above examples, and other models can also be used.
[0061] For the embodiments of the present invention, after obtaining the image frame data of the polarity channel, the image frame data of the polarity channel is input into a preset target detection model for target detection to identify the targets in the image of the polarity channel. At the same time, the image frame data of the polarity channel is input into a preset target segmentation model for target segmentation to obtain a target mask image.
[0062] Step 50: Determine the abnormal target recognition result according to the target detection result and the target segmentation result.
[0063] For the embodiments of the present invention, after determining the target detection result and the target segmentation result, the target detection result and the target segmentation result are rendered to obtain the abnormal target recognition result. For this process, the method includes: determining the center point coordinates, size, and quantity of the detection box according to the target detection result; superimposing the detection box on the polar channel image frame data according to the center point coordinates, size, and quantity of the detection box to obtain the superimposed image frame data; rendering and displaying the superimposed image frame data; determining the mask image according to the target segmentation result; separately displaying the mask image in the form of color blocks; and determining the abnormal target recognition result based on the displayed superimposed image frame data and the mask image in the form of color blocks. The overall detection process of the embodiments of the present invention is as Figure 4 shown. This process receives real-time dynamic vision sensor data, can perform efficient continuous processing by setting the integration time, and can output in the form of a video to obtain a video tracking result.
[0064] The embodiments of the present invention use a dynamic vision sensor and an embedded GPU edge computing platform to form a detection system suitable for the edge side, solving the problems of high data rate, high power consumption, and complex system structure in the prior art. Compared with the prior art, the real-time detection system of the embodiments of the present invention has a lower data rate and system power consumption, simplifies the system structure, has portability, and can support pipeline and outdoor field work. In addition, the embodiments of the present invention adopt a process of event stream data denoising, preprocessing, image conversion, detection and segmentation, and output, and can obtain a processing result in the form of a continuous frame video. Compared with the prior art, the data processing method in the real-time detection system of the embodiments of the present invention has better real-time performance, can realize continuous detection, tracking, and video visualization of high-speed targets, and has better visibility.
[0065] Further, as Figure 1 and Figure 3 a specific implementation of the method shown, this embodiment provides a high-speed abnormal target recognition device, as Figure 5 shown. The device includes: an acquisition unit 101, a denoising unit 102, a mapping unit 103, a detection and segmentation unit 104, and a determination unit 105.
[0066] The acquisition unit 101 can be used to acquire the original event stream data collected by the dynamic vision sensor within a continuous period.
[0067] The denoising unit 102 can be used to detect whether there is a target event within the spatio-temporal neighborhood range of any event in the original event stream data based on a preset spatio-temporal filtering kernel, and perform denoising processing on the original event stream data according to the detection result to obtain the denoised event stream data.
[0068] The mapping unit 103 can be used to determine the polarity corresponding to each denoised event in the event stream data after denoising processing, and map each denoised event to the corresponding image channel according to the polarity to obtain polarity channel image frame data.
[0069] The detection and segmentation unit 104 can be used to input the polarity channel image frame data into a preset target detection model and a preset target segmentation model respectively for processing to obtain a target detection result and a target segmentation result.
[0070] The determination unit 105 can be used to determine an abnormal target recognition result according to the target detection result and the target segmentation result.
[0071] In some embodiments, the acquisition unit 101 can specifically be used to acquire the spatial coordinates, timestamps, and label information of each event collected by the dynamic vision sensor within a continuous time period; and determine the original event stream data according to the spatial coordinates, timestamps, and label information of each event.
[0072] In some embodiments, the denoising unit 102 includes: a determination module and a removal module.
[0073] The determination module can be used to determine whether there is a target event within the spatio-temporal neighborhood range of any one event based on the size corresponding to a preset spatio-temporal filtering kernel.
[0074] The removal module can be used to determine that any one event is noise if there is no target event within the spatio-temporal neighborhood range, and remove any one event from the original event stream data to obtain the event stream data after denoising processing.
[0075] In some embodiments, the determination module can specifically be used to calculate the spatial position distance and time difference between any one event and other events according to the spatial coordinates and timestamp of any one event, and the spatial coordinates and timestamps of other events except any one event in the original event stream data; if the spatial position distance is less than or equal to the spatial size, and the time difference is less than or equal to the time size, it is determined that there is a target event within the spatio-temporal neighborhood range of any one event, and the other event is the target event.
[0076] In some embodiments, the mapping unit 103 includes: a first determination module and a second determination module.
[0077] The first determination module can be used to determine the previous several events corresponding to each noise-reduced event according to the timestamps of the each noise-reduced event. The second determination module can be used to determine the polarity corresponding to each noise-reduced event according to the label information of the previous several events and the label information of each noise-reduced event.
[0078] In some embodiments, the second determination module can specifically be used to determine the number of consistent label information of any two adjacent events among each noise-reduced event and its corresponding previous several events according to the label information of the previous several events and the label information of each noise-reduced event; if the number of consistent label information is less than a preset number, then determine the polarity corresponding to each noise-reduced event according to the label information of the previous event corresponding to each noise-reduced event and the label information of each noise-reduced event; if the number of consistent label information is greater than or equal to the preset number, then determine the number of events corresponding to different label information according to the label information of the previous several events, and determine the target label information corresponding to the largest number of events according to the number of events corresponding to different label information; determine the polarity corresponding to each noise-reduced event according to the target label information and the label information of each noise-reduced event.
[0079] In some embodiments, the mapping unit 103 can specifically be used to create a blank image frame with a size consistent with the field of view of the dynamic vision sensor for each image channel, where different image channels correspond to different polarities; map each noise-reduced event to the blank image frame of the corresponding image channel according to the polarity and spatial coordinates corresponding to each noise-reduced event to obtain the polar channel image frame data.
[0080] In some embodiments, the determination unit 105 can specifically be used to determine the center point coordinates, size and quantity of the detection frame according to the target detection result; superimpose the detection frame on the polar channel image frame data according to the center point coordinates, size and quantity of the detection frame to obtain the superimposed image frame data; render and display the superimposed image frame data; determine the mask image according to the target segmentation result; separately display the mask image in the form of color blocks; determine the abnormal target recognition result based on the displayed superimposed image frame data and the mask image in the form of color blocks.
[0081] It should be noted that other corresponding descriptions of the various functional units involved in a high-speed abnormal target recognition device provided in this embodiment can be referred to Figure 1 and Figure 3 the corresponding descriptions in, and will not be elaborated here.
[0082] Based on the above asFigure 1 and Figure 3 For the method shown in Figure 1 and Figure 3 , correspondingly, this embodiment further provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the high-speed abnormal movement target recognition method as shown in Figure 1 and Figure 3 . Figure 1 and Figure 3 shown above
[0083] Based on such understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions to enable an electronic device (which can be a personal computer, server, or network device, etc.) to execute the methods in various implementation scenarios of this application.
[0084] Based on the method as shown in Figure 1 and Figure 3 above, and the virtual device embodiment shown in Figure 5 , in order to achieve the above object, this embodiment of this application further provides an electronic device, which can specifically be a personal computer, tablet computer, server, or other network device, etc. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the high-speed abnormal movement target recognition method as shown in Figure 1 and Figure 3 above. Figure 1 and Figure 3 shown above, and Figure 5 shown above Figure 1 and Figure 3 shown above
[0085] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0086] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not limit the physical device, and may include more or fewer components, or combine some components, or have different component arrangements.
[0087] The storage medium may further include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the information processing physical device.
[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware.
[0089] The embodiment of the present invention adopts a detection system suitable for the edge side composed of a dynamic vision sensor and an embedded GPU edge computing platform, which solves the problems of high data rate, high power consumption, and complex system structure in the prior art. Compared with the prior art, the real-time detection system of the embodiment of the present invention has a lower data rate and system power consumption, simplifies the system structure, has portability, and can support pipeline and outdoor field work. In addition, the embodiment of the present invention adopts a process of event stream data denoising, preprocessing, image conversion, detection and segmentation, and output, and can obtain a processing result in the form of a continuous frame video. Compared with the prior art, the data processing method in the real-time detection system of the embodiment of the present invention has better real-time performance, can realize continuous detection, tracking, and video visualization of high-speed targets, and has better visibility.
[0090] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.
[0091] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above-disclosed are only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A high-speed moving target recognition method, characterized in that: Applied to edge computing platforms, including: Obtaining raw event stream data collected by dynamic vision sensors in continuous time periods; Based on a preset spatiotemporal filter kernel, detecting whether there is a target event within the spatiotemporal neighborhood of any event in the original event stream data, and performing noise reduction processing on the original event stream data according to the detection result to obtain event stream data after noise reduction processing; Determine the polarity corresponding to each denoised event in the denoised event stream data, and map each denoised event to a corresponding image channel according to the polarity to obtain polarity channel image frame data; Inputting the polarity channel image frame data into a preset target detection model and a preset target segmentation model for processing respectively, to obtain a target detection result and a target segmentation result; Determining a moving target recognition result according to the target detection result and the target segmentation result; The step of determining the polarity corresponding to each denoised event in the denoised event stream data includes: Determine the number of identical label information between any two adjacent events in each denoised event and its corresponding previous events; If the number is less than the preset number, the polarity corresponding to each denoised event is determined according to the label information of the previous event corresponding to each denoised event, wherein if the label information of the previous event is consistent with the label information of each denoised event, the polarity of each denoised event is determined to be positive; If the number is greater than or equal to the preset number, determining the number of events corresponding to different label information according to the label information of the first several events, and determining the target label information corresponding to the maximum number of events according to the number of events corresponding to the different label information; The polarity corresponding to each noise-reduced event is determined according to the target label information.
2. The method according to claim 1, characterized in that The obtaining of raw event stream data collected by the dynamic vision sensor in a continuous period of time includes: Acquire spatial coordinates, timestamps and label information of each event collected by the dynamic vision sensor in a continuous period of time; The original event stream data is determined according to the spatial coordinates, timestamp and tag information of each event.
3. The method according to claim 1, characterized in that The method of detecting whether there is a target event within the spatiotemporal neighborhood of any event in the original event stream data based on a preset spatiotemporal filter kernel, and performing noise reduction processing on the original event stream data according to the detection result to obtain the event stream data after noise reduction processing includes: Based on the size corresponding to the preset spatiotemporal filter kernel, determining whether there is a target event within the spatiotemporal neighborhood of any one of the events; If there is no target event within the space-time neighborhood, it is determined that any one of the events is noise, and the any one of the events is removed from the original event stream data to obtain the event stream data after noise reduction processing.
4. The method according to claim 3, characterized in that The size corresponding to the preset spatiotemporal filter kernel includes a spatial size and a temporal size, and the determining whether there is a target event within the spatiotemporal neighborhood of any one of the events based on the size corresponding to the preset spatiotemporal filter kernel includes: Calculate the spatial position distance and time difference between the any one event and the other events in the original event stream data according to the spatial coordinates and timestamp of the any one event and the spatial coordinates and timestamps of other events except the any one event; If the spatial position distance is less than or equal to the spatial size, and the time difference is less than or equal to the time size, it is determined that there is a target event within the spatiotemporal neighborhood of any one of the events, and the other event is the target event.
5. The method according to claim 1, characterized in that Mapping each noise-reduced event to a corresponding image channel according to the polarity to obtain polarity channel image frame data includes: Creating a blank image frame for each image channel whose size is consistent with the field of view of the dynamic vision sensor, wherein different image channels correspond to different polarities; According to the polarity and spatial coordinates corresponding to each denoised event, each denoised event is mapped to a blank image frame of a corresponding image channel to obtain the polarity channel image frame data; and / or Determining a moving target recognition result according to the target detection result and the target segmentation result includes: Determine the center point coordinates, size and number of the detection frame according to the target detection result; According to the center point coordinates, size and number of the detection frame, superimposing the detection frame on the polarity channel image frame data to obtain superimposed image frame data; Rendering and displaying the superimposed image frame data; Determining a mask image according to the target segmentation result; The mask image is displayed separately in the form of color blocks; The moving target recognition result is determined based on the displayed superimposed image frame data and the mask image in the form of color blocks.
6. A high-speed moving target recognition device, characterized in that: include: An acquisition unit, used for acquiring raw event stream data collected by a dynamic vision sensor in a continuous period of time; A denoising unit, configured to detect whether there is a target event within the spatiotemporal neighborhood of any event in the original event stream data based on a preset spatiotemporal filter kernel, and to perform denoising processing on the original event stream data according to the detection result to obtain the denoised event stream data; A mapping unit, used to determine the polarity corresponding to each denoised event in the denoised event stream data, and map each denoised event to a corresponding image channel according to the polarity to obtain polarity channel image frame data; A detection and segmentation unit, used for inputting the polarity channel image frame data into a preset target detection model and a preset target segmentation model for processing, to obtain a target detection result and a target segmentation result; A determination unit, configured to determine a moving target recognition result according to the target detection result and the target segmentation result; Wherein, the mapping unit includes: a first determination module and a second determination module; The first determination module is used to determine the number of consistent label information between any two adjacent events in each noise-reduced event and its corresponding previous events; The second determination module is used to determine the polarity corresponding to each denoised event according to the label information of the previous event corresponding to each denoised event if the number is less than the preset number, wherein if the label information of the previous event is consistent with the label information of each denoised event, the polarity of each denoised event is determined to be positive; The second determination module is also used to determine the number of events corresponding to different label information based on the label information of the previous events if the number is greater than or equal to the preset number, and determine the target label information corresponding to the maximum number of events based on the number of events corresponding to the different label information; and determine the polarity corresponding to each noise-reduced event based on the target label information.
7. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Event stream denoising method and device
CN119228667A
Aerial target detection method fusing event camera and deep learning
CN119251463A