A method and system for high-speed target tracking based on an event camera

Through the combination of event camera data reconstruction and twin network model, the problem of traditional cameras being difficult to track targets in high-speed motion scenarios is solved, and the clear capture and continuous tracking of high-speed targets is achieved, which is suitable for applications such as unmanned driving and drones.

CN114202564BActive Publication Date: 2025-05-27SHENZHEN LONGGANG INTELLIGENT AUDIOVISUAL RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111543971.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-05-27
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

Traditional cameras are difficult to effectively capture and track targets in high-speed motion scenarios, and the asynchronous independent pixel data of event cameras is difficult to compatible with traditional artificial intelligence processing technologies.

Method used

The event camera is used to capture the data of high-speed motion scenes, and a new event image is formed through data reconstruction. The image features of the target are extracted and compared with the twin network model to achieve continuous tracking of high-speed motion goals.

Benefits of technology

It effectively overcomes the problem of motion blur in high-speed motion scenarios by traditional cameras, achieves clear capture and continuous tracking of high-speed targets, and provides solutions for application scenarios such as unmanned driving and drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202564B_ABST
    Figure CN114202564B_ABST
Patent Text Reader

Abstract

Method and system for tracking high-speed targets based on an event camera, which uses an event camera data acquisition module to collect images of high-speed moving targets; uses an event data reconstruction module to combine a certain number of event pixels, and arranges these combinations according to their position information. If there are repeated positions, the pixel with a larger polarity replaces the pixel with a smaller polarity, and finally a new event image is formed by combination; uses a Siamese network model to extract image features of the target. The method and system of the present invention can reconstruct event data that is not easy to process; and effectively track high-speed targets captured by the event camera using target tracking technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a method and system for high-speed target tracking based on an event camera. Background Art

[0002] The event camera, also known as a dynamic vision sensor, is a new type of sensor inspired by biology. Different from traditional cameras, the event camera captures information by updating asynchronous and independent event pixels, and these pixel information will be activated by changes in light. Since the event camera does not have a global shutter and does not need to be exposed within a certain time to obtain information. Therefore, the event camera can capture high-speed moving targets. While ordinary cameras will produce serious motion blur for high-speed moving targets.

[0003] Target tracking is a technology that uses intelligent algorithms to continuously track targets in videos. Generally speaking, the input of the target tracking algorithm is a sequence of image frames along the time axis. The target tracking method continuously outputs the position of the target in each frame according to the characteristics of the target to be tracked, so as to achieve continuous tracking of the target.

[0004] Although the event camera can effectively capture moving targets, due to the independent and asynchronous distribution of the pixels of the event camera, there are great differences from traditional image frames. At present, traditional artificial intelligence processing technologies cannot fully adapt to this data form.

[0005] The difficulties in solving the above problems and defects are as follows: It is necessary to solve multiple problems such as the inability of traditional cameras to effectively capture high-speed targets and the combination of event camera data and artificial intelligence algorithms, and the solution difficulty is relatively large.

[0006] The significance of solving the above problems and defects is as follows: It provides a new method and system for high-speed target tracking, overcomes the problem of motion blur faced by traditional cameras in obtaining high-speed targets, enables the algorithm to effectively track targets, and provides solutions for application scenarios such as unmanned driving and drones. Summary of the Invention

[0007] The purpose of the method of the present invention is to provide a method and system for tracking high-speed targets based on an event camera, mainly using the event camera to capture high-speed moving targets and processing the original data of the event camera, so that the high-speed moving targets obtained by the event camera can be continuously tracked using target tracking technology. Therefore, the present invention mainly solves two problems: reconstructing and processing difficult-to-process event data; and effectively tracking high-speed targets captured by the event camera using target tracking technology.

[0008] The technical solution of the present invention is as follows:

[0009] According to one aspect of the present invention, a method for high-speed target tracking based on an event camera is provided, including the following steps: S1. Obtain event camera data: Use an event camera to obtain the event camera data recorded in real time in a high-speed moving scene; S2. Generate data packets: Pack the scattered event camera data; S3. Data packet transmission: Transmit the packed data, unpack the transmitted data, and then input the unpacked data into step S4 for event camera data reconstruction; S4. Event camera data reconstruction: Combine a certain number of event pixels, and arrange these combinations according to their position information, and finally combine them to form a new event image; S5. First frame judgment: Judge whether the image frame is the first frame image. If it is the first frame image, input it into step S6 for target initialization processing. If it is not the first frame, input it into step S8 for processing; S6. Initialize the target: Perform an initialization operation on the target, select the target to be tracked in the event frame and input it into step S7; S7. Extract target image features: Use the trained Siamese network model to extract the image features of the target, and input the extracted target image features into step S9 for feature comparison; S8. Extract search image features: If the non-first frame image is input into step S8, extract the features of the searched image, and input the features into step S9 for comparison; S9. Siamese comparison: Fuse and compare the target image features and search image features obtained in steps S7 and S8 in the Siamese network model; and S10. Output results: Output the comparison results, and continuously obtain data for continuous tracking until the tracking exits.

[0010] Preferably, in the above method for high-speed target tracking based on an event camera, in step S1, an event camera is used to collect data on a high-speed moving target scene. For changes in the scene, each event pixel in the event camera is used to record. For each event pixel, it consists of four values: two position information, one information on the magnitude of the polarity, and one time information.

[0011] Preferably, in the above method for high-speed target tracking based on an event camera, in step S2, the four data of each event camera form a data packet.

[0012] Preferably, in the above method for high-speed target tracking based on an event camera, in step S4, when arranging these combinations according to their position information, if there are repeated positions, replace the pixel with a smaller polarity with the pixel with a larger polarity. Among them, the new event image is a spatial image with the same resolution as the event camera, and contains more spatial structure information than a single pixel.

[0013] Preferably, in the above method for high-speed target tracking based on an event camera, in step S6, the selected target to be tracked includes the position information and size information of the target.

[0014] Preferably, in the above method for high-speed target tracking based on an event camera, in step S7, the Siamese network includes a first network and a second network with the same weights. The input of the first network is the image of the selected target to be tracked, and it extracts the target features from the image. The input of the second network is the image of the area to be searched, and it outputs the features of this image; finally, the features are fused and the result is output.

[0015] Preferably, in the above method for high-speed target tracking based on an event camera, in step S7, for the training and application process of the Siamese network, first, feature extraction is performed from the labeled target images and the search space. The extracted feature space is mapped to a higher-dimensional space, and then the distance between them is calculated. The difference between the distance between them and the labeled ground truth is used to calculate the loss function during training; while in the application process, the distance between the output features is calculated, and based on the distance, it is determined whether a suitable target is searched.

[0016] According to another aspect of the present invention, there is provided a high-speed target tracking system based on an event camera, which is characterized in that it includes: an event camera data acquisition module, an event data reconstruction module, and a Siamese network model, wherein: the event camera data acquisition module: is used to collect images of high-speed moving targets; the event data reconstruction module is used to combine a certain number of event pixels, and arrange these combinations according to their position information. If there are repeated positions, the pixel with a larger polarity replaces the pixel with a smaller polarity, and finally a new event image is formed by combination; the Siamese network model is used to extract image features of the target.

[0017] Preferably, in the above high-speed target tracking system based on an event camera, the event camera acquisition module consists of two parts, a data acquisition sub-module and a data transmission sub-module. Among them, the data acquisition module is used to collect data of the high-speed moving target scene by using the event camera. For the changes in the scene, each pixel in the event camera is used for recording. For each event pixel, it consists of four values: two position information, one information about the polarity size, and one time information; among them, the data transmission sub-module is used to, during the transmission process, form a data packet from the four data of each event camera for overall transmission, and transmit the data collected by the event camera one by one in the form of data packets.

[0018] According to the technical solution of the present invention, the beneficial effects are:

[0019] The present invention can track high-speed moving targets. For high-speed moving targets, existing ordinary cameras are unable to capture them clearly because when the frame rate of an ordinary camera is lower than the speed of the moving target, the images obtained by the ordinary camera will have serious motion blur. Aiming at the defect that ordinary cameras cannot acquire and track high-speed targets, the present invention proposes a method for tracking high-speed targets using an event camera, which makes up for the defect of effectively tracking targets in high-speed motion scenarios and effectively realizes the task of target tracking in high-speed motion scenarios.

[0020] When the present invention processes this type of data, it converts the event data into a mode that can be processed by current technologies. By using an event camera to capture high-speed targets and using deep learning technologies to enhance high-speed motion and targets, it can effectively track targets moving at high speed.

[0021] To better understand and illustrate the concept, working principle and invention effect of the present invention, the following will, with reference to the accompanying drawings and through specific embodiments, describe the present invention in detail as follows: BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for describing the specific embodiments or the prior art.

[0023] Figure 1 is a flowchart of the method for tracking high-speed targets based on an event camera of the present invention;

[0024] Figure 2 is a diagram of the twin network architecture of the present invention;

[0025] Figure 3 is a comparison diagram of the images of a high-speed moving object obtained by a traditional camera and the clear vehicle information image obtained by the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical methods and advantages of the present invention clearer, the following will, with reference to the accompanying drawings and specific examples, further describe the present invention in detail. These examples are merely illustrative and not restrictive of the present invention.

[0027] The method of the present invention provides a method for tracking high-speed moving targets. For high-speed moving targets, existing ordinary cameras are unable to acquire their information. Specifically, an event camera is used to acquire the information of high-speed moving targets, and the discrete asynchronous event camera data is reconstructed so that the event data can be effectively applied to target tracking. Then, we use the twin network method to effectively track the moving targets in the event data.

[0028] The principle of the present invention is as follows:

[0029] First, the present invention uses an event camera to obtain data of a high-speed motion scene, and then effectively reconstructs the discrete event data in the scene so that the reconstructed event data has the characteristics of an image. Subsequently, an improved target tracking algorithm is used to track the target in the high-speed motion scene, thereby implementing a high-speed motion target tracking algorithm that cannot be completed by an ordinary camera.

[0030] The method of the present invention proposes a method for tracking high-speed targets for an event camera. By using the advantages of the event camera for the motion scene to obtain data of the high-speed motion scene, effectively reconstructing the event camera data, the target tracking algorithm can be effectively applied to the event data, using a trained Siamese network to extract features from the reconstructed event data, and effectively tracking the high-speed motion target through feature comparison. The specific implementation steps are as follows:

[0031] S1. Obtain event camera data: Use the event camera to obtain the event camera data recorded in real time in the high-speed motion scene. In this step, use the event camera to collect data of the high-speed motion target scene. For the changes in the scene, each pixel in the event camera is used for recording. For each event pixel, it consists of four values: two position information, one information of the polarity magnitude, and one time information.

[0032] S2. Generate data packets: Package the scattered event camera data. Specifically, form a data packet with the four data of each event camera.

[0033] S3. Data packet transmission: Transmit the packaged data, unpack the transmitted data, and then input the unpacked data into step S4 for event camera data reconstruction. During the transmission process, form a data packet with the four data of each event camera for overall transmission, and transmit the data collected by the event camera one by one in the form of data packets.

[0034] S4. Event camera data reconstruction: Combine a certain number of event pixels, and arrange these combinations according to their position information. If there are repeated positions, replace the pixel with a smaller polarity with the pixel with a larger polarity. Finally, combine to form a new event image, and input the event image into step S5. Among them, the new event image is a spatial image with the same resolution as the event camera, and contains more spatial structure information than a single pixel. Since the event camera data is discrete and asynchronous data, it cannot contain complete target spatial features and cannot be further operated. The event camera data reconstruction method in this step can combine multiple single event pixels into a complete event image.

[0035] S5. First Frame Judgment: Determine whether the image frame is the first frame image. If it is the first frame image, input it to step S6 for target initialization processing. If it is not the first frame, input it to step S8 for processing.

[0036] S6. Initialize Target: Perform an initialization operation on the target, that is, select the target to be tracked in the event frame. The selected target should have position information and size information of the target, and input the selected initialized target to step S7.

[0037] S7. Extract Target Image Features: Use the trained Siamese network model to extract image features of the target, and input the extracted target image features to step S9 for feature comparison.

[0038] Among them, after obtaining the event image, it can be input into the Siamese network model one by one for tracking according to the order of generation of the event images. The Siamese network consists of two network structures sharing weight information. The structure of the Siamese network is as Figure 2 shown. The first network 1 and the second network 2 are networks with the same weights, but their inputs and outputs are different. The input of the first network 1 is the image of the selected target to be tracked (i.e., Figure 2 the target template in Figure 2 ), and target feature extraction is performed on this image. The input of the second network 2 is the image of the area to be searched (i.e.,

[0039] the search area in

[0040] ), and the features of this image are output. Finally, the features are fused and the result is output, and this output can determine the position of the target in the search image.

[0041] Siamese Network Model Training and Application: Since the Siamese network consists of two parallel branches, this will affect the training of the model. When designing the training loss function, it is necessary to consider the impact of different network branches on network training. For the training process, first extract features from the labeled target image and the search space, map the extracted feature space to a higher-dimensional space, and then calculate the distance between them. The difference between the distance between them and the labeled true value is used to calculate the loss function during training; while the application process is to calculate the distance between the output features and judge whether the appropriate target is searched according to the distance.

[0042] S8. Extract Search Image Features: If the non-first-frame image is input to step S8, extract features from the searched image, and input the features to step S9 for comparison.

[0041] S9. Siamese Comparison: Fuse and compare the target image features and search image features obtained in steps S7 and S8 in the Siamese network model.

[0042] S10. Output result: Output the comparison result, and continuously obtain data for continuous tracking until the tracking exits.

[0043] A tracking system for high-speed targets based on an event camera, comprising: an event camera data acquisition module, an event data reconstruction module, and a twin network model, wherein: The event camera data acquisition module: is used to collect images of high-speed moving targets. The event camera acquisition module consists of two parts, a data acquisition sub-module and a data transmission sub-module. Among them, the data acquisition module is used to collect data of the high-speed moving target scene by using the event camera. For the changes in the scene, each pixel in the event camera is used for recording. For each event pixel, it consists of four values: two position information, one information of the polarity magnitude, and one time information. Among them, the data transmission sub-module is used to form a data packet from the four data of each event camera for overall transmission during the transmission process, and transmit the data collected by the event camera one by one in the form of data packets. The event data reconstruction module is used to combine a certain number of event pixels, and arrange these combinations according to their position information. If there are repeated positions, replace the pixel with a smaller polarity with the pixel with a larger polarity, and finally combine to form a new event image. The twin network model is used to extract image features of the target.

[0044] Figure 3 It is a comparison chart of the images obtained by a traditional camera for high-speed moving images and the images obtained by the method of the present invention for clear vehicle information. It can be seen from the comparison of the implementation effects: The left traditional camera obtains high-speed moving images, and serious motion blur will occur in the high-speed traditional camera, resulting in the inability to perform target tracking; on the right, the method of the present invention can effectively perform target tracking by obtaining clear vehicle information after visualization in a high-speed scene.

[0045] The method and system of the present invention solve the problem of effectively tracking targets in high-speed moving scenes, realizing a method for tracking high-speed moving targets using an event camera. The effective reconstruction strategy of event camera data solves the problem that event data is discrete and asynchronous and difficult to process. The trained twin network is used to process the data reconstructed by the event camera, realizing a target tracking method in high-speed moving scenes.

[0046] The above description is the best embodiment based on the concept and working principle of the invention. The above embodiments should not be construed as limiting the protection scope of the present claims. Combinations of other implementation manners and implementation modes according to the concept of the present invention all belong to the protection scope of the present invention.

Claims

1. A method for high-speed target tracking based on an event camera, characterized in that, it includes the following steps: S1. Obtain event camera data: Use the event camera to obtain the event camera data recorded in real time in a scene with high-speed movement. Use the event camera to collect data on the target scene with high-speed movement. For changes in the scene, use each event pixel in the event camera to record. For each event pixel, it consists of four values: two position information, one information on the magnitude of the polarity, and one time information; S2. Generate data packets: Pack the scattered event camera data, and form a data packet with the four data of each event camera pixel; S3. Data packet transmission: Transmit the packed data, unpack the transmitted data, and then input the unpacked data into step S4 for event camera data reconstruction; S4. Event camera data reconstruction: Combine a certain number of event pixels, and arrange these combinations according to their position information, and finally combine them to form a new event image; S5. First frame judgment: Judge whether the image frame is the first frame image. If it is the first frame image, input it into step S6 for target initialization processing. If it is not the first frame, input it into step S8 for processing; S6. Initialize the target: Perform an initialization operation on the target, select the target to be tracked in the event frame and input it into step S7; S7. Extract target image features: Use the trained Siamese network model to extract the image features of the target. The extracted target image features are input into step S9 for feature comparison. The Siamese network includes a first network and a second network with the same weights. The input of the first network is the image of the selected target to be tracked, and the target features of the image are extracted. The input of the second network is the image of the area to be searched, and the features of the image are output; finally, the features are fused and the result is output; S8. Extract search image features: If the non-first frame image is input into step S8, extract the features of the searched image and input the features into step S9 for comparison; S9. Siamese comparison: Fuse and compare the target image features and search image features obtained in steps S7 and S8 in the Siamese network model; and S10. Output the result: Output the comparison result, and continuously obtain data for continuous tracking until the tracking exits.

2. The method for high-speed target tracking based on an event camera according to claim 1, characterized in that, in step S4, when arranging these combinations according to their position information, if there are repeated positions, replace the pixel with a smaller polarity with the pixel with a larger polarity. Among them, the new event image is a spatial image with the same resolution as the event camera, and contains more spatial structure information than a single pixel.

3. The method for high-speed target tracking based on an event camera according to claim 1, characterized in that, in step S6, the selected target to be tracked includes the position information and size information of the target.

4. The method for high-speed target tracking based on an event camera according to claim 1, wherein, in step S7, for the training and application process of the Siamese network, first extract features from the labeled target image and the search space, map the extracted feature space to a higher-dimensional space, and then calculate the distance between them. The difference between the distance between them and the labeled ground truth is used to calculate the loss function during training; while in the application process, calculate the distance between the output features and determine whether the appropriate target is searched based on the distance.

5. A tracking system for high-speed targets based on an event camera, which is used to implement the method according to any one of claims 1 to 4, wherein, it includes: an event camera data acquisition module, an event data reconstruction module, and a Siamese network model, where: the event camera data acquisition module: is used to collect images of high-speed moving targets; the event data reconstruction module is used to combine a certain number of event pixels, arrange these combinations according to their position information, and if there are repeated positions, replace the pixel with a smaller polarity with the pixel with a larger polarity, and finally combine to form a new event image; the Siamese network model is used to extract image features of the target.

6. The tracking system for high-speed targets based on an event camera according to claim 5, wherein, the event camera data acquisition module consists of two parts, a data acquisition sub-module and a data transmission sub-module. Among them, the data acquisition sub-module is used to collect data of the high-speed moving target scene by using the event camera, record the changes in the scene with each pixel in the event camera, and each event pixel consists of four values: two position information, one information about the polarity magnitude, and one time information; and the data transmission sub-module is used to form a data packet from the four data of each event camera for overall transmission during the transmission process, and transmit the data collected by the event camera one by one in the form of data packets.

Citation Information

Patent Citations

  • Event camera-based eye detection and tracking method, system, equipment and application

    CN113688733A