Method, apparatus and system for automatic labeling

By combining a stereo camera with a DVS (Disparity Frame) to calculate disparity frames and acquire 3D information, a deep learning model is used to determine the object region, and 3D points are reprojected onto the DVS frame to generate automatic labeling results. This solves the problem of inaccurate DVS frame labeling in existing technologies and achieves efficient and accurate automatic labeling.

CN117121046BActive Publication Date: 2026-07-24HARMAN INT IND INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARMAN INT IND INC
Filing Date
2021-03-26
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing automatic DVS frame labeling methods cannot fully utilize the advantages of DVS, such as low latency, no motion blur, and high dynamic range, and the generated labeling results are inaccurate, thus wasting the advantages of DVS.

Method used

By combining a stereo camera with a DVS (Disparity Frame) to calculate disparity frames and acquire 3D information, a deep learning model is used to determine the object region, and 3D points are reprojected onto the DVS frame to generate automatically labeled results.

Benefits of technology

It achieves efficient and more accurate automatic labeling of DVS frames, fully utilizes the advantages of DVS, generates natural and efficient labeled data, and provides a large amount of labeled data for deep learning training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117121046B_ABST
    Figure CN117121046B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, system, and device for automatically labeling dynamic vision sensor (DVS) frames. The method can include generating, by a pair of cameras, a pair of camera frames within a time interval and generating, by a DVS, at least one DVS frame within the time interval. The method can also calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame. The method can determine an object region for automatic labeling using a deep learning model and can obtain 3D points based on the 3D information and the determined object region. The method can then reproject the 3D points to the at least one DVS frame to generate reprojected points on the at least one DVS frame. The method can also generate at least one automatically labeled result on the at least one DVS by combining the reprojected points on the at least one DVS frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method, apparatus, and system for automatic labeling, and more particularly to a method, apparatus, and system for automatically labeling DVS (Dynamic Visual Sensor) frames. Background Technology

[0002] In recent years, DVS, as a new and cutting-edge sensor, has been widely recognized and applied in many fields such as artificial intelligence, computer vision, autonomous driving, and robotics.

[0003] Compared to conventional cameras, DVS offers advantages such as low latency, no motion blur, high dynamic range, and low power consumption. Specifically, DVS has a latency in the microsecond range, while conventional cameras have a latency in the millisecond range. Therefore, DVS is not affected by motion blur. Consequently, DVS typically has a data rate of 40–180 kB / s (compared to 10 mB / s for conventional cameras), meaning it requires less bandwidth and consumes less power. Furthermore, DVS has a dynamic range of approximately 120 dB, while conventional cameras have a dynamic range of approximately 60 dB. This wider dynamic range is useful in extreme lighting conditions (such as vehicles entering and exiting tunnels, other vehicles in the opposite direction using high beams, changes in sunlight direction, etc.).

[0004] Due to these advantages, DVS has been widely adopted. Efforts have been made to apply DVS to different scenarios. Among all technologies, deep learning is a popular and important direction. When it comes to deep learning, a large amount of labeled data is necessary. However, there may not be enough manual labor to label data manually. Therefore, automatic labeling of DVS frames is needed.

[0005] Currently, there are two methods for automatic labeling of DVS frames. One method is to play regular camera video on the display monitor screen and use DVS to record the screen and label objects. The other method is to use a deep learning model to directly generate labeled DVS frames from camera frames. However, both methods have insurmountable drawbacks. The first method loses accuracy because it is difficult to perfectly match the DVS frames to the display monitor during recording. The second method generates unnatural DVS frames. Different materials have different reflectivities. However, the second method treats them the same because the DVS frames are generated directly from camera frames, which makes the generated DVS frames very unnatural. In addition, both methods fall into the problem of wasting the advantages of DVS because the quality of the camera video limits the final output of the generated DVS frames in several ways. First, the generated DVS frame rate can only reach the camera frame rate at most (although the second method can use zoom methods to obtain more frames, it is still hopeless). Second, motion blur, afterimages, and ghosting recorded by the camera will also exist in the generated DVS frames. This fact is ridiculous because DVS is known for its low latency and lack of motion blur. Third, the high dynamic range of the DVS is wasted because conventional cameras have a lower dynamic range.

[0006] Therefore, improved techniques must be provided to automatically label DVS frames while fully leveraging the advantages of DVS. Summary of the Invention

[0007] According to one or more embodiments of this disclosure, a method for automatically labeling Dynamic Vision Sensor (DVS) frames is provided. The method may include receiving a pair of camera frames generated by a pair of cameras within a time interval and receiving at least one DVS frame generated by the DVS within the time interval. The method may also calculate a disparity frame based on the pair of camera frames and obtain 3D information about the pair of camera frames based on the calculated disparity frame. The method may use a deep learning model to determine object regions for automatic labeling and may obtain 3D points based on the obtained 3D information and the determined object regions. The method may then reproject the 3D points onto at least one DVS frame to generate reprojected points on at least one DVS frame. The method may also generate at least one automatically labeled result on at least one DVS by combining the reprojected points on at least one DVS frame.

[0008] According to one or more embodiments of this disclosure, a system for automatically labeling Dynamic Vision Sensor (DVS) frames is provided. The system may include a pair of cameras, a DVS, and a computing device. The pair of cameras may be configured to generate a pair of camera frames within a time interval. The DVS may be configured to generate at least one DVS frame within the time interval. The computing device may include a processor and a memory unit storing instructions executable by the processor to: receive the pair of camera frames and at least one DVS frame; calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region for automatic labeling using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region, and reproject the 3D points onto at least one DVS frame to generate reprojected points on at least one DVS frame; and generate at least one automatically labeled result on at least one DVS by combining the reprojected points on at least one DVS frame.

[0009] According to one or more embodiments of this disclosure, an apparatus for automatically labeling Dynamic Vision Sensor (DVS) frames is provided. The apparatus may include a computing device comprising a processor and a memory unit storing instructions executable by the processor to: receive a pair of camera frames and at least one DVS frame; calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region for automatic labeling using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region, and reproject the 3D points onto at least one DVS frame to generate reprojected points on at least one DVS frame; and generate at least one automatically labeled result on at least one DVS by combining the reprojected points on at least one DVS frame.

[0010] The methods, apparatus, and systems described in this disclosure enable efficient and more accurate automatic labeling of DVS frames. These methods, apparatus, and systems can bind a pair of cameras to a DVS and simultaneously record the same scene. Based on the combined use of the acquired camera frames and DVS frames, DVS frames can be automatically labeled while recording. Therefore, a large amount of labeled data for DVS deep learning training becomes possible. Compared to existing methods, the methods and systems described in this disclosure fully utilize the advantages of DVS and achieve more accurate and efficient automatic labeling. Attached Figure Description

[0011] Figure 1 A schematic diagram of a system according to one or more embodiments of the present disclosure is shown;

[0012] Figure 2 A flowchart of a method according to one or more embodiments of the present disclosure is shown;

[0013] Figure 3 The parallax principle according to one or more embodiments of this disclosure is illustrated;

[0014] Figure 4 The relationship between parallax and depth information is illustrated in one or more embodiments according to this disclosure;

[0015] Figure 5 Examples of parallax frames calculated from a left camera and a right camera according to one or more embodiments of this disclosure are shown;

[0016] Figure 6 Examples of object detection results on a left camera and a parallax frame according to one or more embodiments of this disclosure are shown;

[0017] Figure 7 An example of reprojecting 3D points onto a DVS frame according to one or more embodiments of this disclosure is shown;

[0018] Figure 8 Exemplary results are shown according to one or more embodiments of this disclosure;

[0019] Figure 9 Another exemplary result is shown according to one or more embodiments of this disclosure.

[0020] For ease of understanding, the same reference numerals are used where possible to denote common elements in the figures. It is anticipated that elements disclosed in one embodiment may be beneficially utilized in other embodiments without specific description. Unless specifically stated otherwise, the figures mentioned herein should not be construed as being drawn to scale. Furthermore, for clarity of presentation and explanation, the figures are generally simplified and details or parts are omitted. The figures and discussion are used to explain the principles discussed below, wherein the same reference numerals denote the same elements. Detailed Implementation

[0021] Examples will be given below. The descriptions of the various examples are for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0022] In general, this disclosure provides a system, apparatus, and method for automatically labeling DVS frames by combining at least one pair of stereo cameras with a DVS (Discrete Visual System). By using a stereo camera to calculate disparity and thus acquire 3D information of the camera frame; using a deep learning model on the camera frame to acquire an object region; reprojecting 3D points corresponding to the object region onto the DVS frame to generate points on the DVS frame; and combining the reprojected points on the DVS frame to generate a final detection result on the DVS frame, the system and method of this disclosure can provide reliably automatically labeled DVS frames due to the combined use of the camera frame and the DVS frame. Based on the obtained combined use of the camera frame and the DVS frame, DVS frames can be automatically labeled simultaneously with recording. Therefore, a large amount of labeled data for deep learning training of DVS becomes possible. Compared with existing methods, the methods, apparatus, and systems described in this disclosure fully utilize the advantages of DVS and can achieve more accurate and efficient automatic labeling.

[0023] Figure 1 A schematic diagram of a system for automatically tagging DVS frames according to one or more embodiments of the present disclosure is shown. Figure 1 As shown, the system may include a recording device 102 and a computer device 104. The recording device 102 may include, but is not limited to, a DVS 102a and a pair of cameras 102b and 102c, such as a left camera 102b and a right camera 102c. Depending on actual needs, the recording device 102 may include more cameras in addition to the left camera 102b and the right camera 102c, but is not limited thereto. For simplicity, only one pair of cameras is shown herein. The term "camera" in this disclosure may include a stereo camera. In the recording device 102, the pair of cameras 102b and 102c and the DVS 102a may be rigidly combined / assembled / integrated together. It should be understood that... Figure 1 This is merely to illustrate the components of the system and is not intended to limit the positional relationships between them. DVS102a can be arranged in any relative positional relationship with the left camera 102b and the right camera 102c.

[0024] The DVS102a employs an event-driven approach to capture dynamic changes in a scene and then creates asynchronous pixels. Unlike conventional cameras, the DVS does not generate images but transmits pixel-level events. When there are dynamic changes in the real scene, the DVS produces some pixel-level output (i.e., events). Therefore, if there is no change, there is no data output. Dynamic changes can include at least one of intensity changes and object movement. Event data is in the form of [x,y,t,p], where x and y represent the coordinates of the event pixel in 2D space, t is the timestamp of the event, and p is the polarity of the event. For example, the polarity of the event can represent changes in scene brightness, such as becoming brighter (positive) or darker (negative).

[0025] The computing device 104 can be any form of device capable of performing calculations, including but not limited to mobile devices, smart devices, laptop computers, tablet computers, in-vehicle navigation systems, etc. The computing device 104 can include, but is not limited to, a processor 104a. The processor 104a can be any technically feasible hardware unit configured to process data and execute software applications, including but not limited to a central processing unit (CPU), microcontroller unit (MCU), application-specific integrated circuit (ASIC), digital signal processor (DSP) chip, etc. The computing device 104 can include, but is not limited to, a memory unit 104b for storing data, code, instructions, etc., executable by the processor. The memory unit 104b can include, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0026] The system for automatically marking DVS frames can be positioned within an operational environment. For example, the system can determine if there are dynamic changes (event-based changes) in the scene, and if such dynamic changes are detected, automatically activate the DVS and the camera pair for operation. The DVS and the camera pair can be synchronized via synchronized timestamps. For the same scene, the left and right cameras can generate at least one left camera frame and at least one right camera frame, respectively, during a time interval. Simultaneously, the DVS can generate at least one DVS frame within the same time interval. Because the time span of a camera frame is greater than the time span of a DVS frame, the number of DVS frames is typically greater than the number of left or right camera frames. For example, the time span of a camera frame is 20 ms, and the time span of a DVS frame is 2 ms. For ease of explanation of the principles of this disclosure, the time interval can be set to be the same as the time span of a camera frame, but is not limited to this. When the time interval is set to the time span of one camera frame, the left and right cameras can generate left and right camera frames, respectively, during the time interval, and the DVS can generate at least one DVS frame within the same time interval. Processor 104a can also perform automatic labeling of DVS frames based on the generated left and right camera frames, which will refer to Figures 2 to 9 Provide a detailed description.

[0027] Figure 2 References to another or more embodiments according to this disclosure are shown. Figure 1 The system method flowchart shown is illustrated. Figure 2As shown, at S201, a determination can be made as to whether there are dynamic changes in the scene. If it is determined that there are no dynamic changes, the process proceeds to S202. At S202, the system can be in a standby state. If it is determined that there are dynamic changes, the method proceeds to S203. At S203, the recording device 102 is activated, meaning that the camera and DVS can operate to generate camera frames and DVS frames respectively. It should be understood that S201 to S203 can be omitted, and the method flow can begin directly from S204.

[0028] At S204, a pair of camera frames generated by a pair of cameras and at least one DVS frame generated by a DVS can be received. For example, the left camera 102b and the right camera 102c can generate a left camera frame and a right camera frame, respectively, within the time interval. At the same time, the DVS 102a can generate at least one DVS frame.

[0029] Furthermore, at S205, a disparity frame can be calculated based on the left and right camera frames, and then 3D information of the left and right camera frames can be obtained based on the calculated disparity frame. The 3D information may include 3D points, each of which represents a spatial location or 3D coordinate corresponding to each pixel within the left and right camera frames.

[0030] For example, triangulation can be used to acquire 3D information from camera frames, and the SGBM (Semi-Global Block Matching) method can be used to calculate the disparity of stereo camera frames. The concept of "disparity" will be described below. The term "disparity" can be understood as 'binocular parallax,' which refers to 'the difference in the position of an object seen by the left and right eyes in the image, caused by the horizontal distance between the eyes (parallax).' In computer vision, it refers to the pixel-level correspondence / matching pair between the left sensor / camera and the right sensor / camera, such as... Figure 3 As described in [the original text]. Reference. Figure 3 Parallax refers to the distance between two corresponding points in the left and right images of a stereo pair. Figure 3 This shows that different 3D points X, X1, X2, and X3 result in different projection positions on the left and right images, where O L Indicates the optical center of the left camera, while O R Indicates the optical center of the right camera. O L With O R The line between them is the baseline; e l This represents the intersection of the left image plane and the baseline, and e r This indicates the intersection of the right image plane and the baseline.

[0031] Taking point X as an example, by following the path from point X to O L The dashed line intersects the left image plane at point X. LThe same principle applies to the right image plane. By moving along from X to O... R The dashed line intersects the right image plane at point X. R This means that point X is projected onto point X in the left camera frame. L And point X in the right camera frame R Then the disparity of the pixels in the frame can be calculated as X. L With X R The difference between them. Therefore, by performing the above calculation on each pixel in the frame, a parallax frame can be obtained based on the left camera frame and the right camera frame.

[0032] Figure 4 This describes the relationship between disparity and depth information for each pixel. Now refer to... Figure 4 This will demonstrate how to obtain 3D information from camera frames based on parallax. Figure 4 This shows a 3D point P(Xp, Yp, Zp), the left camera frame, and the right camera frame. The 3D point is at point p. l (x l ,y l Projected onto the left camera frame at point p, and at point p r (x r ,y r The image at point O is projected onto the right camera frame. l Indicates the optical center of the left camera, while O r Indicates the optical center of the right camera. l This indicates the center of the left camera frame, while c r Indicates the center of the right camera frame. O L With O R The line between them is the baseline. T indicates the distance from O. L To O R The distance. Parameter f represents the camera's focal length, and parameter d represents the parallax, which is equal to x. l With x r The difference between them. Point P in the left camera frame and point p in the right camera frame are respectively... l and pr r The translation between them can be defined by the following formulas (1) to (2), which are also as follows Figure 4 and Figure 7 As shown in the image.

[0033]

[0034]

[0035] According to the disparity-depth equation described above, the position of each pixel in the left and right camera frames can be translated to each 3D point. Therefore, based on the disparity frame, the 3D information of the left and right camera frames can be obtained.

[0036] To promote understanding, Figure 5 The images shown from left to right are the left camera frame, the right camera frame, and the parallax frame calculated from the left and right camera frames. For parallax frames, lighter-colored pixels indicate closer distances, while darker-colored pixels indicate farther distances.

[0037] Returning to the method flowchart, at S206, object regions for automatic labeling can be determined using a deep learning model. Depending on the requirements, various deep learning models capable of (but not limited to) extracting target features can be applied to a camera frame selected from the left and right camera frames. For example, an object detection model can be applied to a single camera frame. Different models can provide different forms of output. For example, some models can output object regions representing the contours of the desired object, where the contours consist of points of the desired object. Other models can output object regions representing the area where the desired object resides, such as rectangular regions. Figure 6 An example is shown for illustrative purposes only, and not as a limitation, in which camera frames can be... Figure 5 The camera frames shown are the same. Figure 6 As shown, for example, the left camera frame is selected. For example, the object detection result on the left camera frame is shown as a rectangular result, and the corresponding result on the parallax frame is also shown as a rectangular result in the parallax frame.

[0038] Then, at S207, based on the 3D information obtained at S205 and the object region determined at S206, the 3D points of the expected object in the object region are obtained. As described with reference to S206, different models can output detection results in different forms. If the detection result is the outline of the expected object composed of points, it can be directly used to obtain the 3D points from the 3D information obtained in S205. If the detection result is the region where the expected object is located, such as a rectangular result, a clustering process needs to be performed. That is, points occupying most of the detection rectangle and points closer to the center of the detection rectangle will be considered as expected objects.

[0039] At S208, the obtained 3D points of the desired object can be reprojected onto at least one DVS frame. Since the stereo camera and DVS are in the same world coordinates and are rigidly combined, the 3D points calculated from the stereo camera frame are also the 3D points seen from the DVS frame. Therefore, a reprojection process can be performed to reproject the 3D points of the desired object onto the DVS frame. It can be understood that triangulation and reprojection can be viewed as opposite processes. The key here is using two stereo camera frames to acquire 3D points and using one camera frame and one DVS frame to acquire matching points on the DVS frame. Figure 7 The 3D point P(X) is shown. p ,Y p Zp Reprojection to the DVS frame. The parallelogram drawn with dashed lines refers to the previous... Figure 4 The right camera frame in the image. Figure 7 The parameters in Figure 4 The parameters are defined the same way. For example... Figure 7 As shown, the formula is the same as Figure 4 The formula is the same. The only difference is that in Figure 4 In the middle, these two frames are stereo camera frames, while... Figure 7 In the image, these two frames are a camera frame and a DVS frame.

[0040] At S209, the reprojected points on the DVS frame can be combined to generate new detection results on the DVS frame, i.e., an automatically labeled DVS frame. After reprojecting the 3D points of the desired object to obtain their positions on the DVS frame, the corresponding detection results can be obtained on the DVS frame. For example, if a rectangular result is needed, a rectangle is created to include all the reprojected points on the DVS frame. For example, if a contour result is needed, each reprojected point on the DVS frame is connected to its nearest neighbor among all points. Figure 8 The example shown describes how to generate automatically labeled results by using reprojected points on a DVS frame. Figure 8 An example of the expected effect of the final result is shown. Figure 8 The left image in the image is the left camera frame, and the right image is the DVS frame. The points on the right image represent the positions of the reprojected 3D points on the DVS frame. The rectangles are the result of automatic labeling on the DVS frame. Figure 8 This is just for illustration, and in reality, there will be many more reprojected points on the DVS frame.

[0041] By using the automatic labeling method described above, because the FPS (frames per second) of a DVS is much higher than that of a conventional camera, a single camera frame can be used to label many DVS frames, which further improves the efficiency of automatic labeling. Figure 10 shows a camera frame and its corresponding automatically labeled DVS frames. These DVS frames are consecutive frames.

[0042] The methods, apparatus, and systems described in this disclosure enable more efficient and accurate automatic labeling of DVS frames. The methods, apparatus, and systems of this disclosure bind a pair of cameras to a DVS and simultaneously record the same scene. Based on the combined use of the acquired camera frames and DVS frames, DVS frames can be automatically labeled while recording. Therefore, a large amount of labeled data for DVS deep learning training becomes possible. Compared to existing methods, the methods, apparatus, and systems described in this disclosure fully utilize the advantages of DVS and can perform more accurate and efficient automatic labeling.

[0043] 1. In some embodiments, a method for automatically labeling dynamic vision sensor (DVS) frames, the method comprising: receiving a pair of camera frames generated by a pair of cameras within a time interval and receiving at least one DVS frame generated by the DVS within the time interval; calculating a disparity frame based on the pair of camera frames and obtaining 3D information of the pair of camera frames based on the calculated disparity frame; determining an object region for automatic labeling using a deep learning model; obtaining 3D points based on the obtained 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generating at least one automatically labeled result on the at least one DVS by combining the reprojected points on the at least one DVS frame.

[0044] 2. The method according to Clause 1, further comprising: wherein the pair of cameras includes a left camera and a right camera; wherein the DVS is arranged to be rigidly combined with the left camera and the right camera.

[0045] 3. The method according to any one of Clauses 1 to 2, wherein determining the object region for automatic labeling further comprises: selecting a camera frame from the pair of camera frames as input to a deep learning model, and determining the object region for automatic labeling based on the output of the deep learning model.

[0046] 4. The method according to any one of clauses 1 to 3, wherein the 3D information comprises 3D points, each of which represents a spatial position / coordinate corresponding to each pixel within a camera frame.

[0047] 5. The method according to any one of clauses 1 to 4, wherein the time interval is predetermined based on the time span between two consecutive camera frames.

[0048] 6. The method according to any one of clauses 1 to 5, generating at least one DVS frame from the DVS within the time interval comprises: integrating pixel events within the time interval to generate the at least one DVS frame.

[0049] 7. The method according to any one of clauses 1 to 6, further comprising: determining whether there is a dynamic change in the scene; and if there is a dynamic change in the scene, activating the DVS and the pair of cameras.

[0050] 8. The method according to any one of Clauses 1 to 7, wherein the dynamic change includes at least one of intensity change and object movement.

[0051] 9. In some embodiments, a system for automatically labeling dynamic vision sensor (DVS) frames, the system comprising: a pair of cameras configured to generate a pair of camera frames within a time interval; a DVS configured to generate at least one DVS frame within the time interval; and a computing device including a processor and a memory unit storing instructions executable by the processor to: calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region for automatic labeling using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region, and reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatically labeled result on the at least one DVS by combining the reprojected points on the at least one DVS frame.

[0052] 10. The system according to Clause 9, wherein the pair of cameras includes a left camera and a right camera, and wherein the DVS is arranged to be rigidly combined with the left camera and the right camera.

[0053] 11. The system according to any one of clauses 9 to 10, wherein the processor is further configured to: select one camera frame from the pair of camera frames as input to a deep learning model, and determine an object region for automatic labeling based on the output of the deep learning model.

[0054] 12. The system according to any one of Clauses 9 to 11, wherein the 3D information comprises 3D points, each of which represents a spatial position / coordinate corresponding to each pixel within the camera frame.

[0055] 13. The system according to any one of clauses 9 to 12, wherein the at least one DVS frame is generated by integrating pixel events within the time interval.

[0056] 14. The system according to any one of clauses 9 to 13, wherein the time interval is predetermined based on the time span between two consecutive camera frames.

[0057] 15. The system according to any one of clauses 9 to 14, wherein the processor is further configured to: determine whether there is a dynamic change in the scene; and if there is a dynamic change in the scene, activate the DVS and the pair of cameras.

[0058] 16. The system according to any one of Clauses 9 to 15, wherein the dynamic change includes at least one of intensity change and object movement.

[0059] 17. In some embodiments, an apparatus for automatically labeling dynamic vision sensor (DVS) frames, the apparatus comprising: a computing device including a processor and a memory unit storing instructions executable by the processor to: receive a pair of camera frames generated by a pair of cameras within a time interval and receive at least one DVS frame generated by the DVS within the time interval; calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region for automatic labeling using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region, and reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatically labeled result on the at least one DVS by combining the reprojected points on the at least one DVS frame.

[0060] The descriptions of various embodiments have been provided for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable those skilled in the art to understand the disclosed embodiments.

[0061] In the foregoing, reference numerals have been used to describe the embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specifically described embodiments. Rather, any combination of the foregoing features and elements, whether or not associated with different embodiments, is contemplated as an embodiment intended for implementation and practice. Furthermore, while the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether a given embodiment achieves a particular advantage does not limit the scope of this disclosure. Therefore, the foregoing aspects, features, embodiments, and advantages are merely illustrative and should not be considered as elements or limitations of the appended claims unless expressly stated in the claims.

[0062] The aspects of this disclosure may take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, which may generally be referred to herein as “circuit,” “module,” or “system.”

[0063] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media will include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store programs used by or in conjunction with instruction execution systems, devices, or apparatuses.

[0064] The foregoing description, with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure, has described various aspects of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executable by the processor of the computer or other programmable data processing apparatus, can perform the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable processors.

[0065] Although the foregoing relates to embodiments of this disclosure, other and additional embodiments of this disclosure may be contemplated without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.

Claims

1. A method for automatically labeling dynamic vision sensor (DVS) frames, the method comprising: Receive a pair of camera frames generated by a pair of cameras within a time interval and receive at least one DVS frame generated by a DVS within the time interval, wherein the pair of cameras includes a left camera and a right camera, and wherein the DVS is arranged to be rigidly combined with the left camera and the right camera such that the DVS and the pair of cameras are in the same world coordinates; Based on the pair of camera frames, calculate the disparity frames and obtain the 3D information of the pair of camera frames based on the calculated disparity frames; Deep learning models are used to determine object regions for automatic labeling based on camera frames; 3D points are obtained based on the acquired 3D information and the determined object region, and the 3D points are reprojected onto the at least one DVS frame to generate reprojected points on the at least one DVS frame, wherein reprojecting the 3D points includes using a camera frame and a DVS frame to obtain matching points on the DVS frame. as well as At least one automatically labeled result is generated on at least one DVS by combining the reprojected points on at least one DVS frame.

2. The method of claim 1, wherein determining the object region for automatic labeling further comprises: Select one camera frame from the pair of camera frames as the input to the deep learning model, and The output of the deep learning model is used to determine the object regions for automatic labeling.

3. The method of claim 1, wherein the 3D information comprises 3D points, each of which represents a spatial position / coordinate corresponding to each pixel within a camera frame.

4. The method according to any one of claims 1 to 3, wherein the time interval is predetermined based on the time span between two consecutive camera frames.

5. The method according to any one of claims 1 to 3, wherein the at least one DVS frame is generated by integrating pixel events within the time interval.

6. The method according to any one of claims 1 to 3, further comprising: Determine if the scene is undergoing dynamic changes; as well as If the scene changes dynamically, the DVS and the pair of cameras are activated.

7. The method of claim 6, wherein the dynamic change includes at least one of intensity change and object movement.

8. A system for automatically labeling dynamic vision sensor (DVS) frames, comprising: A pair of cameras, the pair of cameras being configured to generate a pair of camera frames within a time interval, wherein the pair of cameras includes a left camera and a right camera; DVS, the DVS being configured to generate at least one DVS frame within the time interval, wherein the DVS is arranged to be rigidly combined with the left camera and the right camera such that the DVS and the pair of cameras are in the same world coordinates; as well as A computing device, comprising a processor and a memory unit storing instructions executable by the processor to: Based on the pair of camera frames, a disparity frame is calculated, and the 3D information of the pair of camera frames is obtained based on the calculated disparity frame. Deep learning models are used to determine object regions for automatic labeling based on camera frames; 3D points are obtained based on the acquired 3D information and the determined object region, and the 3D points are reprojected onto the at least one DVS frame to generate reprojected points on the at least one DVS frame, wherein reprojecting the 3D points includes using a camera frame and a DVS frame to obtain matching points on the DVS frame; and At least one automatically labeled result is generated on at least one DVS by combining the reprojected points on at least one DVS frame.

9. The system of claim 8, wherein the processor is further configured to: Choose one camera frame from the pair of camera frames as the input to the deep learning model, and The output of the deep learning model is used to determine the object regions for automatic labeling.

10. The system of claim 8, wherein the 3D information comprises 3D points, each of the 3D points representing a spatial position / coordinate corresponding to each pixel within the camera frame.

11. The system according to any one of claims 8 to 10, wherein the at least one DVS frame is generated by integrating pixel events within the time interval.

12. The system according to any one of claims 8 to 10, wherein the time interval is predetermined based on the time span between two consecutive camera frames.

13. The system according to any one of claims 8 to 10, wherein the processor is further configured to: Determine if the scene is undergoing dynamic changes; and If the scene changes dynamically, the DVS and the pair of cameras are activated.

14. The system of claim 13, wherein the dynamic change includes at least one of intensity change and object movement.

15. An apparatus for automatically labeling dynamic vision sensor (DVS) frames, comprising: A computing device, comprising a processor and a memory unit storing instructions executable by the processor to: Receive a pair of camera frames generated by a pair of cameras within a time interval and receive at least one DVS frame generated by a DVS within the time interval, wherein the pair of cameras includes a left camera and a right camera, and wherein the DVS is arranged to be rigidly combined with the left camera and the right camera such that the DVS and the pair of cameras are in the same world coordinates; Based on the pair of camera frames, calculate the disparity frames and obtain the 3D information of the pair of camera frames based on the calculated disparity frames; Deep learning models are used to determine object regions for automatic labeling based on camera frames; 3D points are obtained based on the acquired 3D information and the determined object region, and the 3D points are reprojected onto the at least one DVS frame to generate reprojected points on the at least one DVS frame, wherein reprojecting the 3D points includes using a camera frame and a DVS frame to obtain matching points on the DVS frame. and At least one automatically labeled result is generated on at least one DVS by combining the reprojected points on at least one DVS frame.