Method, apparatus and system for automated labeling

By combining dual cameras and DVS, calculating the shadow frame and obtaining 3D information, and using deep learning models for automatic annotation, the problems of low labeling accuracy and poor naturalness in the existing technology are solved, and efficient and accurate automatic annotation effect is achieved, making full use of the advantages of DVS.

JP7675834B2Active Publication Date: 2025-05-13HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023551709
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-26
Publication Date
2025-05-13
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

The existing methods of automatically labeling DVS frames have problems with low accuracy and poor naturalness, and cannot fully utilize the advantages of DVS, such as low latency and high dynamic range.

Method used

By combining dual cameras and DVS, the throwing frame between dual camera frames is calculated, 3D information is obtained, and the deep learning model is used to determine the area of ​​the object that needs to be automatically marked, 3D points are obtained, and they are projected onto the DVS frame to generate automatic labeling results.

Benefits of technology

It realizes more accurate and efficient automatic annotation of DVS frames, making full use of the low latency and high dynamic range characteristics of DVS, and the generated annotation results are more natural and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675834000002
    Figure 0007675834000002
  • Figure 0007675834000003
    Figure 0007675834000003
  • Figure 0007675834000004
    Figure 0007675834000004
Patent Text Reader

Abstract

The present disclosure provides a method, system, and apparatus for automatically labeling dynamic vision sensor (DVS) frames. The method may include generating a pair of camera frames by a pair of cameras within an interval, and generating at least one DVS frame by a DVS within the interval. The method may further calculate a disparity frame based on the pair of camera frames, and obtain 3D information of the pair of camera frames based on the calculated disparity frame. The method may determine an object region for automatic labeling using a deep learning model, and obtain 3D points based on the 3D information and the determined object region. And then, the method may reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method, an apparatus and a system for automatic labeling, and in particular to a method, an apparatus and a system for automatic labeling of DVS (Dynamic Vision Sensor) frames. [Background technology]

[0002] In recent years, DVS, a new cutting-edge sensor, has become widely known and is used in many fields, including artificial intelligence, computer vision, autonomous driving, and robotics.

[0003] Compared to conventional cameras, DVS has the advantages of low latency, no motion blur, high dynamic range, and low power consumption. In particular, DVS latency is in the order of microseconds, whereas conventional cameras have latency in the order of milliseconds. Thus, DVS is not affected by motion blur. As a result, DVS data rates are typically 40-180kB / s (usually 10mB / s for conventional cameras), requiring less bandwidth and power consumption. In addition, DVS dynamic range is about 120dB, whereas conventional cameras have a dynamic range of about 60dB. The wider dynamic range is useful in extreme light situations, such as when a vehicle enters or exits a tunnel, when other vehicles in the opposite direction turn on their high beams, or when the direction of sunlight changes.

[0004] Due to these advantages, DVS is widely used. Various efforts have been made to apply DVS to various scenarios. Among all the techniques, deep learning is the main direction that has gained popularity. When talking about deep learning, a large amount of labeled data is essential. However, the manual work of manually labeling data can be insufficient. Therefore, automatic labeling of DVS frames is needed.

[0005] Currently, there are two approaches to automatically labeling DVS frames. One is to play back conventional camera footage on a display monitor screen, record the screen using DVS, and label objects. The other is to use a deep learning model to generate labeled DVS frames directly from camera frames. However, both of these two approaches have insurmountable drawbacks. The first approach loses accuracy because it is difficult to match the DVS frames 100% accurately to the display monitor when recording. The second approach generates unnatural DVS frames. Reflectivity varies depending on the material. However, the second method treats DVS frames the same because they are generated directly from camera frames, and therefore the generated DVS frames are very unnatural. Furthermore, both approaches fall into the problem of wasting the benefits of DVS because the final output of the generated DVS frames is limited by the quality of the camera footage due to the following aspects: First, the generated DVS frame rate can only reach the frame rate of the camera at most (in the second method, we can use upscaling methods to get more frames, but it is still not promising). Second, the motion blur, residual images, and smear recorded by the camera will also be present in the generated DVS frames. This fact is unreasonable and ridiculous, since DVS is known for its low latency and lack of motion blur. Third, the high dynamic range of DVS is wasted because traditional cameras have a low dynamic range.

[0006] Therefore, there is a need to provide an improved technique for automatically labeling DVS frames while fully taking advantage of the advantages of DVS. Summary of the Invention [Means for solving the problem]

[0007] According to one or more embodiments of the present disclosure, a method for automatically labeling dynamic vision sensor (DVS) frames is provided. The method may include receiving a pair of camera frames generated by a pair of cameras within an interval and receiving at least one DVS frame generated by the DVS within the interval. The method may further calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame. The method may determine an object region for automatic labeling using a deep learning model and obtain 3D points based on the obtained 3D information and the determined object region. The method may then reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame. The method may further generate at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0008] According to one or more embodiments of the present disclosure, a system for automatically labeling dynamic vision sensor (DVS) frames is provided. The system may include a pair of cameras, a DVS, and a computing device. The pair of cameras may be configured to generate a pair of camera frames within an interval. The DVS may be configured to generate at least one DVS frame within the interval. The computing device may include a processor and a memory unit storing instructions executable by the processor to: receive the pair of camera frames and at least one DVS frame; calculate a disparity frame based on the pair of camera frames; obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region to be automatically labeled using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region; reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0009] According to one or more embodiments of the present disclosure, an apparatus for automatically labeling a dynamic vision sensor (DVS) frame is provided. The apparatus may include a computing device including a processor and a memory unit storing instructions executable by the processor to: receive a pair of camera frames and at least one DVS frame; calculate a disparity frame based on the pair of camera frames; obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region to be automatically labeled using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region; reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0010] The method, device, and system described in the present disclosure can achieve efficient and more accurate automatic labeling of DVS frames. The method, device, and system of the present disclosure can couple a pair of cameras to a DVS and record the same scene simultaneously. Recording of DVS frames can be performed based on the combined use of the acquired camera frames and DVS frames. At the same time, DVS frames can be automatically labeled. As a result, a large amount of labeled data for DVS deep learning training is possible. Compared with existing approaches, the method and system described in this disclosure can make full use of the advantages of DVS and procure more accurate and efficient automatic labeling. The present specification also provides, for example, the following items: (Item 1) 1. A method for automatically labeling dynamic vision sensor (DVS) frames, comprising: receiving a pair of camera frames generated by a pair of cameras within an interval and at least one DVS frame generated by a DVS within the interval; calculating a parallax frame based on the pair of camera frames, and obtaining 3D information of the pair of camera frames based on the calculated parallax frame; Using a deep learning model to determine object regions for automatic labeling; acquiring 3D points based on the acquired 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; The method comprising: (Item 2) the pair of cameras includes a left camera and a right camera, 2. The method of claim 1, wherein the DVS is arranged to be tightly coupled to the left camera and the right camera. (Item 3) The determining of object regions for automatic labeling further comprises: selecting a camera frame from the pair of camera frames as an input for a deep learning model; determining object regions for automatic labeling based on the output of the deep learning model; The method according to any one of items 1 to 2, comprising: (Item 4) 4. The method according to any one of items 1 to 3, wherein the 3D information comprises 3D points, each of the 3D points representing a spatial position / coordinate corresponding to each pixel in one camera frame. (Item 5) 5. The method according to any one of items 1 to 4, wherein the interval is preset based on a time span between two consecutive camera frames. (Item 6) 6. The method according to any one of items 1 to 5, wherein the at least one DVS frame is generated by integrating pixel events within the interval. (Item 7) Determining whether there is a dynamic change in the scene; activating the DVS and the pair of cameras when there is a dynamic change in the scene; The method according to any one of items 1 to 6, further comprising: (Item 8) 8. The method according to any one of items 1 to 7, wherein the dynamic change includes at least one of an intensity change and a movement of an object. (Item 9) 1. A system for automatically labeling dynamic vision sensor (DVS) frames, comprising: a pair of cameras configured to generate a pair of camera frames within an interval; a DVS configured to generate at least one DVS frame within said interval; 1. A computing device comprising: a processor; calculating a parallax frame based on the pair of camera frames, and obtaining 3D information of the pair of camera frames based on the calculated parallax frame; Using a deep learning model to determine object regions for automatic labeling; acquiring 3D points based on the acquired 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; the computing device comprising: a memory unit storing instructions executable by the processor to perform the steps of: The system comprising: (Item 10) 10. The system of claim 9, wherein the pair of cameras comprises a left camera and a right camera, and the DVS is arranged to be tightly coupled to the left camera and the right camera. (Item 11) The processor, selecting a camera frame from the pair of camera frames as an input for a deep learning model; determining object regions for automatic labeling based on the output of the deep learning model; The system according to any one of items 9 to 10, further configured to perform the following: (Item 12) 12. The system of any one of items 9 to 11, wherein the 3D information includes 3D points, each of the 3D points representing a spatial position / coordinate corresponding to a respective pixel in the camera frame. (Item 13) 13. The system of any one of items 9 to 12, wherein the at least one DVS frame is generated by integrating pixel events within the interval. (Item 14) 14. The system according to any one of items 9 to 13, wherein the interval is preset based on a time span between two consecutive camera frames. (Item 15) The processor, Determining whether there is a dynamic change in the scene; activating the DVS and the pair of cameras when there is a dynamic change in the scene; The system according to any one of items 9 to 14, further configured to perform the following: (Item 16) 16. The system of any one of items 9 to 15, wherein the dynamic change includes at least one of an intensity change and a movement of an object. (Item 17) 1. An apparatus for automatically labeling dynamic vision sensor (DVS) frames, comprising: 1. A computing device comprising: a processor; receiving a pair of camera frames generated by a pair of cameras within an interval and at least one DVS frame generated by a DVS within the interval; calculating a parallax frame based on the pair of camera frames, and obtaining 3D information of the pair of camera frames based on the calculated parallax frame; Using a deep learning model to determine object regions for automatic labeling; acquiring 3D points based on the acquired 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; and a memory unit storing instructions executable by the processor to perform the steps of: The apparatus comprising: [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 shows a schematic diagram of a system according to one or more embodiments of the present disclosure.

[0012] [Diagram 2] 1 illustrates a flowchart of a method according to one or more embodiments of the present disclosure.

[0013] [Diagram 3] 1 illustrates the parallax principle in accordance with one or more embodiments of the present disclosure.

[0014] [Figure 4]1 illustrates a relationship between disparity and depth information in accordance with one or more embodiments of the present disclosure.

[0015] [Diagram 5] 1 illustrates an example of a disparity frame computed from a left camera and a right camera in accordance with one or more embodiments of the present disclosure.

[0016] [Figure 6] 1 illustrates example object detection results and disparity frames for a left camera in accordance with one or more embodiments of the present disclosure.

[0017] [Figure 7] 1 illustrates an example of reprojection of 3D points into a DVS frame, in accordance with one or more embodiments of the present disclosure.

[0018] [Figure 8] 10 illustrates example results according to one or more embodiments of the present disclosure.

[0019] [Figure 9] 13 illustrates another example result according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] For ease of understanding, the same reference numbers have been used, whenever possible, to designate identical elements common to each figure. It is contemplated that elements disclosed in one embodiment may be beneficially utilized in other embodiments without specific mention. The drawings referred to herein should not be understood as being drawn to scale unless specifically noted otherwise. Additionally, the drawings are often simplified and details or components are omitted for clarity of presentation and explanation. The drawings and description serve to explain the principles discussed below, and like names refer to similar elements.

[0021] Examples are provided below. The descriptions of various examples are presented for illustrative purposes, but are not intended to be exhaustive or to be limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0022] In general terms, the present disclosure provides a system, apparatus, and method that can automatically label DVS frames by combining at least one pair of stereo cameras and DVS with each other. The system and method of the present disclosure can provide a reliable automatically labeled DVS frame by using a combination of camera frames and DVS frames by calculating disparity using a stereo camera and acquiring 3D information of the camera frame accordingly, acquiring an object region using a deep learning model for the camera frame, reprojecting the 3D points corresponding to the object region onto the DVS frame to generate points on the DVS frame, and combining the reprojected points on the DVS frame to generate a final detection result on the DVS frame. Based on the combined use of the acquired camera frames and DVS frames, the DVS frames can be automatically labeled at the same time as the recording of the DVS frames. As a result, a large amount of labeled data for deep learning training of the DVS is possible. Compared with existing approaches, the method and system described in the present disclosure can make full use of the advantages of the DVS and procure more accurate and efficient automatic labeling.

[0023] FIG. 1 illustrates a schematic diagram of a system for automatically labeling DVS frames according to one or more embodiments of the present disclosure. As shown in FIG. 1, the system may include a recording device 102 and a computing device 104. The recording device 102 may include, but is not limited to, at least a DVS 102a and a pair of cameras 102b, 102c, for example, a left camera 102b and a right camera 102c. Depending on implementation requirements, more cameras may be included in the recording device 102 in addition to the left camera 102b and the right camera 102c, without limitation. For simplicity, only a pair of cameras is illustrated herein. The term "camera" in this disclosure may include a stereo camera. In the recording device 102, the pair of cameras 102b, 102c and the DVS 102a may be tightly coupled / assembled / integrated with each other. It should be understood that FIG. 1 is merely for illustrating the components of the system, and is not intended to limit the positional relationship of the system components. The DVS 102a can be arranged in any relative positional relationship with the left camera 102b and the right camera 102c.

[0024] The DVS 102a can adopt an event-driven approach to capture dynamic changes in a scene and then create asynchronous pixels. Unlike a traditional camera, the DVS does not generate images but conveys pixel-level events. Dynamic changes in the real scene will cause the DVS to generate some pixel-level output (i.e., events). Thus, if there is no change, there will be no data output. The dynamic changes can include at least one of intensity changes and object motion. The event data is in the format of [x,y,t,p], where x and y are the coordinates of the pixel of the event in 2D space, t is the timestamp of the event, and p represents the polarity of the event. For example, the polarity of the event may represent a change in the brightness of the scene, such as getting brighter (positive) or darker (negative).

[0025] The computing device 104 may be any type of device capable of performing computations, including, but not limited to, a mobile device, a smart device, a laptop computer, a tablet computer, an in-vehicle navigation system, and the like. The computing device 104 may include, but is not limited to, a processor 104a. The processor 104a may be any technically feasible hardware unit configured to process data and execute software applications, including, but not limited to, a central processing unit (CPU), a microcontroller unit (MCU), an application specific integrated circuit (ASIC), a digital signal processor (DSP) chip, and the like. The computing device 104 may include, but is not limited to, a memory unit 104b for storing data, code, instructions, and the like executable by the processor. The memory unit 104b may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0026] The system for automatically labeling DVS frames can be installed in an operating environment. For example, the system can determine whether there is a dynamic change (event-based change) in the scene, and automatically activate and operate the DVS and the pair of cameras when the dynamic change in the scene is detected. The DVS and the pair of cameras can be synchronized by a synchronization timestamp. For the same scene, the left camera and the right camera can generate at least one left camera frame and at least one right camera frame, respectively, during an interval. At the same time, the DVS can generate at least one DVS frame within the same interval. Since the time span of the camera frame is larger than the time span of the DVS frame, the number of DVS frames is usually larger than the number of left or right camera frames. For example, the time span of the camera frame is 20 ms, and the time span of the DVS frame is 2 ms. For the purpose of simply explaining the principles of the present disclosure, but not for the purpose of limitation, the interval can be set to be the same as the time span of the camera frame. If the interval is set as a time span of one camera frame, the left camera and the right camera may generate a left camera frame and a right camera frame, respectively, during the interval, and the DVS may generate at least one DVS frame within the same interval. The processor 104a may further perform automatic labeling of the DVS frames based on the generated left camera frame and right camera frame, which will be described in detail with reference to FIGS. 2-9.

[0027] FIG. 2 illustrates a flowchart of a method associated with the system illustrated in FIG. 1, according to one or more further embodiments of the present disclosure. As illustrated in FIG. 2, in S201, a determination may be made as to whether there is a dynamic change in the scene. If it is determined that there is no dynamic change, the method proceeds to S202. In S202, the system may be in a standby state. If it is determined that there is a dynamic change, the method proceeds to S203. In S203, the recording device 102 is activated. This means that the camera and the DVS can operate to generate camera frames and DVS frames, respectively. It should be understood that S201-S203 may be omitted, and the method flow may immediately start with S204.

[0028] In S204, a pair of camera frames generated by the pair cameras and at least one DVS frame generated by the DVS may be received. For example, the left camera 102b and the right camera 102c may generate a left camera frame and a right camera frame, respectively, within an interval. At the same time, the DVS 102a may generate at least one DVS frame.

[0029] Further, in S205, a disparity frame may be calculated based on the left and right camera frames, and then 3D information of the left and right camera frames may be obtained based on the calculated disparity frame. The 3D information may include 3D points, each of which represents a spatial position or 3D coordinate corresponding to each pixel in the left and right camera frames.

[0030] For example, triangulation can be used to obtain the 3D information of the camera frames, and the SGBM (Semi-Global Block Matching) method can be used to calculate the disparity of the stereo camera frames. The concept of "disparity" is expressed as follows: The term "disparity" can be understood as "binocular disparity", which means "the difference in the position of the image of an object seen by the left and right eyes due to the horizontal separation (parallax) of the eyes". In computer vision, this means a pixel-level correspondence / matching pair between the left sensor / camera and the right sensor / camera, as illustrated in Figure 3. With reference to Figure 3, disparity refers to the distance between two corresponding points in the left and right images of a stereo pair. Figure 3 shows the distance between two corresponding points X, X for different 3D points X, X. 1 , X 2 and X 3 We show that,O,results in different projection positions in the left and right images. L represents the optical center of the left camera, and O R represents the optical center of the right camera. L and O R The line between is the baseline. And e l represents the intersection point between the left image plane and the baseline, and e rは Represents the intersection of the right image plane with the baseline.

[0031] For example, take point X. From X to O L By tracing the dotted line to the left image plane, the intersection point with the left image plane is X L The same principle applies to the right image plane: X to O R By tracing the dotted line to the right, the intersection point with the right image plane is X R In other words, point X is point X in the left camera frame. L and point X in the right camera frame R In this case, the disparity of the pixels in the frame is X L and X R Therefore, a disparity frame can be obtained based on the left and right camera frames by performing the above calculation for each pixel in the frame.

[0032] FIG. 4 shows the relationship between the disparity and depth information of each pixel. Now referring to FIG. 4, a method for obtaining 3D information of a camera frame based on disparity is illustrated. FIG. 4 shows a 3D point P(Xp, Yp, Zp), a left camera frame and a right camera frame. The 3D point is a point p l (x l ,y l ) into the left camera frame, and point p r (x r ,y r ) and projected onto the right camera frame. l represents the optical center of the left camera, and O r represents the optical center of the right camera. l represents the center of the left camera frame, and c r represents the center of the right camera frame. L and O R The line between is the baseline. L From O R The parameter f represents the focal length of the camera, and the parameter d represents the distance from x l and x r The parallax is equal to the difference between the points P and p in the left and right camera frames, respectively. l and p r The translation between can be defined by the following equations (1) and (2), which are also shown in Figures 4 and 7.

number

[0033] According to the above relationship between disparity and depth, the position of each pixel in the left and right camera frames can be converted into a 3D point. Therefore, the 3D information of the left and right camera frames can be obtained based on the disparity frame.

[0034] For ease of understanding, Fig. 5 shows, from left to right, the left camera frame, the right camera frame, and the disparity frame calculated from the left and right camera frames, respectively. In the disparity frame, lighter colored pixels mean closer distances, and darker colored pixels mean farther distances.

[0035] Returning to the method flow chart, in S206, a deep learning model may be used to determine an object region for automatic labeling. According to various requirements, various deep learning models that can extract features of an object may be applied to one camera frame selected from the left and right camera frames, without being limited thereto. For example, an object detection model may be applied to one camera frame. Different models may provide different output formats. For example, in one model, an object region representing the contour of a desired object may be output, and the contour is composed of points of the desired object. For example, in another model, an object region representing an area, such as a rectangular area, in which the desired object is located may be output. FIG. 6 shows an example for illustrative purposes only, but is not limited thereto, and the camera frame may be the same as the camera frame shown in FIG. 5. As shown in FIG. 6, for example, the left camera frame is selected. For example, the object detection result on the left camera frame is shown as a rectangular result, and the corresponding result on the disparity frame is also shown as a rectangular result in the disparity frame.

[0036] Next, in S207, 3D points of the desired object in the object region are obtained based on the 3D information obtained in S205 and the object region determined in S206. As described with respect to S206, different models may output different formats of the detection results. If the detection result is a contour of the desired object consisting of points, it can be used as is to obtain 3D points from the 3D information obtained in S205. If the detection result is a region where the desired object is located, such as a rectangular result, a clustering process needs to be performed. That is, the points that occupy the majority of the detection rectangle and the points that are closer to the center of the detection rectangle will be considered as the desired object.

[0037] In S208, the acquired 3D points of the desired object may be reprojected towards at least one DVS frame. Since the stereo camera and the DVS are in the same world coordinates and they are tightly coupled, the 3D points calculated from the stereo camera frames are also 3D points seen from the DVS frames. Therefore, a reprojection process may be performed to reproject the 3D points of the desired object towards the DVS frames. It can be understood that triangulation and reprojection can be considered as inverse processes to each other. The key here is that two stereo camera frames are used to acquire 3D points, and one camera frame and one DVS frame are used to acquire the matching points on the DVS frames. FIG. 7 shows the projection of a 3D point P(X p , Y p , Z p ) reprojection. The parallelogram drawn with dashed lines refers to the right camera frame in the previous Fig. 4. The parameters in Fig. 7 have the same definitions as those in Fig. 4. As shown in Fig. 7, the equations are the same as those in Fig. 4. The only difference is that in Fig. 4, the two frames are stereo camera frames, while in Fig. 7, the two frames are one camera frame and one DVS frame.

[0038] In S209, the reprojected points on the DVS frame can be combined to generate a new detection result on the DVS frame, i.e., generate an auto-labeled DVS frame. After reprojecting the 3D points of the desired object to obtain the position of the point on the DVS frame, it is possible to obtain the corresponding detection result on the DVS frame. For example, if the result requires a rectangular result, a rectangle containing all the reprojected points on the DVS frame is created. For example, if the result requires a contour result, the reprojected points on the DVS frame are respectively connected to the points that are closest to the point among all the points. As illustrated by the example shown in FIG. 8, the auto-labeled result will be generated by using the reprojected points on the DVS frame. FIG. 8 shows an example of the expected effect of the final result. The left image of FIG. 8 is the left camera frame, and the right image is the DVS frame. The points on the right image represent the positions of the 3D points reprojected on the DVS frame. The rectangle is the auto-labeled result on the DVS frame. FIG. 8 is merely for illustration, and in a real case, there should be more reprojected points on the DVS frame.

[0039] By using the above automatic labeling method, since the FPS (frames per second) of DVS is much higher than that of traditional cameras, one camera frame can be used to label many DVS frames, thus further improving the efficiency of automatic labeling. Figure 10 shows one camera frame and its corresponding automatically labeled DVS frames. These DVS frames are consecutive frames.

[0040] The method, device, and system described in this disclosure can realize more efficient and accurate automatic labeling of DVS frames. The method, device, and system of this disclosure combine a pair of cameras with a DVS to record the same scene simultaneously. Based on the combined use of the acquired camera frames and DVS frames, the DVS frames can be automatically labeled at the same time as the DVS frames are recorded. As a result, a large amount of labeled data for DVS deep learning training is possible. Compared with existing approaches, the method and system described in this disclosure can make full use of the advantages of the DVS and perform more accurate and efficient automatic labeling.

[0041] 1. In some embodiments, a method for automatically labeling dynamic vision sensor (DVS) frames includes receiving a pair of camera frames generated by a pair of cameras within an interval and receiving at least one DVS frame generated by a DVS within the interval; calculating a disparity frame based on the pair of camera frames and acquiring 3D information of the pair of camera frames based on the calculated disparity frame; determining an object region to be automatically labeled using a deep learning model; acquiring 3D points based on the acquired 3D information and the determined object region, reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0042] 2. The method of claim 1, wherein the pair of cameras includes a left camera and a right camera, and the DVS is arranged to be tightly coupled to the left camera and the right camera.

[0043] 3. The method of any one of clauses 1 to 2, wherein determining an object region to be automatically labeled further comprises selecting one camera frame from the pair of camera frames as an input for a deep learning model, and determining an object region to be automatically labeled based on the output of the deep learning model.

[0044] 4. A method according to any one of clauses 1 to 3, wherein the 3D information comprises 3D points, each of the 3D points representing a spatial position / coordinate corresponding to each pixel in one camera frame.

[0045] 5. The method according to any one of clauses 1 to 4, wherein the interval is preset based on a time span between two consecutive camera frames.

[0046] 6. The method of any one of clauses 1 to 5, wherein generating at least one DVS frame by the DVS within the interval includes integrating pixel events within the interval to generate the at least one DVS frame.

[0047] 7. The method of any one of clauses 1 to 6, further comprising determining whether there is a dynamic change in a scene, and activating the DVS and the pair of cameras if there is a dynamic change in the scene.

[0048] 8. The method of any one of clauses 1 to 7, wherein the dynamic changes include at least one of intensity changes and object movement.

[0049] 9. In some embodiments, a system for automatically labeling dynamic vision sensor (DVS) frames, comprising: a pair of cameras configured to generate a pair of camera frames within an interval; a DVS configured to generate at least one DVS frame within the interval; and a computing device comprising a processor and a memory unit storing instructions executable by the processor to: calculate a disparity frame based on the pair of camera frames; obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region to be automatically labeled using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region, reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0050] 10. The system described in clause 9, wherein the pair of cameras comprises a left camera and a right camera, and the DVS is arranged to tightly couple the left camera and the right camera.

[0051] 11. The system of any one of clauses 9 to 10, wherein the processor is further configured to select one camera frame from the pair of camera frames as an input for a deep learning model, and determine an object region for automatic labeling based on the output of the deep learning model.

[0052] 12. A system described in any one of clauses 9 to 11, wherein the 3D information includes 3D points, each of the 3D points representing a spatial position / coordinate corresponding to a respective pixel in the camera frame.

[0053] 13. The system of any one of clauses 9 to 12, wherein the at least one DVS frame is generated by integrating pixel events within the interval.

[0054] 14. A system described in any one of clauses 9 to 13, wherein the interval is preset based on a time span between two consecutive camera frames.

[0055] 15. The system of any one of clauses 9 to 14, wherein the processor is further configured to determine whether there is a dynamic change in the scene, and if there is a dynamic change in the scene, activate the DVS and the pair of cameras.

[0056] 16. A system described in any one of clauses 9 to 15, wherein the dynamic changes include at least one of intensity changes and object movement.

[0057] 17. In some embodiments, an apparatus for automatically labeling dynamic vision sensor (DVS) frames, the apparatus comprising a computing device comprising a processor and a memory unit storing instructions executable by the processor to: receive a pair of camera frames generated by a pair of cameras within an interval and receive at least one DVS frame generated by a DVS within the interval; calculate a disparity frame based on the pair of camera frames and obtain 3D information of the pair of camera frames based on the calculated disparity frame; determine an object region to be automatically labeled using a deep learning model; obtain 3D points based on the obtained 3D information and the determined object region and reproject the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; and generate at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame.

[0058] The description of various embodiments is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications or technical improvements to the technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0059] In the above, reference is made to the embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specific embodiments described. Instead, any combination of the above features and elements, whether related to different embodiments or not, is contemplated for implementing and practicing the contemplated embodiments. Furthermore, the embodiments disclosed herein may achieve advantages over other possible solutions or the prior art, but whether or not a particular advantage is achieved by a given embodiment does not limit the scope of the present disclosure. Thus, the above aspects, features, embodiments, and advantages are merely exemplary and are not considered elements or limitations of the appended claims unless expressly recited in the claim(s).

[0060] Aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "module" or "system."

[0061] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer readable storage media would include an electrical connection with one or more communication lines, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0062] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present methods. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the instructions are executed by the processor of the computer or other programmable data processing apparatus to generate a machine capable of implementing the functions / acts set forth in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable processor.

[0063] While the forgoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which scope is determined by the following claims.

Claims

1. 1. A method for automatically labeling dynamic vision sensor (DVS) frames, the method being executed by a processor, the method comprising: the processor receiving a pair of camera frames generated by a pair of cameras within an interval, and the processor receiving at least one DVS frame generated by a DVS within the interval; The processor calculates a disparity frame based on the pair of camera frames, and the processor obtains 3D information of the pair of camera frames based on the calculated disparity frame; the processor using a deep learning model to determine object regions for automatic labeling; the processor acquiring 3D points based on the acquired 3D information and the determined object region, the processor reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; A method comprising:

2. the pair of cameras includes a left camera and a right camera, The method of claim 1 , wherein the DVS is arranged to be tightly coupled to the left and right cameras.

3. The determining of object regions for automatic labeling further comprises: the processor selecting a camera frame from the pair of camera frames as an input for a deep learning model; determining object regions for automatic labeling based on an output of the deep learning model; The method according to any one of claims 1 to 2, comprising:

4. The method of any one of claims 1 to 3, wherein the 3D information comprises 3D points, each of the 3D points representing a spatial position / coordinate corresponding to a respective pixel in one camera frame.

5. The method according to any one of claims 1 to 4, wherein the interval is preset based on a time span between two consecutive camera frames.

6. The method of any one of claims 1 to 5, wherein the at least one DVS frame is generated by aggregating pixel events within the interval.

7. The processor determining whether there is a dynamic change in the scene; the processor actuating the DVS and the pair of cameras when there is a dynamic change in the scene; The method of any one of claims 1 to 6, further comprising:

8. The method of claim 7 , wherein the dynamic changes include at least one of an intensity change and an object movement.

9. 1. A system for automatically labeling dynamic vision sensor (DVS) frames, comprising: a pair of cameras configured to generate a pair of camera frames within an interval; a DVS configured to generate at least one DVS frame within the interval; 1. A computing device comprising: a processor; calculating a parallax frame based on the pair of camera frames, and obtaining 3D information of the pair of camera frames based on the calculated parallax frame; Using a deep learning model to determine object regions for automatic labeling; acquiring 3D points based on the acquired 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; a memory unit storing instructions executable by the processor to perform the steps of: A system comprising:

10. 10. The system of claim 9, wherein the pair of cameras comprises a left camera and a right camera, and the DVS is arranged to be tightly coupled to the left camera and the right camera.

11. The processor, selecting a camera frame from the pair of camera frames as an input for a deep learning model; determining object regions for automatic labeling based on an output of the deep learning model; The system according to any one of claims 9 to 10, further configured to:

12. The system of any one of claims 9 to 11, wherein the 3D information comprises 3D points, each of the 3D points representing a spatial position / coordinate corresponding to a respective pixel in the camera frame.

13. The system of any one of claims 9 to 12, wherein the at least one DVS frame is generated by integrating pixel events within the interval.

14. The system according to any one of claims 9 to 13, wherein the interval is preset based on a time span between two consecutive camera frames.

15. The processor, Determining whether there is a dynamic change in the scene; activating the DVS and the pair of cameras when there is a dynamic change in the scene; The system according to any one of claims 9 to 14, further configured to:

16. The system of claim 15 , wherein the dynamic changes include at least one of an intensity change and an object movement.

17. 1. An apparatus for automatically labeling dynamic vision sensor (DVS) frames, comprising:

1. A computing device comprising: a processor; receiving a pair of camera frames generated by a pair of cameras within an interval and receiving at least one DVS frame generated by a DVS within the interval; calculating a parallax frame based on the pair of camera frames, and obtaining 3D information of the pair of camera frames based on the calculated parallax frame; Using a deep learning model to determine object regions for automatic labeling; acquiring 3D points based on the acquired 3D information and the determined object region, and reprojecting the 3D points onto the at least one DVS frame to generate reprojected points on the at least one DVS frame; generating at least one automatic labeling result on the at least one DVS frame by combining the reprojected points on the at least one DVS frame; and a memory unit storing instructions executable by the processor to perform the steps of: An apparatus comprising: