Method, device, and computer-readable storage medium having instructions for processing sensor data

By fusing the camera image with 3D measurement points into the data of the virtual sensor, and applying optical flow synchronization and accumulated data fusion, the problems of systematic error and dynamic model error in traditional object tracking systems are solved, the robustness and accuracy of object tracking are improved, and highly automated and autonomous driving are supported.

CN111937036BActive Publication Date: 2025-05-30VOLKSWAGEN AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980026426.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-18
Filing Date
2019-03-27
Publication Date
2025-05-30
Estimated Expiration
2039-03-27

AI Technical Summary

Technical Problem

In traditional object tracking systems, systematic measurement errors of multiple sensors and prediction errors of dynamic models lead to ambiguity in object tracking, which may lead to adverse consequences such as error associations and emergency braking.

Method used

The concept of virtual sensors is introduced, the camera image is fused with 3D measurement points, the data is synchronized through optical flow, and the accumulated sensor data fusion is applied in object tracking to reduce the dependence of systematic errors and dynamic models.

Benefits of technology

Through the data fusion of virtual sensors, the robustness and accuracy of object tracking are improved, the possibility of error association is reduced, and highly automated and autonomous driving functions are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111937036B_ABST
    Figure CN111937036B_ABST
Patent Text Reader

Abstract

Method, device, and computer-readable storage medium having instructions for processing sensor data. In a first step, a camera image is detected (20) by a camera. In addition, 3D measurement points are detected (21) by at least one 3D sensor. Optionally, at least one of these camera images is segmented (22). Then, these camera images are fused (23) with these 3D measurement points by a data fusion unit into virtual sensor data. Finally, the resulting data is output (24) for further processing.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to a method, a device, and a computer-readable storage medium having instructions for processing sensor data. The present invention also relates to a motor vehicle in which the method according to the invention or the device according to the invention is used. Background Art

[0002] Today, modern motor vehicles already have a large number of sensors for various Level-2 (semi-automated) assistance systems.

[0003] For example, DE 102011013776A1 describes a method for detecting or tracking objects in the vehicle's surroundings. These objects are detected from the optical flow based on the determination of corresponding image points in at least two images. Here, the distance of these objects is determined from the optical flow based on the determination of the corresponding image points in the at least two images. Those objects that are within the detection area of the distance sensor and whose distance determined from the optical flow is less than the distance value determined by means of the distance sensor are not considered.

[0004] DE 102017100199 A1 describes a method for detecting pedestrians. In a first step, an image of the area near the vehicle is received. The image is processed using a first neural network to determine the position where a pedestrian might be within the image. Then, the determined position of the image is processed using a second neural network to determine whether a pedestrian is present. If a pedestrian is present, a notification is sent to the driver assistance system or the automated driving system. These neural networks may include deep convolutional networks.

[0005] For Level 3 and higher systems (highly automated and autonomous systems), the number of sensors to be installed will further increase. In this case, for safety reasons, there will be redundant detection areas that are covered by multiple sensors with different measurement principles.

[0006] In this case, camera sensors, radar scanners, and lidar scanners play a crucial role. In particular, it should be assumed that there is at least one camera sensor and at least one 3D sensor in the critical area, which cover these critical areas. Examples of 3D sensors are radar scanners or lidar scanners with elevation measurement.

[0007] In traditional systems, there is so-called object tracking, which establishes object hypotheses that are confirmed and updated by new sensor measurements. Traditionally, so-called "predictive correction filters", such as Kalman filters, are used here. If new measurements arrive, all objects are predicted to the measurement time point of the new measurement with the help of a dynamic model. Immediately afterwards, an attempt is made to assign the measurement to an existing object. If this is successful, the track is updated. If this fails, a new object hypothesis, i.e. a new track, is established.

[0008] In this context, DE 102011119767 A1 describes a method for merging camera data and distance sensor data in order to track at least one external object in a vehicle using a sensor subsystem having a camera and a distance sensor and using an on-board computer. Based on the input received by the vehicle's sensor subsystem, the on-board computer determines that data corresponding to a new object is available. The on-board computer records the data of the new object and estimates the expected position and expected appearance of the object according to a prediction algorithm in order to generate a predicted trajectory for the object. The on-board computer also analyzes the movement of the object, including comparing the predicted trajectory with existing trajectories assigned to the object and stored in the database of the on-board computer.

[0009] In the case of traditional object tracking, especially in the association step, a series of challenges have to be considered in order to avoid ambiguities. For example, the dynamic state cannot always be estimated well: depending on the measurement and state of the track, the Cartesian velocity vector is often not known. Acceleration can only be estimated by long-term observation. This can lead to significant errors in the prediction step. An object may also behave contrary to the dynamic model, for example due to sudden braking. Such abnormal behavior can also lead to prediction errors.

[0010] In addition, there are often systematic measurement errors between different sensors: for example, a laser scanner particularly well perceives strongly reflecting surfaces, such as license plates or cat's eyes, while a black-painted vehicle cannot be detected. A radar sensor, on the other hand, well perceives metal objects with a large radar cross-section, such as taillights, bent sheets, etc. In this case, different points of an object become suitable due to these sensors, which may be far apart from each other if necessary, but should be assigned to the same object. In addition, some sensors, such as radar sensors, have a relatively low selectivity, which exacerbates the ambiguity problem here.

[0011] The incorrect handling of ambiguities can lead to error associations, in which case object tracks are associated with error-prone measurement data and updated. This can have unpleasant consequences. For example, the lateral speed may be incorrectly assigned to a roadside building. As a result, the roadside building appears dynamic and enters the driving envelope (Fahrschlauch). This can cause emergency braking due to "phantom objects". Similarly, it can happen that a roadside building, such as a warning post (Poller) measured by means of a laser scanner, is assigned to a nearby dynamic object, such as a vehicle passing by the warning post just then. This prevents the warning post itself from being recognized in a timely manner, which may lead to a collision with the roadside building. Summary of the Invention

[0012] The object of the present invention is to elucidate solutions for processing sensor data that allow reducing the problems occurring during object tracking.

[0013] This object is solved by the method, computer-readable storage medium with instructions, and device disclosed by the present invention.

[0014] According to a first aspect of the present invention, a method for processing sensor data comprises the following steps:

[0015] - Detecting a camera image by means of a camera;

[0016] - Detecting 3D measurement points by means of at least one 3D sensor; and

[0017] - Fusing these camera images with these 3D measurement points into virtual sensor data.

[0018] According to another aspect of the present invention, a computer-readable storage medium contains instructions which, when implemented by a computer, cause the computer to perform the following steps for processing sensor data:

[0019] - Detecting a camera image by means of a camera;

[0020] - Detecting 3D measurement points by means of at least one 3D sensor; and

[0021] - Fusing these camera images with these 3D measurement points into virtual sensor data.

[0022] The term computer should be understood broadly here. In particular, the computer also includes control devices and other processor-based data processing devices.

[0023] According to another aspect of the present invention, a device for processing sensor data has:

[0024] - an input terminal for receiving a camera image of a camera and 3D measurement points of a 3D sensor; and - a data fusion unit for fusing the camera image with the 3D measurement points into data of a virtual sensor.

[0025] The concept of a virtual sensor is introduced within the framework of a preprocessing step for analyzing sensor data, in particular within the framework of object tracking. The virtual sensor fuses the measurement data of the camera and the 3D sensor at the level of an earlier measurement point and thus abstracts the individual sensors. The data obtained by the virtual sensor can be clustered into high-quality object hypotheses during subsequent object tracking, since these data contain a wide range of information in order to separate the different classes. The solution according to the invention prevents object hypotheses of different sensors with systematic errors from being merged over time into a common model, whereby association errors are prone to occur. This makes it possible to achieve a robust surroundings perception that allows highly automated and autonomous driving functions.

[0026] According to one aspect of the present invention, fusing image data with 3D measurement points into data of a virtual sensor includes:

[0027] - calculating an optical flow from at least one first camera image and at least one second camera image; and - determining, based on the optical flow, a pixel in at least one of the camera images which is to be assigned to one of the 3D measuring points at a point in time of measurement.

[0028] The 3D measuring points are synchronized with the camera images by means of the calculated optical flow. This is particularly advantageous because the optical flow automatically and correctly takes into account external and intrinsic motions. Dynamic models that could cause errors are not stored.

[0029] According to one aspect of the invention, the determination of the pixel in at least one of the camera images which is to be assigned to one of the 3D measuring points at the time of the measurement comprises:

[0030] - camera images that are temporally close to the measurement time points of the 3D sensor are converted based on the optical flow; and - the 3D measurement points are projected into the converted camera images.

[0031] With the help of optical flow, the entire camera image can be converted to the measurement time of the 3D sensor. The 3D measurement points can then be projected from the depth measurement sensor into the camera image. For this purpose, for example, the pixels can be considered as infinitely long light beams that intersect the 3D measurement points.

[0032] According to one aspect of the present invention, the determination of the pixels in at least one of these camera images that should be assigned to one of these 3D measurement points at the time point to be measured includes:

[0033] - Determining, based on the optical flow and a search method, the pixels in the camera image that should be assigned to these 3D measurement points at the time point to be measured; and

[0034] - Projecting these 3D measurement points onto the positions thus determined in the camera image.

[0035] It is possible to determine, by means of the optical flow and a search method, the pixels in the camera image that should be assigned to these 3D measurement points at the time point to be measured. This is particularly useful in the case of a lidar system, where, due to the scanning action, each measurement point has its own time stamp. The solution based on the search method is computationally much less expensive compared to transforming the entire image for each measurement point.

[0036] According to one aspect of the present invention, for the pixels of these camera images, the time until collision is determined from the optical flow. Then, based on the time until collision, the optical flow, and the ranging for the 3D measurement points, the Cartesian velocity vector of the 3D measurement point can be calculated. The Cartesian velocity vector can be used, for example, to distinguish overlapping objects of the same class. To make such a distinction, the sensors used so far have to track the objects over time with the help of dynamic and correlation models, which is relatively error-prone.

[0037] According to one aspect of the present invention, the time until collision is determined from the 3D measurement based on the radial relative velocity and the ranging. Then, based on the time until collision and the optical flow, the Cartesian velocity vector of the 3D measurement point can be calculated. This solution has the advantage that if the radial relative velocity comes from a radar sensor, for example, the measurement of the time until collision is particularly accurate. In addition, the object motion can be observed quite accurately not only horizontally but also vertically (optical flow) in the image. Thus, the resulting velocity vector is generally more accurate than the case where the time until collision is estimated only based on the image.

[0038] According to one aspect of the present invention, the 3D measurement points are made to extend the attributes from at least one of these camera images. These attributes can be, for example, the (average) optical flow or the position in the image space of one or more respective pixels from the camera image. It is also possible to add the velocity vector, the Doppler velocity, the reflectivity, or the radar cross-section, or the confidence. These additional attributes allow for more robust object tracking or also better segmentation.

[0039] According to one aspect of the present invention, the camera images at measurement time points close to 3D measurement are segmented. Optionally, before this segmentation, the measurement points of the 3D sensor are accurately projected into the image by means of optical flow and the measurement attributes of these measurement points are stored in other dimensions. This enables cross-sensor segmentation.

[0040] Herein, this segmentation is preferably performed by a neural network. Through this segmentation, on the one hand, association errors are avoided, and on the other hand, the ambiguity between two categories can be resolved. The class information or identifier obtained from this segmentation is preferably also added as an attribute to these 3D measurement points.

[0041] According to one aspect of the present invention, an object tracking algorithm is applied to the data of the virtual sensor. This algorithm preferably performs accumulative sensor data fusion. This accumulative sensor data fusion enables filtering of data over time and thus enables reliable object tracking.

[0042] Particularly advantageously, the method according to the present invention or the device according to the present invention is used in a vehicle, especially a motor vehicle. Description of the Drawings

[0043] Other features of the present invention are visible from the following description in conjunction with the drawings.

[0044] Figure 1 Schematically shows the process of traditional object tracking;

[0045] Figure 2 Schematically shows the method for processing sensor data;

[0046] Figure 3 Schematically shows the fusion of camera images and 3D measurement points into the data of the virtual sensor;

[0047] Figure 4 Shows a first embodiment of the device for processing sensor data;

[0048] Figure 5 Shows a second embodiment of the device for processing sensor data;

[0049] Figure 6 Schematically shows a motor vehicle in which the solution according to the present invention is implemented;

[0050] Figure 7 Schematically shows the concept of the virtual sensor; and

[0051] Figure 8 Schematically shows the concept of the virtual sensor with a classifier. Detailed Description

[0052] In order to better understand the principle of the present invention, the embodiments of the present invention are described in more detail below according to the accompanying drawings. It is easy to understand that the present invention is not limited to these embodiments and the described features can also be combined or modified without departing from the scope of protection of the present invention.

[0053] Figure 1 The process of conventional object tracking is schematically shown. The input variables for object tracking are sensor data E and track states transformed into the measurement space. In a first step 10, an attempt is made to associate the measurement with the track. Subsequently, a check 11 is made to see whether the association was successful. If so, the corresponding track is updated 12. However, if the association fails, a new track is initialized 13. This procedure is repeated for all measurements. It is also checked 14 for all tracks whether the corresponding track has not been updated for a long time. Tracks to which this is answered in the affirmative are deleted 15. The output variable of object tracking is an object list A. The associated track is predicted 16 to the next measurement time point, and the resulting track state is transformed 17 again into the measurement space for the next continuation of object tracking.

[0054] Figure 2 A method for processing sensor data is schematically shown. In a first step, a camera image is detected 20 by a camera. In addition, 3D measurement points are detected 21 by at least one 3D sensor. Optionally, at least one of the camera images can be segmented 22, for example by means of a neural network. Subsequently, the camera images are fused 23 with the 3D measurement points by means of a data fusion unit to form data of a virtual sensor. In this case, an optical flow is determined, which is used to synchronize the image measurement points and the 3D measurement points. In this case, the 3D measurement points can be extended with properties from at least one of the camera images. Finally, the resulting data are output 24 for further processing. In the case of this further processing, for example, an object tracking algorithm can be applied to the data of the virtual sensor. The algorithm can, for example, perform an accumulation-type sensor data fusion. The data of the virtual sensor can also be segmented. In this case, the segmentation can again be performed by means of a neural network.

[0055] Figure 3Schematically shows the fusion of camera images with 3D measurement points into data of a virtual sensor. In a first step, 30 optical flow is calculated based on at least one first camera image and at least one second camera image. Optionally, for the pixels of these camera images, the time until collision can be determined from the optical flow. In addition, based on the time until collision, the optical flow, and the ranging for the 3D measurement points, the velocity vector of the 3D measurement points can be calculated. Alternatively, the time until collision can be determined from the 3D measurement based on the radial relative velocity and the ranging. Then, based on the time until collision and the optical flow, the Cartesian velocity vector of the 3D measurement points can be calculated. Finally, based on the optical flow, the pixels in at least one of these camera images that are assigned to one of these 3D measurement points are determined. For this purpose, first, the camera images that are temporally close to the measurement time points of the 3D sensor can be transformed based on the optical flow. Then, these 3D measurement points can be projected into the transformed camera images.

[0056] Figure 4 Shows a simplified schematic diagram of a first embodiment of a device 40 for processing sensor data. The device 40 has an input end 41 through which the camera images I1, I2 of a camera 61 and the 3D measurement points MP of at least one 3D sensor 62, 64 can be received. Optionally, the device 40 also has a splitter 42 for segmenting at least one camera image or at least one camera image I1, I2 enriched with other measurements, for example, by means of a neural network. Through a data fusion unit 43, the camera images I1, I2 are fused with the 3D measurement points MP into data VS of a virtual sensor. Here, the 3D measurement points MP can be extended with the attributes from at least one of these camera images I1, I2. For this fusion, the data fusion unit 43 can calculate the optical flow based on at least one first camera image I1 and at least one second camera image I2 in a first step. Optionally, for the pixels of these camera images I1, I2, the time until collision can be determined from the optical flow. Then, based on the time until collision, the optical flow, and the ranging for a given 3D measurement point MP, the velocity vector of the 3D measurement point MP can be calculated. Alternatively, the time until collision can be determined from the 3D measurement based on the radial relative velocity and the ranging. Then, based on the time until collision and the optical flow, the Cartesian velocity vector of the 3D measurement point MP can be calculated. Finally, the data fusion unit 43 determines the pixels in at least one of these camera images I1, I2 that are assigned to one of these 3D measurement points MP based on the optical flow. For this purpose, first, the camera images I1, I2 that are temporally close to the measurement time points MP of the 3D sensors 62, 64 can be transformed based on the optical flow. Then, these 3D measurement points MP can be projected into the transformed camera images.

[0057] Likewise, the optional object tracker 44 can perform object tracking based on the data VS of the virtual sensor. The object tracker 44 can, for example, perform accumulative sensor data fusion. However, this accumulative sensor data fusion can equally well be carried out outside the device 40. Via the output 47 of the device 40, the data VS of the virtual sensor or the results of object tracking or segmentation are output for further processing.

[0058] The splitter 42, the data fusion unit 43, and the object tracker 44 can be controlled by the control unit 45. If necessary, the settings of the splitter 42, the data fusion unit 43, the object tracker 44, or the control unit 45 can be changed via the user interface 48. The data accumulated in the device 40 can be stored in the memory 46 of the device 40 when needed, for example, stored in the memory 26 of the device 20 for later analysis or for use by the components of the device 40. The splitter 42, the data fusion unit 43, the object tracker 44, and the control unit 45 can be implemented as dedicated hardware, for example, implemented as an integrated circuit. However, they can of course also be partially or fully combined or implemented as software running on a suitable processor, such as a GPU or a CPU. The input 41 and the output 47 can be implemented as separate interfaces or can be implemented as a combined bidirectional interface.

[0059] Figure 5 A simplified schematic diagram of a second embodiment of a device 50 for processing sensor data is shown. The device 50 has a processor 52 and a memory 51. For example, the device 50 is a computer or a control device. Instructions are stored in the memory 51, which, when implemented by the processor 52, cause the device 50 to perform the steps of one of the described methods. Thus, the instructions stored in the memory 51 are manifested as a program executable by the processor 52, which program implements the method according to the invention. The device 50 has an input 53 for receiving information, in particular sensor data. The data generated by the processor 52 are provided via the output 54. These data can also be stored in the memory 51. The input 53 and the output 54 can be combined into a bidirectional interface.

[0060] The processor 52 can include one or more processor units, such as a microprocessor, a digital signal processor, or a combination thereof.

[0061] The memories 46, 51 of the described embodiments can have not only a volatile storage area but also a non-volatile storage area, and can include various storage devices and storage media, such as hard disks, optical storage media, or semiconductor memories.

[0062] Figure 6A motor vehicle 50 is schematically shown in which the solution according to the invention is implemented. The motor vehicle 60 has a camera 61 for detecting camera images and a radar sensor 62 for detecting 3D measuring points. The motor vehicle 60 also has a device 40 for processing sensor data, by means of which the camera image and the 3D measuring points are fused into data of a virtual sensor. Other components of the motor vehicle 60 are an ultrasonic sensor 63 and a Lidar system 64 for detecting the surrounding environment, a data transmission unit 65 and a series of auxiliary systems 66, one of which is shown as an example. These auxiliary systems can use the data provided by the device 20, for example for object tracking. With the help of the data transmission unit 65, a connection with a service provider can be established, for example for calling up navigation data. In order to store data, there is a memory 67. Data exchange between different components of the motor vehicle 50 is achieved via a network 68.

[0063] Subsequently, it should be based on Figures 7 to 8 To describe the preferred embodiments of the present invention.

[0064] Instead of fusing the measurement data of different sensors with systematic errors over time in a common model, where correlation errors are prone to occur, the concept of a virtual sensor is introduced, which fuses the measurement data of the camera and 3D sensor at the level of earlier measurement points and thus abstracts the individual sensors.

[0065] Figure 7 The concept of a virtual sensor is schematically shown as the basis for cumulative sensor data fusion. The input variables for sensor fusion by the data fusion unit 43 are the 3D measurement points of the 3D sensor (radar 62) and the camera images of the camera 61. The camera 61 may have already processed these camera images in order to, for example, determine the optical flow, classify the image points within the framework of the segmentation, or extract points from these camera images with the help of an SfM algorithm (SfM: Structure from Motion). However, this processing of these camera images can also be performed by the data fusion unit 43. In addition, the camera 61 can transmit information about the camera position. Other possible data sources are ultrasonic sensors 63 or lidar systems 64. The data are fused in a very short period of time by the data fusion unit 43. Subsequently, the 3D points from the data fusion unit 42 are transferred to the cumulative sensor data fusion 44, which enables filtering over time.

[0066] The main challenge in data fusion is that sensors 61, 62 make measurements at different time points. Therefore, accurate synchronization of the data from different sensors 61, 62 is required. For the synchronization of sensors 61, 62, preferably the optical flow determined from these camera images is used. Subsequently, the basis of synchronization should first be elaborated. How to handle the various coordinate systems that occur is described in more detail below.

[0067] The following 3D measurement points are given, which are recorded at time point t. Now, at least two camera images are used, for example, the camera images before and after the measurement time point t, in order to first calculate the optical flow o.

[0068] Preferably, the following image is used, the shooting time point t of which is closest to the measurement time point of the 3D sensor. The time difference between the shooting time point of this image and the measurement is Δt. The optical flow o is measured in the image space (polar space).

[0069] Pixels with position p and optical flow o are now predicted as follows:

[0070] p′ = p + o·△t (1)

[0071] In consideration of the time to collision (TTC) (the determination of this time to collision is described in more detail below), the formula can also be refined:

[0072]

[0073] Using this approach, the entire image can be converted to the measurement time point t of the 3D sensor. Subsequently, these 3D measurement points can be easily projected from the depth measurement sensor into this image. For this purpose, these pixels can be regarded as infinitely long light beams intersecting the 3D measurement points.

[0074] However, in the case of a lidar system, due to the scanning effect, each measurement point has its own timestamp. In this case, the entire image can be converted for each measurement point, but this is computationally expensive. An alternative possibility is to search for those pixels with position p in the original image, which satisfy the above equation (1) for the 3D measurement points with image coordinates p'.

[0075] To this end, a variety of algorithms can be used. On the one hand, linear algorithms can be utilized to render all optical flow vectors such that the bounding box of the vector is specified in each pixel. If multiple flow vectors intersect in a pixel, the bounding box is correspondingly enlarged so that both vectors are contained within the box. The subsequent search algorithm then only has to consider the bounding box that must contain the pixel being searched.

[0076] Another possibility lies in implementing a search tree, such as Quadtrees, similar to collision recognition.

[0077] Most 3D measurement points have angular uncertainties, such as those caused by beam expansion. Thus, it is preferable to consider all pixels in the vicinity of this uncertainty in order to extend the attributes of the 3D measurement points from the image. These attributes can be, for example, the average optical flow o(o x ,o y ) or the position p(p x ,p y ) in the image space.

[0078] Due to the latest advances in the field of image processing using "[Deep] Convolutional Neuronal Networks (CNN)" ([deep] convolutional neural networks), it is possible to perform pixel-precise segmentation of images with the corresponding computing power. If at least one of these camera images is segmented by such a neural network, the 3D measurement points can additionally be extended with the classes and the associated identifiers obtained from the segmentation.

[0079] The points obtained from the virtual sensor can be clustered into high-quality object hypotheses because these points contain a wide range of information for separating the classes. This is especially the class information and identifiers from the segmentation and the Cartesian velocity vector, which is useful, for example, in the case of overlapping objects of the same class.

[0080] The extended 3D measurement points or clusters from the virtual sensor or these clusters are then transferred to an accumulative sensor data fusion, which enables filtering over time. In the case of some current neural networks, it is possible that these neural networks form so-called entities. As an example, a parking lot with parked vehicles can be given, and these parked vehicles are detected obliquely by a camera. Then, even though there is overlap in the image, the latest methods can still separate the different vehicles. If the neural network forms entities, then of course these entities can be used as clustering information in the accumulative sensor data fusion.

[0081] If there is information about the image segments due to the segmentation of these camera images, the complete calculation of the optical flow can be dispensed with if necessary. As an alternative, with suitable algorithms, it is also possible to determine the change of the individual image segments over time, which can be implemented particularly efficiently.

[0082] The time until collision can be determined from the optical flow o within the image space. This time until collision describes when a point penetrates the principal plane of the camera optics.

[0083] Using two corresponding points p 1 、p 2 in the image at two time points t 1 、t 2 with the distance b = p 1 -p 2 or using the distance at one time point and the associated optical flows o 1 、o 2 , the TTC can be calculated:

[0084]

[0085] In the following, for the mathematical representation, the pinhole camera model is used. From the image positions p x 、p y (in pixels), the TTC (in s), the optical flow o (in pixels / s), and the ranging d (in m) in the direction towards the image plane of the camera sensor, the Cartesian velocity vector v (in m / s) for 3D measurement can be determined, which is with respect to the self-motion in the camera coordinate system. It should be noted that the optical flow o and the pixel position p are specified within the image space, while the velocity v is determined in the camera coordinate system x,y,z .

[0086] In addition to the measurement attributes, the camera constant K is required, which takes into account the image width b (in m) and the resolution D (pixels per m) of the imaging system. Then, the velocity is obtained as follows:

[0087]

[0088] If the 3D measurement points should come from a radar sensor, the radial relative velocity (Doppler velocity) can additionally be used to stabilize the measurement: With the aid of this relative velocity and the ranging, an alternative TTC can be determined by taking the quotient. This is particularly useful in the case of features near the extended points of the camera, because there is only a small optical flow there. That is, this concerns objects within the driving envelope. However, this driving envelope is mostly covered by a particularly large number of sensors, so that the information is usually available.

[0089] Figure 8 Schematically shows the concept of a virtual sensor with a classifier. This concept largely corresponds to the concept known from Figure 7 the known concepts. Currently, convolutional neural networks are often used for image classification. If possible, these convolutional neural networks require locally correlated data, which undoubtedly exists in an image. Adjacent pixels often belong to the same object and are described as adjacent in the polar image space.

[0090] However, these neural networks preferably do not rely solely on image data, which provides little data under poor lighting conditions and also makes distance estimation generally difficult. Therefore, in other dimensions, measurement data from other cameras, especially from laser measurements and radar measurements, is projected into the state space. Here, for good performance, it is useful to synchronize the measurement data by means of optical flow so that these neural networks can make good use of data locality.

[0091] Here, the synchronization can be carried out in the following way. The starting point is the camera image, the capture time point of which is as close as possible to all sensor data. In addition to the pixel information, other data is now annotated: in the first step, this includes the offset in the image, for example, the offset by means of optical flow. By means of the prediction step described above, the pixels are identified again, which are associated with the available 3D measurement data, such as the available 3D measurement data from laser measurements or radar measurements, according to the pixel offset. Since there is beam expansion in these measurements, this mostly involves multiple pixels. The associated pixels are extended in other dimensions and the measurement attributes are entered correspondingly. Possible attributes are, for example: in addition to the ranging according to laser, radar or ultrasound, there is also the Doppler velocity, reflectivity or radar cross-section of the radar or also the confidence level.

[0092] Now, the synchronized camera image with extended measurement attributes is classified using a classifier or a segmenter 42, preferably using a convolutional neural network. In this case, all the information can now be generated as described above in connection with Figure 7 what has been described.

[0093] Subsequently, the mathematical background that is necessary for the synchronization of these camera images and these 3D measurement points should be elaborated in detail. Here, it is assumed that the camera can be modeled as a pinhole camera. This assumption is only used to make the transformation easier to operate. If the camera used cannot be appropriately modeled as a pinhole camera, a distortion model can be used as an alternative to generate a view that satisfies the pinhole camera model. In these cases, the parameters of the virtual pinhole camera model must be used in the subsequent equations.

[0094] First, the coordinate systems and the transformations between these coordinate systems must be defined. A total of five coordinate systems are defined:

[0095] - C W is the 3D world coordinate system

[0096] - C V is the 3D coordinate system of the vehicle itself

[0097] - C C is the 3D coordinate system of the camera

[0098] - C S is the 3D coordinate system of the 3D sensor

[0099] - C I is the 2D image coordinate system.

[0100] The coordinate systems of the camera, 3D sensor, image, and the vehicle itself are closely related to each other. Since the vehicle moves relative to the world coordinate system, the following four transformations between these coordinate systems are defined:

[0101] - T V←W (t) is the transformation that transforms a 3D point in the world coordinate system into the 3D coordinate system of the vehicle itself. Since the vehicle moves over time, this transformation depends on time t. - T S←V is the transformation that transforms a 3D point in the 3D coordinate system of the vehicle itself into the 3D coordinate system of the 3D sensor. - P C←V is the transformation that transforms a 3D point in the 3D coordinate system of the vehicle itself into the 3D coordinate system of the camera. - P I←C is the transformation that projects a 3D point in the 3D coordinate system of the camera onto the 2D image coordinate system.

[0102] A world point moving in the world coordinate system, such as a point on a vehicle, can be described by x w (t).

[0103] This point is detected by the camera at time point t 0 and detected by the 3D sensor at time point t 1 The camera observes the corresponding image point xi(t 0 ) in homogeneous coordinates:

[0104] x i (t 0 ) = P I←C · T C←V · T V←W (t 0 )· x w (t 0 ) (6)

[0105] The point x observed by the 3D sensor s :

[0106] x s (t 1 ) = T S←V ·T V←W (t 1 )·x w (t 1 ) (7)

[0107] Equations (6) and (7) are related to each other through the motion of the vehicle itself and the motion of the world point. Information about the motion of the vehicle itself is available, while the motion of the world point is unknown.

[0108] Therefore, it is necessary to determine information about the motion of the world point.

[0109] A second measurement of the camera at time point t 2 should be given:

[0110] x i (t 2 ) = P 1←C ·T C←V ·T V←W (t 2 )·x w (t 2 ) (8)

[0111] Now, equations (6) and (8) can be combined with each other:

[0112]

[0113] In the coordinate system of the vehicle itself, the observed point x v (t) is given by the following equation:

[0114] x v (t) = T V←W (t)·x w (t) (10)

[0115] If this is applied to equation (10), then we obtain:

[0116]

[0117] Equation (11) establishes the relationship between the optical flow vector and the motion vector of the world point. Δx i (t 0 , t 2 ) is the optical flow vector between the camera images taken at time t 0 and t 2 , and Δx v (t0 , t 2 ) is the world point in C V The corresponding motion vector expressed by . Therefore, the optical flow vector is the projection of the motion vector in 3D space.

[0118] The measurements of the camera and the 3D sensor cannot be combined directly with each other. First, the following additional assumption must be introduced: the motion in the image plane at time t 0 With t 2 Under this assumption, the image point belonging to the world point is determined by the following equation:

[0119]

[0120] From equation (11), it is clear that the motion of the world point as well as the motion of the host vehicle must be linear.

[0121] The transformation for the 3D sensor according to equation (7) can be used so that in the camera coordinate system C C Determine in C S At time point t 1 Measured 3D measuring points:

[0122]

[0123] The following pixel coordinates of the world point at time t can also be determined with the aid of equation (6): 1 The captured virtual camera image will have these pixel coordinates:

[0124]

[0125] If equation (13) is applied to equation (14), we obtain:

[0126]

[0127] Otherwise, x can be determined according to equation (12) i (t 1 ):

[0128]

[0129] Equation (16) establishes the relationship between the camera's measurements and the 3D sensor's measurements. If the world point is well defined in the world coordinate system, the time point t 0 ,t 1 and t 2 If the image coordinates in the two camera images and the measurements of the 3D sensor are known, equation (16) establishes a complete relationship, ie there are no unknown variables.

[0130] Even if the correct correspondence between the measurements of the 3D sensor and the camera is not known, this situation can be utilized. If there is a measurement of the 3D sensor at time point t 1 , this measurement can be transformed into the virtual camera image, that is, the camera image that the camera would detect at time point t 1 . The virtual pixel coordinates for this are x i (t 1 ). Now, with the help of the optical flow vector v i (t 0 , t 2 ), it is possible to search for the pixel x that is equal to or at least very close to x i (t 1 ). i (t 0 )

[0131] List of Reference Numerals

[0132] 10 Associate the measurement with the track

[0133] 11 Check whether the association is successful

[0134] 12 Update the corresponding track

[0135] 13 Initialize a new track

[0136] 14 Check the track with respect to the time point of the last update

[0137] 15 Delete the track

[0138] 16 Predict the track to the next measurement time point

[0139] 17 Transform the track state into the measurement space

[0140] 20 Detect the camera image

[0141] 21 Detect the 3D measurement point

[0142] 22 Segment at least one camera image

[0143] 23 Fuse the camera image with the 3D measurement point

[0144] 24 Output the data of the virtual sensor

[0145] 30 Calculate the optical flow

[0146] 31 Determine the time until collision

[0147] 32 Determine the pixel assigned to one of these 3D measurement points

[0148] 40 Device

[0149] 41 Input Terminal

[0150] 42 Splitter

[0151] 43 Data Fusion Unit

[0152] 44 Object Tracker

[0153] 45 Monitoring Unit

[0154] 46 Memory

[0155] 47 Output Terminal

[0156] 48 User Interface

[0157] 50 Device

[0158] 51 Memory

[0159] 52 Processor

[0160] 53 Input Terminal

[0161] 54 Output Terminal

[0162] 60 Motor Vehicle

[0163] 61 Camera

[0164] 62 Radar Sensor

[0165] 63 Ultrasonic Sensor

[0166] 64 Lidar System

[0167] 65 Data Transmission Unit

[0168] 66 Auxiliary System

[0169] 67 Memory

[0170] 68 Network

[0171] A Object List

[0172] E Sensor Data

[0173] FL Optical Flow

[0174] I1, I2 Camera Images

[0175] MP Measurement Point

[0176] TTC Time to Collision

[0177] VS Data of Virtual Sensor

Claims

1. A method for processing sensor data, the method having the following steps: - Detecting (20) camera images (I1, I2) by means of a camera (61); - Detecting (21) 3D measurement points (MP) by means of at least one 3D sensor (62, 64); and - Fusing (23) the camera images (I1, I2) with the 3D measurement points (MP) into virtual sensor data (VS), wherein fusing the image data with the 3D measurement points into virtual sensor data comprises: - Calculating (30) an optical flow (FL) based on at least one first camera image (I1) and at least one second camera image (I2) respectively before and after the measurement time point of the 3D measurement points (MP); and - Determining (32) a pixel in at least one of the camera images (I1, I2) that should be assigned to one of the 3D measurement points (MP) at the measurement time point based on the optical flow (FL) according to the following formula: p′ = p + o·Δt where p' is the image coordinate of the 3D measurement point (MP), p is the position of the pixel, o is the optical flow (FL), and Δt is the time difference between the shooting time points of the camera images (I1, I2) and the measurement time point, and the shooting time points of the camera images (I1, I2) are closest to the measurement time point of the 3D sensor.

2. The method according to claim 1, wherein the determination (32) of a pixel in at least one of the camera images (I1, I2) that should be assigned to one of the 3D measurement points (MP) at the measurement time point comprises: - Transforming camera images (I1, I2) that are temporally close to the measurement time point of the 3D sensor (62, 64) based on the optical flow (FL); and - Projecting the 3D measurement points (MP) into the transformed camera images.

3. The method according to claim 1, wherein the determination (32) of a pixel in at least one of the camera images (I1, I2) that should be assigned to one of the 3D measurement points (MP) at the measurement time point comprises: - Determining, based on the optical flow (FL) and a search method, the pixels in the camera images (I1, I2) that should be assigned to the 3D measurement points (MP) at the measurement time point; and - Projecting the 3D measurement points (MP) onto the positions thus determined in the camera images (I1, I2).

4. The method according to any one of claims 1 to 3, wherein for the pixels of the camera images, a time to collision (TTC) is determined (31) from the optical flow (FL), and a Cartesian velocity vector of the 3D measurement points (MP) is calculated based on the time to collision (TTC), the optical flow (FL), and the ranging for the 3D measurement points (MP).

5. The method according to claim 4, wherein the time until collision is determined (31) based on measurements of the 3D sensors (62, 64), rather than being determined from the optical flow.

6. The method according to any one of claims 1 - 3, wherein expanding the 3D measurement points attributes from at least one of the camera images (I1, I2).

7. The method according to any one of claims 1 - 3, wherein segmenting at least one of the camera images (I1, I2) at a measurement time point close to the 3D sensors (62, 64).

8. The method according to claim 7, wherein the segmentation takes into account measurements of the 3D sensors (62, 64) in addition to image information.

9. The method according to any one of claims 1 - 3, wherein applying an object tracking algorithm to the data (VS) of the virtual sensor.

10. The method according to claim 9, wherein the object tracking algorithm performs accumulative sensor data fusion.

11. A computer - readable storage medium having instructions which, when executed by a computer, cause the computer to perform the steps of the method for processing sensor data according to any one of claims 1 to 10.

12. A device (20) for processing sensor data, the device having: - an input end (41) for receiving camera images (I1, I2) of a camera (61) and 3D measurement points (MP) of 3D sensors (62, 64); and - a data fusion unit (43) for fusing (23) the camera images (I1, I2) with the 3D measurement points (MP) into data (VS) of a virtual sensor, wherein fusing the image data with the 3D measurement points into data of a virtual sensor comprises: - calculating (30) an optical flow (FL) based on at least one first camera image (I1) and at least one second camera image (I2) respectively before and after a measurement time point of the 3D measurement points (MP); and - determining (32) based on the optical flow (FL) according to the following formula the pixels in at least one of the camera images (I1, I2) that should be assigned to one of the 3D measurement points (MP) at the measurement time point: p′ = p + o·Δt where p' is the image coordinate of the 3D measurement point (MP), p is the position of the pixel, o is the optical flow (FL), and Δt is the time difference between the shooting time point of the camera images (I1, I2) and the measurement time point, and the shooting time point of the camera images (I1, I2) is closest to the measurement time point of the 3D sensor.

13. A motor vehicle (60), characterized in that the motor vehicle (60) has the device (40) according to claim 12 or is configured to implement the method for processing sensor data according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for acquisition and / or tracking of objects e.g. static objects, in e.g. front side of vehicle, involves disregarding objects having optical flow distance smaller than distance value determined by sensor from detection range of sensor

    DE102011013776A1

  • APPEARANCE-BASED UNION OF CAMERA AND DISTANCE DATA FOR MULTI-OBJECTS

    DE102011119767A1

  • Pedestrian detection WITH ALERT MAPS

    DE102017100199A1

  • Systems and methods for braking a vehicle based on a detected object

    US20150336547A1