Improved process for detecting objects using a neural network

The improved process for detecting objects using neural networks in vehicle driving assistance systems addresses the limitations of existing technologies by supplying rectified optical flow cards, enhancing detection accuracy across various vehicle movements.

DE112023002835T5Pending Publication Date: 2025-05-08AUMOVIO AUTONOMOUS MOBILITY GERMANY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112023002835
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-28
Filing Date
2023-06-20
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing neural networks used in vehicle driving assistance systems are less effective in detecting objects when the vehicle's movement includes components other than longitudinal transfer, such as lateral translation, vertical translation, or rotations, due to their training data primarily focusing on longitudinal movements.

Method used

An improved process for detecting objects using a neural network that involves supplying the network with rectified optical flow cards. This process includes three steps: appreciating the translation and rotation parameters of the camera's movement, applying transformation matrices to rectify the optical flow, and standardizing the optical flow to ignore parasitic movements, thereby enhancing the network's performance across various vehicle movements.

Benefits of technology

The improved process enhances the performance of neural networks in detecting objects by compensating for non-longitudinal movements of the vehicle, thereby improving detection accuracy and reliability in diverse driving conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to an improved process for detecting objects (16) using a neural network, wherein the neural network is supplied at the input with at least one original first optical flow map representing the computerized tracking of moving objects (16) in a scene, by analyzing the differences in content between a first image captured by an image capture device (14) at a first position (P1) at an earlier time and a subsequent second image captured by the image capture device (14) at a second position (P2) at a current time, characterized in that it consists of rectifying the original optical flow map using the estimation information of the ego movement of the image capture device (14).
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] The invention relates to the field of detecting objects using a neural network and, in particular, to an improved process for detecting objects using a neural network.

[0002] The present invention is particularly designed to be implemented by a computer with which a vehicle is equipped, in particular a motor vehicle comprising a driver assistance system. State of the art

[0003] It is known to equip a motor vehicle with a driver assistance system, which is generally known by the acronym FAS for “driver assistance system”.

[0004] Such an assistance system comprises, as is known, at least one image capture device attached to the vehicle which makes it possible to generate a series of images representing the vehicle's surroundings.

[0005] The image capture device is, for example, a lidar (acronym for “light detection and ranging”), a camera or a radar.

[0006] The captured images are used by a computer to assist the driver, for example by detecting “objects” such as a pedestrian, a stationary vehicle or any other object on the road, and by calculating, for example, the time before collision with the detected object.

[0007] The information provided by the images captured by the image acquisition device makes it possible to implement simultaneous positioning and mapping (known by the acronym SLAM for “Simultaneous Localization And Mapping”) in order to enable the scene representing the environment of the motor vehicle to be constructed and enriched simultaneously, and also to enable the motor vehicle to be located in the scene.

[0008] Thus, the use of neural networks, also known by the acronym CNN for “Convolutional Neural Networks”, is known.

[0009] Neural networks are used to perform a scene perception function and provide information about various objects present in the vehicle's environment.

[0010] Neural networks are particularly used to provide semantic, detection, motion and location information about objects in the scene.

[0011] For this purpose, the neural network should be trained in learning mode by providing it with numerous previously labeled images at the input to teach the neural network to recognize objects.

[0012] Once the neural network has been trained, it can be used in detection mode to detect and recognize objects.

[0013] For this purpose, the previously trained neural network is supplied with images, for example images captured by an image capture device attached to a motor vehicle.

[0014] Thus, it is known to supply a neural network with optical flow maps.

[0015] An optical flow map describes and represents the computerized tracking of moving objects by analyzing differences in content between successive video frames.

[0016] A computer can locate reference frames that mark the boundaries, edges, and territories of individual still images.

[0017] Detecting their movement allows the computer to track an object in time and space.

[0018] In other words, an optical flow map represents the movement of moving areas, or moving objects, between a first image captured by an image capture device at a first position at an earlier time t-1 and a subsequent second image captured by the image capture device at a second position at a current time t.

[0019] The direction and force of movement of the moving objects are illustrated, for example, by motion vectors on the optical flow map.

[0020] In the field of the invention, namely driver assistance, the image capture device is, in terms of movement, in sync with the motor vehicle that drives it.

[0021] The image capture device is mounted on a camera reference frame having a transverse axis, a vertical axis, and a longitudinal axis extending from rear to front along a main direction of movement of the vehicle.

[0022] The movement of the image capture device, or its ego motion, includes six degrees of freedom, namely a transverse translation, a vertical translation from top to bottom, a longitudinal translation, a rotation around the transverse axis called pitch, a rotation around the vertical axis called yaw, and a rotation around the longitudinal axis called roll.

[0023] During the lifetime of the vehicle, the vehicle and the image capture device will move essentially in the main direction of travel of the vehicle on the road, i.e., in longitudinal translation.

[0024] Thus, neural networks are trained with optical flow maps that mainly correspond to a movement in longitudinal translation and, in very rare cases, to the other movements described above.

[0025] Consequently, neural networks will be effective in detecting objects when the vehicle moves into longitudinal translation.

[0026] As a result, neural networks will be less effective at detecting objects when the vehicle's motion includes a component other than longitudinal translation, such as a slide corresponding to lateral translation, a speed bump corresponding to up-down translation and pitching, a sharp turn corresponding to yaw rotation, a pothole corresponding to roll rotation, or otherwise a sharp acceleration of the vehicle corresponding to pitch rotation.

[0027] Thus, it is observed that the performance of neural networks is limited by the data they are fed. Brief description of the invention

[0028] The present invention aims to propose an improved process for detecting objects by means of a neural network, which increases the performance of the neural network, in particular under conditions corresponding to movements of the vehicle other than a longitudinal translational movement in the main direction of travel of the vehicle.

[0029] This object, as well as others which will become apparent from reading the following description, is achieved with an improved process for detecting objects by means of a neural network, the neural network being supplied at input with at least one original first optical flow map representing the computerized tracking of moving objects in a scene, by analyzing the differences in content between a first image acquired by an image acquisition device at a first position at a previous time and a subsequent second image acquired by the image acquisition device at a second position at a current time, characterized in that it comprises at least the following: • a first step of estimating the three translation parameters and the three rotation parameters of the ego movement of the image capture device between the first position and the second position, wherein the three translation parameters comprise two secondary translation parameters and one main translation parameter along a longitudinal axis corresponding to the main movement of the image capture device, • a second step consisting in estimating a first rotation matrix allowing to move from the first position to a first virtual rectified position and a second rotation matrix allowing to move from the second position to a second virtual rectified position, such that the two secondary translation parameters and the three rotation parameters allowing to move from the first rectified position to the second rectified position are equal to zero, • a third normalization step consisting in calculating a rectified second optical flow map obtained by applying the first rotation matrix and the second rotation matrix to the original first optical flow map and representing the computerized tracking of moving objects between the first rectified position and the second rectified position.

[0030] Thus, the process according to the invention makes it possible to obtain rectified optical flow maps that ignore parasitic movements of the image acquisition sensor, such as pitch, roll and yaw movements, when the image acquisition sensor is mounted on a motor vehicle.

[0031] According to other optional features of the invention, taken alone or in combination: - the process includes a learning step consisting of providing a neural network with a plurality of rectified optical flow maps according to the third step of normalization of the process to train the neural network. This feature allows for improving the learning performance of the neural network supplied with rectified optical flow maps; - the process includes a detection step consisting of supplying a neural network with a plurality of rectified optical flow maps according to the third normalization step to perform object detection. This feature allows for improving the detection performance of the neural network supplied with rectified optical flow maps.

[0032] The invention also relates to a computer designed to implement the process described above.

[0033] Furthermore, the invention relates to a motor vehicle comprising a computer of the type described above and at least one image capture device which is in synchronization with the motor vehicle in terms of movement, the motor vehicle moving mainly along a longitudinal axis. Description of the drawings

[0034] Other features and advantages of the invention will become apparent from reading the following description with reference to the attached figures, which illustrate: [ Fig. 1] a schematic plan view of a motor vehicle equipped with an image capture device, at a first, previous position and at a second, current position; [ Fig. 2] a schematic view similar to that of Fig. 1 of the motor vehicle in a first rectified position and in a second rectified position after the process according to the invention; [ Fig. 3] a schematic view of an original first optical flow map; [ Fig. 4] a schematic view of a second optical flow map rectified by the process according to the invention.

[0035] In the description and in the claims, the terminology transverse, vertical and longitudinal is adopted in a non-limiting manner with reference to the transverse axis Xc, the vertical axis Yc and the longitudinal axis Zc of the camera reference frame Rc indicated in the figures, respectively, taking into account that the motor vehicle extends longitudinally and moves forward in the longitudinal direction. Description of the embodiments

[0036] Fig. 1 shows a motor vehicle 10 equipped with a computer 12 and an image capture device 14.

[0037] A camera reference frame Rc, which is the reference frame attached to the image capture device 14, is taken into account.

[0038] A target reference frame (not shown) corresponding to the target position of the image capture device 14, that is, the theoretical ideal position that the image capture device 14 should assume on the associated motor vehicle 10, is also taken into account.

[0039] The camera reference frame Rc includes an axis Xc extending transversely, an axis Yc extending vertically, and an axis Zc extending longitudinally along a main direction of movement of the motor vehicle 10 from rear to front.

[0040] It is noted that the image capture device 14 is linked to the motor vehicle 10 with respect to movement.

[0041] Thus, the movement of the image capture device 14, or its ego movement, comprises six degrees of freedom, namely one degree of freedom corresponding to a transverse translation along the axis Xc, one degree of freedom corresponding to a vertical translation from top to bottom along the axis Yc, one degree of freedom corresponding to a longitudinal translation along the axis Zc, one degree of freedom about the axis Xc corresponding to a rotation called pitch, one degree of freedom about the axis Yc corresponding to a rotation called yaw, and one degree of freedom about the axis Zc corresponding to a rotation called roll.

[0042] The motor vehicle 10 is equipped with a driver assistance system intended to assist the driver, for example by analyzing the data provided by the image capture device 14, in particular to detect “objects” such as a pedestrian, a stationary vehicle or any other obstacle on the road.

[0043] For this purpose, the computer 12 of the motor vehicle 10 implements an improved process for detecting objects by means of a neural network according to the invention.

[0044] The neural network is supplied with a multitude of optical flow maps at the input.

[0045] For the sake of clarity, the flow of the process according to the invention will be described below with a single optical flow map, called the “original first optical flow map C1”.

[0046] The original first optical flow map C1, which was Fig. 3, represents the computerized tracking of moving objects in a scene by analyzing the differences in content between two successive images captured by the image capture device 14.

[0047] With reference to Fig. 1, the images that make it possible to obtain the original first optical flow map C1 include a first image taken by the image acquisition device 14 at a first position P1 at an earlier time, and a subsequent second image taken by the image acquisition device 14 at a second position P2 at a current time.

[0048] The process according to the invention comprises a first step of estimating the ego motion of the image capture device 14.

[0049] The first step consists of estimating the three translation parameters and the three rotation parameters of the ego movement of the image capture device 14 between the first position P1 and the second position P2.

[0050] The three translation parameters include a main translation parameter along the longitudinal axis Zc of the main movement of the image capture device 14 and two secondary translation parameters along the vertical axis Yc and along the transverse axis Xc of the camera reference frame Rc.

[0051] The estimation of the ego motion of the image capture device 14 may be obtained by various methods, such as a simultaneous positioning and mapping method, known by the acronym SLAM for "Simultaneous Localization And Mapping", or by means of the data provided by an inertial sensor guided by the motor vehicle 10, or otherwise by means of a geolocation system.

[0052] After the first step, the process performs a second step consisting of estimating a transformation matrix that allows to pass from a first virtual rectified position P3 to a second virtual rectified position P4 of the image acquisition device 14.

[0053] As in Fig. 2, the transition from the first rectified position P3 to the second rectified position P4 requires only a longitudinal translation along the longitudinal axis Zc, which in Fig. 2 is illustrated by a dashed line Tz.

[0054] The longitudinal translation along the longitudinal axis Zc corresponds to the main direction of movement of the assembly formed by the motor vehicle 10 and the image capture device 14.

[0055] Thus, the second step consists of estimating a first rotation matrix allowing to pass from the first position P1 to the first virtual rectified position P3 and a second rotation matrix allowing to pass from the second position P2 to the second virtual rectified position P4, such that the two secondary translation parameters and the three rotation parameters allowing the transition from the first rectified position P3 to the second rectified position P4 are equal to zero.

[0056] Following the second step, the process performs a third normalization step, which consists of calculating a rectified second optical flow map C2, illustrated in Fig. 4, which is obtained by applying the first rotation matrix and the second rotation matrix to the original first optical flow map C1.

[0057] The rectified second optical flow map C2 represents the computerized tracking of moving objects between the first rectified position P3 and the second rectified position P4 of the image capture device 14.

[0058] Thus, the rectified second optical flow map C2 has the advantage of rectifying the optical flows generated by the motion vectors in the Fig. 3 and Fig. 4, so that the transition from the first rectified position P3 to the second rectified position P4 is limited to a longitudinal translation along the longitudinal axis Zc of travel of the motor vehicle 10.

[0059] An exemplary implementation of the process according to the invention is described below.

[0060] With reference to Fig. 1, the image capture device 14 draws a curved trajectory for the movement from the first position P1 to the second position P2, which results in a rotation of the image capture device 14 and the motor vehicle 10 about the vertical axis Yc.

[0061] This curved trajectory is in the Fig. 1 and Fig. 2 is illustrated by a dashed axis line.

[0062] The ego motion of the image capture device 14 is estimated during the first estimation step of the process.

[0063] Next, during the second step of the process, a first rotation matrix allowing to transition from the first position P1 to the first virtual rectified position P3 and a second rotation matrix allowing to transition from the second position P2 to the second virtual rectified position P4 are estimated such that the two secondary translation parameters and the three rotation parameters allowing the transition from the first rectified position P3 to the second rectified position P4 are equal to zero.

[0064] As in Fig. 3, which illustrates the original first optical flow map C1, a plurality of motion vectors represent the tracking of an object 16 by analyzing the differences in content between the first image captured by the image capture device 14 at the first position P1 and the second image captured by the image capture device 14 at the second position P2.

[0065] Fig. Figure 4 shows the rectified second optical flow map C2 calculated during the third step of the normalization process.

[0066] It is noted that the movement induced by the curved trajectory of the motor vehicle 10, which corresponds to a rotation of the detection sensor 14 about the vertical axis Yc, is compensated.

[0067] The compensation is reflected as a reduction in the length of the motion vectors from Fig.4, ie, a suppression of the distortion caused by the curved trajectory of the motor vehicle 10.

[0068] It will be understood that the process according to the invention is designed to correct distortion caused by any movement other than longitudinal translation along the axis Zc, such as a slide corresponding to lateral translation along the axis Xc, a speed bump corresponding to up-down translation along the axis Yc and / or rotation about the axis Xc, a sharp turn corresponding to yaw rotation about the axis Yc, a pothole corresponding to roll rotation about the axis Zc, or otherwise a sharp acceleration of the vehicle corresponding to pitch rotation about the axis Xc.

[0069] The process according to the invention can be used in object detection mode and in learning mode.

[0070] In detection mode, the process includes a detection step consisting of supplying a neural network with a plurality of rectified optical flow maps according to the third normalization step.

[0071] The detection mode allows to improve the performance of the neural network for detecting objects in numerous situations, especially in the pitch, roll, yaw and other situations mentioned above.

[0072] In learning mode, the process includes a learning step consisting of providing a neural network with a plurality of rectified optical flow maps according to the third normalization step to train the neural network.

[0073] The learning mode makes it possible to improve the performance of the neural network trained by the process according to the invention.

[0074] Of course, the invention has been described in the foregoing by way of example. It is understood that one skilled in the art will be able to devise various alternative embodiments of the invention without thereby departing from the scope of the invention.

[0075] For example, it is possible to compensate and correct a significant inclination of the image capture device 14 in certain elevated vehicles, such as trucks, or to correct a mounting error of the image capture device 14 with respect to its theoretical target position.

[0076] For this purpose, a suitable transformation between the camera reference frame Rc, which is the reference frame attached to the image capture device 14, and the target reference frame corresponding to the target position of the image capture device 14 of the process is applied following the second step.

[0077] This transformation consists, for example, in applying an additional rotation matrix to the original first optical flow map C1 during the third step in order to correct a mounting error of the image acquisition device 14 with respect to its theoretical target position.

Claims

[1] An improved process for detecting objects (16) by means of a neural network, the neural network being supplied at the input with at least one original first optical flow map (C1) representing the computerized tracking of moving objects (16) in a scene by analyzing the differences in content between a first image captured by an image capture device (14) at a first position (P1) at an earlier time and a subsequent second image captured by the image capture device (14) at a second position (P2) at a current time, characterized by that it includes at least the following: • a first step of estimating the three translation parameters and the three rotation parameters of the ego movement of the image capture device (14) between the first position (P1) and the second position (P2), wherein the three translation parameters comprise two secondary translation parameters and one main translation parameter along a longitudinal axis (Zc) corresponding to the main movement of the image capture device (14), • a second step consisting in estimating a first rotation matrix allowing to pass from the first position (P1) to a first virtual rectified position (P3) and a second rotation matrix allowing to pass from the second position (P2) to a second virtual rectified position (P4), such that the two secondary translation parameters and the three rotation parameters allowing to pass from the first rectified position (P3) to the second rectified position (P4) are equal to zero, • a third normalization step consisting in calculating a rectified second optical flow map (C2) obtained by applying the first rotation matrix and the second rotation matrix to the original first optical flow map (C1) and representing the computerized tracking of moving objects (16) between the first rectified position (P3) and the second rectified position (P4). [2] An improved process for detecting objects (16) using a neural network according to claim 1, characterized by that it includes a learning step consisting of supplying a neural network with a plurality of rectified optical flow maps (C2) according to the third step of normalization of the process in order to train the neural network. [3] An improved process for detecting objects (16) using a neural network according to any one of the preceding claims, characterized bythat it comprises a detection step consisting in supplying a neural network with a plurality of rectified optical flow maps (C2) according to the third normalization step in order to perform object detection. [4] A computer (12) adapted to implement the process of any preceding claim. [5] A motor vehicle (10) comprising a computer (12) according to claim 4 and at least one image capture device (14) which is in motion in sync with the motor vehicle (10), the motor vehicle (10) moving primarily along a longitudinal axis (Zc).