Improved Process for Object Detection by Neural Network

The improved process for object detection in neural networks corrects optical flow maps using ego-motion estimation and rotation matrices to address limitations in vehicle movements, enhancing detection accuracy and training efficiency.

JP2025520803APending Publication Date: 2025-07-03CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024576536
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-28
Filing Date
2023-06-20
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Neural networks used in vehicle driving assistance systems are ineffective in detecting objects when vehicle movements include components other than translation in the longitudinal direction, such as skidding, turning, or sudden acceleration, due to limited training data.

Method used

An improved process for object detection by neural networks that includes estimating ego-motion parameters and applying rotation matrices to correct optical flow maps, compensating for parasitic movements like pitch, roll, and yaw, and using corrected optical flow maps for training and detection.

Benefits of technology

Enhances neural network performance in various vehicle movements by correcting for non-longitudinal movements, improving object detection accuracy and training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520803000001_ABST
    Figure 2025520803000001_ABST
Patent Text Reader

Abstract

The present invention relates to an improved process for detecting an object (16) by a neural network, the neural network being supplied, as input, with at least one first initial optical flow map representing a computer tracking of a moving object (16) in a scene by analyzing the difference in content between a first image taken by an image acquisition device (14) at a first position (P1) at a previous time and a successive second image taken by the image acquisition device (14) at a second position (P2) at the current time, the process being characterized in that it comprises correcting the initial optical flow map using ego-motion estimation information from the image acquisition device (14).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection by neural networks, and more particularly to an improved process for object detection by neural networks.

[0002] The present invention is particularly designed to be executed by a computer equipped in a vehicle, particularly an automobile, including a driving assistance system.

Background Art

[0003] It is known to equip an automobile with a driving assistance system generally known as ADAS, which is an abbreviation for "Advanced Driver Assistance System".

[0004] As is well known, such a support system includes at least one image acquisition device, which is mounted on the vehicle and can thereby generate a series of images representing the environment of the automobile.

[0005] The image acquisition device is, for example, Lidar (an abbreviation for "Light Detection And Ranging"), a camera, or a radar.

[0006] The captured images are used by a computer to assist the driver by detecting "objects", such as pedestrians, parked vehicles, or any other object on the road, etc., and by calculating, for example, the time until a collision with the detected object.

[0007] Based on the information obtained from the images captured by the image acquisition device, it is possible to perform self-position estimation and map creation (known as SLAM, an abbreviation for "Simultaneous Localization And Mapping") to simultaneously construct a scene representing the environment of the automobile, incorporate information therein, and identify the position of the automobile within the scene.

[0008] Therefore, the use of neural networks, also known as CNN, an abbreviation of "Convolutional Neural Networks", is known.

[0009] Neural networks are used to perform scene recognition functions and provide information about various objects in the vehicle's environment.

[0010] Neural networks are particularly used to provide semantic, detection, motion, and position information about objects in a scene.

[0011] For this purpose, neural networks should be trained by supplying many previously labeled images as their input in the learning mode to teach the neural networks object recognition.

[0012] Once the training of the neural network is completed, it can be used to detect and recognize objects in the detection mode.

[0013] To do this, an image, for example, an image captured by an image acquisition device mounted on a vehicle, is supplied to the already trained neural network.

[0014] Therefore, it is known to supply an optical flow map to a neural network.

[0015] The optical flow map explains and represents the computer tracking of moving objects by analyzing the differences in content between consecutive video images.

[0016] The computer can position a reference frame that marks the boundaries, edges, and regions of individual still images.

[0017] By detecting these movements, the computer can track objects in time and space.

[0018] In other words, the optical flow map represents the moving area between the first image captured by the image acquisition device at the first position at the previous time t-1 and the subsequent second image captured by the image acquisition device at the second position at the current time t, or the movement of the moving object.

[0019] The direction and force of the movement of the moving object are indicated, for example, by movement vectors on the optical flow map.

[0020] In the field of the present invention, namely driving assistance, the image acquisition device is integral with the motor vehicle on which it is mounted in terms of movement.

[0021] The image acquisition device is attached to the reference frame of the camera, which includes the horizontal axis, the vertical axis, and the longitudinal direction extending along the main moving direction of the motor vehicle from the rear to the front.

[0022] The movement of the image acquisition device, or its ego motion, includes six degrees of freedom, namely lateral translation, vertical translation from top to bottom, translation in the longitudinal direction, rotation called pitch around the horizontal axis, rotation called yaw around the vertical axis, and rotation called roll around the longitudinal axis.

[0023] During the use period of the vehicle, the vehicle and the image acquisition device basically move in the main driving direction of the vehicle on the road, that is, perform translation in the longitudinal direction.

[0024] Therefore, the neural network is trained with an optical flow map that mostly corresponds to the movement of translation in the longitudinal direction and only partly corresponds to the other movements described above.

[0025] As a result, the neural network is effective in object detection when the vehicle is moving by translation in the longitudinal direction.

[0026] Conversely, a neural network is not very effective in object detection when the movement of the vehicle includes components other than translation in the longitudinal direction, such as skidding corresponding to translation in the lateral direction, steps on the road corresponding to translation from top to bottom, sharp turning corresponding to yaw rotation, potholes corresponding to roll rotation, or sudden acceleration of the vehicle corresponding to pitch rotation.

[0027] Therefore, it can be seen that the performance of the neural network is limited by the data supplied to it.

Summary of the Invention

Problems to be Solved by the Invention

[0028] The present invention aims to propose an improved process for object detection by a neural network that enhances the performance of the neural network particularly in situations corresponding to vehicle movements other than translational movement in the longitudinal direction in the main traveling direction of the vehicle.

Means for Solving the Problems

[0029] In addition to this object, other objects that will become apparent from reading the following description are achieved by an improved process for object detection by a neural network, wherein the neural network is supplied, as input, with at least one initial first optical flow map representing computer tracking of a moving object in a scene by analyzing the difference in content between a first image taken by an image acquisition device at a first position at a previous time and a successive second image taken by the image acquisition device at a second position at a current time. The process includes at least, · A first step of estimating three translational parameters and three rotational parameters of the ego motion of the image acquisition device between the first position and the second position, the three translational parameters including two secondary translational parameters and one main translational parameter along the longitudinal axis corresponding to the main movement of the image acquisition device. · A second step including estimating a first rotation matrix enabling movement from a first position to a first virtual correction position and a second rotation matrix enabling movement from a second position to a second virtual correction position such that two secondary translation parameters and three rotation parameters enabling passage from the first correction position to the second correction position are equal to zero, · A third step including calculating a corrected second optical flow map obtained by applying the first rotation matrix and the second rotation matrix to an initial first optical flow map and representing computer tracking of a moving object between the first correction position and the second correction position, characterized by including the above.

[0030] Therefore, by the process according to the present invention, it is possible to obtain a corrected optical flow map that ignores parasitic movements of an image acquisition sensor, such as pitch, roll, and yaw movements when the image acquisition sensor is mounted on a vehicle.

[0031] According to another optional feature of the present invention, considered alone or in combination,

[0032] - The process includes a learning step of supplying a plurality of corrected optical flow maps from a third normalization step of the process to a neural network to train the neural network. By this feature, the learning performance of the neural network supplied with the optical flow map corrected in this way can be improved.

[0033] - The process includes a detection step of supplying a plurality of corrected optical flow maps from a third normalization step to a neural network to perform object detection. By this feature, the detection performance of the neural network supplied with the optical flow map corrected in this way can be improved.

[0034] The present invention also relates to a computer configured to execute the aforementioned process.

[0035] Furthermore, the present invention relates to a motor vehicle, which relates to a motor vehicle including a computer of the above-described type and at least one image acquisition device that is integral with the motor vehicle in terms of movement, and the motor vehicle moves mainly along the longitudinal axis.

[0036] Other features and advantages of the present invention will become apparent upon reading the following description with reference to the accompanying drawings, which illustrate the following.

Brief Description of the Drawings

[0037]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0038] In the description and claims, the terms horizontal, vertical, and longitudinal are used non-limitingly with respect to the horizontal axis Xc, vertical axis Yc, and longitudinal axis Zc of the reference frame Rc of the camera shown in the figures, considering that the motor vehicle extends in the longitudinal direction and moves forward in the longitudinal direction.

[0039] FIG. 1 shows a motor vehicle 10 equipped with a computer 12 and an image acquisition device 14.

[0040] The reference frame Rc of the camera, which is the reference frame attached to the image acquisition device 14, is considered.

[0041] The reference position of the image acquisition device 14, that is, a reference reference frame (not shown) corresponding to the theoretically ideal position that the image acquisition device 14 should occupy on the related automobile 10 is also considered.

[0042] The reference frame Rc of the camera includes an axis Xc extending in the horizontal direction, an axis Yc extending in the vertical direction, and an axis Zc extending in the longitudinal direction from the rear to the front along the main moving direction of the automobile 10.

[0043] It should be noted that the image acquisition device 14 is connected to the automobile 10 at a point that moves.

[0044] Therefore, the movement of the image acquisition device 14, or its ego motion, includes six degrees of freedom, which are, namely, one degree of freedom corresponding to the horizontal translation along the axis Xc, one degree of freedom corresponding to the vertical translation from top to bottom along the axis Yc, one degree of freedom corresponding to the longitudinal translation along the axis Zc, one degree of freedom around the axis Xc corresponding to a rotation called pitch, one degree of freedom around the axis Yc corresponding to a rotation called yaw, and one degree of freedom around the axis Zc corresponding to a rotation called roll.

[0045] The automobile 10 is provided with a driving assistance system that aims to assist the driver by analyzing, for example, the data provided by the image acquisition device 14 and detecting "objects" such as pedestrians, parked vehicles, or any other obstacles on the road.

[0046] For this purpose, the computer 12 of the automobile 10 executes an improved process for object detection by a neural network according to the present invention.

[0047] A plurality of optical flow images are supplied as inputs to the neural network.

[0048] For clarity, hereinafter, the operation of the process according to the present invention will be described with one optical flow map, which is called the "initial first optical flow map C1".

[0049] The initial first optical flow map C1 is shown in FIG. 3 and represents the computer tracking of moving objects in a scene by analyzing the difference in content between two consecutive images taken by the image acquisition device 14.

[0050] Regarding FIG. 1, the images that enable the acquisition of the initial first optical flow map C1 include a first image taken by the image acquisition device 14 at the first position P1 at a previous time and a second consecutive image taken by the image acquisition device 14 at the second position P2 at the current time.

[0051] The process according to the present invention includes a first step of estimating the ego - motion of the image acquisition device 14.

[0052] The first step includes estimating three translational parameters and three rotational parameters of the ego - motion of the image acquisition device 14 between the first position P1 and the second position P2.

[0053] The three translational parameters include a main translational parameter along the longitudinal axis Zc of the main translational movement of the image acquisition device 14 and two secondary translational parameters along the vertical axis Yc and the horizontal axis Xc of the camera's reference frame Rc.

[0054] The estimation of the ego - motion of the image acquisition device 14 can be obtained by various methods such as a simultaneous localization and mapping method known as SLAM for short, or by data provided by an inertial sensor provided in the vehicle 10, or by a geolocation system.

[0055] After the first step, the process performs a second step, which includes estimating a transformation matrix that enables the movement of the image acquisition device 14 from the first virtual corrected position P3 to the second virtual corrected position P4.

[0056] As can be seen from FIG. 2, passing from the first correction position P3 to the second correction position P4 requires only a translation in the longitudinal direction along the longitudinal axis Zc, which is shown by the dotted line Tz in FIG. 2.

[0057] The translation in the longitudinal direction along the longitudinal axis Zc corresponds to the main movement direction of the assembly formed by the vehicle 10 and the image acquisition device 14.

[0058] Therefore, the second step includes estimating the first rotation matrix that enables movement from the first position P1 to the first virtual correction position P3 and the second rotation matrix that enables movement from the second position P2 to the second virtual correction position P4 such that two secondary translation parameters and three rotation parameters that enable passage from the first correction position P3 to the second correction position P4 are equal to zero.

[0059] After the second step, the process performs a third normalization step, which includes calculating a corrected second optical flow map C2 shown in FIG. 4, obtained by applying the first rotation matrix and the second rotation matrix to the initial first optical flow map C1.

[0060] The corrected second optical flow map C2 represents the computer tracking of an object moving between the first correction position P3 and the second correction position P4 of the image acquisition device 14.

[0061] Therefore, the corrected second optical flow map C2 has the advantage that the optical flow represented by the motion vectors in FIGS. 3 and 4 is corrected such that the passage from the first correction position P3 to the second correction position P4 is limited to a translation in the longitudinal direction along the longitudinal axis Zc of the travel of the vehicle 10.

[0062] An exemplary implementation of the process according to the present invention will be described below.

[0063] Referring to FIG. 1, the image acquisition device 14 depicts a curved trajectory for moving from the first position P1 to the second position P2, which is converted into the rotation of the image acquisition device 14 and the automobile 10 around the vertical axis Yc.

[0064] This curved trajectory is indicated by an axis with dots in FIGS. 1 and 2.

[0065] The ego-motion of the image acquisition device 14 is estimated during the first estimation step of the process.

[0066] Next, during the second step of the process, a first rotation matrix that enables movement from the first position P1 to the first virtual correction position P3 and a second rotation matrix that enables movement from the second position P2 to the second virtual correction position P4 are estimated such that two secondary translation parameters and three rotation parameters that enable passage from the first correction position P3 to the second correction position P4 are equal to zero.

[0067] As can be seen from FIG. 3 showing the initial first optical flow map C1, a plurality of motion vectors represent the tracking of the object 16 by analyzing the content difference between the first image captured by the image acquisition device 14 at the first position P1 and the second image captured by the image acquisition device 14 at the second position P2.

[0068] FIG. 4 shows the corrected second optical flow map C2 calculated during the third normalization step of the process.

[0069] Note that the movement induced by the curved trajectory of the automobile 10 corresponding to the rotation of the acquisition sensor 14 around the vertical axis Yc is compensated.

[0070] This compensation will shorten the length of the motion vectors in FIG. 4, that is, the bias caused by the curved trajectory of the automobile 10 is suppressed.

[0071] The process according to the invention is designed to correct for any movement other than translational movement in the longitudinal direction along axis Zc, such as skidding corresponding to translational movement in the lateral direction along axis Xc, vertical translational movement along axis Yc and / or steps on the road corresponding to rotation around axis Xc, sharp turns corresponding to yaw rotation around axis Yc, potholes corresponding to roll rotation around axis Zc, or biases resulting from sudden acceleration of the vehicle corresponding to pitch rotation around axis Xc.

[0072] The process according to the invention can be used in object detection mode and learning mode.

[0073] In the detection mode, the process includes a detection step, which includes supplying a plurality of corrected optical flow maps by a third normalization step to a neural network.

[0074] The detection mode makes it possible to improve the performance of a neural network for detecting objects in various situations, particularly in the aforementioned pitch, roll, yaw, and other situations.

[0075] In the learning mode, the process includes a learning step, which includes supplying a plurality of corrected optical flow maps by a third normalization step to a neural network to train the neural network.

[0076] The learning mode makes it possible to improve the performance of a neural network trained by the process according to the invention.

[0077] Of course, the present invention has been described by way of example. It should be understood that those skilled in the art can create various modified embodiments of the present invention without departing from the scope of the present invention thereby.

[0078] For example, it is possible to compensate for and correct a large inclination of the image acquisition device 14 in a specific vehicle with a certain height such as a truck, or to correct a mounting error of the image acquisition device 14 with respect to its theoretical reference position.

[0079] For this purpose, an appropriate transformation is applied between the reference frame Rc of the camera, which is a reference frame attached to the image acquisition device 14, and the reference reference frame corresponding to the reference position of the image acquisition device 14 after the second step of the process.

[0080] This transformation includes, for example, applying an additional rotation matrix to the initial first optical flow map C1 during the third step to correct a mounting error of the image acquisition device 14 with respect to its theoretical reference position.

Claims

1. In an improved process for detecting an object (16) by a neural network, the neural network is supplied, as input, with at least one initial first optical flow map (C1) representing a computer-based tracking of a moving object (16) in a scene by analyzing the difference in content between a first image taken by an image acquisition device (14) at a first position (P1) at a previous time and a successive second image taken by the image acquisition device (14) at a second position (P2) at the current time. At least, - A first step of estimating three translational parameters and three rotational parameters of the egomotion of the image acquisition device (14) between the first position (P1) and the second position (P2), wherein the three translational parameters include two secondary translational parameters and one main translational parameter along the longitudinal axis (Zc) corresponding to the main movement of the image acquisition device (14), - A second step including estimating a first rotation matrix enabling movement from the first position (P1) to a first virtual corrected position (P3) and a second rotation matrix enabling movement from the second position (P2) to a second virtual corrected position (P4) such that the two secondary translational parameters and the three rotational parameters enabling passage from the first corrected position (P3) to the second corrected position (P4) are equal to zero, - A third step including calculating a corrected second optical flow map (C2) representing a computer-based tracking of the moving object (16) between the first corrected position (P3) and the second corrected position (P4), obtained by applying the first rotation matrix and the second rotation matrix to the initial first optical flow map (C1). An improved process for detecting an object (16) by a neural network, characterized by comprising the above steps.

2. An improved process for detecting an object (16) by a neural network according to Claim 1, characterized by comprising a learning step including supplying a plurality of corrected optical flow maps (C2) obtained by the third normalization step of the process to the neural network to train the neural network.

3. An improved process for object (16) detection by a neural network according to claim 1 or 2, comprising a detection step including supplying a plurality of corrected optical flow maps (C2) by the third normalization step to the neural network to perform object detection.

4. A computer (12) configured to execute the process according to any one of claims 1 to 3.

5. An automobile (10) including the computer (12) according to claim 4 and at least one image acquisition device (14) integral with the automobile (10) in terms of movement, the automobile (10) moving mainly along the longitudinal axis (Zc).