Image processing method and method for predicting collision

By using trained machine learning models to generate unrotated parallax maps in automotive applications, the problem of stereo camera calibration failure is solved, and accurate depth information acquisition and 3D point cloud construction in the case of camera error is achieved, supporting object detection and collision prevention.

CN120092261APending Publication Date: 2025-06-03CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073921.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-17
Filing Date
2023-10-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In automotive applications, the calibration of the stereo camera is prone to failure due to the vibration or regular movement of the vehicle, resulting in the inability to accurately obtain the depth information required by the three-dimensional model, which in turn affects object detection and collision prevention of the surrounding environment of the vehicle.

Method used

The problem of stereo camera calibration failure is solved by inputting the image set to the trained machine learning model, a rotation disparity map is generated, and the rotation angle and unrotated disparity of each pixel is determined based on the diagram, and an unrotated disparity map is determined.

Benefits of technology

It realizes that the point depth of the image center can be accurately determined in the case of camera error, and an accurate 3D point cloud can be built to support object detection and collision prevention of the vehicle's surrounding environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092261A_ABST
    Figure CN120092261A_ABST
Patent Text Reader

Abstract

According to various embodiments, a computer-implemented image processing method may include inputting a stereoscopic image set to a machine learning model. The stereoscopic image set may include a first image and a second image. The image processing method may further include generating a rotational disparity map of the stereoscopic image set using the machine learning model. The image processing method may further include, for each pixel in the first image, identifying a corresponding pixel in the second image based on the rotational disparity map, determining a respective rotational angle between the pixel in the first image and the corresponding pixel in the second image relative to a common image center, and determining a respective non-rotated disparity based on the respective rotation angle. The image processing method may further include generating an unrotated disparity map based on a respective unrotated disparity for each pixel in the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments relate to methods and devices for processing stereoscopic images, methods and devices for predicting collisions, and methods for training machine learning models. Background Art

[0002] Stereoscopic vision technology typically uses two cameras to view the same object. The two cameras can be spaced apart by a baseline distance. Assume that the baseline distance is accurately known. The two cameras can simultaneously capture two images, also referred to as stereoscopic images. The stereoscopic images can be analyzed to identify the differences between the two images. The difference between the two images can be referred to as parallax. The parallax between the two images can be used to determine the depth of points in the image. The depth information can be used to project points in a three-dimensional model. Such a three-dimensional (3D) model can be used to facilitate various driving assistance or autonomous driving functions. For example, the 3D model can provide information about the positions of various objects near the vehicle, thus assisting in navigation or obstacle avoidance.

[0003] For automotive applications, a stereoscopic camera can be mounted on a vehicle to capture images of the vehicle's surrounding environment. The stereoscopic camera needs to be precisely calibrated so that the parallax obtained from the images is accurate. However, regular movements of the vehicle, such as crossing a pothole or a road hump on the road, may cause vibrations that displace the stereoscopic camera, thereby invalidating the calibration of the stereoscopic camera. When this occurs, it is not possible to rely on the stereoscopic images to generate the accurate depth information required for the 3D model. Therefore, the 3D model will not be suitable for use as a means for detecting objects present in the vehicle's surrounding environment or for preventing collisions with objects. Summary of the Invention

[0004] According to various embodiments, a computer-implemented image processing method is provided. The image processing method may include inputting an image set into a machine learning model. The image set may include a first image and a second image. The image processing method may further include: generating a rotated parallax map of the image set using the trained machine learning model. The image processing method may further include: for each pixel in the first image, determining, based on the rotated parallax map, the rotation angle of the pixel in the first image relative to its corresponding position in the second image with respect to a common image center; and determining a corresponding unrotated parallax based on the respective rotation angle. The image processing method may further include: generating an unrotated parallax map based on the respective unrotated parallax of each pixel in the first image.

[0005] According to various embodiments, an image processing device is provided, the image processing device including a processor. The processor may be configured to execute the above-described image processing method.

[0006] According to various embodiments, a computer-implemented method for predicting a collision is provided. The method may include inputting a sequence of image sets into a motion detection model. Each image set in the sequence may include a first image and a second image. The method may further include: for each image set, determining a corresponding depth map by the motion detection model based on the first image and the second image in the stereo image set, thereby generating a sequence of depth maps. The method may further include: determining an optical flow by the motion detection model based on at least one of the first image from the sequence of image sets and the second image from the sequence of stereo image sets. The method may further include: determining the motion of an object in the image set by the motion detection model based on the optical flow. The method may further include: determining a collision time with the object based on the sequence of depth maps and the determined motion of the object.

[0007] According to various embodiments, a device for predicting a collision is provided. The device may include a processor configured to execute the above-described method for predicting a collision.

[0008] Additional features of advantageous embodiments are provided in the dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the drawings, like reference numerals generally refer to like parts throughout the different views. The drawings are not necessarily to scale, but generally focus on showing the principles of the invention. In the following description, various embodiments are described with reference to the following drawings, in which:

[0010] Figure 1 A block diagram of a computer-implemented method for training a machine learning model to determine image disparity according to various embodiments is shown.

[0011] Figure 2 A block diagram of a computer-implemented method for generating a 3D model using stereo images according to various embodiments is shown.

[0012] Figure 3 An example of an image set according to various embodiments is shown.

[0013] Figure 4 An example of how the positions of feature points can be different in a left image and a right image is shown.

[0014] Figure 5A An example of an image set according to various embodiments is shown.

[0015] Figure 5B Disparity images generated by a prior art machine learning model and a trained machine learning model according to various embodiments are shown.

[0016] Figure 6Shows an image processing method according to various embodiments.

[0017] Figure 7A Shows a simplified block diagram of an image processing apparatus according to various embodiments.

[0018] Figure 7B Shows the operations of an image processing apparatus according to various embodiments.

[0019] Figure 8A Shows a simplified functional block diagram of a device for predicting a collision according to various embodiments.

[0020] Figure 8B Shows a simplified hardware block diagram of a device according to various embodiments.

[0021] Figure 9 Shows a schematic diagram of a device that performs a method for predicting a collision according to various embodiments.

[0022] Figure 10 Shows a schematic diagram of a neural network according to various embodiments.

[0023] Figure 11 Shows a schematic diagram of a neural network according to various embodiments, the neural network receiving an input different from the input in Figure 10 and the input in.

[0024] Figure 12 Shows a flowchart of a computer-implemented method for predicting a collision according to various embodiments. Detailed Description

[0025] The embodiments described below in the context of a device are similarly valid for the corresponding methods, and vice versa. In addition, it should be understood that the embodiments described below can be combined, for example, a part of one embodiment can be combined with a part of another embodiment.

[0026] It should be understood that any property described herein for a particular device can also apply to any device described herein. It should be understood that any property described herein for a particular method can also apply to any method described herein. In addition, it should be understood that for any device or method described herein, not all of the described components or steps necessarily have to be included in that device or method, but rather some (but not all) of the components or steps can be included.

[0027] The term "coupled" (or "connected") herein can be understood as electrically coupled or mechanically coupled (e.g., attached or fixed), or merely in contact without any fixation, and it should be understood that both direct coupling and indirect coupling (in other words: coupling without direct contact) can be provided.

[0028] In this context, a device as described in this specification may include a memory, which is used, for example, in a process executed in the device. The memory used in an embodiment may be a volatile memory such as DRAM (Dynamic Random Access Memory), or a non-volatile memory such as PROM (Programmable Read-Only Memory), EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or flash memory such as floating-gate memory, charge-trapping memory, MRAM (Magnetoresistive Random Access Memory), or PCRAM (Phase-Change Random Access Memory).

[0029] To facilitate understanding of the present invention and to put it into practice, various embodiments will now be described by way of example and not limitation with reference to the accompanying drawings.

[0030] Figure 1 A block diagram of a computer-implemented method 100 for training a machine learning model 102 to determine image disparity in accordance with various embodiments is shown. The method of training the machine learning model 102 may include providing training image data 112 to the machine learning model 102. The training image data 112 may include left and right stereo images from an uncalibrated stereo camera. In other words, the training image data 112 may include uncalibrated stereo images. Ground truth disparity data 114 corresponding to each pair of left and right stereo images of the training image data 112 may also be provided to the machine learning model 102 as a training signal. The machine learning model 102 may be configured to extract features from the training image data 112 and may learn patterns between the extracted features and the ground truth disparity data 114. As a result of training using the training image data 112 and the ground truth disparity data 114, the machine learning model 102 learns to compute a disparity map for each pair of left and right stereo images. The resulting machine learning model is referred to herein as the trained machine learning model 104. The trained machine learning model 104 may include a neural network 1000 as Figure 10 further described.

[0031] In a general stereo camera setup, feature points of the left stereo image and corresponding feature points of the right stereo image are aligned on the same x-axis (also referred to as the horizontal axis). Due to vibration or mechanical setup, one of the stereo cameras in a stereo camera may be misaligned with the other stereo camera due to at least one of the following factors: roll, i.e., the rotation angle (a), pitch, yaw, and translation (x t , y t ), where x t refers to the movement of an image pixel along the horizontal axis, and y tRefers to the movement of image pixels along the vertical axis. The trained machine learning model 104 can be configured to determine the misaligned rotational parallax as described above because the training image data 112 includes uncalibrated stereo images. The method 100 can further include: rotating the stereo images to artificially invalidate the calibration of the training image data 112.

[0032] According to an embodiment that can be combined with any of the above embodiments or any of the embodiments further described below, the training image data 112 can include images obtained from a plurality of cameras that may not be stereo cameras. These images can capture a similar scene from different angles and can thus be similarly used to determine the depth of an object in the scene, just like stereo images.

[0033] According to an embodiment that can be combined with any of the above embodiments or any of the embodiments further described below, the camera or stereo camera that captures the training image data 112 can be coupled to a vehicle. The training image data 112 can capture the scene around the vehicle.

[0034] According to an embodiment that can be combined with any of the above embodiments or any of the embodiments further described below, the method 100 can further include: classifying the disparity range into close and far while training the machine learning model 102. For example, the ground truth disparity data 114 can be divided into close ground truth disparity data and far ground truth disparity data. The machine learning model 102 can be trained to use the close ground truth disparity data as a training signal to calculate the close ground truth disparity. The machine learning model 102 can be further trained to use the far ground truth disparity data as a training signal to calculate the far ground truth disparity in a process separate from training with the close ground truth disparity. By training the machine learning model 102 separately for close and far disparities, the accuracy of the trained machine learning model 104 can be improved. The time taken to train the machine learning model 102 can also be reduced.

[0035] Figure 2FIG. 0 shows a block diagram of a computer-implemented method 200 for generating a 3D model using stereo images according to various embodiments. The method 200 may include providing image data to a trained machine learning model 104. The image data may include an image set 212. The image set 212 may include a left stereo image and a right stereo image. The trained machine learning model 104 may extract feature information from the left and right stereo images and may generate a rotational disparity map 214 based on the extracted feature information. A rotation correction module 202 may receive the rotational disparity map 214 and correct the rotational disparity map 214 for a relative rotation between the images in the image set 212. The rotation correction module 202 may generate an unrotated disparity map 216 based on the rotational disparity map 214. The rotation correction module 202 may provide the unrotated disparity map 216 to a 3D reconstruction module 206. The 3D reconstruction module 206 may generate a 3D point cloud 208 based on the unrotated disparity map 216, camera parameters 204, and the image set 212. The rotational disparity map 214 may include an actual disparity for each pixel between an uncalibrated left stereo image and a right stereo image. The unrotated disparity map 214 may include a corrected disparity for each pixel between an uncalibrated left stereo image and a right stereo image as if the misaligned stereo images have been corrected back to their calibrated positions.

[0036] Compared to lidar sensors, the method 200 may be able to generate a 3D point cloud with more accurate and lower-cost depth measurements. The method 200 may also be able to determine the depth in a scene captured in an image regardless of dynamic misalignment of the camera (including rotation, horizontal shift, and vertical shift). Even if the image set 212 includes distorted images and modified intrinsic parameters, and even if the camera parameters are inaccurate (e.g., incorrect focal lens information), the method 200 may be able to generate a 3D point cloud.

[0037] According to an embodiment that may be combined with any of the above embodiments or any of the embodiments further described below, the image set 212 may include images obtained from a plurality of cameras that may not be stereo cameras. These images may capture a similar scene from different angles and may thus be similarly used to determine the depth of an object in the scene, just like stereo images.

[0038] According to an embodiment that may be combined with any of the above embodiments or any of the embodiments further described below, the camera or stereo camera that captures the image set 212 may be coupled to a vehicle. The image set 212 may capture a scene around the vehicle.

[0039] Figure 3Shows an example of an image set 212 according to various embodiments. The image set 212 may include a left image 302 and a right image 304. The image set 212 may be uncalibrated. The right image 304 is rotated relative to the left image 302. This may occur when at least one of the stereo cameras is not properly installed or displaced from its original position. For example, when a vehicle drives over a road bump or makes a sharp movement, the stereo cameras may experience vibrations that cause a slight displacement. The rotation correction module 202 may determine an unrotated disparity map 216 based on the rotated disparity map 214 such that the unrotated disparity map 216 provides information about the displacement between each pixel in the left image 302 and its corresponding right image 304 as if the left image 302 and the right image 304 were aligned on the horizontal axis. Regarding Figure 4 Describes the process of determining the unrotated disparity map 216.

[0040] Figure 4 Shows an example of how the positions of the feature points may be different in the left image 302 and the right image 304. In this example, the image center 402 may be considered (0,0). The feature point 404 in the left image 302 is represented as (x 2 , y 2 ). The rotated corresponding position 406 of the feature point 404 in the right image 304 is represented as (x' 2 , y' 2 ). The line connecting the corresponding position 406 and the image center 402 may define the first hypotenuse 420 of a right triangle 410, indicated herein as "h 1 ". The corresponding position 406 of the feature point 404 in the right image 304 may be determined based on the rotated disparity map 214.

[0041] The rotated disparity map may include the rotated disparity values (dx 2 , dy 2 ) of the feature point 404. The corresponding position 406 and the first hypotenuse 420 may be determined based on the feature point 404 and the rotated disparity values according to the following equations (1) to (3). x' 2 = x 2 + dx 2 Equation (1) y' 2 = y 2 + dy 2 Equation (2)

[0042] The unrotated disparity values of the feature point 404 may be represented as (dx 2 , dy 2)。When the right image 304 is adjusted to have a zero rotation angle relative to the left image 302, the unrotated corresponding position 408 of the feature point 404 in the right image 304 is represented as (x 2r , y 2r ).

[0043] The corresponding position 406, the unrotated corresponding position 408, and the image center 402 can define the vertices of an isosceles triangle 422. The distance between the corresponding position 406 and the image center 402 can be at least substantially equal to the distance between the unrotated corresponding position 408 and the image center 402, in other words, equal to the first hypotenuse 420. The vertex angle 426 of the isosceles triangle 422 is represented as a. The vertex angle 426 can also be referred to as the rotation angle 426 herein.

[0044] The line connecting the corresponding position 406 and the unrotated corresponding position 408 can define the second hypotenuse 424 of another right triangle 412, represented as "h 2 " in this text. The second hypotenuse 424 can be the base of the isosceles triangle 422. The second hypotenuse 424 can be determined according to equation (4).

[0045] The base 428 of the another right triangle 412 is represented as d. The length of the base 428 can be determined according to equation (5).

[0046] The unrotated parallax d 2 can be determined based on the base 428 and the rotated parallax dx 2 . The unrotated parallax can be determined according to the following equation (6). d 2 = d + dx 2 Equation (6)

[0047] The rotation correction module 202 may determine the unrotated disparity map 216 based on the above calculations. The rotation correction module 202 may determine the corresponding position 406 based on the rotated disparity values in the rotated disparity map 214. The rotation correction module 202 may determine the first hypotenuse 420 based on the corresponding position 406. The rotation correction module 202 may determine the vertex angle 426. Determining the vertex angle 426 may include determining the direction of the epipolar line of the image set 212 in the image plane without using 3D space. Then, the vertex angle 426 may be obtained by projecting the rotated disparity measurement results toward the direction of the epipolar line. The rotation correction module 202 may determine the second hypotenuse 424 based on the vertex angle 426 and the first hypotenuse 420. The rotation correction module 202 may determine the base 428 based on the second hypotenuse 424 and the rotated disparity values. The rotation correction module 202 may determine the unrotated disparity values based on the base 428 and the rotated disparity values.

[0048] According to an embodiment that may be combined with any of the above embodiments or with any of the embodiments further described below, the rotation correction module 202 may be configured to determine the direction of the epipolar line of the image set 212 in the image plane without using 3D space. The rotation correction module 202 may determine the epipolar line direction without prior knowledge of the camera internal parameters and the camera external parameters. The rotation correction module 202 may determine the epipolar line direction by processing the images in the image set 212 region by region. In other words, the image may be segmented into multiple regions, and the epipolar line direction may be determined in each of the multiple regions.

[0049] The rotation correction module 202 may be configured to check for infinity-negative disparity in the vertical disparity. The true distance represented in the image may be calculated by projecting the disparity in the direction of the epipolar line. The actual distance may be determined region by region, in other words, in each of the multiple regions.

[0050] According to an embodiment that may be combined with any of the above embodiments or with any of the embodiments further described below, the rotation correction module 202 is further configured to determine the translation (x t , y t ), and is further configured to correct the right image relative to the left image based on the determined translation. The horizontal translation x t may be calculated using the tracking and filtering of the input images over time. The vertical translation yt may be determined based on the pixel information of the y-disparity center in the y-disparity array.

[0051] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the 3D reconstruction module 206 can be configured to generate an unscaled 3D point cloud based on the unrotated disparity map 216. The trained machine learning model 104 can receive a sequence of the image sets 212 over time to generate a sequence of the rotated disparity maps 214. Thus, the rotation correction module 202 can generate a sequence of the unrotated disparity maps. The 3D reconstruction module 206 can correct the 3D point cloud for camera factors such as scale and yaw angle based on comparing at least two consecutive 3D point clouds. Points in the 3D point cloud should move at least at substantially the same speed because these points move relative to the vehicle to which the camera is coupled. In other words, the points in the 3D point cloud can move at a speed that is at least substantially equal to the longitudinal speed of the vehicle. The 3D reconstruction module 206 can correct the 3D point cloud for camera factors based on the speed deviation of the points in the 3D point cloud.

[0052] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the method 200 can further include: correcting negative disparity in the image set 212. The correction of the negative disparity can be performed before performing 3D reconstruction. The negative disparity may be caused by a yaw angle error between the cameras that capture the image set 212. A change in the yaw angle of at least one camera may result in a horizontal shift. In a stereo image, the horizontal disparity decreases as the distance (i.e., depth) increases. Thus, the horizontal disparity of an object at infinity (also referred to herein as an infinite object) should be at least substantially zero. However, when there is a negative yaw angle between the cameras, the infinite object will show a negative horizontal disparity instead of zero disparity. The method 200 can further include: applying a low-pass filter to the unrotated disparity map 216 to obtain the maximum negative disparity value in the unrotated disparity map 216, i.e., the negative disparity value with the largest magnitude. The method 200 can include using the maximum negative disparity value to correct the disparity in the entire image region. The correction of the disparity can be performed region by region, where each image can be divided into multiple regions.

[0053] Figure 5A An example of the image set 212 according to various embodiments is shown. In this example, the image set 212 includes a first image 502 and a second image 504. The first image 502 is taken by a first camera, while the second image 504 is taken by a second camera. Both the first camera and the second camera capture images of the same scene from slightly offset positions. In other words, the second camera is positioned at a calibrated distance from the first camera. In the case where the second camera fails to be calibrated in terms of rotation, the second image 504 is rotated relative to the first image 502.

[0054] Figure 5BShows disparity images generated by a prior art machine learning model and a trained machine learning model 104 according to various embodiments. The disparity images include a first disparity image 510 and a second disparity image 520, both of which are generated by the machine learning model based on Figure 5A an example image set 212. The first disparity image 510 is generated by a prior art machine learning model. The prior art machine learning model fails to handle the misalignment of the second image 504, and thus, the objects in the image set 212 are not visible in the first disparity image 510. In contrast, the second disparity image 520 clearly shows the objects in the image set 212, indicating that despite the misalignment of the second image 504, the trained machine learning model 104 is still able to correctly match features between the first image 502 and the second image 504.

[0055] Figure 6 Shows an image processing method 600 according to various embodiments. The image processing method 600 may include processes 602, 604, 606, 608, and 610. Process 602 may include inputting the image set 212 into the trained machine learning model 104. The image set 212 may include a first image and a second image. Process 604 may include using the trained machine learning model 104 to generate a rotated disparity map 214 of the image set 212. Process 606 may include, for each pixel in the first image, determining a rotation angle of the pixel in the first image relative to a common image center with respect to its corresponding position in the second image based on the rotated disparity map 214. Process 608 may include, for each pixel in the first image, determining a corresponding unrotated disparity based on the corresponding rotation angle. Process 610 may include generating an unrotated disparity map 216 based on the corresponding unrotated disparities of each pixel in the first image. The image processing method 600 may determine an unrotated disparity map, which may be used to accurately determine the depth of points in the image set 212, for example, to construct an accurate 3D point cloud without having to calibrate the cameras that captured the image set 212. These cameras may be mounted on a vehicle to capture the surrounding environment of the vehicle. The cameras mounted on the vehicle may be displaced from their initial calibrated positions due to vibrations or movement of the vehicle. Using the image processing method 600, the depth of points in the image set 212 can be accurately determined without being affected by camera misalignment.

[0056] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image set 212 can include stereoscopic images. The first image can be a left stereoscopic image, and the second image can be a right stereoscopic image. The stereoscopic image set can capture a similar scene from slightly different positions such that the images within the stereoscopic image set can be compared to obtain depth information about each point in the images. This depth information can be used to provide situational awareness to the vehicle, such as indicating the distance of the vehicle to various objects in its surroundings.

[0057] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the trained machine learning model 104 can be configured to determine both the horizontal disparity and the vertical disparity of the image set 212 such that the rotational disparity map 214 can include both the horizontal disparity values and the vertical disparity values of the image set. By determining both the horizontal disparity and the vertical disparity of the image set 212, the trained machine learning model 104 can be capable of handling misalignment of the cameras in both the vertical and horizontal directions. In this way, even if at least one of the cameras rotates or is displaced out of position in both directions, the image processing method 600 can produce an accurate disparity map.

[0058] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the common image center can be any point in the image and can be close to the image center of at least one of the first image and the second image.

[0059] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the center of the first image can be used as the common image center. The first image can be used as a reference image, and the translation of the second image relative to the first image can be estimated.

[0060] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing method 600 can further include: dividing the first image into a plurality of first regions, and dividing the second image into a plurality of second regions corresponding to the plurality of first regions. The image processing method 600 can further include: for each of the plurality of first regions, determining an epipolar line and the direction of the epipolar line based on the pixels in the first image and the corresponding positions of the pixels in the second image.

[0061] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing method 600 can further include: for each pixel of the first image, determining a corresponding horizontal disparity based on the projection of the corresponding unrotated disparity through the direction of the epipolar line. Thereby, the horizontal disparity in the image set 212 caused by, for example, misalignment of at least one camera can be corrected.

[0062] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing method 600 may further include: generating a 3D point cloud 208 based on the image set 212 and further based on the unrotated disparity map 216. The 3D point cloud may provide detailed information about the surrounding environment of the vehicle and may be used to assist various functions of the vehicle, such as navigation, positioning, and obstacle avoidance.

[0063] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing method 600 may further include: inputting another image set into the trained machine learning model 104. The another image set may include another first image and another second image captured at time frames consecutive to the first image and the second image in the image set 212. The image processing method 600 may further include: using the trained machine learning model 104 to generate another rotated disparity map of the another image set. The image processing method 600 may further include: for each pixel in the another first image, identifying a corresponding pixel in the another second image based on the another rotated disparity map; determining a corresponding rotation angle of the pixel in the another first image and the corresponding pixel in the another second image with respect to another common image center; and determining a corresponding unrotated disparity based on the corresponding rotation angle. The image processing method 600 may further include: generating another unrotated disparity map based on the corresponding unrotated disparity of each pixel in the another first image. The image processing method 600 may further include: generating another 3D point cloud based on the another image set and further based on the another unrotated disparity map; determining the distance that each point in the 3D point cloud moves from the 3D point cloud 208 to the another 3D point cloud; and correcting at least one of the 3D point cloud 208 and the another 3D point cloud based on the determined distance. The 3D reconstruction module 206 may perform the above correction on the 3D point cloud 208 or the another 3D point cloud. The points in the 3D point cloud should move at least at substantially the same speed because these points move relative to the vehicle to which the camera is coupled. The 3D reconstruction module 206 may correct the 3D point cloud for camera factors such as scale and yaw angle based on the speed deviation of the points in the 3D point cloud.

[0064] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the trained machine learning model 104 may be trained using a training data set 112 including a plurality of image pairs generated by rotating a pair of calibrated images to a corresponding plurality of different angles, and using a ground truth disparity map 114 generated based on the pair of calibrated images as a training signal. When generating the training data set 112 using the calibrated images, an accurate ground truth disparity map 114 can be obtained.

[0065] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the trained machine learning model 104 can be trained to determine near parallax and further trained separately to determine far parallax. Doing so can optimize the training process, thereby reducing the required training time and improving the accuracy of the trained machine learning model 104 in determining parallax.

[0066] Figure 7A A simplified block diagram of an image processing device 700 according to various embodiments is shown. The image processing device 700 may include a processor 702. The processor 702 may be configured to execute an image processing method 600.

[0067] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing device 700 may be any one of a server, a computer, a vehicle, or a robot. The image processing device 700 may be an autonomous vehicle (such as a self-driving car) or a drone.

[0068] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the image processing device 700 may further include at least one camera 704. The camera 704 may be configured to capture an image set 212. The image processing device 700 may further include an engine 706 configured to drive the image processing device 700. The image processing device 700 may further include a steering module 708 configured to steer the image processing device 700.

[0069] Figure 7B The operation of the image processing device 700 according to various embodiments is shown. The input 720 of the image processing device 700 may include a sequence of image sets. The sequence of image sets may include continuously captured images. For example, the sequence of image sets may include an image set 720a captured at t - 1 and another image set 720b captured at t, where t represents time as a variable. In other words, the image set 720b may be a subsequent frame of the image set 720a. Each image set may include a plurality of images, each captured at a corresponding location. These locations may be offset from each other such that the combination of the plurality of images may provide depth information of an object shown in the image. For example, each of the image sets 720a and 720b may include a pair of stereo images. The image set 720a may include a left image 722a and a right image 724a. The image set 720b may include a left image 722b and a right image 724b.

[0070] A sequence of image sets 720a, 720b can be provided to a trained machine learning model 104. The trained machine learning model 104 can be trained to compute a corresponding rotational disparity map 214 for each image set. The trained machine learning model 104 can be trained to identify features in the images and further configured to determine the spatial offset of the features between the images (also referred to herein as disparity data). The spatial offset between the images within the same image set can provide depth information of the features. The rotation correction module 202 can compute a corresponding unrotated disparity map 216 for each rotational disparity map 214, e.g., as described with respect to Figure 4 as described.

[0071] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the device 700 can further include a point cloud generator 726. The point cloud generator 726 can include a 3D reconstruction module 206. The point cloud generator 726 can be configured to generate a 3D point cloud 208 based on the input 720 and the unrotated disparity map 216. The generated 3D point cloud 208 can be a 3D reconstruction of the environment in which the vehicle is traveling. Since the 3D point cloud 250 can be generated in real time, the device 100 can provide environmental data that was previously unavailable to the vehicle, such as previously uncharted terrain. Further, the 3D point cloud 208 can indicate the presence of dynamic objects such as other traffic participants to the vehicle. The 3D point cloud 208 can be further generated based on camera intrinsics (such as the focal length and axis skew offset from the camera principal point along the x, y The point cloud generator 726 can also improve the 3D point cloud 208 to remove outliers and incorrect predicted values such that the 3D point cloud 208 can be used as an accurate dense 3D map. The point cloud generator 726 can remove outliers and incorrect predicted values based on the probabilities of those data points.

[0072] The various aspects described with respect to the method 200 and the image processing method 600 can be applicable to the image processing device 700.

[0073] Collision prediction plays an important role in improving traffic safety. Most past studies have focused on determining the time to collision (TTC) during the day when oncoming vehicles can be seen. However, due to less visual information of vehicles and complex lighting conditions, traffic accidents are more frequent at night. At night, the visual feature information is almost reduced to the headlights and taillights of vehicles. Traditional collision prediction systems rely on the width and motion information of vehicles, which is difficult to evaluate in insufficient light. Existing deep learning models trained for night vehicle light recognition usually cannot generate reliable TTC values based on pitch-black images. In advanced driver assistance systems, several algorithms rely only on these features for vehicle detection. TTC is the most commonly used safety metric, and it indicates the time it takes for a vehicle to collide with another vehicle.

[0074] According to various embodiments, a computer-implemented method 1200 for predicting a collision is provided. Method 1200 may estimate the TTC based on stereo disparity information and optical flow. The stereo disparity information may be obtained using an image processing method 600. Also regarding Figure 12 Method 1200 is described.

[0075] Method 1200 may include receiving a sequence of image sets. The sequence of image sets in input 720 may be captured by sensors mounted on a vehicle. Examples of suitable sensors include stereo vision cameras, thermal stereo cameras, or dense lidars. A lidar sensor may provide a dense point cloud of the surrounding environment and the distance to each detected object in the scene. Each image set may include a plurality of images, and each image in the plurality of images may be captured from a different position on the vehicle such that the plurality of images in each image set may be offset from each other along at least one axis. The plurality of images may be captured by corresponding pluralities of sensors respectively. The plurality of images may include a first image and a second image.

[0076] In an example, the sensors may be a pair of stereo cameras. Thus, the image set includes a pair of stereo images. The two cameras may be separated by a baseline, assuming its distance is accurately known. The two cameras may capture two consecutive images simultaneously. The images may be analyzed to identify the differences between the images, and the positions of corresponding pixels in the two images may be identified during a stereo matching process. The disparity between the corresponding pixels in the two images may be used to estimate the depth. In an example, the pair of stereo cameras may be mounted on the side mirrors of the vehicle. The stereo cameras may be FSC231 stereo cameras with an 8.3M pixel resolution and a 30° field of view.

[0077] Method 1200 may include using a motion detection module including a machine learning model to determine the disparity between the images in the image set. The machine learning model may be, for example, a trained machine learning model 104. The machine learning model may include a neural network, such as a graph neural network, a transformer, or a recurrent neural network. The machine learning model may include the neural network 1000 Figure 10 described. Using the estimated disparity, the distance and TTC of each given pixel or object in the scene captured by the images may be calculated. The distance may be measured in meters, while the TTC may be measured in seconds. The extracted features for performing stereo matching may be used to estimate the TTC of each pixel in the image using a deep neural network.

[0078] Method 1200 may include identifying a moving object by identifying moving pixels in a sequence of image sets. In other words, identifying a moving object may involve finding corresponding pixels in a sequence of image sets over time. The identification of the moving object may be achieved by passing two consecutive stereo image pairs through an optical flow estimation algorithm.

[0079] Method 1200 may include using a neural network to combine the resulting flow information with the current stereo frame to segment the moving objects in the scene. A point cloud generator may generate a 3D point cloud of the vehicle's surroundings based on the estimated depth (i.e., the distance of each pixel).

[0080] According to various embodiments, method 1200 may include providing stereo camera data to a motion detection module. The motion detection module may include a machine learning model, also referred to herein as a motion detection network. The motion detection network may include a trained machine learning model 104. The motion detection network may include, for example, a convolutional neural network (CNN). The stereo camera data may include a first stereo image pair and a second stereo image pair. The first stereo image pair and the second stereo image pair may be captured sequentially. The motion detection network may be trained using the stereo camera data, ground truth disparity, and ground truth TTC values. The stereo camera data may be images captured at night such that the motion detection network can learn to compute disparity values and TTC values for night images. The motion detection network may also be trained with images associated with other environmental conditions such as daylight, rain, or snow. The motion detection network may be trained with uncalibrated stereo images and may generate a rotated disparity map of the stereo camera data. The motion detection module may include a rotation correction module 202 that converts the rotated disparity map to an unrotated disparity map. The motion detection network may perform stereo matching by extracting feature information from the left and right stereo images and mapping each image to a dense disparity map. The motion detection network may be trained to determine the rotated disparity, optical flow, and TTC image output for two consecutive stereo image pairs. To segment the moving objects in the scene, the obtained flow information may be combined with the current stereo frame to train the motion detection network to predict a mask of the moving objects. The motion detection module may determine the relative depth of the moving objects based on the unrotated disparity map and may further determine a TTC image based on the determined relative depth. The TTC image may include the TTC value for each pixel in the image. The motion detection module may further perform trajectory planning based on the TTC estimation.

[0081] According to various embodiments, the disparity map, input image, and flow information may further be used to estimate the pose and perform simultaneous localization and mapping (SLAM) of the scene for vehicle navigation.

[0082] Figure 8AA simplified functional block diagram of a device 800 for predicting collisions according to various embodiments is shown. The device 800 may be configured to receive an input 720 and may be configured to generate an output. The input 720 may include a sequence of image sets. The output may include a predicted time to collision (TTC) 820 of the vehicle with another object or vehicle. The device 800 may be capable of determining the TTC 820 based on processing visual data (i.e., images). The device 800 may be configured to execute a method 1200.

[0083] The sequence of image sets in the input 720 may be captured by sensors mounted on the vehicle. Each image set may include a plurality of images, and each of the plurality of images may be captured from a different location on the vehicle such that the plurality of images in each image set may be offset from each other along at least one axis. The plurality of images may be captured by corresponding plurality of sensors respectively. The plurality of images may include a first image and a second image. In some embodiments, the first image and the second image may be a pair of stereo images.

[0084] The device 800 may include a motion detection module 802. The motion detection module 802 may include a machine learning model, e.g., a trained machine learning model 104. The motion detection module 802 may be configured to generate a corresponding depth map 812 of the image set based on the first image and the second image of each image set. Thus, the motion detection module 802 may output a sequence of depth maps 812 based on the received sequence of image sets. Each depth map 812 may be an image or an image channel that contains information related to the distance of the surface of an object from the viewpoint. The viewpoint may be the vehicle, or more specifically, the sensors mounted on the vehicle.

[0085] The motion detection module 802 may also be configured to: determine an optical flow 814 based on at least one of the first image from the sequence of image sets and the second image from the sequence of image sets. The motion detection module 802 may determine the motion of an object in the image set based on the optical flow 814 to generate motion data 816.

[0086] The machine learning model in the motion detection module 802 may be used to determine both disparity and motion. The machine learning model may generate a segmentation mask of a moving object based on the detected motion.

[0087] The device 800 may further include a prediction module 804. The prediction module 804 may be configured to receive the sequence of depth maps 812 and the motion data 816 from the motion detection module 802. The prediction module 804 may be configured to generate a predicted TTC 820 based on the received sequence of depth maps 812 and the motion data 816.

[0088] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the motion detection module 802 can include the device 700. The motion detection module 802 can determine the depth map 812, for example, according to the image processing method 600 by generating a corresponding unrotated disparity map 216 for each image set.

[0089] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the device 800 can further include a point cloud generator 726 that generates a 3D point cloud 208 based on at least one image set in a sequence of image sets. The device 800 can further be configured to detect an object in the 3D point cloud 208.

[0090] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the point cloud generator 726 can generate a corresponding 3D point cloud based on each image set in a sequence of image sets, thereby generating a sequence of 3D point clouds 208. The device 800 can compare the distances by which each point in the sequence of 3D point clouds 208 moves, and can correct at least one 3D point cloud in the sequence of 3D point clouds 208 based on the determined distances.

[0091] Figure 8B A simplified hardware block diagram of the device 800 according to various embodiments is shown. The device 800 can include at least one processor 830. The at least one processor 830 can be configured to execute the functions of the machine learning model 802 and the prediction module 804.

[0092] According to various embodiments, the device 800 can be a driving assistance system.

[0093] According to various embodiments, the device 800 can be a vehicle.

[0094] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the device 800 can further include a plurality of sensors 832. The plurality of sensors 832 can be configured to generate a sequence of image sets 816. The plurality of sensors 832 can include, for example, a stereo camera, a surround view camera, an infrared camera, an event camera, or a dense lidar. The device 800 can include at least one memory 834. The at least one memory 834 can store the machine learning model 802 and the prediction module 804. The at least one memory 834 can include a non-transitory computer-readable medium. The at least one processor 830, the plurality of sensors 832, and the at least one memory 834 can be coupled to each other via a coupling wire 840, for example, mechanically or electrically.

[0095] Figure 9A schematic diagram of a device 800 that executes a method for predicting a collision according to various embodiments is shown. The motion detection module 802 of the device 800 may include a machine learning model that may generate a sequence of depth maps 812, an optical flow 814, and motion data 816 based on a sequence of received image sets. The sequence of image sets may include at least a first image set 720a and a second image set 720b. Each of the image sets may include at least a first image and a second image, such as a left image 722a and a right image 724a of the first image set 720a, and a left image 722b and a right image 724b of the second image set 720b. The motion detection module 802 may also receive a moving object mask 914, which may be generated by another machine learning model.

[0096] The motion detection module 802 may output a segmented moving object image 910 based on the moving object mask 914 applied to the sequence of image sets. The device 800 may also include a moving object tracking module 906 configured to determine the motion data 816 based on the segmented moving object image 910.

[0097] The prediction module 804 of the device 800 may include a relative depth estimation module 902. The relative depth estimation module 902 may be configured to receive the sequence of depth maps 812, the optical flow 814, and the motion data 816. The relative depth estimation module 902 may determine the relative depth of pixels in the image based on the received inputs. The prediction module 804 may further include a TTC determination module 904. The TTC determination module 904 may determine the TTC of each pixel based on the determined relative depth and the motion data 816.

[0098] Figure 10 A schematic diagram of a neural network 1000 according to various embodiments is shown. The neural network 1000 may be part of at least one of a trained machine learning model 104 and the motion detection module 802. The neural network 1000 may include a feature extraction network 1310 and a disparity calculation network 1320. The feature extraction network 1310 may be configured to extract features from images in the sequence of image sets to generate a feature map. The disparity calculation network 1320 may be configured to determine a two-dimensional offset (also referred to herein as displacement or disparity) between the images based on the generated feature map. Thus, the neural network 1000 may use a single common neural network set to determine both optical flow and depth information, since both optical flow and depth information are related to the two-dimensional offset between images. In this way, compared to training two separate neural networks, the neural network 1000 may be trained in less time with fewer resources.

[0099] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the neural network 1000 can be trained using scene flow stereo images via supervised training. Training the neural network 1000 may require, for example, 25,000 scene flow stereo images as training data. Stereo images from the Karlsruhe Institute of Technology and Toyota Technological Institute at Chicago (KITTI) dataset can be used to fine-tune the neural network 1000. As an example, about 400 stereo images from the KITTI dataset can be used.

[0100] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the feature extraction network 1310 can include a plurality of neural network branches. For example, the neural network branches can include a first branch 1350 and a second branch 1360. Each neural network branch can include a corresponding convolutional stack and a pooling module connected to the convolutional stack. The convolutional stack can include, for example, a CNN 1312. The pooling module can include, for example, an SPP module 1314. The plurality of neural network branches can share the same weights. By having the neural network branches share the same weights, the feature extraction network 1310 can be trained in less time with fewer resources compared to training the neural network branches separately.

[0101] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the disparity calculation network 1320 can include a 3D convolutional neural network (CNN) 1324 that is configured to generate three disparity maps based on the generated feature map 1318. The feature extraction network 1310 can extract features at different levels. To aggregate feature information along the disparity dimension as well as the spatial dimension, the 3D CNN 1324 can be configured to perform cost volume regularization on the extracted features. The 3D CNN 11324 can include an encoder-decoder architecture that includes a plurality of 3D convolutional layers and 3D deconvolutional layers with intermediate supervision. The 3D CNN 324 can have a stacked hourglass architecture that includes three hourglasses, thereby producing three disparity outputs. The architecture of the 3D CNN 1324 can enable it to generate accurate disparity outputs and predict optical flow. The filter size in the 3D CNN 1324 can be 3×3.

[0102] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the neural network 1000 can be configured to determine a corresponding depth map 812 for each image set based on the determined 2D offsets between multiple images of the image set. The depth map 812 can provide information about the distances between objects in the images from the vehicle. This information can improve the positioning accuracy. Each pixel in the depth map can be determined based on the following equation: Herein, the baseline refers to the horizontal distance between viewpoints that capture images within the same image set, the focal length refers to the distance between the lens and the image sensor, and the parallax refers to a 2D offset. The 2D offset can be obtained from the unrotated parallax map 216.

[0103] According to an embodiment that can be combined with any of the above embodiments or with any of the embodiments further described below, the neural network 1000 can be configured to determine the optical flow 814 based on a determined 2D offset between at least one of a first image in an adjacent image set and a second image in the adjacent image set. The optical flow 814 can provide information about the position change of an object in the image.

[0104] The neural network 1000 can be configured to perform a stereo matching operation. The stereo matching operation can include determining an estimated per-pixel displacement map between input images. The input images can include multiple images within the same image set. For example, when the input includes images captured by a stereo camera, the input images can include a left image 722b and a right image 724b.

[0105] Generally, a stereo image can be a rectified stereo image or an unrectified stereo image. A rectified stereo image is a stereo image in which the displacement of each pixel is constrained to a horizontal line. The displacement map can be referred to as parallax herein. To obtain a rectified stereo image, it is necessary to accurately calibrate the sensor or camera used to capture the images.

[0106] On the other hand, an unrectified stereo image may exhibit both vertical parallax and horizontal parallax. Vertical parallax can be defined as the vertical displacement between corresponding pixels in the left and right images. Horizontal parallax can be defined as the horizontal displacement between corresponding pixels in the left and right images. Obtaining a rectified stereo image from a sensor mounted on a vehicle is challenging because the sensor may shift over time or rotate in place due to the movement and vibration of the vehicle. Thus, the input images captured by a sensor mounted on a vehicle can be regarded as unrectified stereo images.

[0107] The neural network 1000 can be trained to be robust to rotation, shift, vibration, and distortion of the input images. The neural network 1000 can be trained to handle both horizontal (x-parallax) displacement and vertical (y-parallax) displacement. In other words, the neural network 1000 can be configured to determine both the horizontal parallax and the vertical parallax of the input images. This makes the neural network 1000 robust to vertical translation and rotation between cameras or sensors.

[0108] The neural network 1000 can have a two-branch neural network architecture, which includes a first branch 1350 and a second branch 1360, such that each branch can be configured to determine the disparity on the corresponding axis. Each of the feature extraction network 310 and the disparity calculation network 1320 can include components of the first branch 350 and the second branch 360.

[0109] For each of the first branch 11350 and the second branch 1360, the feature extraction network 1310 can include a convolutional neural network (CNN) 1312, a spatial pyramid pooling (SPP) module 1314, and a convolutional layer 1316. The CNN 1312 can extract feature information from the input image. The CNN 1312 can include three small convolutional filters with a kernel size of (3×3), which are cascaded to build a deeper network with the same receptive field. The CNN 1312 can include the conv1_x, conv2_x, conv3_x, and conv4_x layers that form a basic residual block for learning single feature extraction. For conv3_x and conv4_x, dilated convolutions can be applied to further expand the receptive field. The output feature map size can be (1 / 4×1 / 4) of the input image size. Then, the SPP module 1314 can be applied to collect context information from the output feature map. The SPP module 1314 can learn the relationship between an object and its sub-regions to combine hierarchical context information. The SPP module 1314 can include four fixed-size average pooling blocks with sizes of 64×64, 32×32, 16×16, and 8×8. The convolutional layer 1316 can be a 1×1 convolutional layer for reducing the feature dimension. The feature extraction network 1310 can upsample the feature map to the same size as the original feature map using bilinear interpolation. The size of the original feature map can be 1 / 4 of the input image size. The feature extraction network 1310 can concatenate different levels of feature maps extracted by various convolutional filters into a left SPP feature map 1318 and a right SPP feature map 1319.

[0110] The disparity calculation network 1320 can receive the left SPP feature map 1318 and the right SPP feature map 1319. The disparity calculation network 1320 can concatenate the left SPP feature map 1318 and the right SPP feature map 1319 into separate cost volumes 1322 for x and y displacements, respectively. Each cost volume can have four dimensions, i.e., height × width × disparity × feature size. The disparity calculation network 1320 can include 3D-CNNs 1322 in each branch. The 3D-CNN 1322 can include a stacked hourglass (encoder-decoder) architecture configured to generate three disparity maps. The disparity calculation network 1320 can further include an upsampling module 1326 and a regression module 1328 in each branch. The upsampling module 1326 can upsample the three disparity maps such that their resolutions match the resolution of the input image size. The regression module 1328 can apply regression to the upsampled disparity maps to calculate the output disparity map. The disparity calculation network 1320 can calculate the probability of each disparity based on the predicted cost via a SoftMax operation. The predicted disparity can be calculated as the sum of each disparity weighted by its probability. Next, a smooth loss function can be applied between the ground truth disparity and the predicted disparity. The smooth loss function can measure how close the predicted value is to the ground truth disparity value. The smooth loss function can be a combination of l1 and l2 losses. It is used in deep neural networks due to its robustness and low sensitivity to outliers. The disparity calculation network 1320 then outputs a horizontal displacement 1332 at the first branch 1350 and a vertical displacement 1334 at the second branch 1360. Then, the machine learning model 102 can use a known stereo calculation method (such as semi-global matching) to determine the depth map 112 based on the horizontal displacement 1332 and the vertical displacement 1334.

[0111] Figure 11 A schematic diagram of a neural network 1000 according to various embodiments is shown, which receives inputs different from those Figure 10 in the input. The neural network 1000 can also be configured to perform optical flow calculation. The optical flow calculation can include predicting a per-pixel displacement field such that for each pixel in a frame, the neural network 1000 can estimate its corresponding pixel in the next frame. For the optical flow calculation operation, the input images used by the neural network 1000 can be two consecutive images taken from the same location. In this example, the input images are the left image 722b captured at time = t and the left image 722a captured at time = t - 1. The output of the neural network 1000 for the optical flow calculation operation is an optical flow 814 including displacements in the x direction and the y direction.

[0112] The first branch 1350 can process the left image 722b, while the second branch 1360 can process the earlier left image 722a. Similar to that regarding Figure 10For the described stereo matching operation, the optical flow calculation operation may include using the CNN 1312 of the feature extraction network 1310 to extract features. The SPP module 1314 may collect context information from the output feature maps generated by the CNN 11312. The feature extraction network 1310 generates a final SPP feature map that is provided to the disparity calculation network 1320. The difference from the stereo matching operation is that the generated SPP feature maps are the earlier SPP feature map (for time = t - 1) and the subsequent SPP feature map (for time = t). The disparity calculation network 1320 may concatenate the SPP feature maps into separate cost volumes 1322 for time = t - 1 and time = t, respectively. Each cost volume may have four dimensions, i.e., height × width × disparity × feature size. The 3D-CNN 1322 of each branch may generate three disparity maps based on the corresponding cost volume 1322. The upsampling module 1326 may upsample the three disparity maps. The regression module 1328 may apply regression to the upsampled disparity maps to calculate the output disparity map. The disparity calculation network 1320 may calculate the probability of each disparity based on the predicted cost via a SoftMax operation. The predicted disparity may be calculated as the sum of each disparity weighted by its probability. Next, a smooth loss function may be applied between the ground truth disparity and the predicted disparity. The disparity calculation network 320 then outputs the x-direction displacement and y-direction displacement determined to occur between time = t - 1 and time = t.

[0113] According to various embodiments, suitable deep learning models for the machine learning model 102 may include, for example, pyramid stereo matching and RAFTNet.

[0114] According to an embodiment that may be combined with any of the above embodiments or with any of the embodiments further described below, the training of the neural network 1000 may be performed, for example, according to gradient descent based on standard backpropagation. As an example of how the neural network 1000 may be trained, a training data set may be provided to the neural network 1000, and the following training process may be performed:

[0115] An example of a suitable training data set for training the neural network 1000 may be the KITTI data set.

[0116] Before training the neural network 1000, the weights may be randomly initialized to a number between 0.01 and 0.1, while the biases may be randomly initialized to a number between 0.1 and 0.9.

[0117] Subsequently, the first observation of the data set may be loaded into the input layer of the neural network, and (multiple) output values may be generated by the forward propagation of the input values of the input layer. Thereafter, the following loss function may be used to calculate the loss using the (multiple) output values:

[0118] Mean Squared Error (MSE): where n represents the number of neurons in the output layer, and y represents the true output value, and represents the predicted output. In other words, represents the difference between the actual output and the predicted output.

[0119] The weights and biases can then be updated by the AdamOptimizer with a learning rate of 0.001. Other parameters of the AdamOptimizer can be set to their default values. For example: beta_1 = 0.9 beta_2 = 0.999 eps = 1e-08 weight_decay = 0

[0120] The above steps can be repeated with the next set of observations until all observations have been used for training. This can represent the first training epoch, and can be repeated until 10 epochs are completed.

[0121] Figure 12FIG. 1200 is a flowchart of a computer-implemented method for predicting a collision according to various embodiments. Method 1200 may include processes 1202, 1204, 1206, 1208, and 1210. Process 1202 may include inputting a sequence of image sets into a motion detection module, where each image set in the sequence includes a first image and a second image. The sequence of image sets may include, for example, image sets 720a, 720b. The motion detection module may be, for example, motion detection module 802. The motion detection module may include a neural network 1000. Process 1204 may include, for each image set, determining a corresponding depth map by the motion detection module based on the first image and the second image of the image set, thereby generating a sequence of depth maps. In an example, the image set may include a pair of stereo images, where the first image may be, for example, left image 722a or 722b, and the second image may be, for example, right image 724a or 724b. The depth map may include, for example, depth map 812. Process 1206 may include determining an optical flow by the motion detection module based on at least one of the first image from the sequence of image sets and the second image from the sequence of image sets. The optical flow may be, for example, optical flow 814. Process 1208 may include determining the motion of an object in the image set by the motion detection module based on the optical flow. Process 1210 may include determining the TTC of the object based on the sequence of depth maps and the determined motion of the object. Method 1200 may be able to use camera images to determine the TTC even when the road lighting is dim. Thus, adopting method 1200 in a vehicle may avoid traffic accidents at night. The various aspects described with respect to device 800 may be applicable to method 1200.

[0122] According to an embodiment that may be combined with any of the above embodiments or with any of the embodiments further described below, process 1204 may include: generating a corresponding unrotated disparity map for each image set using image processing method 600, and determining a corresponding depth map based on the corresponding unrotated disparity map. Using the unrotated disparity map generated according to image processing method 600 to determine the depth map allows for accurate depth determination even when the camera of the images in the output image set is not calibrated. The uncalibration of the camera may occur due to the vibration or movement of the vehicle on which the camera is mounted.

[0123] According to an embodiment that may be combined with any of the above embodiments or with any of the embodiments further described below, method 1200 may further include: generating a 3D point cloud based on at least one image set in the sequence of image sets, and detecting an object in the three-dimensional point cloud. The 3D point cloud is data-intensive. The 3D point cloud provides detailed information about the surrounding environment of the vehicle, such that object detection in the 3D point cloud is accurate.

[0124] According to an embodiment that can be combined with any of the above embodiments or any of the embodiments further described below, method 1200 may further include: generating a corresponding three-dimensional point cloud based on each image set in a sequence of image sets, thereby producing a sequence of three-dimensional point clouds; comparing the distances by which each point in the sequence of three-dimensional point clouds is moved; and correcting at least one three-dimensional point cloud in the sequence of three-dimensional point clouds based on the determined distances. By comparing the distances by which each point is moved, method 1200 can correct for distortions caused, for example, by camera intrinsic or extrinsic parameters.

[0125] Although embodiments of the present invention have been specifically shown and described with reference to particular embodiments, those skilled in the art should understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present invention as defined by the appended claims. Accordingly, the scope of the present invention is indicated by the appended claims and is thus intended to cover all changes falling within the meaning and scope of the equivalents of the claims. It should be understood that common numerals used in the relevant drawings refer to components for like or the same purpose.

[0126] Those skilled in the art should understand that the terms used herein are for the purpose of various embodiments only and are not intended to limit the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms "a / an" and "the" are intended to also include the plural forms. It should be further understood that the terms "comprising" and / or "including" used in this specification specify the presence of the stated features, integrations, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integrations, steps, operations, elements, components, and / or groups thereof.

[0127] It should be understood that the specific order or hierarchy of the blocks in the disclosed process / flowchart is an illustration of an exemplary method. It should be understood that the specific order or hierarchy of the blocks in the process / flowchart can be rearranged based on design preferences. Further, some blocks may be combined or omitted. The appended method claims present the elements of the various blocks in a sample order and are not intended to be limited to the specific order or hierarchy presented.

[0128] The foregoing description is provided to enable a person of ordinary skill in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, where the mention of an element in the singular is not intended to mean "one and only one" (unless specifically so stated) but rather "one or more." The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term "some" means one or more. Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "any combination of A, B, C, or their combinations" include any combination of A, B, and / or C, and may include multiple As, multiple Bs, or multiple Cs. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "any combination of A, B, C, or their combinations" can be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combination can include one or more members of A, B, or C. All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later become known to a person of ordinary skill in the art are hereby expressly incorporated by reference herein and are intended to be covered by the claims.

Claims

1. A computer-implemented image processing method (600), comprising: inputting an image set (212) including a first image and a second image into a trained machine learning model (104); using the trained machine learning model (104) to generate a rotational disparity map (214) of the image set (212); for each pixel in the first image, determining a rotational angle (426) between the pixel in the first image and its corresponding position in the second image with respect to a common image center (402) based on the rotational disparity map (214), and determining a corresponding unrotated disparity (426) based on the corresponding rotational angle; and generating an unrotated disparity map (216) based on the corresponding unrotated disparities of each pixel in the first image.

2. The image processing method (600) according to any one of the preceding claims, wherein, the image set (212) includes stereo images, wherein the first image is a left stereo image and the second image is a right stereo image.

3. The image processing method (600) according to any one of the preceding claims, wherein, the trained machine learning model (104) is configured to determine both a horizontal disparity and a vertical disparity of the image set such that the rotational disparity map (214) includes both a horizontal disparity and a vertical disparity.

4. The image processing method (600) according to any one of the preceding claims, wherein, the center of the first image serves as the common image center (402).

5. The image processing method (600) according to any one of the preceding claims, further comprising: segmenting the first image into a plurality of first regions; segmenting the second image into a plurality of second regions corresponding to the plurality of first regions; for each first region of the plurality of first regions, determining an epipolar line and a direction of the epipolar line based on a pixel in the first image and its corresponding position in the second image.

6. The image processing method (600) according to claim 5, further comprising: for each pixel of the first image, determining a corresponding horizontal disparity based on the corresponding unrotated disparity through a projection of the direction of the epipolar line.

7. The image processing method (600) according to any one of the preceding claims, further comprising: generating a three-dimensional point cloud (208) based on the image set (212) and further based on the unrotated disparity map (216).

8. The image processing method (600) according to claim 7, further comprising: inputting another image set including another first image and another second image captured at a time frame consecutive to the first image and the second image in the image set (212) into the trained machine learning model (104); using the trained machine learning model (104) to generate another rotational disparity map of the another image set, for each pixel in the another first image, identifying a corresponding pixel in the another second image based on the another rotational disparity map, and Determine a corresponding rotation angle between a pixel in the other first image and a corresponding pixel in the other second image with respect to another common image center, and Determine a corresponding unrotated disparity based on the corresponding rotation angle; Generate another unrotated disparity map based on the corresponding unrotated disparity of each pixel in the other first image; Generate another three-dimensional point cloud based on the other image set and further based on the other unrotated disparity map; Determine the distance that each point in the three-dimensional point cloud moves from the three-dimensional point cloud (208) to the other three-dimensional point cloud; And Correct at least one of the three-dimensional point cloud (208) and the other three-dimensional point cloud based on the determined distance.

9. The image processing method (600) according to any one of the preceding claims, Wherein, The trained machine learning model (104) is trained using a training data set including a plurality of image pairs generated by rotating a pair of calibrated images to a corresponding plurality of different angles, and using a ground truth disparity map (114) generated based on the pair of calibrated images as a training signal.

10. The image processing method (600) according to any one of the preceding claims, Wherein, The trained machine learning model (104) is trained to determine a close-range disparity and is further separately trained to determine a long-range disparity.

11. The image processing method (600) according to any one of the preceding claims, Wherein, The trained machine learning model (104) includes A feature extraction network (1310) configured to extract features from the image set to generate a feature map (1318), and Wherein, the trained machine learning model (104) further includes a disparity calculation network (1320) configured to determine the rotated disparity map (214) based on the generated feature map (1318).

12. The image processing method (600) according to claim 11, Wherein, The feature extraction network (1310) includes a plurality of neural network branches (1350, 1360), wherein each neural network branch includes a corresponding convolutional stack (1312) and a pooling module (1314) connected to the convolutional stack (1312), and wherein the plurality of neural network branches (1350, 1360) share the same weights.

13. The image processing method (600) according to any one of claims 11 to 12, Wherein, The disparity calculation network (1320) includes a three-dimensional convolutional neural network (1324) configured to generate three disparity maps based on the generated feature map (1318).

14. An image processing device (700), Comprising: A processor (702) configured to execute the image processing method (600) according to any one of the preceding claims.

15. The image processing device (700) according to claim 14, further Comprising: A camera (704); An engine (706) for driving the image processing device (700); and / or a steering module (708) configured to steer the image processing device (700), wherein the processor (702) is configured to drive the image processing device (700) and / or steer the image processing device using the unrotated disparity maps.

16. A computer-implemented method (1200) for predicting a collision, the method (1200) comprising: inputting a sequence of image sets (720a, 720b) into a motion detection module (802), wherein each image set in the sequence includes a first image (722a, 722b) and a second image (724a, 724b); for each image set, determining a corresponding depth map (812) by the motion detection module (802) based on the first image (722a, 722b) and the second image (724a, 724b) in the image set, thereby generating a sequence of depth maps (812); determining an optical flow (814) by the motion detection module (802) based on at least one of the first images (722a, 722b) from the sequence of image sets (720a, 720b) and the second images (724a, 724b) of the sequence of image sets; and determining the motion (814) of an object in the image set by the motion detection module (802) based on the optical flow; and determining a collision time with the object based on the sequence of depth maps (812) and the determined motion of the object.

17. The method (1200) according to claim 16, wherein determining a corresponding depth map (812) by the motion detection module (802) for each image set includes generating a corresponding unrotated disparity map (216) for each image set using the image processing method (600) according to claim 1, and determining the corresponding depth map (812) based on the corresponding unrotated disparity map (216).

18. The method (1200) according to any one of claims 16 to 17, further comprising: generating a three-dimensional point cloud (208) based on at least one image set in the sequence of image sets (720a, 720b); and detecting the object in the three-dimensional point cloud (208).

19. The method (1200) according to claim 18, further comprising: generating a corresponding three-dimensional point cloud (208) for each image set in the sequence of image sets (720a, 720b), thereby generating a sequence of three-dimensional point clouds (208); comparing the distances by which each point in the sequence of three-dimensional point clouds (208) moves; and correcting at least one three-dimensional point cloud (208) in the sequence of three-dimensional point clouds (208) based on the determined distances.

20. A device (800) for predicting a collision, the device (800) comprising: a processor (830) configured to execute the method (1200) according to any one of claims 16 to 19.