Method and system for determining the distance of an object
Through the combination of machine learning and sensors, the object bounding box in the two-dimensional image is projected into the three-dimensional bird's-eye view, solving the problem of accurate positioning of object distance and position in autonomous driving vehicles, and improving the real-time and accuracy of the autonomous driving system.
Patent Information
- Application Number
- CN202210141121.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-19
- Filing Date
- 2022-02-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-02-16
AI Technical Summary
The prior art is difficult to effectively and reliably determine the distance of objects around a vehicle, especially in autonomous driving environments, where traditional methods cannot accurately provide three-dimensional position and distance information of objects.
Using machine learning methods, camera images are used to classify objects, and through distance sensors and correction verification, combined with machine learning models such as SSD-type neural networks, the bounding box of the object on the two-dimensional image plane is reprojected into the three-dimensional bird's-eye view coordinate system, and the distance-oriented object description is constructed, and the depth information is provided by lidar or radar sensors for correction.
Real-time, accurate distance and position determination of objects around autonomous driving vehicles is achieved, and the safety and reliability of autonomous driving systems are improved.
Smart Images

Figure CN114973182B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods and systems for determining the distance to an object. Background Art
[0002] Knowing the distance to objects around a vehicle is an important aspect of (at least partially) autonomous vehicles.
[0003] Therefore, there is a need to provide effective and reliable methods and systems for determining the distance to objects around a vehicle. Summary of the Invention
[0004] The present disclosure provides computer-implemented methods, computer systems, and non-transitory computer-readable media. Embodiments are given in the description and the drawings.
[0005] In one aspect, the present disclosure relates to a computer-implemented method for determining the distance to an object, the method comprising steps performed (in other words: carried out) by computer hardware components: determining an image containing (in other words: showing) the object; determining the class of the object based on the image; determining a rough estimate of the distance based on a distance sensor; and determining the distance to the object based on the rough estimate and based on the class of the object.
[0006] The distance can be a scalar distance or a vector having more than one component (e.g., a two-dimensional or three-dimensional vector such that the distance can be used to determine the position relative to the vehicle provided with the sensor in a two-dimensional or three-dimensional space).
[0007] The determination of the class can be understood as classifying the object into one of a plurality of classes based on the image.
[0008] For example, a technique can be provided that is used for training with an ML (machine learning) model and reprojects an object of interest imaged in a 2D camera view onto a bird's-eye view (BEV) coordinate system with distance compensation correction to construct a distance-oriented object description.
[0009] A real-time object detection system that can be applied in a self-driving vehicle can be provided. The system can detect multiple object classes at once and can provide positioning related to the distance between them and the vehicle itself. A description of the environment can be provided along with knowledge of the positions of the detected objects based on the reprojection plane.
[0010] According to another aspect, the class of the object is selected from a plurality of traffic participant types. The traffic participant types can be selected such that all objects of a particular type have a similar size.
[0011] According to another aspect, the plurality of traffic participant types includes pedestrians, and / or bicycles, and / or cars, and / or trucks.
[0012] According to another aspect, the class of the object is selected from a plurality of orientations.
[0013] According to another aspect, the plurality of orientations includes left, and / or upper left, and / or up, and / or upper right, and / or right, and / or lower right, and / or down, and / or lower left.
[0014] For example, a method for constructing a correction vector according to a training class assignment can be provided. The object classes can be semantically related, and / or related to size (e.g., width and height), and / or orientation specific (e.g., having granularity: down, lower left, left, upper left, up, upper right, right, lower right).
[0015] According to another aspect, machine learning methods, such as artificial neural networks, can be used to determine the class. The training method of the machine learning method (e.g., artificial neural network) can allow accurate reprojection of the detected class.
[0016] According to another aspect, the artificial neural network can be of the SSD (Single Shot Detector) type or the YOLO (You Only Look Once) type. It will be understood that the YOLO method is one of the SSD type methods. The method for training an SSD type (or YOLO type) neural network can allow accurate reprojection of image detection to obtain knowledge of the BEV positions between objects. The SSD type (or YOLO type) machine learning model can be trained only on images. It has been found that the marking of image 2D bounding boxes is fast and an open training set for image object detection is available.
[0017] According to another aspect, the computer-implemented method can further include the following steps performed by the computer hardware component: determining a bounding box of an object in the image. The bounding box of the object can be determined based on the image.
[0018] According to another aspect, a rough estimate can also be determined based on the bounding box.
[0019] According to another aspect, the rough estimate is determined based on the matching of the measurement results of the distance sensor with the bounding box.
[0020] According to another aspect, the distance is determined based on a hash table for the class of the object.
[0021] According to another aspect, the hash table includes a correction for determining the distance based on the rough estimate.
[0022] In another aspect, the present disclosure relates to a computer system that includes a plurality of computer hardware components configured to perform multiple or all of the steps of the computer-implemented methods described herein. The computer system can be part of a vehicle.
[0023] The computer system can include a plurality of computer hardware components (such as a processor (e.g., a processing unit or a processing network), at least one memory (e.g., a memory unit or a memory network), and at least one non-transitory data storage). It should be understood that additional computer hardware components can be provided and used to perform the steps of the computer-implemented methods in the computer system. The non-transitory data storage and / or the memory unit can include a computer program for instructing the computer to perform multiple or all of the steps or aspects of the computer-implemented methods described herein, for example, using the processing unit and at least one memory unit.
[0024] In another aspect, the present invention relates to a vehicle that includes a computer system as described herein, a camera configured to acquire images, and a distance sensor.
[0025] In another aspect, the present disclosure relates to a non-transitory computer-readable medium that includes instructions for performing multiple or all of the steps or aspects of the computer-implemented methods described herein. The computer-readable medium can be configured as: an optical medium, such as a compact disc (CD) or a digital versatile disc (DVD); a magnetic medium, such as a hard disk drive (HDD); a solid-state drive (SSD); a read-only memory (ROM), such as a flash memory; and so on. Additionally, the computer-readable medium can be configured as a data storage accessible via a data connection such as an Internet connection. The computer-readable medium can be, for example, an online data repository or cloud storage.
[0026] The present disclosure also relates to a computer program for instructing a computer to perform multiple or all of the steps or aspects of the computer-implemented methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The exemplary embodiments and functions of the present disclosure are described herein in conjunction with the following schematically illustrated drawings:
[0028] Figure 1 is an illustration of an image and a corresponding bounding box;
[0029] Figure 2 is in Figure 1 is an illustration of the distance determined by the distance sensor in the scenario of;
[0030] Figure 3Schematic diagram of a coarsely reprojected image, in which an object is illustrated;
[0031] Figure 4 Flowchart illustrating a method for determining the distance of an object according to various embodiments;
[0032] Figure 5 Illustrates a distance determination system according to various embodiments; and
[0033] Figure 6 Illustrates a computer system having a plurality of computer hardware components according to various embodiments, the plurality of computer hardware components being configured to perform steps of a computer-implemented method for determining the distance of an object. Detailed Description
[0034] It may be desirable for a self-driving (in other words: autonomous) vehicle to be able to determine the distance to other objects and, after finding their positions, determine a set of possible actions for the self-driving vehicle. It should be understood that different objects have different meanings or significances, and while some object classes may be ignorable, other objects may need to be located to determine the best and safest driving manner for a self-driving car.
[0035] Machine learning methods may allow for accurately finding the localization of objects of interest in the form of 2D (two-dimensional) bounding boxes on the camera image plane. The determination of 2D bounding boxes may include:
[0036] 1) Determining the coordinates of the bounding box that encloses the object (in the image);
[0037] 2) Determining the type of the object (e.g., sedan, truck, pedestrian, construction vehicle, etc.); and
[0038] 3) Determining the certainty score (or probability) of the detection.
[0039] Determining 2D bounding boxes can be very effective. 2D bounding box determination can work in real time on embedded devices and can have very good accuracy in differentiating many classes. It should be understood that the number of classes to be differentiated may depend on the training data. For example, up to 80 classes can be modeled.
[0040] However, such a planar camera representation may not give information about how far the object of interest is actually from the vehicle (on which the camera for acquiring the image is set) and its 3D (three-dimensional) position related to the actual distance therebetween. According to various embodiments, an interchangeable and reliable system is provided that returns such information in real time and in a very accurate manner.
[0041] According to various embodiments, a fast and accurate construction of a representation can be provided that allows the identification of the positions of all objects of interest found by a detection algorithm operating on an image relative to the host vehicle. The architecture according to various embodiments can be applied to various sensor setups.
[0042] According to various embodiments, by training a machine learning method (such as an artificial neural network (e.g., of the SSD type)) together with an accurate method for re-projecting all objects of interest found in a 2D planar camera view onto a BEV coordinate system using a "sensor with depth" (e.g., radar, lidar, 3D camera, time-of-flight camera) (in other words: determining 3D coordinates based on the 2D coordinates projected by the camera onto the image sensor), a distance-oriented object description can be constructed. Additionally, according to various embodiments, further correction can be provided by compensating for the distance of the object, which can be implied by selecting an appropriate training strategy that depends on specifying the object position (heading direction) and the size (width, length) that encapsulates it into a training class type, thereby making the class more granular.
[0043] According to various embodiments, a coarse re-projection method can be provided that approximates the positioning of objects in the environment. Additionally, a method for performing the training of an ML (machine learning) method can be provided, which is beneficial and allows for an overall increase in the accuracy of the positions of the re-projected objects.
[0044] To perform the re-projection itself, a sensor with depth or the depth implicitly indicated by the sensor (as single-camera / stereo-camera depth) can be used. For example, the sensor can be a lidar sensor or a radar sensor. The method of re-projecting using a sensor with depth can utilize operations opposite to those used during projection (i.e., projecting a 3D point cloud onto a camera image). Additionally, a method for estimating the depth of an image point can be provided, as it is most likely that a particular point in the image plane is not directly associated with any projection point from a sensor with depth (e.g., because the point is lost due to the sparsity of the sensor points).
[0045] To have a reasonably accurate view of the scene and be able to identify positions in the BEV, i.e., give the distances from the host vehicle to all objects of interest, it can be assumed that there is at least one "sensor with depth" point projected onto the camera view in the vicinity of the 2D bounding box detection of each object in the image sample, where the vicinity can be understood as the sensor points projected by at least one of the respective bounding boxes. For example, if additional objects need to be detected, a sensor can be provided that can project points as far and as densely as needed, such that there are close projection points for the objects of our interest. Thus, an environment including a reasonable response to all objects of interest can be provided.
[0046] According to various embodiments, the reprojection can start by finding the centers of the respective 2D bounding boxes. Then, given the center points, for each center point, an efficient method using the k-d tree approach can be utilized to determine the k closest projected "depth sensor" points (where k is an integer). Then, the average of their depths can be used as an estimated depth for the bounding box center. Using this depth, the center can be reprojected onto 3D using the inverse operation of the operation used when projecting the "depth sensor" points onto the camera plane. Then, the center points (and / or the bounding boxes) can be projected onto a bird's-eye view.
[0047] However, only after these steps, the obtained localization may be rough. It can identify that in the BEV, the object of interest occupies the reprojected points, without giving an accurate representation of their positions.
[0048] Based on the information obtainable from the standard training of a 2D bounding box detection neural network, the type of the object can be determined and the probability of its existence can be determined (e.g., using a machine learning method of the SSD type). According to various embodiments, enhancing this information can be used to make the localization (in other words: the determination of the position) more robust. This can be provided by a specific way of training the detection neural network (e.g., a specific way of providing the training classes).
[0049] According to various embodiments, a specific partitioning of the objects into training classes (e.g., pedestrians, cars, bicycles...) can be provided, where it can indicate a fixed size for all objects in a specific class. By specifying the class, it can be assumed that its elements have similar sizes, such that each class contains elements having at least substantially the same or similar sizes (e.g., width and / or length).
[0050] According to various embodiments, a neural network (e.g., of the SSD type) can be trained to subdivide the main class into classes with similar widths and heights, such as truck_type_a, truck_type_b...
[0051] According to various embodiments, the orientation of the object can be determined. For the reprojection setup, a discrete representation of the orientation can be used (e.g.: left, upper left, up, upper right, right, lower right, down, lower left can be used. However, it will be understood that a more granular or coarser representation can be used. The orientation can take the form of further subclasses in the neural network. Thus, for example, the class of the object can be "car_bottom_right" (i.e., the object type of the class "car" and the object orientation of "bottom_right").
[0052] In the case of having the orientation and type of an object, compensation can be provided for the re-projection of the detection center from a planar camera onto the BEV. Default size values can be used for each class; for example, most sedans have similar dimensions (from the BEV perspective, we are only concerned with their width and height). Then, we can apply the default orientation to the BEV bounding box with an indication of its front.
[0053] According to various embodiments, compensation for the center point of re-projection can be specified in a hash table, and for each class (such as truck_type_a_front_left) in the hash table, a BEV correction vector can be provided. Using these vectors, compensation can be performed when re-projecting each class. It will be understood that the known front and rear orientations of the object obtained through a learning strategy according to various embodiments can be used for other purposes.
[0054] According to various embodiments, each identified 2D bounding box can be re-projected to achieve an accurate description of the environment in the BEV along with the knowledge of the position of the detected object. The architecture according to various embodiments can be used for different sensor setups. Instead of various lidar sensors that result in different point cloud densities, to provide depth, depth predicted based on a single camera or a stereo camera can be provided, or depth from radar readings.
[0055] Figure 1 Illustration 100 shows an image and a corresponding bounding box (such as bounding box 102). Each bounding box can be associated with the class of the object determined in the bounding box and a confidence value (or probability value) indicating how reliable the determined information (including the bounding box of the class) is.
[0056] Figure 2 Illustration 200 shows Figure 1 the distances determined by a distance sensor in the scene.
[0057] Figure 3 Illustration 300 shows a coarsely re-projected image, in which an object 302 is illustrated. By applying the correction vector and the orientation information implied by the identified class type, the positioning of the object (the bounding box and its orientation) can be further corrected.
[0058] Figure 4 Flowchart 400 shows a method for determining the distance of an object according to various embodiments. At 402, an image containing the object can be determined. At 404, the class of the object can be determined based on the image. At 406, a rough estimate of the distance can be determined based on the distance sensor. At 408, the distance of the object can be determined based on the rough estimate and the class of the object.
[0059] According to various embodiments, the class of the object can be selected from multiple traffic participant types.
[0060] According to various embodiments, the plurality of traffic participant types may include pedestrians, and / or bicycles, and / or cars, and / or trucks.
[0061] According to various embodiments, the class of an object can be selected from a plurality of orientations.
[0062] According to various embodiments, the plurality of orientations may include left, and / or upper left, and / or up, and / or upper right, and / or right, and / or lower right, and / or down, and / or lower left.
[0063] According to various embodiments, an artificial neural network can be used to determine the class.
[0064] According to various embodiments, the artificial neural network can be of the SSD type.
[0065] According to various embodiments, the method may further include determining a bounding box of an object in the image.
[0066] According to various embodiments, a rough estimate can be further determined based on the bounding box.
[0067] According to various embodiments, a rough estimate can be determined based on matching the measurements of a distance sensor with the bounding box.
[0068] According to various embodiments, the distance can be determined based on a hash table of object classes.
[0069] According to various embodiments, the hash table may include (for each category) a correction for determining the distance based on the rough estimate.
[0070] Each of the steps 402, 404, 406, 408 and the further steps described above can be performed by computer hardware components.
[0071] Figure 5 A distance determination system 500 according to various embodiments is shown. The distance determination system 500 may include an image determination circuit 502, a class determination circuit 504, a rough estimate determination circuit 506, and a distance determination circuit 508.
[0072] The image determination circuit 502 can be configured to determine an image containing an object.
[0073] The class determination circuit 504 can be configured to determine the class of an object based on the image.
[0074] The rough estimate determination circuit 506 can be configured to determine a rough estimate of the distance based on a distance sensor.
[0075] The distance determination circuit 508 can be configured to determine the distance of an object based on the rough estimate and based on the class of the object.
[0076] The image determination circuit 502, the category determination circuit 504, the rough estimate determination circuit 506, and the distance determination circuit 508 can be coupled to each other to exchange electrical signals, for example, via an electrical connection 510 (such as a cable or a computer bus) or via any other suitable electrical connection.
[0077] "Circuit" can be understood as any type of logic implementation entity, which can be a dedicated circuit or a processor that executes a program stored in a memory, firmware, or any combination thereof.
[0078] Figure 6 A computer system 600 with multiple computer hardware components is shown, and the multiple computer hardware components are configured to execute steps of a computer-implemented method for determining the distance of an object according to various embodiments. The computer system 600 can include a processor 602, a memory 604, and a non-transitory data storage 606. A camera 608 and a distance sensor 610 can be provided as part of the computer system 600 (as Figure 6 illustrated), or can be provided external to the computer system 600.
[0079] The processor 602 can execute instructions provided in the memory 604. The non-transitory data storage 606 can store a computer program, including instructions that can be transferred to the memory 604 and then executed by the processor 602. The camera 608 can be used to determine an image containing an object. The distance sensor 610 can be used to determine a rough estimate.
[0080] The processor 602, the memory 604, and the non-transitory data storage 606 can be coupled to each other to exchange electrical signals, for example, via an electrical connection 612 (such as a cable or a computer bus) or via any other suitable electrical connection. The camera 608 and / or the distance sensor 610 can be coupled to the computer system 600 via an external interface, for example, or can be provided as part of the computer system (in other words: inside the computer system, for example, coupled via the electrical connection 612).
[0081] The terms "coupled" or "connected" are intended to include direct "coupling" (e.g., via a physical link) or direct "connection" and indirect "coupling" or indirect "connection" (e.g., via a logical link), respectively.
[0082] It should be understood that what has been described for one of the above methods can similarly hold for the distance determination system 500 and / or the computer system 600.
[0083] List of reference numerals
[0084] Illustration of 100 images and corresponding bounding boxes
[0085] 102 bounding box
[0086] 200 Figure 1 Illustration of the distance determined by a distance sensor in a scenario
[0087] Schematic diagram of a 300 rough reprojection image, in which an object is illustrated
[0088] 302 Object
[0089] Flowchart illustrating a method for determining the distance of an object according to various embodiments
[0090] Step 402 of determining an image containing an object;
[0091] Step 404 of determining the class of the object based on the image
[0092] Step 406 of determining a rough estimate of the distance based on a distance sensor
[0093] Step 408 of determining the distance of the object based on the rough estimate and based on the class of the object
[0094] 500 Distance determination system
[0095] 502 Image determination circuit
[0096] 504 Class determination circuit
[0097] 506 Rough estimate determination circuit
[0098] 508 Distance determination circuit
[0099] 510 Connection
[0100] 600 Computer system according to various embodiments
[0101] 602 Processor
[0102] 604 Memory
[0103] 606 Non-transitory data storage unit
[0104] 608 Camera
[0105] 610 Distance sensor
[0106] 612 Connection
Claims
1. A computer-implemented method for determining the distance of an object, the method comprising the following steps performed by computer hardware components: Determine (402) an image containing the object; Determine (404) the class of the object based on the image; Determine the bounding box of the object in the image; Determine (406) a coarse estimate of the distance based on matching the measurement results of a distance sensor with the bounding box; And Determine (408) the distance of the object based on the coarse estimate and based on the class of the object.
2. The computer-implemented method according to claim 1, wherein, The class of the object is selected from a plurality of traffic participant types.
3. The computer-implemented method according to claim 2, wherein, The plurality of traffic participant types includes pedestrians, and / or bicycles, and / or cars, and / or trucks.
4. The computer-implemented method according to any one of claims 1 to 3, wherein, The class of the object is selected from a plurality of orientations.
5. The computer-implemented method according to claim 4, wherein, The plurality of orientations includes left, and / or upper left, and / or up, and / or upper right, and / or right, and / or lower right, and / or down, and / or lower left.
6. The computer-implemented method according to claim 1, wherein, The class is determined using an artificial neural network.
7. The computer-implemented method according to claim 6, wherein, The artificial neural network is of the single-shot detector type.
8. The computer-implemented method according to claim 1, wherein, The distance is determined based on a hash table of the class of the object.
9. The computer-implemented method according to claim 8, wherein, The hash table includes a correction for determining the distance based on the coarse estimate.
10. A computer system (600), the computer system (600) comprising a plurality of computer hardware components configured to perform the steps of the computer-implemented method according to any one of claims 1 to 9.
11. A vehicle, the vehicle comprising: The computer system (600) according to claim 10; A camera (608) configured to acquire the image; And The distance sensor (610).
12. A non-transitory computer-readable medium comprising instructions for performing the computer-implemented method according to any one of claims 1 to 9.