Object detection using cross-modal sensors in vehicles
By using a cross-modal sensor system combining lidar and cameras in the vehicle, performing data downsampling and staged processing, the problem of high computational intensity in autonomous and semi-autonomous vehicles is solved, achieving fast and accurate object detection.
Patent Information
- Application Number
- CN202110522082.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-16
- Filing Date
- 2021-05-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-05-13
AI Technical Summary
Object detection in modern vehicles, especially autonomous and semi-autonomous vehicles, is computationally intensive using multiple sensors, making it difficult to maintain detection accuracy while reducing computational complexity.
A cross-modal sensor system that uses a combination of lidar and cameras generates lower-resolution data by downsampling the data and uses a controller for light processing to combine the data from the two sensors to identify objects, and processes unconfirmed proposals in stages to improve accuracy.
With limited computing resources, it achieves fast and accurate detection of objects, reduces computational intensity, and improves the speed and reliability of object detection, especially the performance of long-distance and small object detection.
Smart Images

Figure CN114509773B_ABST
Abstract
Description
[0001] introduction
[0002] The information provided in this section is for the purpose of generally presenting the context of the present disclosure. The work of the presently named inventors, to the extent it is described in this section and in all aspects of the description that may not otherwise qualify as prior art at the time of filing, is not admitted, either explicitly or implicitly, to be prior art against the present disclosure. Technical Field
[0003] The present disclosure relates generally to object detection and, more particularly, to object detection using cross-modal sensors in a vehicle. Background Art
[0004] Modern vehicles, especially autonomous and semi-autonomous ones, increasingly rely on object detection capabilities. Objects can be detected using a variety of sensors, such as cameras, radar, and lidar. However, accurate object detection using these sensors is often computationally intensive. Sacrificing object detection accuracy by reducing computational complexity by using lower-resolution sensors is unacceptable. Summary of the Invention
[0005] A system includes a first sensor, a second sensor, and a controller. The first sensor is of a first type and is configured to sense objects around a vehicle and capture first data about the objects in a frame. The second sensor is of a second type and is configured to sense objects around the vehicle and capture second data about the objects in a frame. The controller is configured to downsample the first and second data to generate downsampled first and second data having a lower resolution than the first and second data. The controller is configured to identify a first group of objects by processing the downsampled first and second data having a lower resolution. The controller is configured to identify a second group of objects by selectively processing the first and second data from the frame.
[0006] In other features, the controller is configured to detect a first set of objects based on processing the downsampled second data; generate proposals regarding identities of the objects based on processing the downsampled first data; and confirm the identities of the detected first set of objects based on the first set of proposals.
[0007] In other features, the controller is configured to: process a second set of proposals using corresponding data from the first and second data from the framework; and identify a second set of objects based on the processing of the second set of proposals using corresponding data from the first and second data from the framework.
[0008] In another feature, the controller is configured to display the identified first and second groups of objects on a display in the vehicle.
[0009] In another feature, the controller is configured to navigate the vehicle based on the identified first and second groups of objects.
[0010] In other features, the first data is three-dimensional and the second data is two-dimensional or three-dimensional.
[0011] In other features, the first sensor is a lidar sensor and the second sensor is a camera.
[0012] In another feature, the proposals include N1 proposals for a first object within a first range of the vehicle and N2 proposals for a second object within a second range of the vehicle, the second range being beyond the first range, where N1 and N2 are integers greater than 1 and N1>N2.
[0013] In other features, the controller is further configured to detect a first set of objects based on processing the downsampled second data; confirm the identities of the detected first set of objects based on a first set of N1 proposals that match the detected first set of objects; and identify the second set of objects by processing the second set of N1 proposals using corresponding data from the first data and the second data of the framework.
[0014] In other features, the controller is further configured to detect a first set of objects based on processing the downsampled second data; confirm the identities of the detected first set of objects based on a first set of N2 proposals that match the detected first set of objects; and identify the second set of objects by processing the second set of N2 proposals using corresponding data from the first data and the second data of the framework.
[0015] In yet other features, a method includes: sensing first data regarding objects around a vehicle in a framework using a first sensor of a first type; and sensing second data regarding objects around the vehicle in the framework using a second sensor of a second type. The method includes: downsampling the first and second data to generate downsampled first and second data having a lower resolution than the first and second data; identifying a first group of objects by processing the downsampled first and second data having a lower resolution; and identifying a second group of objects by selectively processing the first and second data from the framework.
[0016] In other features, the method further includes detecting a first set of objects based on processing the downsampled second data; generating proposals regarding identities of the objects based on processing the downsampled first data; and confirming the identities of the detected first set of objects based on the first set of proposals.
[0017] In other features, the method further comprises processing a second set of proposals using corresponding data from the first data and the second data of the framework; and identifying a second set of objects based on the processing of the second set of proposals using corresponding data from the first data and the second data of the framework.
[0018] In another feature, the method further includes displaying the identified first and second group of objects on a display in the vehicle.
[0019] In another feature, the method further includes navigating the vehicle based on the identified first and second groups of objects.
[0020] In other features, the first data is three-dimensional and the second data is two-dimensional or three-dimensional.
[0021] In other features, the first sensor is a lidar sensor and the second sensor is a camera.
[0022] In another feature, the proposals include N1 proposals for a first object within a first range of the vehicle and N2 proposals for a second object within a second range of the vehicle, the second range being beyond the first range, where N1 and N2 are integers greater than 1 and N1>N2.
[0023] In other features, the method further includes: detecting a first set of objects based on processing the downsampled second data; confirming identities of the detected first set of objects based on a first set of N1 proposals that match the detected first set of objects; and identifying the second set of objects by processing the second set of N1 proposals using corresponding data from the first data and the second data of the framework.
[0024] In other features, the method further includes: detecting the first set of objects based on processing the downsampled second data; confirming the identities of the detected first set of objects based on a first set of N2 proposals that match the detected first set of objects; and identifying the second set of objects by processing the second set of N2 proposals using corresponding data from the first data and the second data of the framework.
[0025] The present invention includes the following technical solutions:
[0026] Solution 1. A system comprising:
[0027] a first sensor of a first type configured to sense objects surrounding the vehicle and capture first data regarding the objects in a frame;
[0028] a second sensor of a second type configured to sense objects surrounding the vehicle and capture second data about the objects in the frame; and
[0029] A controller configured to:
[0030] downsampling the first data and the second data to generate downsampled first data and second data having a lower resolution than the first data and the second data;
[0031] identifying a first set of objects by processing the downsampled first and second data having the lower resolution; and
[0032] A second set of objects is identified by selectively processing the first and second data from the framework.
[0033] Option 2. The system according to Option 1, wherein the controller is configured to:
[0034] detecting the first set of objects based on processing the downsampled second data;
[0035] generating a proposal regarding the identity of the object based on processing the downsampled first data; and
[0036] The identities of the first set of detected objects are confirmed based on the first set of proposals.
[0037] Option 3. The system according to Option 2, wherein the controller is configured to:
[0038] processing a second set of proposals using corresponding data from the first and second data of the framework; and
[0039] The second set of objects is identified based on processing the second set of proposals using corresponding data of the first and second data from the framework.
[0040] Option 4. The system of Option 1, wherein the controller is configured to display the identified first and second groups of objects on a display in the vehicle.
[0041] Option 5. The system of Option 1, wherein the controller is configured to navigate the vehicle based on the identified first and second groups of objects.
[0042] Option 6. The system according to Option 1, wherein the first data is three-dimensional and the second data is two-dimensional or three-dimensional.
[0043] Option 7. The system according to Option 1, wherein the first sensor is a lidar sensor and the second sensor is a camera.
[0044] Option 8. A system according to Option 2, wherein the proposals include N1 proposals regarding a first object within a first range of the vehicle, and N2 proposals regarding a second object within a second range of the vehicle, the second range exceeding the first range, wherein N1 and N2 are integers greater than 1, and N1>N2.
[0045] Option 9. The system according to Option 8, wherein the controller is further configured to:
[0046] detecting the first set of objects based on processing the downsampled second data;
[0047] confirming the identities of the first set of detected objects based on the first set of N1 proposals that match the first set of detected objects; and
[0048] The second set of objects is identified by processing a second set of N1 proposals using corresponding data from the first data and the second data of the framework.
[0049] Option 10. The system according to Option 8, wherein the controller is further configured to:
[0050] detecting the first set of objects based on processing the downsampled second data;
[0051] confirming the identities of the first set of detected objects based on the first set of N2 proposals that match the first set of detected objects; and
[0052] The second set of objects is identified by processing a second set of N2 proposals using corresponding data from the first data and the second data of the framework.
[0053] Scheme 11. A method comprising:
[0054] sensing first data regarding objects surrounding the vehicle in the framework using a first sensor of a first type;
[0055] sensing second data regarding objects surrounding the vehicle in the framework using a second sensor of a second type;
[0056] downsampling the first data and the second data to generate downsampled first data and second data having lower resolution than the first data and the second data;
[0057] identifying a first set of objects by processing the downsampled first and second data having the lower resolution; and
[0058] A second set of objects is identified by selectively processing the first data and the second data from the framework.
[0059] Scheme 12. The method according to Scheme 11, further comprising:
[0060] detecting the first set of objects based on processing the downsampled second data;
[0061] generating a proposal regarding the identity of the object based on processing the downsampled first data; and
[0062] The identities of the first set of detected objects are confirmed based on the first set of proposals.
[0063] Scheme 13. The method according to Scheme 12, further comprising:
[0064] processing a second set of proposals using corresponding data of the first data and the second data from the framework; and
[0065] The second set of objects is identified based on processing the second set of proposals using corresponding data of the first data and the second data from the framework.
[0066] Embodiment 14. The method of embodiment 11 further comprises displaying the identified first group of objects and second group of objects on a display in the vehicle.
[0067] Embodiment 15. The method of embodiment 11, further comprising navigating the vehicle based on the identified first and second groups of objects.
[0068] Option 16. The method according to Option 11, wherein the first data is three-dimensional and the second data is two-dimensional or three-dimensional.
[0069] Option 17. The method according to Option 11, wherein the first sensor is a lidar sensor and the second sensor is a camera.
[0070] Option 18. A method according to Option 12, wherein the proposal includes N1 proposals regarding a first object within a first range of the vehicle, and N2 proposals regarding a second object within a second range of the vehicle, the second range exceeding the first range, wherein N1 and N2 are integers greater than 1, and N1>N2.
[0071] Option 19. The method according to Option 18, further comprising:
[0072] detecting the first set of objects based on processing the downsampled second data;
[0073] confirming the identities of the first set of detected objects based on the first set of N1 proposals that match the first set of detected objects; and
[0074] The second set of objects is identified by processing a second set of N1 proposals using corresponding data from the first data and the second data of the framework.
[0075] Option 20. The method according to Option 18, further comprising:
[0076] detecting the first set of objects based on processing the downsampled second data;
[0077] confirming the identities of the first set of detected objects based on the first set of N2 proposals that match the first set of detected objects; and
[0078] The second set of objects is identified by processing a second set of N2 proposals using corresponding data from the first data and the second data of the framework.
[0079] Further areas of applicability of the present disclosure will become apparent from the detailed description, claims, and accompanying drawings.The detailed description and specific examples are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The present disclosure will become more fully understood from the detailed description and accompanying drawings, in which:
[0081] Figure 1 is a functional block diagram of a system for detecting objects around a vehicle using two different types of sensors according to the present disclosure;
[0082] Figure 2 and Figure 3 Shown for use Figure 1 A flowchart of a method for a system to detect short-range and long-range objects around a vehicle;
[0083] Figure 4 A flow chart showing a method that combines Figure 2 and Figure 3 For use Figure 1 A method for detecting objects around a vehicle by a system; and
[0084] Figure 5 Shown for Figure 2-4 For use Figure 1Flowchart of a method for performing high-resolution processing on portions of data from two sensors in a method for detecting objects around a vehicle using a system.
[0085] In the drawings, reference numerals may be repeated to identify similar and / or identical elements. DETAILED DESCRIPTION
[0086] The present disclosure relates to detecting objects using cross-modal sensors. For example, objects can be detected using a combination of a lidar sensor and a camera, although any other type of sensor may be used instead. Object detection using cross-modal sensors can be performed in two stages. In a first stage, referred to as a glance stage, data about objects around the vehicle is captured in a frame by a first sensor (e.g., a lidar sensor) and a second sensor (e.g., a camera). The data captured by each sensor is downsampled and rapidly processed in parallel at a lower resolution. Objects are detected from the downsampled camera data, and proposals about the detected objects are generated from the downsampled lidar data. The objects detected from the downsampled camera data and the proposals generated from the downsampled lidar data are combined. Objects that can be definitively identified (i.e., verified) by both sensors from the combination of the detected objects and the proposals are confirmed as being correctly identified.
[0087] In the second phase, referred to as the focus phase, proposals from the combination that can be verified by one sensor but not the other (i.e., unconfirmed proposals) are processed at a higher resolution than in the first phase, which uses only the corresponding data from both sensors in the framework. Data from the second sensor (e.g., a camera) used only for the unconfirmed proposals is processed at a higher resolution than in the first phase to detect objects in the unconfirmed proposals. Furthermore, data from the first sensor (e.g., a lidar) used only for the unconfirmed proposals is processed at a higher resolution than in the first phase and used to confirm the identities of detected objects in the unconfirmed proposals.
[0088] The two-stage processing system disclosed above offers numerous advantages over existing technologies. Typically, limited computing resources are available to process data from high-resolution sensors (such as cameras, lidar, etc.). If low-resolution data is used due to computational constraints, these sensors perform poorly at detecting distant or small objects at the lower resolution. If high / raw resolution data is used, most systems are unable to process all the data in real time due to limited computing resources. In contrast, the two-stage processing system disclosed above can be implemented with limited computing resources without sacrificing accuracy and reliability in object detection, regardless of the distance from the vehicle or the size of the object. Specifically, in the two-stage processing system, because the processing in the first stage is performed at a lower resolution, it is less computationally intensive. As a result, the processing in the first stage can be performed quickly and with relatively low power consumption. In the second stage, only a very limited amount of raw data corresponding to unconfirmed proposals is processed at its original resolution, which is higher than the majority of the downsampled data processed at a lower resolution in the first stage. Thus, a very limited amount of high-resolution and power-intensive processing is performed as needed. These and other features of the two-stage processing system disclosed above are described in detail below.
[0089] Figure 1 A block diagram of a system 100 for detecting objects around a vehicle using two different types of sensors according to the present disclosure is shown. By way of example only, throughout this disclosure, the first type of sensor is a lidar sensor, and the second type of sensor is a camera. Alternatively, any other type of sensor of a different modality may be used instead.
[0090] Camera data (whether captured by a 2D or 3D camera) has limitations in estimating distant objects (for example, 2D camera data lacks depth information, and 2D / 3D camera data has very few pixels for distant objects), but it captures color and texture information from objects. LiDAR data lacks color and texture information but provides depth information about objects that camera data lacks. Therefore, these two types of sensors together provide data that can be combined to accurately identify objects both relatively close to and far from the vehicle, as explained in detail below.
[0091] System 100 includes: a first sensor 102 of a first type (e.g., a lidar sensor); a second sensor 104 of a second type (e.g., a camera); a controller 106 for processing data from the first and second sensors 102, 104 and detecting objects around the vehicle; a display 108 (e.g., located in the vehicle's dashboard) for displaying detected objects; a navigation module 110 (e.g., of an autonomous or semi-autonomous vehicle) for navigating the vehicle based on the detected objects; and one or more vehicle control subsystems 112 (e.g., a braking subsystem, a cruise control subsystem, etc.) controlled by navigation module 110 based on the detected objects. Controller 106 includes memory 120 and a processor 122. Controller 106 processes data from both sensors 102, 104 in parallel, as described below.
[0092] The first sensor 102 (e.g., a lidar sensor) generates 3D point cloud data 130 regarding objects sensed around the vehicle in a frame. The second sensor 104 (e.g., a camera) generates 2D pixel data 132 regarding objects sensed around the vehicle in a frame. The memory 120 stores the point cloud data 130 and the pixel data 132. Furthermore, the memory 120 stores instructions that are executed by the processor 122 to process the data from the first sensor 102 and the second sensor 104 as follows.
[0093] The processor 122 downsamples the pixel data 132 and generates downsampled pixel data 140 having a lower resolution than the original, higher resolution pixel data 132 captured by the camera 104. The processor 122 includes a lightweight object detector 142 that detects objects by processing the downsampled pixel data 140. The object detector 142 is lightweight; that is, not computationally intensive because the object detector 142 processes the downsampled pixel data 140 at a lower resolution and does not process the original, higher resolution pixel data 132 captured by the camera 104. For example, the object detector 142 may detect some objects surrounding the vehicle (in the Figure 1 is shown as N3).
[0094] Processor 122 also downsamples point cloud data 130 and generates downsampled point cloud data 150 having a lower resolution than the original, higher resolution point cloud data 130 captured by lidar sensor 102. Processor 122 extracts feature vectors 152 from downsampled point cloud data 150. Processor 122 includes a trained neural network 154 that generates proposals for objects surrounding the vehicle based on feature vectors 152 extracted from downsampled point cloud data 150.
[0095] These proposals may include N1 proposals for short-range objects (i.e., objects located within a relatively short range (e.g., 0-40 m) from the vehicle) and N2 proposals for long-range objects (i.e., objects located within a relatively long range (i.e., beyond the short range) (e.g., >40 m) from the vehicle). Thus, the N1 proposals may be referred to as short-range proposals, and the N2 proposals may be referred to as long-range proposals. For example, N1>N2. Again, the neural network 154 is lightweight; that is, not computationally intensive because the feature vectors 152 used by the neural network 154 are extracted from the downsampled point cloud data 150 at a lower resolution, and not from the original, higher resolution point cloud data 130 captured by the lidar sensor 102.
[0096] Processor 122 projects the 3D proposals onto the 2D objects detected by object detector 142. Processor 122 processes a combination 160 of the projected N1 and N2 proposals generated based on the downsampled point cloud data 150 and the N3 objects detected based on the downsampled pixel data, as described below with reference to Figure 2 and Figure 3 Just as described.
[0097] Figure 2 A method 200 is shown for processing a combination 160 of N1 proposals (ie, short-range proposals) generated based on projections of downsampled point cloud data 150 and N3 objects detected based on downsampled pixel data 140. Figure 2-5 , the term control generally refers to the controller 106 and specifically refers to the processor 122.
[0098] At 202, control combines the projected N1 proposals (i.e., short-range proposals) generated based on the downsampled point cloud data 150 with the N3 objects detected based on the downsampled pixel data 140. At 204, control performs maximum bipartite matching between the projected N1 proposals and the N3 objects based on the intersection-over-union (IoU) ratio of the bounding boxes. IoU is the ratio of the overlapping area (i.e., intersection) between two bounding boxes to the area of the union of the two bounding boxes. The projected N1 proposals with an IoU greater than 0.5 are selected as valid candidates, and maximum bipartite matching is used to find the best matching pair between the selected N1 proposals and the N3 detected objects.
[0099] At 206, the control determines whether the matching proposals can be verified by the camera data from the N1 proposals that match the N3 detected objects. At 208, the control disregards or ignores those matching proposals that cannot be verified by the camera data as false positives detected by the lidar sensor 102. Not processing false positives detected by the lidar sensor 102 also results in computational savings.
[0100] For those match proposals verified by the camera data, the control determines whether these match proposals can also be verified by the lidar data at 210. If these match proposals are also verified by the lidar data, then at 212, the control confirms the identity of the object detected by the camera 104.
[0101] For those match proposals that are verified by the camera data but not by the lidar data, the control processes these match proposals at 214, which represent potential false positives from the camera and Figure 1 The control processes these unconfirmed proposals 162 using the corresponding higher resolution raw pixel data 132 from the camera 104 and the corresponding higher resolution raw point cloud data 130 from the lidar sensor 102. Figure 5 To describe Figure 1 The high resolution processing shown at 170 in FIG.
[0102] Figure 3 A method 250 is shown for processing a combination 160 of N2 proposals (i.e., long-range proposals) of projections generated based on downsampled point cloud data 150 and N3 objects detected based on downsampled pixel data 140. Methods 200 and 250 are shown separately for illustrative purposes only; control executes methods 200 and 250 in parallel.
[0103] At 252, the control combines the projected N2 proposals generated based on the downsampled point cloud data 150 and the N3 objects detected based on the downsampled pixel data 140. At 254, the control performs maximum bipartite matching between the projected N2 proposals and the N3 objects based on the IoU ratio of the bounding boxes. The projected N2 proposals with IoU>0.5 are selected as valid candidates, and maximum bipartite matching is used to find the best matching pair between the selected N2 proposals and the detected N3 objects.
[0104] At 256, from the N2 proposals that match the N3 detected objects, the control determines whether these matching proposals can be verified by the lidar data. At 258, the control disregards or ignores those matching proposals that cannot be verified by the lidar data as false positives detected by the camera 104. Not processing false positives detected by the camera 104 also results in computational savings.
[0105] For those match proposals verified by the lidar data, control determines whether these match proposals can also be verified by the camera data at 260. If these match proposals are also verified by the camera data, control confirms the identity of the object detected by the camera 104 at 262.
[0106] For those match proposals that are verified by the lidar data but not by the camera data, the control processes these match proposals at 264, which represent potential false positives from the camera and Figure 1 The control processes these unconfirmed proposals 162 using the corresponding higher resolution raw pixel data 132 from the camera 104 and the corresponding higher resolution raw point cloud data 130 from the lidar sensor 102. Figure 5 To describe Figure 1 The high resolution processing shown at 170 in FIG.
[0107] Figure 4 A combined method 300 (ie, a combination of methods 200 and 250) is shown for processing data from Figure 1 The two different types of sensors shown in FIG are used to collect data from the vehicle and detect objects around the vehicle (primarily at a lower resolution and secondly, if necessary, partially at a higher resolution).
[0108] At 402, the control captures 3D point cloud data of an object from a lidar sensor in the frame. At 404, the control captures 2D pixel data of the object from a camera in the frame. At 406, the control downsamples the point cloud data and the pixel data. At 408, the control detects N3 objects in 2D from the downsampled pixel data.
[0109] At 410, the control extracts features from the downsampled point cloud data. At 412, the control inputs the extracted features into the trained neural network and generates N1 3D proposals for short-range objects and N2 3D proposals for long-range objects based on the downsampled point cloud data.
[0110] At 414, the control projects the 3D proposal onto the 2D objects detected from the downsampled pixel data. At 416, when the match of the projected proposal to the detected objects is verified by both the camera data and the lidar data, the control confirms the identities of the N3 detected objects to which the projected 3D proposal matches.
[0111] At 418, from the N1 short-range proposals that match the N3 detected objects, the control ignores those proposals that cannot be verified by the camera data as false positives from the lidar data. In addition, the control processes those proposals from the N1 short-range proposals that match the N3 detected objects and can be verified by the camera data but not by the lidar data (false positives from the camera data) using the corresponding high-resolution data from the framework. Figure 5 To describe high resolution processing.
[0112] At 420, from the N2 short-range proposals that match the N3 detected objects, the control ignores those proposals that cannot be verified by the lidar data as false positives from the camera data. In addition, the control processes those proposals from the N2 long-range proposals that match the N3 detected objects and can be verified by the lidar data but not by the camera data (false positives from the camera data) using the corresponding high-resolution data from the framework. Figure 5 To describe high resolution processing.
[0113] Figure 5 A method 450 is shown for processing a proposal that matches an object detected by a camera, is verified by one of the two sensors but not by the other of the two sensors. Figure 1 is shown as unconfirmed proposal 162 in Figure 1 Their processing at a higher resolution is shown at 170 in FIG, which is described below.
[0114] At 452, the control obtains partial pixel data from the framework for only the proposals to be processed at high resolution. The partial pixel data obtained from the framework is the original high resolution raw data captured by the camera 104. At 454, the control detects objects in the proposals by processing the partial pixel data at high resolution.
[0115] At 456, the control obtains only the partial point cloud data for the proposals to be processed at high resolution from the framework. The partial point cloud data obtained from the framework is the original high resolution raw data captured by the lidar sensor 102. At 458, the control processes the partial point cloud data for these proposals at high resolution.
[0116] At 460, the control uses the processing performed at 458 to confirm the identity of the object detected at 454. For example, the control uses depth information obtained from the partial point cloud data processed at high resolution to confirm the identity of the object detected by processing the partial pixel data at high resolution. Figure 4 418 in the low resolution) is performed by Figure 1 The only computationally intensive object detection performed by the system 100 shown in FIG. Figure 4 This partial high-resolution object detection is performed when the low-resolution object detection performed at 418 using downsampled data from the two sensors does not identify all objects.
[0117] exist Figure 1 In the example, the display 108 (e.g., of an infotainment module in a vehicle) displays Figure 4 418 of them and Figure 5 460 in the detected objects. These detected objects are also input to the navigation module 110. The navigation module 110 can control one or more vehicle control subsystems 112 based on the detected objects.
[0118] Thus, the systems and methods of the present disclosure significantly improve the technical field of object detection in autonomous and semi-autonomous vehicles in general and in particular. In particular, the systems and methods significantly improve the speed at which objects can be detected using significantly reduced and simplified processing resources (due to low-resolution processing of data from sensors of different modalities, as explained above), without sacrificing accuracy, which can be important in autonomous and semi-autonomous vehicles.
[0119] The foregoing description is merely illustrative in nature and is not intended to limit the present disclosure, its application, or uses. The broad teachings of the present disclosure can be implemented in many forms. Therefore, although the present disclosure includes specific examples, the true scope of the present disclosure should not be so limited, as other modifications will become apparent upon studying the drawings, the description, and the following claims. It should be understood that one or more steps within the method can be performed in a different order (or simultaneously) without changing the principles of the present disclosure. In addition, although each of the embodiments is described above as having certain features, any one or more of those features described with respect to any embodiment of the present disclosure can be implemented in any of the other embodiments and / or combined with features thereof, even if the combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and permutations and combinations of one or more embodiments with each other remain within the scope of the present disclosure.
[0120] Spatial and functional relationships between elements (e.g., between modules, circuit elements, semiconductor layers, etc.) are described using various terms, including "connected," "engaged," "coupled," "adjacent," "immediately adjacent," "on top of," "above," "below," and "positioned." Unless explicitly described as "directly," when describing a relationship between a first element and a second element in the above disclosure, the relationship can be a direct relationship with no other intervening elements between the first element and the second element, but can also be an indirect relationship with one or more intervening elements (either spatially or functionally) between the first element and the second element. As used herein, the phrase at least one of A, B, and C should be interpreted to mean a logical (A OR B OR C) using a non-exclusive logical OR, and should not be interpreted to mean "at least one of A, at least one of B, and at least one of C."
[0121] In the accompanying drawings, the direction of an arrow, as indicated by the arrow head, generally illustrates the flow of information (such as data or instructions) of interest. For example, when component A and component B exchange various types of information, but the information transmitted from component A to component B is relevant to the illustration, an arrow may point from component A to component B. This unidirectional arrow does not imply that no other information is transmitted from component B to component A. Furthermore, for information sent from component A to component B, component B may send a request for the information or an acknowledgment of receipt to component A.
[0122] In this application (including the definitions below), the term "module" or the term "controller" may be replaced with the term "circuit". The term "module" may refer to, be part of, or include: an application-specific integrated circuit (ASIC); a digital, analog, or mixed analog / digital discrete circuit; a digital, analog, or mixed analog / digital integrated circuit; a combinational logic circuit; a field-programmable gate array (FPGA); a processor circuit (shared, dedicated, or group) that executes code; a memory circuit (shared, dedicated, or group) that stores code executed by the processor circuit; other suitable hardware components that provide the described functionality; or a combination of some or all of the above, such as in a system-on-chip.
[0123] A module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module of the present disclosure may be distributed across multiple modules connected via the interface circuits. For example, multiple modules may allow for load balancing. In another example, a server (also referred to as a remote or cloud) module may perform some functions on behalf of a client module.
[0124] As used above, the term code may include software, firmware, and / or microcode, and may refer to programs, routines, functions, classes, data structures, and / or objects. The term shared processor circuit encompasses a single processor circuit that executes some or all code from multiple modules. The term group processor circuit encompasses a processor circuit that is combined with additional processor circuits to execute some or all code from one or more modules. Reference to multiple processor circuits encompasses multiple processor circuits on discrete dies, multiple processor circuits on a single die, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or combinations of the above. The term shared memory circuit encompasses a single memory circuit that stores some or all code from multiple modules. The term group memory circuit encompasses a memory circuit that is combined with additional memory to store some or all code from one or more modules.
[0125] The term memory circuit is a subset of the term computer-readable medium. As used herein, the term computer-readable medium does not encompass transient electrical or electromagnetic signals propagating through a medium (such as on a carrier wave); thus, the term computer-readable medium may be considered to be tangible and non-transitory. Non-limiting examples of non-transitory, tangible computer-readable media are nonvolatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital magnetic tape or hard drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs).
[0126] The apparatus and methods described in this application may be implemented in part or in whole by a special-purpose computer, wherein the special-purpose computer is constructed by configuring a general-purpose computer to perform one or more specific functions implemented in a computer program. The functional blocks, flow chart components, and other elements described above serve as software descriptions that can be converted into computer programs through routine work by a skilled technician or programmer.
[0127] A computer program includes processor-executable instructions stored on at least one non-transitory, tangible computer-readable medium. A computer program may also include or rely on stored data. A computer program may include a basic input / output system (BIOS) that interacts with the hardware of a special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, and the like.
[0128] A computer program may include: (i) descriptive text to be parsed, such as HTML (Hypertext Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation), (ii) assembly code, (iii) object code generated by a compiler from source code, (iv) source code executed by an interpreter, (v) source code compiled and executed by a just-in-time compiler, etc. By way of example only, the source code may be written using syntax from languages including C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Java®, Fortran, Perl, Pascal, Curl, OCaml, Javascript®, HTML5 (Hypertext Markup Language Version 5), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Flash®, Visual Basic®, Lua, MATLAB, SIMULINK, and Python®.
Claims
1. An object detection system, comprising: a first sensor of a first type configured to sense objects surrounding the vehicle and capture first data regarding the objects in a frame; a second sensor of a second type configured to sense objects surrounding the vehicle and capture second data about the objects in the frame; as well as A controller configured to: downsampling the first data and the second data to generate downsampled first data and second data having a lower resolution than the first data and the second data; identifying a first set of objects by processing the downsampled first and second data having the lower resolution; as well as identifying a second set of objects by selectively processing the first data and the second data from the framework; Wherein, the controller is configured as follows: detecting the first set of objects based on processing the downsampled second data; generating a proposal regarding the identity of the object based on processing the downsampled first data; and confirming the identities of the first set of detected objects based on the first set of proposals; The proposals include N1 proposals for a first object within a first range of the vehicle and N2 proposals for a second object within a second range of the vehicle, the second range being outside the first range, wherein N1 and N2 are integers greater than 1, and N1>N2; The controller is further configured to do any of the following: detecting the first set of objects based on processing the downsampled second data; confirming the identities of the first set of detected objects based on the first set of N1 proposals that match the first set of detected objects; and identifying the second set of objects by processing a second set of N1 proposals using corresponding data from the first data and the second data of the framework; or: detecting the first set of objects based on processing the downsampled second data; confirming the identities of the first set of detected objects based on the first set of N2 proposals that match the first set of detected objects; and The second set of objects is identified by processing a second set of N2 proposals using corresponding data from the first data and the second data of the framework.
2. The system according to claim 1, wherein: The controller is configured to: processing a second set of proposals using corresponding data from the first and second data of the framework; as well as The second set of objects is identified based on processing the second set of proposals using corresponding data of the first and second data from the framework.
3. The system according to claim 1, wherein: The controller is configured to display the identified first and second groups of objects on a display in the vehicle.
4. The system according to claim 1, wherein: The controller is configured to navigate the vehicle based on the identified first and second groups of objects.
5. The system according to claim 1, wherein: The first data is three-dimensional, and the second data is two-dimensional or three-dimensional.
6. The system according to claim 1, wherein: The first sensor is a lidar sensor and the second sensor is a camera.
7. A method for object detection, comprising: sensing first data regarding objects surrounding the vehicle in the framework using a first sensor of a first type; sensing second data regarding objects surrounding the vehicle in the framework using a second sensor of a second type; downsampling the first data and the second data to generate downsampled first data and second data having lower resolution than the first data and the second data; identifying a first set of objects by processing the downsampled first and second data having the lower resolution; as well as identifying a second set of objects by selectively processing the first data and the second data from the framework; It further includes: detecting the first set of objects based on processing the downsampled second data; generating a proposal regarding the identity of the object based on processing the downsampled first data; and confirming the identities of the first set of detected objects based on the first set of proposals; wherein the proposals include N1 proposals regarding a first object within a first range of the vehicle and N2 proposals regarding a second object within a second range of the vehicle, the second range being outside the first range, wherein N1 and N2 are integers greater than 1, and N1>N2; The method further comprises any one of the following steps: detecting the first set of objects based on processing the downsampled second data; confirming the identities of the first set of detected objects based on the first set of N1 proposals that match the first set of detected objects; and identifying the second set of objects by processing a second set of N1 proposals using corresponding data from the first data and the second data of the framework; or detecting the first set of objects based on processing the downsampled second data; confirming the identities of the first set of detected objects based on the first set of N2 proposals that match the first set of detected objects; and The second set of objects is identified by processing a second set of N2 proposals using corresponding data from the first data and the second data of the framework.
8. The method according to claim 7, further comprising: processing a second set of proposals using corresponding data of the first data and the second data from the framework; as well as The second set of objects is identified based on processing the second set of proposals using corresponding data of the first data and the second data from the framework. 9 . The method of claim 7 , further comprising displaying the identified first and second groups of objects on a display in the vehicle. 10 . The method of claim 7 , further comprising navigating the vehicle based on the identified first and second groups of objects.
11. The method according to claim 7, wherein: The first data is three-dimensional, and the second data is two-dimensional or three-dimensional.
12. The method according to claim 7, wherein: The first sensor is a lidar sensor and the second sensor is a camera.
Citation Information
Patent Citations
High resolution 3D point clouds generation based on CNN and CRF models
CN109215067A