System and method for obstacle detection and association based on neural network for mobile platform
By using multiple sensors to collect data on the mobile platform and using the hierarchical structure of ANN and region-based methods for processing, the problem of insufficient efficiency and accuracy of obstacle detection in the prior art is solved, and more efficient and accurate obstacle detection is achieved.
Patent Information
- Application Number
- CN201980087253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-01-22
AI Technical Summary
The prior art has problems with insufficient detection efficiency and accuracy when detecting obstacles in the three-dimensional environment around the mobile platform, especially when the 3D information reconstructed using stereo camera data is not accurate enough.
Data is collected using sensors carried by mobile platforms such as optical radar, radar, time-of-flight camera, stereo camera or monocular camera, and depth information, feature maps and candidate areas are determined through the hierarchical structure of artificial neural networks (ANNs) and region-based methods, and these data are fed to the obstacle detection neural network to predict the state attributes of the obstacle.
The accuracy and efficiency of obstacle detection are improved, and the attributes such as the location, posture, type and other attributes of obstacles can be determined more accurately, thereby improving the performance of automatic or unmanned navigation systems.
Smart Images

Figure CN113228043B_ABST
Abstract
Description
Technical Field
[0001] The techniques of the present disclosure are generally directed to detecting obstacles in a three-dimensional (3D) environment adjacent to a mobile platform, such as one or more pedestrians, vehicles, buildings, or other obstacle types. Background Art
[0002] One or more sensors may be used to scan or otherwise detect the environment surrounding the mobile platform. For example, the mobile platform may be equipped with a stereoscopic vision system (e.g., a "stereo camera") to sense its surroundings. A stereo camera is typically a type of camera with two or more lenses each having a separate imaging sensor or film frame. When two or more lenses are used simultaneously but photos / videos are taken from different angles, the differences between the corresponding photos / videos provide a basis for calculating depth information (e.g., the distance between an object in the scene and the stereo camera). As another example, a mobile platform may be equipped with one or more light radar sensors, which typically emit pulse signals (e.g., laser signals), detect pulse signal reflections, and determine depth information about the environment to facilitate object detection and / or recognition. Automatic or unmanned navigation typically requires determining various attributes of obstacles, such as position, orientation, or size. There is still a need for more efficient obstacle detection technology that can help improve the performance of various more advanced applications. Summary of the invention
[0003] The following summary is provided for the reader's convenience and points to some representative embodiments of the disclosed technology.
[0004] In one aspect, a computer-implemented method for detecting obstacles using one or more sensors carried by a mobile platform includes: obtaining sensor data indicating at least a portion of an environment surrounding the mobile platform from the one or more sensors; determining depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment.
[0005] In some embodiments, the one or more sensors include at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0006] In some embodiments, the depth information comprises a point cloud.
[0007] In some embodiments, the point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0008] In some embodiments, the depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
[0009] In some embodiments, the depth information is generated by feeding the feature map to a depth estimation neural network separate from the obstacle detection neural network.
[0010] In some embodiments, the sensor data comprises image data, and wherein the feature map is generated at least by feeding the image data to a base neural network separate from the obstacle detection neural network.
[0011] In some embodiments, the sensor data comprises a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from the image data.
[0012] In some embodiments, the projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the image data.
[0013] In some embodiments, the projected data includes at least one of a height, distance, or angle measurement for each grid block of the 2D grid.
[0014] In some embodiments, the feature map is also generated by feeding the projected data to the base neural network.
[0015] In some embodiments, the feature map is smaller in size than the image data.
[0016] In some embodiments, at least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
[0017] In some embodiments, the method further comprises: feeding the feature map to an intermediate neural network to generate the plurality of candidate regions defined according to the image data.
[0018] In some embodiments, each candidate region is a two-dimensional (2D) region that includes a corresponding target pixel of the image data.
[0019] In some embodiments, the corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
[0020] In some embodiments, at least two of the base neural network, the intermediate neural network, and the obstacle detection neural network are jointly trained.
[0021] In some embodiments, the obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate area in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate areas.
[0022] In some embodiments, the obstacle detection neural network includes a second subnetwork configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
[0023] In some embodiments, the obstacle detection neural network includes a third subnetwork, which is configured to predict, for each candidate region in the subset, at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to the candidate region based at least in part on the initial 3D position and the one or more region features.
[0024] In some embodiments, the mobile platform includes at least one of the following: an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, a robot, a smart wearable device, a virtual reality (VR) head-mounted display, or an augmented reality (AR) head-mounted display.
[0025] In some embodiments, the method further comprises controlling a mobility function of the mobile platform based at least in part on the control command.
[0026] In some embodiments, the method further comprises enabling navigation of the mobile platform based at least in part on at least one of the predicted type, posture, orientation, three-dimensional position, or three-dimensional size of the one or more obstacles.
[0027] In another aspect, a non-transitory computer-readable medium storing computer-executable instructions that, when executed, cause one or more processors associated with a mobile platform to perform actions including: obtaining sensor data indicating at least a portion of an environment surrounding the mobile platform from one or more sensors carried by the mobile platform; determining depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment.
[0028] In some embodiments, the one or more sensors include at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0029] In some embodiments, the depth information includes the point cloud.
[0030] In some embodiments, the point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0031] In some embodiments, the depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
[0032] In some embodiments, the depth information is generated by feeding the feature map to a depth estimation neural network separate from the obstacle detection neural network.
[0033] In some embodiments, the sensor data comprises image data, and wherein the feature map is generated at least by feeding the image data to a base neural network separate from the obstacle detection neural network.
[0034] In some embodiments, the sensor data comprises a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from the image data.
[0035] In some embodiments, the projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the image data.
[0036] In some embodiments, the projected data includes at least one of a height, distance, or angle measurement for each grid block of the 2D grid.
[0037] In some embodiments, the feature map is also generated by feeding the projected data to the base neural network.
[0038] In some embodiments, the feature map is smaller in size than the image data.
[0039] In some embodiments, at least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
[0040] In some embodiments, the action further comprises: feeding the feature map to an intermediate neural network to generate the plurality of candidate regions defined according to the image data.
[0041] In some embodiments, each candidate region is a two-dimensional (2D) region that includes a corresponding target pixel of the image data.
[0042] In some embodiments, the corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
[0043] In some embodiments, at least two of the base neural network, the intermediate neural network, and the obstacle detection neural network are jointly trained.
[0044] In some embodiments, the obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate area in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate areas.
[0045] In some embodiments, the obstacle detection neural network includes a second subnetwork configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
[0046] In some embodiments, the obstacle detection neural network includes a third subnetwork, which is configured to predict, for each candidate region in the subset, at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to the candidate region based at least in part on the initial 3D position and the one or more region features.
[0047] In some embodiments, the mobile platform includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, or a robot.
[0048] In some embodiments, the actions further include enabling navigation of the mobile platform based at least in part on at least one of the predicted type, posture, orientation, three-dimensional position, or three-dimensional size of the one or more obstacles within the environment.
[0049] In another aspect, a mobile platform includes a programmed controller that at least partially controls one or more movements of the mobile platform, wherein the programmed controller includes one or more processors that are configured to: obtain sensor data indicating at least a portion of an environment surrounding the mobile platform from one or more sensors; determine depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feed the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment.
[0050] In some embodiments, the one or more sensors include at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0051] In some embodiments, the depth information includes the point cloud.
[0052] In some embodiments, the point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, or a monocular camera.
[0053] In some embodiments, the depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
[0054] In some embodiments, the depth information is generated by feeding the feature map to a depth estimation neural network separate from the obstacle detection neural network.
[0055] In some embodiments, the sensor data comprises image data, and wherein the feature map is generated at least by feeding the image data to a base neural network separate from the obstacle detection neural network.
[0056] In some embodiments, the sensor data comprises a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from the image data.
[0057] In some embodiments, the projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the image data.
[0058] In some embodiments, the projected data includes at least one of a height, distance, or angle measurement for each grid block of the 2D grid.
[0059] In some embodiments, the feature map is also generated by feeding the projected data to the base neural network.
[0060] In some embodiments, the feature map is smaller in size than the image data.
[0061] In some embodiments, at least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
[0062] In some embodiments, the one or more processors are further configured to feed the feature map to an intermediate neural network to generate the plurality of candidate regions defined based on the image data.
[0063] In some embodiments, each candidate region is a two-dimensional (2D) region that includes a corresponding target pixel of the image data.
[0064] In some embodiments, the corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
[0065] In some embodiments, at least two of the base neural network, the intermediate neural network, and the obstacle detection neural network are jointly trained.
[0066] In some embodiments, the obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate area in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate areas.
[0067] In some embodiments, the obstacle detection neural network includes a second subnetwork configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
[0068] In some embodiments, the obstacle detection neural network includes a third subnetwork, which is configured to predict, for each candidate region in the subset, at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to the candidate region based at least in part on the initial 3D position and the one or more region features.
[0069] In some embodiments, the mobile platform includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, or a robot.
[0070] In some embodiments, the one or more processors are further configured to enable navigation of the mobile platform based at least in part on at least one of the predicted type, posture, orientation, three-dimensional position, or three-dimensional size of the one or more obstacles within the environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is a schematic diagram of a representative system 100 having elements configured in accordance with some embodiments of the disclosed technology.
[0072] Figure 2 is a flow chart illustrating a method of using a hierarchy of artificial neural networks (ANNs) for detecting obstacles for a mobile platform, according to some embodiments of the presently disclosed technology.
[0073] Figure 3 is a flow chart illustrating another method of using a hierarchy of ANNs for detecting obstacles for a mobile platform in accordance with some embodiments of the presently disclosed technology.
[0074] Figure 4A and Figure 4B An example of a 2D grid and candidate regions identified thereon is shown in accordance with some embodiments of the presently disclosed technology.
[0075] Figure 5 is a flow chart illustrating an obstacle detection process using an obstacle detection network in accordance with some embodiments of the presently disclosed technology.
[0076] Figure 6 An example of a mobile platform configured according to various embodiments of the technology of this disclosure is shown.
[0077] Figure 7 is a block diagram illustrating an example of the architecture of a computer system or other control device that can be used to implement various portions of the techniques of this disclosure.
[0078] Figure 8 is a flow chart illustrating a candidate area determination process using a candidate area network according to some embodiments of the presently disclosed technology.
[0079] Fig. 9 An example process for detecting obstacles using one or more sensors carried by a mobile platform in accordance with some embodiments of the presently disclosed technology is shown.
[0080] Fig.10 An example process of generating a feature map according to some embodiments of the presently disclosed technology is shown.
[0081] Fig.11A and Fig. 11BAn example of cascaded convolution and pooling layers used in a base neural network according to some embodiments of the presently disclosed technology and example data involved therewith is shown.
[0082] Fig.12 An example implementation of generating a feature map according to some embodiments of the presently disclosed techniques is shown. DETAILED DESCRIPTION
[0083] 1. Overview
[0084] Obstacle detection is an important aspect of automatic or unmanned navigation technology. Image data and / or point cloud data collected by sensors (e.g., cameras or light radar sensors) carried by a mobile platform (e.g., an unmanned car, ship, or aircraft) can be used as the basis for detecting obstacles in an environment that is otherwise observable around the mobile platform or from the mobile platform. The 2D position (e.g., on the image), orientation, attitude, 3D position and size, and / or other attributes of the obstacle can be useful in various advanced navigation applications. The accuracy and efficiency of obstacle detection and 3D positioning can determine the safety and reliability of the corresponding navigation system to some extent.
[0085] Typically, 3D information reconstructed from stereo camera data (e.g., images) is less accurate than 3D point clouds generated by light radar or other emission / detection sensors. Therefore, obstacle detection methods using point clouds generated by light radar may not be applicable to stereo camera data. On the other hand, image-based obstacle detection methods typically only output the 2D position of obstacles and are unable to determine the precise 3D position of obstacles in the physical world. Aspects of the technology disclosed herein use a hierarchical structure of artificial neural networks (ANNs) and a region-based approach to detect obstacles and determine their state attributes (e.g., positioning information) based on various data collected from sensors. In some embodiments, a specific structure of an ANN hierarchical structure in which multiple ANNs are interconnected by specifically defined inputs and outputs contributes to various advantages and improvements (e.g., computational efficiency, detection accuracy, system robustness, etc.) of the technology disclosed herein, etc. As will be appreciated by those skilled in the art, an ANN is a computing system that "learns" a task (i.e., gradually improves performance thereon) by considering examples, generally without the need for task-specific programming. For example, in image recognition, an ANN can learn to recognize images containing cats by analyzing example images that have been manually labeled as "cat" or "non-cat" and using the results to recognize cats in other images.
[0086] ANN is usually based on a collection of connected units or nodes called artificial neurons. Each connection between artificial neurons can send a signal from one artificial neuron to another artificial neuron. The artificial neuron that receives the signal can process the signal and then signal the artificial neuron connected to it. Usually, in ANN implementations, the signal at the connection between artificial neurons is a real number, and the output of each artificial neuron is calculated by a nonlinear function of the sum of its inputs. Artificial neurons and connections usually have weights that are adjusted as learning proceeds. The weight increases or decreases the intensity of the signal at the connection. Artificial neurons can have a threshold value so that a signal is sent only when the aggregate signal crosses the threshold value. Usually, artificial neurons are organized in layers. Different layers can perform different types of transformations on their inputs. The signal may travel from the first (input) layer to the last (output) layer after passing through each layer many times.
[0087] In some embodiments, one or more ANNs used by the technology of the present disclosure include convolutional neural networks (CNNs or ConvNets). Typically, CNNs use variants of multilayer perceptrons designed to require minimal preprocessing. CNNs can also be translation-invariant or spatially invariant artificial neural networks (SIANNs) based on their shared weight architecture and translation-invariant properties. As an illustration, CNNs are inspired by biological processes, in which the connection patterns between neurons are similar to the organization of the visual cortex of animals. Neurons in each cortex respond to excitations only in a restricted area of the visual field called a receptive field. The receptive fields of different neurons partially overlap, so that they cover the entire visual field.
[0088] In some embodiments, the technology of the present disclosure implements various ANNs in a hierarchical structure and interconnects the ANNs to achieve more accurate and / or efficient obstacle detection. In some aspects, the base neural network can receive 3D point cloud data, stereo camera image data, and / or monocular image data, and generate intermediate features (e.g., feature maps) for feeding into one or more other neural networks. In some aspects, the candidate area neural network can at least receive the intermediate features and determine a 2D candidate area indicating at least a portion of an obstacle within the environment. In some aspects, the obstacle detection neural network can receive environmental depth information (e.g., 3D point cloud or depth map data), intermediate features, and candidate areas, and predict various attributes of the detected obstacles.
[0089] For the sake of clarity, some details of the following description structures and / or processes are not set forth in the following description: these structures or processes are well known and commonly associated with mobile platforms (e.g., UAVs, automobiles, or other types of mobile platforms) and corresponding systems and subsystems, but may unnecessarily obscure some important aspects of the technology of the present disclosure. In addition, although the following disclosure sets forth several embodiments of different aspects of the technology of the present disclosure, some other embodiments may have different configurations or different components than those described herein. Accordingly, the technology of the present disclosure may have other embodiments that have additional elements and / or do not have the following references. Figures 1 to 9 Several elements are described.
[0090] Provided Figures 1 to 9 The accompanying drawings are not intended to limit the scope of the claims in this application unless otherwise specified.
[0091] Many embodiments of the present technology described below may take the form of computer or controller executable instructions, including routines executed by a programmable computer or controller. A programmable computer or controller may or may not reside on a corresponding mobile platform. For example, a programmable computer or controller may be an onboard computer of a mobile platform, or a separate but dedicated computer associated with a mobile platform, or a part of a network-based or cloud-based computing service. Those skilled in the relevant art will appreciate that, in addition to those shown and described below, the technology may also be implemented on a computer or controller system. The technology may be embodied in a dedicated computer or data processor that is specially programmed, configured or constructed to execute one or more computer executable instructions described below. Therefore, the terms "computer" and "controller" generally used herein refer to any data processor, and may include Internet devices and handheld devices (including palmtop computers, wearable computers, cellular or mobile phones, multiprocessor systems, processor-based or programmable consumer electronics, network computers, microcomputers, etc.). The information processed by these computers and controllers may be presented on any appropriate display medium including an LCD (liquid crystal display). Instructions for performing computer or controller executable tasks may be stored in or on any suitable computer readable medium, including hardware, firmware, or a combination of hardware and firmware. Instructions may be contained in any suitable storage device, including, for example, a flash drive, a universal serial bus (USB) device, and / or other suitable media. In certain embodiments, instructions are correspondingly non-transitory.
[0092] 2. Representative Examples
[0093] Figure 11 is a schematic diagram of a representative system 100 having elements configured according to some embodiments of the disclosed technology. System 100 includes a mobile platform 110 (e.g., an autonomous vehicle) and a control system 120. Mobile platform 110 may be any suitable type of movable object that may be used in various embodiments, such as an unmanned aerial vehicle, a manned aircraft, an autonomous vehicle, a self-balancing vehicle, or a robot.
[0094] The mobile platform 110 may include a body 112 that may carry a load 114. Many different types of loads may be used according to the embodiments described herein. In some embodiments, the load includes one or more sensors, such as an imaging device or an optoelectronic scanning device. For example, the load 114 may include an optical radar, a radar, a time-of-flight (ToF) camera, a stereo camera, a monocular camera, a video camera, and / or a still camera. The camera may be sensitive to wavelengths in any of a variety of suitable bands, including visible light, ultraviolet light, infrared light, and / or other bands. The load 114 may also include other types of sensors and / or other types of cargo (e.g., packages or other deliverables). In some embodiments, the load 114 is supported relative to the body 112 using a carrying mechanism 116 (e.g., a gimbal, a luggage rack, or a pole). The carrying mechanism 116 may allow the load 114 to be independently positioned relative to the body 112.
[0095] The mobile platform 110 may be configured to receive control commands from the control system 120 and / or send data to the control system 120. Figure 1 In the embodiment shown in FIG. 1 , the control system 120 includes some components carried on the mobile platform 110 and / or some components located outside the mobile platform 110. For example, the control system 120 may include a first controller 122 carried by the mobile platform 110 and / or a second controller 124 (e.g., a manually operated remote control) located away from the mobile platform 110 and connected via a communication link 128 (e.g., a wireless link such as a radio frequency (RF) based link). The first controller 122 may include a computer-readable medium 126 that executes instructions to direct the actions of the mobile platform 110, including but not limited to the operation of various components of the mobile platform including a payload 162 (e.g., a camera). The second controller 124 may include one or more input / output devices, such as display buttons and control buttons. In some embodiments, the operator at least partially manipulates the second controller 124 to remotely control the mobile platform 110 and receives feedback from the mobile platform 110 via a display interface and / or other interface on the second controller 124. In some embodiments, the mobile platform 110 operates automatically, in which case the second controller 124 may be eliminated or used only by the operator for invalid functions.
[0096] To provide safe and efficient operation, being able to automatically or semi-automatically detect obstacles and / or engage in avoidance maneuvers to avoid obstacles can be beneficial for autonomous vehicles, UAVs, and other types of unmanned vehicles. Additionally, sensing environmental objects can be useful for mobile platform functions such as navigation, target tracking, and mapping, particularly when the mobile platform is operating in a semi-autonomous or fully autonomous manner.
[0097] Thus, the mobile platforms described herein may include one or more sensors configured to detect objects in the environment surrounding the mobile platform (e.g., separate and independent from the load-type sensors). In some embodiments, the mobile platform includes one or more sensors configured to measure the distance between an object and the mobile platform (e.g., Figure 1 The distance measuring device 140 can be carried by the mobile platform in various ways, for example, above, below, on the side or within the body of the mobile platform. Optionally, the distance measuring device can be coupled to the mobile platform via a gimbal or other carrying mechanism that allows the device to translate and / or rotate relative to the mobile platform. In some embodiments, the distance measuring device is an optical distance measuring device that uses light to measure the distance to an object. The optical distance measuring device can be a light radar system or a laser rangefinder. In some embodiments, the distance measuring device is a camera that can image data from which depth information can be determined. The camera can be a stereo camera or a monocular camera.
[0098] Fig. 9 An example process 900 for detecting obstacles using one or more sensors carried by a mobile platform in accordance with some embodiments of the technology disclosed herein is shown. At box 910, the process 900 includes obtaining sensor data (e.g., a point cloud, a depth map, an image, etc.) indicating at least a portion of an environment surrounding the mobile platform from one or more sensors (e.g., a lidar, a radar, a time-of-flight (ToF) camera, a stereo camera, a monocular camera, etc.). At box 920, the process 900 includes determining depth information (e.g., a depth map, a point cloud, etc.), a feature map (e.g., a 2D grid-based feature), and a plurality of candidate regions (e.g., regions defined on a 2D grid such as an image) based at least in part on the sensor data. As an illustration, each candidate region indicates at least a portion of an obstacle within the environment.
[0099] In the context of blocks 910 and 920, Fig.12 An example implementation of generating a feature map according to some embodiments of the technology disclosed herein is shown. Fig.12 , the image and preliminary features derived from the point cloud can be fed to the base neural network. As discussed above, the image and point cloud can be obtained at block 910 of process 900. As will be described below with reference to Figure 2 As discussed in detail, a pre-processing module (e.g., a portion of a controller associated with a mobile platform) can project 3D points in the point cloud onto a 2D grid (e.g., a 2D plane) defined according to the image, thereby generating a set of preliminary features based on the 2D grid including, for example, height values, angle values, and / or distance values. As will be described below with reference to Figure 2 , Fig.10 , Fig.11A and Fig. 11B As discussed in detail, the base neural network may include multiple transformation layers that may transform the image and preliminary features into one or more feature maps for further processing. In some embodiments, as will be described below with reference to Figure 3 , Fig.10 , Fig.11A and Fig. 11B As discussed in detail, the base neural network may include multiple transformation layers that can transform an image into one or more feature maps for further processing.
[0100] Go back for reference Fig. 9 At block 930, process 900 includes feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment. The state attribute includes at least one of a type, a pose, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles within the environment.
[0101] More specifically, Figure 2 2 is a flowchart illustrating a method 200 for using a hierarchical structure of an ANN for detecting obstacles for a mobile platform according to some embodiments of the technology disclosed herein. The method 200 may be implemented by a controller (e.g., an onboard computer of a mobile platform, an associated computing device, and / or an associated computing service).
[0102] refer to Figure 2, the controller may use one or more sensors carried by the mobile platform to obtain point cloud data 202 (or another form of depth information) and image data 204. As discussed above, a stereo camera, a lidar, a radar, a time of flight (ToF) camera, a stereo camera, or a monocular camera, or other sensor may provide data for obtaining depth information (e.g., distance measurements between different parts of the scene and the sensor) of an environment surrounding or otherwise adjacent to the mobile platform but not necessarily adjacent to the mobile platform. The point cloud data 202 may be obtained directly from a depth detection sensor (e.g., a lidar, a radar, or a time of flight (ToF) camera) or directly from reconstruction using a stereo camera or a monocular camera. As an illustration, the point cloud data (or another form of depth information) is obtained based on sensor data generated by a lidar, a radar, a time of flight (ToF) camera, a stereo camera, a monocular camera, etc. In some embodiments, the depth information may include a depth map determined at least in part based on disparity data generated directly or indirectly from the image data 204.
[0103] As an illustration, the image data 204 may be provided by a stereo camera or by a monocular camera. In some embodiments, the controller obtains a temporally continuous series of point clouds and images (e.g., point cloud frames and image frames). In some embodiments, point clouds and images corresponding to the same moment are used in the method 200.
[0104] The controller feeds the point cloud data 202 to a pre-processing module 210 (e.g., a neural network of one or more layers), which outputs a preliminary feature 212 based on a 2D grid. As an illustration, the pre-processing module 210 can be implemented to project the point cloud data 202 onto a 2D grid defined according to the image data to obtain projected data, wherein the 2D grid has the same size as the obtained image data 204. Figure 4A An example of such a 2D grid is shown. Projecting the point cloud data 202 may be performed based on extrinsic calibration parameters and / or intrinsic calibration parameters associated with the camera that generated the image data 204. For example, if the image data 204 has a size of 720x1280 pixels, the pre-processing module 210 may project the scanned points of the point cloud onto each 720x1280 grid block 402 corresponding to the pixels of the image data 204. In other words, each pixel may correspond to a grid block 402 of the 2D grid. Features such as height (e.g., z coordinate of 3D coordinates), depth (e.g., distance to the mobile platform), angle (e.g., normal vector), etc. may be calculated for each grid block based on projecting the point cloud data 202 onto the 2D grid defined according to the image data. In this way, the preliminary features 212 based on the 2D grid may have 720x1280 grid blocks, and each grid block may include one or more features (e.g., height, distance, or angle measurement) derived from the point cloud data.
[0105] According to an example implementation, point cloud data 202 is projected from a 3D coordinate system associated with a camera that generated image data 204 to a 2D grid of image data 204. If the number of points in point cloud data 202 is N, and the 3D coordinates of each point are p = (x, y, z), then the points may be projected based on:
[0106]
[0107]
[0108] Among them, f x , f y is the focal length, c x , c y is the optical center coordinate (e.g., f x , f y and c x , c y (can be obtained from the intrinsic calibration parameters associated with the camera), and (u, v) are the pixel coordinates after the point is projected. Thus, a correspondence or mapping between a three-dimensional point (x, y, z) and a pixel coordinate (u, v) is established.
[0109] For each pixel coordinate (u, v), in order to generate preliminary features 212, the controller can perform feature encoding based on its corresponding 3D coordinates (x, y, z). For example, according to some encoding schemes, the controller can encode the point cloud into a set of 3-channel preliminary features. The 3-channel features can represent distance, height and angle respectively. Feature encoding can be based on:
[0110]
[0111]
[0112]
[0113] Among them, c1, c2, c3 are distance values, height values, and angle values respectively, and α z , α y is a normalization coefficient (which may be predetermined). Therefore, the projection of the point cloud data 202 may generate a set of 2D grid-based preliminary features each including a corresponding distance value, a height value, and an angle value.
[0114] Go back for reference Figure 2, the controller may feed the preliminary features 212 (e.g., the projected data) and the image data 204 to a basic neural network 220 (e.g., including one or more CNNs). In various embodiments, the basic neural network 220 may include one or more convolution operation layers and / or pooling (e.g., downsampling) operation layers. As an illustration, using multiple layers of feature transformation, the basic neural network 220 may output a feature map 222 based on the preliminary features 212 and the image data 204. In some embodiments, the feature map 222 may take the form of a 2D grid-based feature map that is smaller in size than the preliminary features 212. For example, the basic neural network may include a plurality of modules cascaded in sequence (e.g., each including one or more convolution layers and / or pooling layers), each of the plurality of modules performing nonlinear feature transformation (e.g., convolution) and / or pooling (e.g., downsampling) on the input features (the preliminary features 212 and the image data 204). After multiple levels of feature transformation and / or pooling operations, the basic neural network may output a feature map 222 of a smaller size than the preliminary features 212.
[0115] In this regard, Fig.10 An example process for generating a feature map is shown. In general, a CNN may include two main types of network layers, namely, convolutional layers and pooling layers. Convolutional layers may be used to extract various features from input (e.g., an image) to the convolutional layers. Pooling layers may be used to compress features input to the pooling layers, thereby reducing the number of training parameters of the neural network and alleviating the degree of model overfitting. Fig.10 , if the input preliminary features are 32*32 in size, then after the convolution operation, the preliminary features can be transformed into a first set of 6 feature maps. The first set of feature maps each has a size of 28*28. After performing a pooling operation on the first set of feature maps, a second set of 6 feature maps is generated. The second set of feature maps each has a size of 14*14.
[0116] Fig.11AAn example of cascaded convolutional layers and pooling layers used in a basic neural network according to some embodiments of the technology disclosed herein is shown. As shown, the basic neural network includes a network having 3 convolutional layers (e.g., resulting in operations of C1, C3, and C5, respectively) and 3 pooling layers (e.g., resulting in operations of S2, S4, and S6, respectively) connected in series in a cascaded manner. Using the cascaded convolutional layers and pooling layers, the input (64*64 size) is transformed into a first set (C1) of 6 feature maps (60*60 size), and then transformed into a second set (S2) of 6 feature maps (30*30 size), a third set (C3) of 16 feature maps (26*26 size), a fourth set (S4) of 16 feature maps (13*13 size), a fifth set (C5) of feature maps (10*10 size), and a sixth set (C6) of feature maps (5*5 size) as output. In some embodiments, the sixth set of feature maps (C6) may also be transformed by a fully connected layer and / or a Gaussian layer into a vector, for example, as an output from a basic neural network. Fig.11A Example of input and feature map collection of the cascade structure.
[0117] Go back for reference Figure 2 , the feature map 222 may be a common input to multiple neural networks or applicable components for implementing obstacle detection and / or 3D positioning according to various embodiments of the technology disclosed herein. Figure 2 , the controller may feed the feature map 222 to a candidate region neural network 230 (e.g., including one or more CNNs) to generate a plurality of candidate regions 204 defined according to the image data. In various embodiments, the candidate region neural network 230 may include one or more layers of convolution and / or pooling operations.
[0118] Figure 8 is a diagram illustrating the use of a candidate region neural network 830 (e.g., with Figure 2 Flow chart of candidate region determination process 800 corresponding to the candidate region neural network 230 in FIG. Figure 8 The candidate region neural network 830 may include one or more modules 840-870 for feature transformation, likelihood estimation, 2D grid regression, and / or redundancy filtering.
[0119] As an illustration, the feature transformation module 840 receives the feature map 822 as input, which is further transformed to be fed to (a) a likelihood estimation module 850 that can predict the probability that each pixel in the image data 204 (or each grid block in the corresponding 2D grid) "belongs" to an obstacle, and (b) a 2D grid regression module 860 that can determine the corresponding 2D region representing the obstacle to which the pixel (or grid block) "belongs". The outputs from the likelihood estimation module 850 and the 2D grid regression module 860 are fed to a redundancy filtering module 870 that can filter out redundant 2D regions based on the predicted probabilities and / or overlap of the 2D regions. The redundancy filtering module 870 can then output a smaller number of candidate regions to be included in the candidate region data 832 (e.g., the candidate region data 232).
[0120] Go back for reference Figure 2 , as an illustration, the candidate region neural network 230 may include one or more CNNs. The candidate region neural network 230 may output candidate region data 232 including or indicating candidate regions. Each candidate region may indicate at least a portion of an obstacle within the environment (e.g., a portion of an image showing at least some portion of an obstacle). For example, Figure 4B An example candidate region 410 identified on a 2D grid is shown. As an illustration, the 2D grid may be defined based on the image data 204 (eg, smaller in size), and in some embodiments, the 2D grid used to identify the candidate region may be the image data 204.
[0121] The candidate region may be a 2D region including a set of connected or unconnected grid blocks having a basis block 412. Each grid block may correspond to a respective pixel or block (e.g., 3x3 in size) of pixels in the image data 204. Using each grid block of the 2D grid as a basis block 412, the candidate region neural network 230 may output (1) a likelihood (e.g., an estimated probability) that the basis block 412 indicates a portion of an obstacle, and (2) a candidate region 410 including the basis block 412 and likely corresponding to at least a portion of the obstacle. In some embodiments, various criteria may be applied to the candidate regions and / or the likelihood associated therewith for selecting a subset of data for output by filtering out redundant candidates. For example, candidate regions that exceed a threshold level of overlap with one or more other candidate regions may be excluded from the output. As another example, if the likelihood that the corresponding basis block of the candidate region belongs to an obstacle falls below a threshold, they may be filtered out.
[0122] According to the context generally described above, the technology of the present disclosure may include (1) data acquisition and preprocessing aspects, and (2) feature map and candidate region determination aspects. By way of implementation example, the data acquisition and preprocessing aspects may include: obtaining a stereo color image (e.g., with a resolution of 720*1280) acquired by a stereo camera carried by a moving mobile platform; generating a 3D point cloud via 3D reconstruction based on the stereo image; and determining input features based on the 3D point cloud and the image.
[0123] In order to determine the input features, preliminary features can be obtained from the 3D point cloud. As an illustration, the point cloud is projected onto a 2D grid (e.g., a plane) of the image using the calibration parameters of the stereo camera. The projection can result in preliminary features (e.g., projected data) of the same size as the image (e.g., 720*1280). For each grid block of the 2D grid, various features (e.g., height, distance, or angle measurements) can be calculated based on actual needs and / or computational efficiency. For example, 3 features associated with each grid block (e.g., height, distance, and angle) can be calculated, so the preliminary features have a dimension of 720*1280*3.
[0124] Next, (a) preliminary features (e.g., with dimensions of 720*1280*3) and (b) the left eye (or right eye) image of the stereo image can be concatenated or otherwise combined. As an illustration, the left eye image is an RGB image also with dimensions of 720*1280*3. The combination of (a) and (b) thus generates an input feature with dimensions of 720*1280*6, which can be used as input for determining feature maps and candidate regions.
[0125] According to a specific example, the feature map and candidate region determination aspects may include the use of a basic neural network and a candidate neural network. Based on actual needs and / or computational efficiency, the neural network may include layers of convolution operations and pooling operations.
[0126] As an illustration, the basic neural network receives input features (e.g., with a dimension of 720*1280*6) as input. The basic neural network may include 4 modules cascaded in sequence, each of which performs nonlinear feature transformation and 2x downsampling. Therefore, after 4 rounds of downsampling, the basic neural network may output a feature map with a resolution of 45*80. The candidate region neural network receives the feature map as input, predicts the probability that each pixel in the left eye image "belongs" to an obstacle and the corresponding 2D region represents or indicates the obstacle. Based on the predicted probability, the candidate region neural network filters out redundant 2D regions and outputs the remaining 2D regions as candidate regions representing or indicating obstacles. The number of candidate regions may be several hundred, for example, 500, 400, 300 or less.
[0127] Go back for reference Figure 2 , the controller can feed the point cloud data 202, the feature map 222, and the candidate region data 232 to the obstacle detection neural network 240 (e.g., including one or more CNNs). Figure 5 As discussed in more detail, the obstacle detection neural network 240 may output one or more state attributes 242 of a detected obstacle including a predicted type, pose, orientation, 3D position, 3D size, and / or other attributes of a detected obstacle among one or more obstacles within the environment. The controller may output commands or instructions based on the state attributes 242 of the detected obstacle to control at least some movement of the mobile platform (e.g., acceleration, deceleration, steering, etc.) to avoid contact with the detected obstacle.
[0128] Figure 3 3 is a flow chart illustrating a method 300 of using a hierarchical structure of an ANN for detecting obstacles for a mobile platform according to some embodiments of the technology disclosed herein. The method 300 may be implemented by a controller (e.g., an onboard computer of a mobile platform, an associated computing device, and / or an associated computing service). In some embodiments, the method 300 may use various combinations of data obtained by a stereo camera or a monocular camera at different layers or levels of the hierarchical structure of the ANN to achieve obstacle detection and 3D positioning.
[0129] refer to Figure 3 , the controller may obtain image data 302 using a camera or other visual sensor carried by the mobile platform. As discussed above, the image data generated by the camera may provide a basis for obtaining depth information (e.g., measurements of distances between different parts of the scene and the sensor) of an environment surrounding the mobile platform or otherwise adjacent to the mobile platform but not necessarily adjacent to the mobile platform. In various embodiments, the image data 302 may be provided by a stereo camera and / or a monocular camera. In some embodiments, the controller obtains a stereo image and / or a monocular image corresponding to a particular moment. In some embodiments, the controller obtains a temporally continuous series of images (e.g., frame images).
[0130] More specifically, the controller may feed the image data 302 into a base neural network 320 (e.g., including one or more CNNs). The base neural network 320 may be structurally equivalent to, similar to, or different from the one described in reference Figure 2The base neural network 220 used in the method 200 described above. In various embodiments, the base neural network 320 may include one or more layers of convolution and / or pooling operations. As an illustration, the base neural network 320 may output a feature map 322 based on the image data 302 using multiple layers of convolution and / or pooling operations. The feature map 322 may take the form of a 2D grid-based feature map that is smaller in size than each image included in the image data 302.
[0131] Continue to refer Figure 3 , the controller may feed the feature map 322 into a depth estimation neural network 310 (e.g., including one or more CNNs), which outputs depth information 312 (e.g., a depth map). As an illustration, the depth estimation neural network 310 may be implemented to estimate depth information corresponding to different locations defined within the image data 302. For example, for a target image included in the image data 302, the depth estimation neural network 310 may analyze multiple frames of images before and / or after the target image, and output depth information 312 including an estimated depth value (e.g., a distance from the mobile platform) for each pixel of the target image. Alternatively, the depth information (e.g., a depth map) may be determined based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
[0132] The controller may feed the feature map 322 to an intermediate neural network such as a candidate region neural network 330 (e.g., a candidate region neural network 830 that may include one or more CNNs). The candidate region neural network 330 may be structurally equivalent to, similar to, or different from the one described in reference Figure 2 The candidate region neural network 230 used in the method 200 described above. In various embodiments, the candidate region neural network 330 may include one or more layers of convolution and / or pooling operations. For example, each neuron may apply a corresponding convolution operation to their input, and the output of a cluster of neurons at one layer may be combined into a single neuron at the next layer. As an illustration, the candidate region neural network 330 may output candidate region data 332 including or indicating candidate regions. Each candidate region may indicate at least a portion of an obstacle within the environment.
[0133] As mentioned above Figure 4BAs discussed, the candidate region may be a set of connected or unconnected grid blocks including the basis block 412. Each grid block may correspond to a respective pixel or block of pixels (e.g., 2x4 in size) in the obtained image. Using each grid block of the 2D grid as the basis block 412, the candidate region neural network 330 may output (1) a likelihood (e.g., an estimated probability) that the basis block 412 indicates at least some portion of an obstacle, and (2) a corresponding candidate region 410 that includes the basis block 412 and may indicate at least a portion of an obstacle. As described above with reference to Figure 4B As discussed, various criteria may be applied to the candidate regions, their associated basis blocks, and / or possibilities for selecting a subset of data for output.
[0134] By reference Figure 8 In an implementation example of , the obtained image has a resolution of 100*50 (ie, the image has 5000 pixels) and the image includes a 2D representation of one or more obstacles including obstacle A. Figure 8 , the feature transformation module 840 receives as input the feature map 822 corresponding to the obtained image, and transforms it to feed to (a) the likelihood estimation module 850 which can predict the probability that each pixel in the image "belongs" to an obstacle. As an illustration, the likelihood estimation module 850 predicts that 100 pixels in the image "belong" to an obstacle with corresponding probabilities.
[0135] Continue to refer Figure 8 , the feature transformation module 840 also feeds the transformed feature map to the 2D grid regression module 860, which can determine the corresponding 2D regions representing the obstacles to which the pixels (or grid blocks) "belong". As an illustration, because 100 pixels "belong" to obstacle A, the 2D grid regression module 860 can determine 100 corresponding 2D regions (e.g., 2D frames) representing or indicating obstacle A.
[0136] Based on the estimated probability that each of the 100 pixels "belongs" to obstacle A, a non-maximum suppression method (or other suitable filtering method) can be used to remove a subset of regions (e.g., those that overlap each other more than a threshold degree) from the 100 2D regions. The remaining 2D regions can be retained as output candidate regions for obstacle A.
[0137] Go back for reference Figure 3 , the controller can feed the depth information 312, the feature map 322, and the candidate region data 332 into an obstacle detection neural network 340 (e.g., including one or more CNNs). The obstacle detection neural network 340 can be structurally equivalent to, similar to, or different from the one described in reference Figure 2The obstacle detection neural network 240 used in the method 200 described above. As will be discussed in more detail below, the obstacle detection neural network 340 can output one or more state attributes 342 of the detected obstacle including a prediction of the type, pose, orientation, 3D position, 3D size, and / or other attributes of the detected obstacle. The controller can output commands or instructions based on the state attributes 342 of the detected obstacle to control at least some movement of the mobile platform (e.g., acceleration, deceleration, steering, etc.) to avoid contact with the detected obstacle.
[0138] Figure 5 is to illustrate the use of some embodiments of the technology according to the present disclosure (for example, in reference Figure 2 The obstacle detection neural network 240 used in the method 200 described above or in reference Figure 3 Flow chart of obstacle detection process 500 of obstacle detection neural network 540 corresponding to obstacle detection neural network 340 used in method 300 described above. Figure 5 The obstacle detection neural network 540 may include an initial position subnetwork 510 (e.g., a first subnetwork including one or more ANNs), a region feature subnetwork 520 (e.g., a second subnetwork including one or more ANNs), and a 3D prediction subnetwork 530 (e.g., including one or more ANNs).
[0139] The initial position subnetwork 510 may receive as input the depth information 502 (e.g., the point cloud data 202 as in the method 200 or the depth information 312 as in the method 300) and the candidate region data 504 (e.g., the candidate region data 232 as in the method 200 or the candidate region data 332 output from the candidate region neural network 330 as in the method 300). If the depth information 502 is not in the form of a point cloud, embodiments of the disclosed techniques include converting the depth information 502 into point cloud data based on, for example, extrinsic calibration parameters and / or intrinsic calibration parameters of an associated camera that generated the image data 204 or 302. For each candidate region included in the candidate region data 504 (e.g., a 2D region on the obtained image or an associated 2D grid), the initial position subnetwork 510 can (a) use the depth information 312 to identify a 3D region corresponding to the candidate region (e.g., a subset of the point cloud), and (b) calculate and output an initial 3D position of possible obstacles including the identified 3D region based on various statistics characterizing the 3D region (e.g., the mean of the median values of the 3D coordinates of the corresponding scan points).
[0140] The region feature subnetwork 520 (e.g., a fully connected neural network) can receive candidate region data 504 (e.g., candidate region data 232 as in method 200 or candidate region data 332 as in method 300) and a feature map 506 (e.g., feature map 222 as in method 200 or feature map 322 as in method 300) as input; perform one or more layers of linear and / or nonlinear feature transformations; and output region features of each candidate region included in the candidate region data 504. The size of the region feature of each candidate region can be determined based on actual needs and / or computing resource limitations. As an example, the candidate regions are normalized to a fixed size, and then a fixed-length feature vector is obtained based on one or more pooling operations before performing multiple layers of feature transformations to obtain the region features of each candidate region.
[0141] As an illustration, each candidate region may correspond to a corresponding 2D region in an original image, which is used as a basis for generating a feature map 506. Using the relationship (e.g., ratio) between the size of the original image and the size of the feature map, the controller may identify a corresponding reduced 2D region on the feature map corresponding to each candidate region. Various operations may be performed on the feature map in the reduced 2D region to generate regional features. In various embodiments, the operations may include pooling and / or feature transformations (e.g., fully connected and / or convolutions). As an example, each reduced 2D region may be normalized so that each regional feature may be a fixed-length feature vector calculated based on the corresponding reduced 2D region identified on the feature map.
[0142] The 3D prediction subnetwork 530 can receive outputs from the initial position subnetwork 510 and the regional feature subnetwork 520, and output state attributes 542 of the detected obstacles. For example, the 3D prediction subnetwork 530 can predict and output the type, posture, orientation, 3D position, 3D size and / or other attributes of the detected obstacles. The 3D prediction subnetwork 530 can determine and output a confidence level for each candidate region indicating the probability that the candidate region belongs to an obstacle. In some embodiments, the output is filtered based on the confidence level. For example, candidate regions whose confidence levels fall below a threshold can be excluded from the output. As an illustration, one or more controllers of the mobile platform can use various state attributes of the detected obstacles to perform automatic or semi-automatic mapping, navigation, emergency maneuvers, or other actions to control certain movements of the mobile platform.
[0143] In some embodiments, the 3D prediction subnetwork 530 includes one or more submodules (e.g., neural network branches) that predict at least one of the following: semantic category (e.g., type of obstacle), 2D region (e.g., 2D region of image data), orientation, 3D size, and 3D position of the obstacle. Each submodule may include mapping the output of the region feature subnetwork 520 (e.g., region features of each candidate region) to a corresponding dimension of the submodule output.
[0144] As an illustration, the semantic category prediction submodule may predict a confidence level indicating the probability that a candidate region "belongs" to a semantic category (eg, the probability of "belonging to" a vehicle, a pedestrian, a bicycle, or background).
[0145] As an illustration, the 2D region prediction submodule may use the center point, length and width to represent the 2D region corresponding to the obstacle. The 2D region prediction submodule may estimate the deviation from each candidate region to the 2D region corresponding to the obstacle, thereby obtaining the position of the 2D region indicating the obstacle.
[0146] As an illustration, the directional prediction submodule can divide the range from -180 degrees to +180 degrees into multiple intervals (e.g., two intervals of [-180°, 0°] and [0°, 180°]), and calculate the center of each interval. The directional prediction submodule can predict the specific interval to which the directional angle of the obstacle belongs, and calculate the difference between the directional angle of the obstacle and the center of the interval to which it belongs, thereby obtaining the directional angle of the obstacle.
[0147] As an illustration, the 3D size prediction submodule may perform prediction using the average length, width, and height (or other measurements related to 3D size) of the 3D representations (e.g., frames) of obstacles for each semantic category. The average measurements may be obtained from training data collected offline. During the prediction process, the 3D size prediction submodule may predict the ratio of the 3D size of the obstacle to the average 3D size of the corresponding category, thereby obtaining the 3D size attribute of the obstacle.
[0148] As an illustration, the 3D position prediction submodule may predict the deviation between the 3D position of the obstacle and the initial 3D position of the corresponding input candidate region, thereby obtaining the 3D position of the obstacle.
[0149] In some embodiments, based on the confidence level of the predicted semantic category, the 3D prediction subnetwork 530 can also filter its output to retain only those outputs with a confidence level greater than a certain threshold. Various suitable filtering methods (e.g., non-maximum suppression) can be used based on actual needs and / or computational efficiency.
[0150] According to the above description, state attributes such as semantic category, 2D area, orientation, 3D size and 3D position of obstacles in the current road scene of the mobile platform can be obtained. This output can be provided to downstream applications of the mobile platform, such as route planning and control, to facilitate automatic navigation, automatic driving or other functions.
[0151] The various ANN components used in the embodiments of the technology according to the present disclosure can be trained in various ways deemed appropriate by those skilled in the art. As an illustration, training samples can be collected in advance, each sample containing input data (e.g., point clouds and corresponding images for method 200, stereo images for method 300) and its associated 3D region that is manually identified as representing an obstacle. The parameters of the neural network can be learned by a sufficiently large number of training samples.
[0152] As an illustration, the base neural network 220, the candidate region neural network 230, and the obstacle detection neural network 240 can be trained separately or jointly. When trained separately, the training data for different neural networks can be independent of each other (e.g., based on different times and contexts). When trained jointly, the training data for different neural networks correspond to each other (e.g., associated with the same series of point clouds and / or image frames). In some embodiments, a portion of the ANN hierarchy used in method 200 (e.g., the base neural network 220 and the candidate region neural network 230) is jointly trained, while at least another portion of the ANN hierarchy (e.g., the obstacle detection neural network 240) is trained separately.
[0153] Similarly, the depth estimation neural network 310, the base neural network 320, the candidate region neural network 330, and the obstacle detection neural network 340 can be trained separately or jointly. For example, a suitable training method may include collecting different images and light radar point clouds corresponding to the images as training data. Because the light radar point cloud provides a depth measurement of the environment depicted by the image, the base neural network 320 and the depth estimation neural network 310 can be jointly trained based on the image (as input to the base neural network 320) and its associated depth measurement (as output from the depth estimation neural network 310). When training is performed on sufficient data samples, appropriate network parameters can be obtained.
[0154] Moreover, the obstacle detection neural network 540 can be trained alone or jointly with other neural networks disclosed herein. For example, different images and light radar point clouds corresponding to the images can be collected. The light radar point cloud can include manually marked 3D regions representing obstacles. The joint training of the neural network of method 200 can be based on the image and its corresponding light radar point cloud (as input), and various attributes of the corresponding marked 3D region (as output). When training is performed on sufficient data samples, appropriate network parameters can be obtained.
[0155] Figure 6 An example of a mobile platform configured according to various embodiments of the technology disclosed herein is shown. As shown, a representative mobile platform as disclosed herein may include at least one of the following: an unmanned aerial vehicle (UAV) 602, a manned aircraft 604, an autonomous car 606, a self-balancing vehicle 608, a ground robot 610, a smart wearable device 612, a virtual reality (VR) head-mounted display 614, or an augmented reality (AR) head-mounted display 616.
[0156] Figure 7 7 is a block diagram showing an example of the architecture of a computer system 700 or other control device that can be used to implement various portions of the technology of the present disclosure. Figure 7 In the embodiment, computer system 700 includes one or more processors 705 and memory 710 connected via interconnect 725. Interconnect 725 can represent any one or more separate physical buses, point-to-point connections, or both connected by appropriate bridges, adapters, or controllers. Thus, interconnect 725 can include, for example, a system bus, a peripheral component interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a small computer system interface (SCSI) bus, a universal serial bus (USB), an IIC (I2C) bus, or an Institute of Electrical and Electronics Engineers (IEEE) standard 674 bus (sometimes referred to as "Firewire").
[0157] The processor 705 may include a central processing unit (CPU) to control the overall operation of, for example, a host computer. In some embodiments, the processor 705 implements this by executing software or firmware stored in the memory 710. The processor 705 may be or may include one or more programmable general or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), etc., or a combination of these devices.
[0158] The memory 710 may be or include the main memory of the computer system. The memory 710 represents any suitable form of random access memory (RAM), read-only memory (ROM), flash memory, etc., or a combination of these devices. When in use, the memory 710 may contain a machine instruction set that, when executed by the processor 705, causes the processor 705 to perform operations to implement embodiments of the technology disclosed herein. In some embodiments, the memory 710 may contain an operating system (OS) 730 that manages computer hardware and software resources and provides common services for computer programs.
[0159] Also connected to processor 705 via interconnect 725 is an (optional) network adapter 715. Network adapter 715 provides computer system 700 with the ability to communicate with remote devices such as storage clients and / or other storage servers, and may be, for example, an Ethernet adapter or a Fibre Channel adapter.
[0160] The techniques described herein can be implemented, for example, by programmable circuits (e.g., one or more microprocessors) programmed with software and / or firmware, or entirely in dedicated hardwired circuits, or in a combination of these forms. The dedicated hardwired circuits can be in the form of, for example, one or more application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), and the like.
[0161] Software or firmware for implementing the techniques described herein may be stored on a machine-readable storage medium and may be executed by one or more general or special purpose programmable microprocessors. The term "machine-readable storage medium" as used herein includes any mechanism that can store information in a form accessible to a machine (a machine may be, for example, a computer, a network device, a cellular phone, a personal digital assistant (PDA), a manufacturing tool, any device with one or more processors, etc.). For example, a machine-accessible storage medium includes recordable / non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.), etc.
[0162] The term "logic" as used herein may include, for example, programmable circuitry programmed with specific software and / or firmware, dedicated hardwired circuitry, or a combination thereof.
[0163] In addition to or in place of the foregoing, some embodiments of the present disclosure have other aspects, elements, features and / or steps. These potential additions and replacements are described in the remainder of this specification. References to "various embodiments" and "certain embodiments" in this specification indicate that specific features, structures or characteristics described in conjunction with the embodiments are included in at least one embodiment of the present disclosure. These embodiments, even alternative embodiments (e.g., referred to as "other embodiments") are not mutually exclusive with other embodiments. In addition, various features that can be exhibited by some embodiments rather than by other embodiments are described. Similarly, various requirements are described, which may be requirements of some embodiments but not requirements of other embodiments. For example, some embodiments use depth information generated by a stereo camera, while other embodiments may use depth information generated by a light radar, 3D-ToF or RGB-D. Some other embodiments may use depth information generated by a combination of sensors. As used herein, phrases such as "and / or" in "A and / or B" refer to separate A, separate B, and both A and B.
[0164] To the extent any material incorporated by reference herein conflicts with the present disclosure, the present disclosure controls.
Claims
1. A computer-implemented method for detecting obstacles using both a laser unit and a camera unit carried by a common autonomous vehicle, the method comprising: determining preliminary features based at least in part on projecting a point cloud obtained by the laser unit onto a two-dimensional grid corresponding to an image obtained by the camera, wherein the point cloud includes a three-dimensional measurement of at least a portion of an environment surrounding the autonomous vehicle, and wherein the image includes a two-dimensional representation of the portion of the environment surrounding the autonomous vehicle; Feeding the preliminary features and the image to a base neural network to generate a feature map; feeding the feature map to an intermediate neural network to generate a plurality of candidate regions of the image indicating at least a portion of obstacles within the environment surrounding the autonomous vehicle; feeding the point cloud, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one of a type, pose, orientation, three-dimensional position, or three-dimensional size of one or more obstacles within the environment; and Navigation of the autonomous vehicle is achieved based at least in part on at least one of the predicted type, attitude, orientation, three-dimensional position, or three-dimensional size of the one or more obstacles.
2. A computer-implemented method for detecting obstacles using a stereo camera unit carried by an autonomous vehicle, the method comprising: feeding image data obtained by the stereo camera unit to an underlying neural network to generate a feature map, wherein the image data comprises a two-dimensional representation of at least a portion of an environment surrounding the autonomous vehicle; feeding the feature map to an intermediate neural network to generate a plurality of candidate regions of the image indicating at least a portion of obstacles within the environment surrounding the autonomous vehicle; Feeding the feature map to a depth estimation neural network to generate a depth map corresponding to the image data; Feeding the depth map, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one of a type, pose, orientation, three-dimensional position, or three-dimensional size of one or more obstacles in the environment; and Navigation of the autonomous vehicle is achieved based at least in part on at least one of the predicted type, attitude, orientation, three-dimensional position, or three-dimensional size of the one or more obstacles.
3. A computer-implemented method for detecting an obstacle using one or more sensors carried by a mobile platform, the method comprising: obtaining sensor data indicative of at least a portion of an environment surrounding the mobile platform from the one or more sensors; determining depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment; wherein the state attribute comprises at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment; and based at least in part on the predicted at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment, navigation of the mobile platform is implemented; wherein the sensor data comprises one or more images, and wherein the feature map is generated at least by feeding the one or more images to a base neural network separate from the obstacle detection neural network.
4. The method according to claim 3, wherein: The one or more sensors include at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
5. The method according to claim 3, wherein: The depth information is determined at least in part based on the point cloud.
6. The method according to claim 5, wherein: The point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
7. The method according to claim 3, wherein: The depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
8. The method according to claim 3, wherein: The depth information is generated by feeding the feature map into a depth estimation neural network separate from the obstacle detection neural network.
9. The method according to claim 3, wherein: The sensor data includes a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from at least one of the one or more images.
10. The method according to claim 9, wherein: The projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the one or more images.
11. The method according to claim 9, wherein: The projected data includes at least one of a height, a distance, or an angle measurement for each grid block of the 2D grid.
12. The method according to claim 9, wherein: The feature map is generated by further feeding the projected data to the base neural network.
13. The method according to claim 3, wherein: The size of the feature map is smaller than the size of at least one of the one or more images.
14. The method according to claim 3, wherein: At least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
15. The method according to claim 3, further comprising: The feature map is fed to an intermediate neural network to generate the plurality of candidate regions defined according to the image data.
16. The method according to claim 15, wherein: Each candidate region is a two-dimensional 2D region including a corresponding target pixel of the image data.
17. The method according to claim 16, wherein: The corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
18. The method according to claim 15, wherein: At least two of the basic neural network, the intermediate neural network and the obstacle detection neural network are jointly trained.
19. The method according to claim 3, wherein: The obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate region in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate regions.
20. The method according to claim 19, wherein: The obstacle detection neural network includes a second sub-network configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
21. The method according to claim 20, wherein: The obstacle detection neural network includes a third subnetwork, which is configured to: predict at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to each candidate region in the subset based at least in part on the initial 3D position and the one or more region features.
22. The method according to claim 3, wherein: The mobile platform includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, or a robot.
23. A non-transitory computer-readable medium storing computer-executable instructions that, when executed, cause one or more processors associated with a mobile platform to perform actions comprising: obtaining sensor data indicative of at least a portion of an environment surrounding the mobile platform from one or more sensors carried by the mobile platform; determining depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment; wherein the state attribute comprises at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment; and based at least in part on the predicted at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment, navigation of the mobile platform is implemented; wherein the sensor data comprises one or more images, and wherein the feature map is generated at least by feeding the one or more images to a base neural network separate from the obstacle detection neural network.
24. The computer readable medium of claim 23, wherein: The one or more sensors include at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
25. The computer-readable medium of claim 23, wherein: The depth information is determined at least in part based on the point cloud.
26. The computer-readable medium of claim 25, wherein: The point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
27. The computer-readable medium of claim 23, wherein: The depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
28. The computer-readable medium of claim 23, wherein: The depth information is generated by feeding the feature map into a depth estimation neural network separate from the obstacle detection neural network.
29. The computer-readable medium of claim 23, wherein: The sensor data includes a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from at least one of the one or more images.
30. The computer-readable medium of claim 29, wherein: The projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the one or more images.
31. The computer-readable medium of claim 29, wherein: The projected data includes at least one of a height, a distance, or an angle measurement for each grid block of the 2D grid.
32. The computer-readable medium of claim 29, wherein: The feature map is generated by further feeding the projected data to the base neural network.
33. The computer readable medium of claim 23, wherein: The size of the feature map is smaller than the size of at least one of the one or more images.
34. The computer-readable medium of claim 23, wherein: At least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
35. The computer-readable medium of claim 23, wherein: The actions also include: feeding the feature map to an intermediate neural network to generate the plurality of candidate regions defined according to the image data.
36. The computer readable medium of claim 35, wherein: Each candidate region is a two-dimensional 2D region including a corresponding target pixel of the image data.
37. The computer readable medium of claim 36, wherein: The corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
38. The computer-readable medium of claim 35, wherein: At least two of the basic neural network, the intermediate neural network and the obstacle detection neural network are jointly trained.
39. The computer-readable medium of claim 26, wherein: The obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate region in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate regions.
40. The computer readable medium of claim 39, wherein: The obstacle detection neural network includes a second sub-network configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
41. The computer readable medium of claim 40, wherein: The obstacle detection neural network includes a third subnetwork, which is configured to: predict at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to each candidate region in the subset based at least in part on the initial 3D position and the one or more region features.
42. The computer-readable medium of claim 23, wherein: The mobile platform includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, or a robot.
43. A mobile platform comprising a programmed controller that at least partially controls one or more movements of the mobile platform, wherein: The programmed controller includes one or more processors configured to: obtaining sensor data indicative of at least a portion of an environment surrounding the mobile platform from one or more sensors carried by the mobile platform; determining depth information, a feature map, and a plurality of candidate regions based at least in part on the sensor data, wherein each candidate region indicates at least a portion of an obstacle within the environment; and feeding the depth information, the feature map, and at least a subset of the plurality of candidate regions to an obstacle detection neural network to predict at least one state attribute of one or more obstacles within the environment; wherein the state attribute comprises at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment; and based at least in part on the predicted at least one of a type, a posture, an orientation, a three-dimensional position, or a three-dimensional size of the one or more obstacles in the environment, navigation of the mobile platform is implemented; wherein the sensor data comprises one or more images, and wherein the feature map is generated at least by feeding the one or more images to a base neural network separate from the obstacle detection neural network.
44. The mobile platform of claim 43, wherein: The one or more sensors include at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
45. The mobile platform of claim 43, wherein: The depth information is determined at least in part based on the point cloud.
46. The mobile platform of claim 45, wherein: The point cloud is obtained based on sensor data generated by at least one of a lidar, a radar, a time-of-flight ToF camera, a stereo camera, or a monocular camera.
47. The mobile platform of claim 43, wherein: The depth information includes a depth map determined at least in part based on disparity data generated directly or indirectly from at least one of a stereo camera or a monocular camera.
48. The mobile platform of claim 43, wherein: The depth information is generated by feeding the feature map into a depth estimation neural network separate from the obstacle detection neural network.
49. The mobile platform of claim 43, wherein: The sensor data includes a point cloud, and wherein the feature map is generated based at least in part on projecting the point cloud onto a 2D grid defined from at least one of the one or more images.
50. The mobile platform of claim 49, wherein: The projection is based at least in part on extrinsic calibration parameters and / or intrinsic calibration parameters regarding at least one sensor generating the one or more images.
51. The mobile platform of claim 50, wherein: The projected data includes at least one of a height, a distance, or an angle measurement for each grid block of the 2D grid.
52. The mobile platform of claim 50, wherein: The feature map is generated by further feeding the projected data to the base neural network.
53. The mobile platform of claim 43, wherein: The size of the feature map is smaller than the size of at least one of the one or more images.
54. The mobile platform of claim 43, wherein: At least one of the base neural network or the obstacle detection neural network includes one or more convolutional layers and / or pooling layers.
55. The mobile platform of claim 43, wherein: The one or more processors are further configured to feed the feature map to an intermediate neural network to generate the plurality of candidate regions defined based on the image data.
56. The mobile platform of claim 55, wherein: Each candidate region is a two-dimensional 2D region including a corresponding target pixel of the image data.
57. The mobile platform of claim 56, wherein: The corresponding target pixel is associated with a probability of being indicative of at least a portion of the obstacle within the environment.
58. The mobile platform of claim 55, wherein: At least two of the basic neural network, the intermediate neural network and the obstacle detection neural network are jointly trained.
59. The mobile platform of claim 43, wherein: The obstacle detection neural network includes a first subnetwork configured to determine an initial 3D position of an obstacle for each candidate region in the subset based at least in part on the depth information and the at least one subset of the plurality of candidate regions.
60. The mobile platform of claim 59, wherein: The obstacle detection neural network includes a second sub-network configured to generate one or more region features for each candidate region in the subset based at least in part on the at least one subset of the plurality of candidate regions and the feature map.
61. The mobile platform of claim 60, wherein: The obstacle detection neural network includes a third subnetwork, which is configured to: predict at least one of a type, a posture, an orientation, a 3D position, or a 3D size of an obstacle corresponding to each candidate region in the subset based at least in part on the initial 3D position and the one or more region features.
62. The mobile platform of claim 43, wherein: The mobile platform includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, or a robot.
Citation Information
Patent Citations
Method, apparatus, device and computer storage medium for obtaining obstacle information
CN109145680A