Dense mapping using a distance sensor with multiple scans and multi-view geometry
By combining the distance sensor and camera data to generate dense three-dimensional maps of the environment, the problems of drift and loop closure in three-dimensional reconstruction in the prior art are solved, and high-quality three-dimensional map generation is achieved in challenging environments.
Patent Information
- Application Number
- CN202010589260.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-26
- Filing Date
- 2020-06-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-06-24
AI Technical Summary
The prior art faces the problems of proportional factor cumulative drift and loop closure in outdoor three-dimensional reconstruction, and the visual-based method outputs inaccurately under conditions such as poor lighting, lack of texture, occlusion or moving objects.
By combining continuous scanning of distance sensors and continuous images of the camera, a dense three-dimensional map of the environment is generated. The processing circuit is based on multi-view geometry and continuous scanning, performs depth estimation and image matching, generates fine three-dimensional maps, and updates in real time.
Generating high-quality three-dimensional maps in challenging environments is achieved, improving the accuracy and stability of positioning and mapping, and overcoming the shortcomings of visual methods under adverse conditions.
Smart Images

Figure CN112150620B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to three-dimensional mapping. Background Art
[0002] Outdoor three-dimensional reconstruction can be used in many applications such as autonomous navigation, positioning of aircraft, mapping, obstacle avoidance, and many other applications. However, three-dimensional reconstruction can be challenging due to the large scale of the three-dimensional map, unformed features in the map, and poor lighting conditions. Creating a fully automated and real-time modeling process with high-quality results poses difficulties because the processes of collecting, storing, and matching data are expensive both in terms of memory and time.
[0003] Vision-based methods for three-dimensional reconstruction have relatively low cost and high spatial resolution. However, vision-based simultaneous localization and mapping (vSLAM) solutions for scene reconstruction suffer from scale factor cumulative drift and loop closure problems. Due to poor image quality, the output of the vSLAM process may be inaccurate, which may be caused by external factors such as poor lighting, lack of texture, occlusion, or moving objects.
[0004] Millimeter-wave (MMW) radar-based solutions offer the advantage of higher reliability independent of lighting and weather conditions. However, MMW radar cannot identify the height, shape, and size of targets. In addition, the depth output of MMW radar is very sparse.
[0005] LiDAR-based solutions provide a large number of accurate three-dimensional points for scene reconstruction. However, the alignment of a large amount of data requires a large number of processing algorithms, which can be memory-intensive and time-consuming. Scenes reconstructed using point cloud-based methods usually have an unstructured representation and cannot be directly represented as a connected surface. Compared with radar, LiDAR is generally more expensive and affected by external lighting and weather conditions (e.g., raindrops, dust particles, and extreme sunlight), which can result in noisy measurements. Summary of the Invention
[0006] Generally speaking, the present disclosure relates to systems, devices, and techniques for generating a three-dimensional map of an environment using continuous scans performed by a distance sensor and continuous images captured by a camera. The system can generate a dense three-dimensional map of the environment based on an estimate of the depth of an object in the environment. The system can fuse the distance sensor scans and camera images to generate a three-dimensional map and can continuously update the three-dimensional map based on newly acquired scans and camera images.
[0007] In some examples, the system includes a distance sensor configured to receive signals reflected from objects in the environment and generate two or more consecutive scans of the environment at different times. The system also includes a camera configured to capture two or more consecutive camera images of the environment, where each of the two or more consecutive camera images of the environment is captured by the camera at different positions within the environment. The system also includes a processing circuit configured to generate a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0008] In some examples, a method includes receiving, by a processing circuit, from a distance sensor, two or more consecutive scans of the environment performed by the distance sensor at different times, where the two or more consecutive scans represent information derived from signals reflected from objects in the environment. The method also includes receiving, by the processing circuit, two or more consecutive camera images of the environment captured by a camera, where each of the two or more consecutive camera images of the object is captured by the camera at different positions within the environment. The method also includes generating, by the processing circuit, a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0009] In some examples, a device includes a computer-readable medium having executable instructions stored thereon, the executable instructions configured to be executed by a processing circuit to cause the processing circuit to receive from a distance sensor: two or more consecutive scans of the environment performed by the distance sensor at different times, where the two or more consecutive scans represent information derived from signals reflected from objects in the environment. The device also includes instructions for causing the processing circuit to receive two or more consecutive camera images of the environment from a camera, where each of the two or more consecutive camera images is captured by the camera at different positions within the environment. The device also includes instructions for causing the processing circuit to generate a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0010] Details of one or more examples of the present disclosure are set forth in the following drawings and description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a conceptual block diagram of a system including a distance sensor and a camera according to some examples of the present disclosure.
[0012] Figure 2 is a conceptual block diagram of a drone according to some examples of the present disclosure, the drone including a system for generating a three-dimensional map of the environment.
[0013] Figure 3Is a diagram showing the generation of a three - dimensional map based on multi - view geometry and a distance sensor according to some examples of the present disclosure.
[0014] Figure 4 Is a flowchart for determining a fine depth estimate based on a spatial cost volume according to some examples of the present disclosure.
[0015] Figure 5 Is a diagram showing the geometry of a system including a distance sensor and a camera according to some examples of the present disclosure.
[0016] Figure 6 Is a flowchart showing an exemplary process for generating a three - dimensional map of an environment based on consecutive images and consecutive scans according to some examples of the present disclosure.
[0017] Figure 7 Is a flowchart showing an exemplary process for multi - view geometry processing using consecutive images and consecutive scans according to some examples of the present disclosure. Detailed Description
[0018] Various examples for generating a three - dimensional map of an environment by combining camera images with consecutive scans performed by a distance sensor are described below. To generate a three - dimensional map, the system can generate a multi - view geometry of the environment based on sequential images captured by the camera and based on consecutive scans performed by the distance sensor. The system can combine the multi - view geometry and consecutive distance sensor scans to form a dense map of the environment around the distance sensor and the camera. When the system generates a three - dimensional map, the system takes into account the translational and rotational motions of the distance sensor and the camera in the three - dimensional environment.
[0019] The system can use sequential camera images to perform visual simultaneous localization and mapping (vSLAM) to form a multi - view geometry as the camera moves throughout the three - dimensional environment. The system can determine a depth estimate of an object in three - dimensional space based on the reflections received by the distance sensor. The system can use the distance sensor returns as a constraint to fine - tune the depth estimate from the vSLAM process. When the system receives a new camera image and return information from the distance sensor, the system can update the dense map of the environment.
[0020] Compared with a system that performs a one - time fusion of one image and one scan performed by a distance sensor, a system that combines multi - views and multiple scans can form and update a dense point cloud of the surrounding environment. The system can rely on the complementarity of the distance sensor (e.g., depth detection ability and robustness to environmental conditions) and the camera (e.g., high spatial resolution and high angular resolution).
[0021] The system can use a dense point cloud to track objects within an environment and / or determine the positions of cameras and distance sensors within the environment. A three-dimensional map can be used for obstacle detection and terrain object avoidance and for landing clearances for loop pilots as well as autonomous navigation and landing operations.
[0022] Figure 1 FIG. 4 is a conceptual block diagram of a system 100 including a distance sensor 110 and a camera 120 in accordance with some examples of the present disclosure. System 100 includes a distance sensor 110, a camera 120, a processing circuit 130, a positioning device 140, and a memory 150. System 100 can be mounted on a vehicle that moves throughout a three-dimensional environment such that the distance sensor 110 and the camera 120 can have translational and rotational motion. The distance sensor 110 and the camera 120 can each move in six degrees of freedom (e.g., pitch, roll, and yaw), as well as translational motion.
[0023] System 100 can be mounted on a vehicle or non-vehicle moving object, attached to a vehicle or non-vehicle moving object, and / or built into a vehicle or non-vehicle moving object. In some examples, system 100 can be mounted on an aircraft such as an airplane, helicopter, or weather balloon, or on a space vehicle such as a satellite or spacecraft. In other examples, system 100 can be mounted on a land vehicle (such as an automobile) or a water vehicle (such as a boat or submarine). System 100 can be mounted on a manned vehicle or an unmanned vehicle such as an unmanned aerial vehicle, a remotely controlled vehicle, or any suitable vehicle without any pilot or crew. In some examples, a portion of system 100 (e.g., the distance sensor 110 and the camera 120) can be mounted on a vehicle and another portion of system 100 (e.g., the processing circuit 130) can be external to the vehicle.
[0024] The distance sensor 110 emits a signal into the environment 160 and receives a reflected signal 112 from the environment 160. The signal emitted by the distance sensor 110 can reflect off the object 180 and return to the distance sensor 110. The processing circuit 130 can determine the distance (e.g., depth 190) from the distance sensor 110 to the object 180 by processing the reflected signal 112 received by the distance sensor 110. The distance sensor 110 can include a radar sensor (e.g., millimeter-wave radar and / or phased array radar), a lidar sensor, and / or an ultrasonic sensor. Exemplary details of the distance sensor can be found in the commonly assigned U.S. Patent Application Publication 2018 / 0246200, titled "Integrated Radar and ADS-B," filed on November 9, 2017, and the commonly assigned U.S. Patent Application Publication 2019 / 0113610, titled "Digital Active Phased Array Radar," filed on February 5, 2018, the entire contents of which are incorporated herein. For example, the distance sensor 110 can include a radar sensor configured to perform a full scan of the field of view in less than five seconds or in some examples in less than three seconds using electronic scanning.
[0025] The distance sensor 110 can determine the range or distance (e.g., depth 190 to the object 180) with higher accuracy than the camera 120. The measurement of the depth 190 to the object 180 obtained by the distance sensor 110 based on the reflected signal 112 can have a constant range error as the distance increases. As described in further detail below, the processing circuit 130 can determine a first estimate of the depth 190 based on the image captured by the camera 120. Then, the processing circuit 130 can determine a second estimate of the depth 190 based on the reflected signal 112 received by the distance sensor 110 and use the second estimate to supplement the first estimate of the depth 190 based on the camera image.
[0026] The distance sensor 110 can perform a scan by emitting a signal into a part or all of the environment 160 and receiving a reflected signal from an object in the environment. The distance sensor 110 can perform a continuous scan by transmitting a signal across a part or all of the environment 160 for a first scan and then repeating the process by transmitting a signal across a part or all of the environment 160 for a second scan.
[0027] As the camera 120 moves within the environment 160, the camera 120 captures successive or sequential images of the environment 160 and the object 180. Thus, the camera 120 captures images at different locations within the environment 160 and provides the captured images to the processing circuit 130 for generating a three-dimensional map and / or to the memory 150 for storage and later use by the processing circuit 130 for mapping the environment 160. The camera 120 may include a visual camera and / or an infrared camera. The processing circuit 130 may store the position and pose information (e.g., translation and rotation) of the camera 120 for each image captured by the camera 120. The processing circuit 130 may use the position and pose information to generate a three-dimensional map of the environment 160. The camera 120 may have a lighter weight and lower power consumption than the distance sensor 110. In addition, compared to the distance sensor 110, the camera 120 is capable of sensing angular information with higher accuracy.
[0028] The processing circuit 130 may use the images captured by the camera 120 to perform vSLAM to simultaneously map the environment 160 and track the position of the system 100. vSLAM is an image-based mapping technique that uses a moving camera and multi-view geometry. vSLAM includes simultaneously tracking the movement of the system 100 and mapping the environment 160. In the vSLAM method, the processing circuit 130 may use the estimated depth of an object in the environment 160 to track the position of the system 100 within the environment 160. During the tracking step, the processing circuit 130 may be configured to use the pose information from the inertial sensor to track the position of the system within the environment 160. Then, the processing circuit 130 uses the position, orientation, and pose of the camera 120 for each image to generate a map of the environment 160. During the mapping step, the processing circuit 130 may build a three-dimensional map by extracting key points from multiple images fused with the movement information from the tracking step.
[0029] Unlike other systems that perform vSLAM using only images, the system 100 and the processing circuit 130 may use successive scans from the distance sensor 110 and the multi-view geometry from sequential image frames captured by the camera 120 to compute the pixel-level uncertainty confidence of the spatial cost volume given the rotation and translation of the camera 120. The processing circuit 130 may also utilize the depth constraint distribution generated from multiple scans of the distance sensor 110 to fine-tune the vSLAM depth estimation to improve depth accuracy.
[0030] The processing circuit 130 may be configured to warp the returns from successive scans from the distance sensor 110 to an intermediate view of the scan with a known camera pose. The processing circuit 130 may compare the camera image and / or the distance sensor image by warping one view to another. Warping the returns from successive scans to an intermediate view of the scan based on the known pose of the camera 120 may improve the point density around the intermediate view. The processing circuit 130 may be configured to calculate a spatially cost volume that is adaptively spaced within a depth range based on pixel-level uncertainty confidence, and use the depth distribution output from the distance sensor 110 to calibrate the vSLAM depth output and improve the density and accuracy of depth measurements based on the map generated by vSLAM.
[0031] The processing circuit 130 receives return information based on the reflected signal 112 from the distance sensor 110 and receives images from the camera 120. The processing circuit 130 may generate a distance sensor image based on the return information received from the distance sensor 110. The distance sensor images may each include a rough map of the environment 160, which includes depth information of the objects in the environment 160. The processing circuit 130 may generate a multi-view geometry based on the images received from the camera 120 and combine the multi-view geometry with the rough map of the environment 160.
[0032] The processing circuit 130 may match points in the distance sensor image and the camera image to determine the depth of an object in the environment 160. For example, the processing circuit 130 may identify key points in the camera image and then detect corresponding points in the distance sensor image. The processing circuit 130 may extract features from the camera image and match the extracted features with points in the distance sensor image. Example details of key point detection and matching can be found in the commonly assigned U.S. Patent Application Serial No. 16 / 169,879, titled "Applying an Annotation to an Image Based on Keypoint", filed on October 24, 2018, the entire content of which is incorporated herein.
[0033] The processing circuit 130 may be mounted on a vehicle together with other components of the system 100, and / or the processing circuit 130 may be located outside the vehicle. For example, if the distance sensor 110 and the camera 120 are mounted on an unmanned aerial vehicle (UAV), the processing circuit 130 may be located on the UAV and / or in a ground system. The system 100 may perform scans and capture images of the environment during an inspection. After the inspection, the UAV may send data to a ground-based computer including the processing circuit 130 that generates a three-dimensional map of the environment 160. However, the processing circuit 130 may also be co-located with the distance sensor 110 and the camera 120 on the UAV.
[0034] Whether the processing circuit 130 is co-located or remotely located from the distance sensor 110 and the camera 120, the processing circuit 130 can generate a three-dimensional map of the environment 160 as the vehicle moves through the environment 160. The processing circuit 130 can generate a travel path through the environment 160 for the vehicle based on the three-dimensional map. The processing circuit 130 can navigate the vehicle and control the movement of the vehicle based on the travel path generated as the vehicle moves along the travel path.
[0035] The positioning device 140 determines the position or location of the system 100 and provides this information to the processing circuit 130. The positioning device 140 can include satellite navigation devices, such as a Global Navigation Satellite System (GNSS) configured to receive positioning signals from satellites and other transmitters. An example of a GNSS is the Global Positioning System (GPS). The positioning device 140 can be configured to deliver the received positioning signals to the processing circuit 130, which can be configured to determine the position of the system 100. The processing circuit 130 can determine the positions of the distance sensor 110 and the camera 120 based on the positioning data from the positioning device 140. The processing circuit 130 can also determine the position and orientation based on information from a navigation system, a heading system, a gyroscope, an accelerometer, and / or any other device for determining the orientation and heading of a moving object. For example, the system 100 can include an inertial system having one or more gyroscopes and accelerometers.
[0036] The memory 150 stores the three-dimensional map of the environment 160 generated by the processing circuit 130. In some examples, the memory 150 can store program instructions, which can include one or more program modules executable by the processing circuit 130. When executed by the processing circuit 130, such program instructions can cause the processing circuit 130 to provide the functions that belong to it herein. The program instructions can be embodied in software and firmware. The memory 150 can include any volatile, non-volatile, magnetic, optical, or electronic medium, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), electrically erasable programmable ROM (EEPROM), flash memory, or any other digital media.
[0037] The environment 160 includes Figure 1 objects 180 and other objects not shown. The system 100 can move through the entire environment 160 as the processing circuit 130 generates and updates the three-dimensional map. Additionally, objects in the environment 160 can move as the distance sensor 110 and the camera 120 collect data. For example, the system 100 can be mounted on an unmanned aerial vehicle and used to inspect a structure in the environment 160. The system 100 can perform scans and capture images of the environment 160 during a detection, and the processing circuit 130 can generate a three-dimensional map of the environment 160 during or after the detection.
[0038] The object 180 is located at a depth 190 from the system 100, but the distance from the distance sensor 110 to the object 180 may be different from the distance from the camera 120 to the object 180. Thus, the processing circuitry 130 can use the position and orientation of each of the distance sensor 110 and the camera 120 to determine the position of the object 180 within the environment 160.
[0039] Another system can perform a one-time fusion of a single scan performed by a distance sensor and a single image captured by a camera. This other system can use the one-time fusion to determine the position of an object, for example, to avoid a collision between a vehicle and the object. The system can identify the object in the image and then find the object in the scan information to determine the distance between the object and the vehicle. To determine an updated position of the object, the system can later perform another fusion of another camera image and another scan performed by the distance sensor. The system can identify the object in the later image and then find the object in the scan information. This system cannot use the one-time fusion of the image and the scan to build and update a dense map of the environment. Additionally, the system cannot use the one-time fusion to continuously determine the position of the system within the environment.
[0040] According to the techniques of the present disclosure, the processing circuitry 130 can generate a dense map of the environment 160 based on two or more consecutive scans performed by the distance sensor 110 and two or more images captured by the camera 120. Using the scans received from the distance sensor 110, the processing circuitry 130 can determine an accurate estimate of the depth of the object within the environment 160. Using the images received from the camera 120, the processing circuitry 130 can determine the angular information of the object within the environment 160. By using consecutive scans and consecutive camera images, the processing circuitry 130 can generate and update a dense three-dimensional map of the environment 160.
[0041] The processing circuitry 130 is capable of generating a high-resolution map of the environment 160. For example, the system 100 may have a field of view with a horizontal dimension of sixty degrees and a vertical dimension of forty-five degrees. The resolution of the distance sensor 110 can be less than one-tenth of a degree, such that the processing circuitry 130 generates a three-dimensional map of depth values with a horizontal dimension of 640 pixels and a vertical dimension of 480 pixels. This example is for illustrative purposes only, and other examples are possible, including examples with a wider field of view and / or a higher angular resolution.
[0042] System 100 can be implemented in various applications. In an example where the distance sensor 110 includes a phased array radar, the processing circuit 130 can perform sensor fusion on the radar echoes received by the distance sensor 110 and the images captured by the camera 120. The processing circuit 130 can also perform obstacle detection and avoidance for a UAV operating outside the visual line of sight. The enhanced depth estimation accuracy of the system 100 can be used to determine the depth of obstacles and prevent collisions with obstacles. Additionally, when GPS is not working (e.g., in a GPS-denied area), the system 100 can provide navigation.
[0043] In an example where GPS, GNS, or cellular services are fully available or reliable, a UAV or urban air mobility (UAM) can use the techniques of the present disclosure for internal guidance during takeoff and landing. In some examples, a UAV can use a radar-enhanced vision system that supports deep learning to generate an estimate of depth based on a single inference, regardless of the multi-view geometry from camera images. The UAV can also use dense tracking and mapping (DTAM) with a single camera. The present disclosure describes a new method for dense depth prediction that calculates the pixel-level uncertainty confidence of a spatial cost volume for a given set of rotational and translational motions of the camera 120 by leveraging consecutive radar scans and multi-view geometry from sequential image frames. The processing circuit 130 can use multiple scans performed by the distance sensor 110 to fine-tune the vSLAM depth estimate to improve depth density and estimation accuracy.
[0044] Figure 2 is a conceptual block diagram of a UAV 202 including a system 200 for generating a three-dimensional map of an environment 260 according to some examples of the present disclosure. The system 200 includes a processing circuit configured to determine the position and orientation of the system 200 based on translational and rotational motions 210, yaw 212, roll 214, and pitch 216. The system 200 has six degrees of freedom because the three-dimensional mapping process takes into account the yaw, roll, and pitch of the distance sensor as well as the yaw, roll, and pitch of the camera.
[0045] The system 200 can determine depth estimates of objects 280 and 282 and vehicle 284 within the environment 260. Objects 280 and 282 are ground-based objects such as buildings, trees or other vegetation, terrain, light poles, utility poles, and / or cellular transmission towers. Vehicle 284 is an airborne object depicted as an airplane, but vehicle 284 can also be a helicopter, UAV, and / or weather balloon. In some examples, the environment 260 can include other objects such as birds. The system 200 is configured to generate a three-dimensional map of the positions of the objects including objects 280 and 282, vehicle 284, and the ground surface 270 within the environment 260.
[0046] Figure 3It is an illustration showing the generation of a three-dimensional map 330 based on a multi-view geometry 320 and continuous scans 310 by a distance sensor according to some examples of the present disclosure. Figure 1 The processing circuit 130 of the illustrated system 100 can identify points 312 in the continuous scan 310. The processing circuit 130 can also identify points 322 in the continuous images of the multi-view geometry 320. The processing circuit 130 can be configured to match the points 312 with the points 322 and use the matched points to determine a fine depth estimate of an object in the environment.
[0047] The processing circuit 130 is configured to generate a three-dimensional map 330 based on the continuous scan 310 and the multi-view geometry 320. The processing circuit 130 can determine the position of a point 332 within the three-dimensional map 330 based on the position of the points 312 within the continuous scan 310 and the position of the points 322 within the multi-view geometry 320. The multi-view geometry 320 can provide a smooth depth estimate, and the continuous scan 310 can provide constraints on the depth estimate based on the multi-view geometry 320. The processing circuit 130 can use adaptive range slices instead of fixed range slices to adjust the depth estimate in the three-dimensional map.
[0048] Based on the continuous scan 310, the processing circuit 130 can determine a depth estimate of an object within the environment 160. However, the depth estimate based on the continuous scan 310 can be sparse and may have errors due to noise. Additionally, there may be little or no correlation between the data from each of the continuous scans 310, especially for distance sensors with a slow scan speed and for vehicles moving at a fast speed. However, when paired with the multi-view geometry 320 based on the images captured by the camera 120, the continuous scans 310 can be used to generate a dense three-dimensional map. The three-dimensional mapping process described herein goes beyond a one-time fusion of the images captured by the camera 120 and the scans performed by the distance sensor 110 to combine the continuous scans 310 and sequential image frames to generate a dense three-dimensional map.
[0049] Figure 4 It is a flowchart for determining a fine depth estimate based on a spatial cost volume according to some examples of the present disclosure. Refer to Figure 1 the illustrated system 100 for description Figure 4 of an exemplary process, but other components can illustrate similar techniques. Although Figure 4 is described as including a processing circuit 130 that performs spatial cost volume processing, in addition to or as an alternative to using a spatial cost volume, the processing circuit 130 can use a conditional (Markov) random field model or an alpha matte.
[0050] In Figure 4In the example of, the distance sensor 110 performs multiple scans (410). The distance sensor 110 receives the reflected signal 112 as part of the successive scans and sends the information to the processing circuit 130. This information may include altitude data, azimuth data, velocity data, and distance data. For example, the distance sensor 110 may send data indicating the elevation angle of the object 180 (e.g., the angle from the distance sensor 110 to the object 180 relative to the horizontal plane), the azimuth angle of the object 180, the velocity of the object 180 (e.g., the Doppler velocity of the object 180), and the depth 190.
[0051] In Figure 4 the example of, the camera 120 captures sequential image frames (420). The camera 120 sends the captured images together with the translational and rotational motion information to the processing circuit 130. Additionally or alternatively, the processing circuit 130 may receive the translational and rotational motion information from the inertial sensors.
[0052] In Figure 4 the example of, the processing circuit 130 performs spatial cost volume processing (430). The processing circuit 130 may perform spatial cost volume processing by determining the multi-view geometry based on the cumulative measurements from multiple sensors (e.g., the distance sensor 110 and the camera 120). For example, the processing circuit 130 may perform spatial cost volume processing based on multiple scans from a millimeter wave (MMW) radar and the multi-view geometry of an image sequence with corresponding rotations and translations of the camera 120.
[0053] The processing circuit 130 may be configured to feed the fast Fourier transform (FFT) range cells and angle cells as spatial features. The range cells and angle cells represent the return power of the reflected signals received by the distance sensor 110 for each azimuth angle and elevation angle in the environment 160. The processing circuit 130 may also utilize the multi-view geometry to warp successive distance sensor scans to the intermediate scan of the distance sensor 110 with the known pose of the distance sensor 110. The processing circuit 130 may also calculate the spatially cost volume adaptively spaced within the depth range based on the multi-view geometry and based on the pixel-level uncertainty confidence. As further described below, the processing circuit 130 may be configured to utilize the depth output from the distance sensor 110 as a constraint to calibrate the vSLAM depth output and improve the density and accuracy of the depth map measurements.
[0054] Example details of spatial cost volume processing can be found in "Fast Cost-Volume Filtering for Visual CorResponse and Beyond" published by Asmaa Hosni et al. in IEEE Transactions on Pattern Analysis and Machine Intelligence in February 2013, the entire content of which is incorporated herein by reference. A spatial cost volume is a multi-view geometry where the space or environment is divided into multiple blocks. The processing circuit 130 can construct a spatial cost volume and warp a new image to the same view as a previous image to obtain a spatial cost volume. The processing circuit 130 can construct the spatial cost volume in an adaptive manner, where range slices are adjusted based on errors. The spatial cost volume is similar to a four-dimensional model (e.g., space and time).
[0055] The processing circuit 130 can use a first FFT applied to the continuous scan 310 to extract range information of an object within the environment 160. The processing circuit 130 can use a second FFT to extract velocity information based on the Doppler effect. The processing circuit can also use a third FFT to extract angle information of an object within the environment 160. The processing circuit 130 can use the output of the first FFT as an input to perform the second FFT. Then, the processing circuit 130 can use the output on the second FFT as an input to perform the third FFT.
[0056] In Figure 4 the example, the processing circuit 130 corresponds points in the distance sensor scan to pixels (440) in the camera image. The processing circuit 130 can use key point recognition techniques to identify points in the distance sensor image and the camera image. The processing circuit 130 can generate a constraint distribution based on the matching of points in the distance sensor image and points in the camera image. The processing circuit 130 can perform pixel-level matching using direction or angle, color information (e.g., red-green-blue information), and range information.
[0057] In Figure 4 the example, the processing circuit 130 determines a rough estimate (450) of the depth 190 of the object 180 based on the spatial cost volume. In some examples, the processing circuit 130 can determine rough estimates of the depths of all objects within the environment 160. As part of determining the rough estimate of the depth 190, the processing circuit 130 can calculate a depth error based on the depth range in the scene, which can be divided into multiple range slices.
[0058] The pixel-level uncertainty of the spatial cost volume can be measured by the uncertainty of the generated depth map. The processing circuit 130 can calculate the pixel-level uncertainty of the spatial cost volume based on the translational and rotational motions of the camera 120. There may be sensors embedded in the camera 120 that allow the processing circuit 130 to extract the translation and orientation. The processing circuit 130 can calculate the depth error based on the depth range in the scene, which can be divided into multiple slices.
[0059] Three techniques for measuring depth error are labeled L1-rel, L1-inv, and Sc-inv. The processing circuit 130 can measure L1-rel by the absolute difference of depth values in the logarithmic space averaged over multiple pixels to normalize the depth error. The difference in depth values can refer to the difference between the predicted depth and the ground truth depth. The processing circuit 130 can measure L1-inv by the absolute difference of the reciprocals of depth values in the logarithmic space averaged over multiple pixels (n), which emphasizes more on the near depth values. Sc-inv is a scale-invariant metric that allows the processing circuit 130 to measure the relationship between points in the scene regardless of the absolute global scale. The processing circuit 130 can be configured to switch the depth error evaluation method based on the scene depth range to better reflect the uncertainty in depth calculation.
[0060] In Figure 4 the example, the processing circuit 130 determines a refined estimate (460) of the depth 190 based on the correspondence between the distance sensor and the image pixels. The correspondence between the distance sensor and the image pixels may include matching points in the camera image with points in the distance sensor image, as described with respect to Figure 3 above. The processing circuit 130 can be configured to refine the depth based on the constraint distribution generated by the processing circuit 130 from multiple scans. The data from the distance sensor 110 becomes an auxiliary channel to constrain the multi-view geometry based on the images captured by the camera 120. Even without ground truth depth data, the processing circuit 130 can check the temporal / spatial pixel consistency between the distance sensor 110 and the camera 120 to calculate the depth estimation error.
[0061] Figure 5 is a diagram showing the geometry of a system including a distance sensor and a camera according to some examples of the present disclosure. Oc and Or are Figure 5 the relative positions of the camera and the distance sensor in the example geometry. Rc and Rr are the camera frame and the distance sensor frame respectively. The distance sensor data (azimuth and elevation) provides the polar coordinates of the target. Figure 5Shows the projection points on the camera plane shown together with the horizontal distance sensor plane. The processing circuit can use GPS, an inertial system, a predetermined distance between the camera and the distance sensor, and a SLAM tracking algorithm to determine the positions and orientations of the camera and the distance sensor.
[0062] The consistency error C between the predicted depth of the camera multi-view geometry and the predicted depth of multiple scans of the distance sensor can be evaluated using Equation (1). error . In Equation (1), C error is the spatial consistency error of the selected pixel, N is defined as the number of selected pixels used for evaluation, evaluated with reference to the depth output from multiple scans of the distance sensor, and evaluated with reference to the depth output from the camera multi-view geometry.
[0063] As Figure 5 shown, the camera and the distance sensor are not located at the same position. Even though both the camera and the distance sensor can be mounted on the same vehicle, the camera and the distance sensor can have different positions and different orientations. The generation of the three-dimensional map can be based on the relative positions and orientations of the camera for each camera image captured by the camera. The generation of the three-dimensional map can be based on the relative positions and orientations of the distance sensor for each scan performed by the distance sensor.
[0064] Figure 6 is a flowchart showing an exemplary process for generating a three-dimensional map of an environment based on consecutive images and consecutive scans according to some examples of the present disclosure. The exemplary process is described with reference Figure 1 to the system 100 shown, Figure 6 but other components can illustrate similar techniques.
[0065] In Figure 6 the example, the processing circuit 130 receives two or more consecutive scans of the environment 160 performed by the distance sensor 110 at different times from the distance sensor 110, where the two or more consecutive scans represent information (600) derived from signals reflected from objects in the environment 160. In an example where the distance sensor 110 includes a radar sensor, the distance sensor 110 can perform a scan by transmitting a radar signal into the field of view and receiving the reflected radar signal. The processing circuit 130 can use digital beamforming techniques to generate scan information at each altitude and azimuth angle within the field of view. When the radar sensor continuously transmits and receives signals to determine the depth in each direction within the field of view, the processing circuit 130 can form and move the beam throughout the field of view. In an example where the distance sensor 110 includes a lidar sensor, the distance sensor 110 can perform a scan by transmitting signals in each direction within the field of view.
[0066] In Figure 6 the example of, processing circuit 130 receives two or more consecutive camera images of the environment captured by camera 120, where each of the two or more consecutive camera images of the object is captured by camera 120 at different positions within environment 160 (602). Processing circuit 130 may use a key point detection algorithm such as an edge detection algorithm to identify key points in each image. Then, processing circuit 130 may match the key points across sequential images. Processing circuit 130 may also determine the position and orientation of camera 120 for each image captured by camera 120. Processing circuit 130 may store the position and orientation of camera 120 into memory 150 for generating a three-dimensional map.
[0067] In Figure 6 the example of, processing circuit 130 generates a three-dimensional map of environment 160 based on two or more consecutive scans and two or more consecutive camera images (604). Processing circuit 130 may generate a dense three-dimensional map with a pixel resolution in elevation and azimuth less than 1 degree, less than 0.5 degree, less than 0.2 degree, or less than 0.1 degree. Processing circuit 130 may generate a dense three-dimensional map based on a rough depth estimate for an object in environment 160 using consecutive images captured by camera 120. Processing circuit 130 may generate a fine depth estimate for an object in environment 160 based on consecutive scans performed by distance sensor 110.
[0068] Figure 7 is a flowchart showing an exemplary process for multi-view geometry processing using consecutive images and consecutive scans according to some examples of the present disclosure. Referring to Figure 1 the system 100 shown describes Figure 7 the exemplary process, but other components may illustrate similar techniques.
[0069] In Figure 7 the example of, processing circuit 130 performs correspondence of the distance sensor with the image (700). Using the geometric layout of distance sensor 110 and camera 120, processing circuit 130 may transform coordinates between the distance sensor image and the camera image. Processing circuit 130 may map the distance sensor targets to the image frame based on a coordinate transformation matrix.
[0070] In Figure 7In an example, the processing circuit 130 adaptively spaces N depth tags within a depth slice to interpolate a first rough depth to construct a spatial cost volume (702). The spatial cost volume can be defined as a function SCV(x,d), where x represents a pixel position and d represents a depth tag. The processing circuit 130 can calculate hyperparameters for constructing the spatial cost volume function from a set of images, camera poses, a set of distance sensor images, and distance sensor poses through a feature learning-based solution.
[0071] In Figure 7 an example, the processing circuit 130 warps multiple scans of the distance sensor image to a keyframe image centered at the first rough depth based on the pose of the camera 120 (704). Warping multiple scans of the distance sensor image to a keyframe image centered at the first rough depth can be done using relative pose and depth. An intermediate scan of the MMW distance sensor with a known pose can be selected as a reference to calculate a spatially spaced-apart cost volume adaptively within a depth range based on multi-view geometry.
[0072] In Figure 7 an example, the processing circuit 130 refines the rough prediction of the surrounding depth based on a distance sensor depth lookup table according to a constraint distribution generated from multiple scans of the distance sensor (706). The processing circuit 130 uses the depth output from the multiple scans of the distance sensor geometry as a constraint distribution to shape the rough estimate to improve depth accuracy. Range finding based on the distance sensor can be achieved through consecutive multiple scans of the distance sensor with known translation and rotation parameters. Since the distance resolution output by the distance sensor is better than the distance resolution of vSLAM, the depth values generated from multiple scans can be used as a constraint file through a lookup table to remove outliers of the depth values from vSLAM.
[0073] In Figure 7 an example, the processing circuit 130 refines the depth based on a confidence score from the rough prediction (708). This is a coarse-to-fine approach. We extract a rough depth estimate and utilize the constraint distribution for better regularization to refine the depth prediction. The confidence score can be calculated based on the spatial consistency error of the selected pixels.
[0074] The following numbered embodiments illustrate one or more aspects of the present disclosure.
[0075] Example 1. The present invention discloses a method, which includes a processing circuit receiving, from a distance sensor, two or more consecutive scans of an environment performed by the distance sensor at different times, where the two or more consecutive scans represent information derived from signals reflected by an object in the environment. The method further includes the processing circuit receiving two or more consecutive camera images of the environment captured by a camera, where each of the two or more consecutive camera images of the object is captured by the camera at different positions within the environment. The method further includes the processing circuit generating a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0076] Example 2. The method according to Example 1 further includes matching points in the two or more consecutive scans with points in the two or more consecutive camera images.
[0077] Example 3. The method according to Example 1 or Example 2, wherein generating the three-dimensional map of the environment includes determining an estimate of the depth of a first object in the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0078] Example 4. The method according to Example 1 - 3 or any combination thereof, wherein generating the three-dimensional map of the environment includes determining a refined estimate of the depth of a first object based on matching points on consecutive distance sensor images with consecutive camera images.
[0079] Example 5. The method according to Example 1 - 4 or any combination thereof further includes estimating the depth of a second object in the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0080] Example 6. The method according to Example 1 to 5 or any combination thereof, wherein generating the three-dimensional map of the environment is based on the depth of the first object and the depth of the second object.
[0081] Example 7. The method according to Example 1 to 6 or any combination thereof further includes measuring the rotational movement of the camera relative to the distance sensor.
[0082] Example 8. The method according to Example 1 - 7 or any combination thereof, wherein generating the three-dimensional map is based on the two or more consecutive scans, the two or more consecutive camera images, and the rotational movement of the camera.
[0083] Example 9. The method according to Example 1 - 8 or any combination thereof further includes measuring the translational movement of the camera relative to the distance sensor.
[0084] Example 10. The method according to Example 1-9 or any combination thereof, wherein generating the three-dimensional map is based on two or more consecutive scans, two or more consecutive camera images, and the translational movement of the camera.
[0085] Example 11. The method according to Example 1-10 or any combination thereof, wherein receiving two or more consecutive scans includes receiving signals reflected from an object by a radar sensor.
[0086] Example 12. The method according to Example 1-11 or any combination thereof, further comprising performing simultaneous localization and mapping based on two or more consecutive camera images using two or more consecutive scans as depth constraints.
[0087] Example 13. The method according to Example 1-12 or any combination thereof, wherein estimating the depth of an object includes performing a first fast Fourier transform on two or more consecutive scans to generate an estimate of the depth of the object.
[0088] Example 14. The method according to Example 1-13 or any combination thereof, further comprising performing a second fast Fourier transform on two or more consecutive scans to generate the relative speed of the object.
[0089] Example 15. The method according to Example 1-14 or any combination thereof, further comprising performing a third fast Fourier transform on two or more consecutive scans to generate an estimate of the angle from the distance sensor to the object.
[0090] Example 16. The method according to Example 1-15 or any combination thereof, further comprising constructing a spatial cost volume of the environment based on two or more consecutive camera images.
[0091] Example 17. The method according to Example 1 to 16 or any combination thereof, further comprising determining the pixel-level uncertainty of the spatial cost volume based on two or more consecutive scans and two or more consecutive camera images.
[0092] Example 18. The method according to Example 17, wherein determining the pixel-level uncertainty of the spatial cost volume is further based on the rotational movement and the translational movement of the camera.
[0093] Example 19. The present invention discloses a system that includes a distance sensor configured to receive signals reflected from objects in the environment and generate two or more consecutive scans of the environment at different times. The system further includes a camera configured to capture two or more consecutive camera images of the environment, where each of the two or more consecutive camera images of the environment is captured by the camera at different positions within the environment. The system further includes a processing circuit configured to generate a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0094] Example 20. The system according to Example 19, wherein the processing circuit is configured to perform the method according to Example 1-18 or any combination thereof.
[0095] Example 21. The system according to Example 19 or Example 20, wherein the processing circuit is further configured to match points in the two or more consecutive scans with points in the two or more consecutive camera images.
[0096] Example 22. The system according to Example 19-21 or any combination thereof, wherein the processing circuit is configured to generate a three-dimensional map of the environment at least in part by estimating the depth of a first object in the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0097] Example 23. The system according to Example 19-22 or any combination thereof, wherein the processing circuit is configured to generate a three-dimensional map of the environment at least in part by determining a refined estimate of the depth of a first object by matching points on consecutive distance sensor images with consecutive camera images.
[0098] Example 24. The system according to Example 19-23 or any combination thereof, wherein the processing circuit is further configured to estimate the depth of a second object in the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0099] Example 25. The system according to Example 19-24 or any combination thereof, wherein the processing circuit is configured to generate a three-dimensional map of the environment based on the depth of the first object and the depth of the second object.
[0100] Example 26. The system according to Example 19-25 or any combination thereof, wherein the processing circuit is further configured to measure the rotational movement of the camera relative to the distance sensor.
[0101] Example 27. The system according to any one of Examples 19-26 or any combination thereof, wherein the processing circuit is configured to generate a three-dimensional map based on two or more consecutive scans, two or more consecutive camera images, and the rotational movement of the camera.
[0102] Example 28. The system according to any one of Examples 19-27 or any combination thereof, wherein the processing circuit is further configured to measure the translational movement of the camera relative to the distance sensor.
[0103] Example 29. The system according to any one of Examples 19-28 or any combination thereof, wherein the processing circuit is configured to generate a three-dimensional map based on two or more consecutive scans, two or more consecutive camera images, and the translational movement of the camera.
[0104] Example 30. The system according to any one of Examples 19-29 or any combination thereof, wherein the distance sensor includes a radar sensor configured to emit radar signals and receive signals reflected from objects in the environment.
[0105] Example 31. The system according to any one of Examples 19-30 or any combination thereof, wherein the processing circuit is further configured to perform simultaneous localization and mapping based on two or more consecutive camera images using two or more consecutive scans as depth constraints.
[0106] Example 32. The system according to any one of Examples 19-31 or any combination thereof, wherein the processing circuit is further configured to perform a first fast Fourier transform on two or more consecutive scans to generate an estimate of the depth of a first object in the environment.
[0107] Example 33. The system according to any one of Examples 19-32 or any combination thereof, wherein the processing circuit is further configured to perform a second fast Fourier transform on two or more consecutive scans to generate the relative speed of a first object; Example 34. The system according to any one of Examples 19-33 or any combination thereof, wherein the processing circuit is further configured to perform a third fast Fourier transform on two or more consecutive scans to generate an estimate of the angle from the distance sensor to a first object.
[0108] Example 35. The system according to any one of Examples 19-34 or any combination thereof, wherein the processing circuit is further configured to construct a spatial cost volume of the environment based on two or more consecutive camera images.
[0109] Example 36. The system according to any one of Examples 19-35 or any combination thereof, wherein the processing circuit is further configured to determine the pixel-level uncertainty of the spatial cost volume based on two or more consecutive scans and two or more consecutive camera images.
[0110] Example 37. The system according to Example 36, wherein the processing circuit is configured to determine pixel-level uncertainty of a spatial cost volume based on rotational motion of the camera and translational motion of the camera.
[0111] Example 38. An apparatus includes a computer-readable medium having executable instructions stored thereon, the executable instructions being configured to be executable by a processing circuit to cause the processing circuit to receive from a distance sensor two or more consecutive scans of an environment performed by the distance sensor at different times, wherein the two or more consecutive scans represent information derived from signals reflected from an object in the environment. The apparatus further includes instructions for causing the processing circuit to receive from a camera two or more consecutive camera images of the environment, wherein each of the two or more consecutive camera images is captured by the camera at different positions within the environment. The apparatus further includes instructions for causing the processing circuit to generate a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0112] Example 39. The apparatus according to Example 38, further includes instructions for performing the method according to any combination of Examples 1-18.
[0113] Example 40. The present disclosure discloses a system that includes means for receiving signals reflected from an object in an environment and generating two or more consecutive scans of the environment at different times. The system further includes means for capturing two or more consecutive camera images of the environment, wherein each of the two or more consecutive camera images of the environment is captured at different positions within the environment. The system further includes means for generating a three-dimensional map of the environment based on the two or more consecutive scans and the two or more consecutive camera images.
[0114] Example 41. The apparatus according to Example 40, further includes means for performing the method according to Example 1-18 or any combination thereof.
[0115] The present disclosure contemplates a computer-readable storage medium including instructions that cause a processor to perform any of the functions and techniques described herein. The computer-readable storage medium may take any exemplary form of volatile, non-volatile, magnetic, optical, or dielectric, such as random access memory (RAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), or flash memory. The computer-readable storage medium may be referred to as non-transitory. A programmer (such as a patient programmer or a clinician programmer) or other computing device may also include a more portable removable memory type to enable simple data transfer or offline data analysis.
[0116] The techniques described in this disclosure, including those attributable to systems 100 and 200, distance sensor 110, camera 120, processing circuitry 130, positioning device 140, and / or memory 150, and various component parts, may be implemented at least in part in hardware, software, firmware, or any combination thereof. For example, aspects of these techniques may be implemented within one or more processors (including one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuitry, as well as any combination of such components), embodied in a programmer (such as a physician programmer or patient programmer), a stimulator, a remote server, or other devices. The term "processor" or "processing circuitry" generally may refer to any of the foregoing logic circuitry alone or in combination with other logic circuitry, or any other equivalent circuitry.
[0117] As used herein, the term "circuitry" refers to an ASIC, an electronic circuit, a processor (shared, dedicated, or group), memory that executes one or more software or firmware programs, combinational logic circuitry, or other suitable components that provide the described functionality. The term "processing circuitry" refers to one or more processors distributed across one or more devices. For example, "processing circuitry" may include a single processor or multiple processors on a device. "Processing circuitry" may also include processors on multiple devices, where the operations described herein may be distributed across the processors and devices.
[0118] Such hardware, software, firmware may be implemented within the same device or within separate devices to support the various operations and functions described in this disclosure. For example, any of the techniques or processes described herein may be executed within one device or at least partially distributed between two or more devices, such as between systems 100 and 200, distance sensor 110, camera 120, processing circuitry 130, positioning device 140, and / or memory 150. Additionally, any of the described units, modules, or components may be implemented together or separately as discrete but interoperable logic devices. Describing different features as modules or units is intended to highlight different functional aspects and does not necessarily imply that such modules or units must be implemented by separate hardware or software components. Instead, the functions associated with one or more modules or units may be performed by separate hardware or software components, or integrated within common or separate hardware or software components.
[0119] The techniques described in this disclosure may also be embodied or encoded in an article of manufacture that includes a non-transitory computer-readable storage medium encoded with instructions. The instructions embedded or encoded in the article of manufacture (including the encoded non-transitory computer-readable storage medium) may cause one or more programmable processors or other processors to implement one or more of the techniques described herein, such as when the instructions included or encoded in the non-transitory computer-readable storage medium are executed by one or more processors. Exemplary non-transitory computer-readable storage media may include RAM, ROM, programmable ROM (PROM), EPROM, EEPROM, flash memory, a hard disk, a CD-ROM, a floppy disk, magnetic media, optical media, or any other computer-readable storage device or tangible computer-readable medium.
[0120] In some examples, the computer-readable storage medium includes a non-transitory medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. In certain examples, the non-transitory storage medium may store data that can vary over time (e.g., in RAM or a cache). The elements of the devices and circuits described herein, including but not limited to systems 100 and 200, distance sensor 110, camera 120, processing circuit 130, positioning device 140, and / or memory 150, may be programmed with various forms of software. For example, one or more processors may be at least partially implemented as or include one or more executable applications, application modules, libraries, classes, methods, objects, routines, subroutines, firmware, and / or embedded code.
[0121] Various examples of the present disclosure have been described. Any combination of the described systems, operations, or functions is contemplated. These examples and other examples are within the scope of the following claims.
Claims
1. A system for generating a three-dimensional map of an environment, comprising: a distance sensor configured to receive signals reflected from objects in the environment and generate two or more consecutive scans of the environment at different times; a camera configured to capture two or more consecutive camera images of the environment, wherein each of the two or more consecutive camera images of the environment is captured by the camera at different positions within the environment; and a processing circuit configured to: construct a spatial cost volume of the environment based on the two or more consecutive scans of the distance sensor and the two or more consecutive camera images; determine the pixel-level uncertainty of the spatial cost volume based on the two or more consecutive scans and the two or more consecutive camera images; determine a rough estimate of the depth of an object in the environment based on the spatial cost volume; match points in the two or more consecutive scans with points in the two or more consecutive camera images; determine a fine estimate of the depth of the object based on the rough estimate of the depth of the object and further based on matching points on consecutive distance sensor images with consecutive camera images; and generate a three-dimensional map of the environment based on the fine estimate of the depth of the object, the two or more consecutive scans, and the two or more consecutive camera images.
2. The system according to claim 1, wherein the object is a first object, wherein the processing circuit is further configured to estimate the depth of a second object in the environment based on the two or more consecutive scans and the two or more consecutive camera images, and wherein the processing circuit is configured to generate the three-dimensional map of the environment based on the fine estimate of the depth of the first object and the depth of the second object.
3. The system according to claim 1, wherein the processing circuit is further configured to: measure the rotational movement of the camera relative to the distance sensor; and measure the translational movement of the camera relative to the distance sensor, wherein the processing circuit is configured to generate the three-dimensional map based on the two or more consecutive scans, the two or more consecutive camera images, the rotational movement of the camera, and the translational movement of the camera.
4. The system according to claim 1, wherein the distance sensor includes a radar sensor configured to transmit radar signals and receive the signals reflected from the objects in the environment.
5. The system according to claim 1, wherein the processing circuit is further configured to perform simultaneous localization and mapping based on the two or more consecutive camera images using the two or more consecutive scans as depth constraints.
6. The system according to claim 1, wherein the processing circuit is further configured to: perform a first fast Fourier transform on the two or more consecutive scans to generate an estimate of the depth of an object in the environment; Perform a second fast Fourier transform on the two or more consecutive scans to generate the relative velocity of the object; Perform a third fast Fourier transform on the two or more consecutive scans to generate an estimate of the angle from the distance sensor to the object.
7. The system according to claim 1, wherein the processing circuit is configured to determine the pixel-level uncertainty of the spatial cost volume based on the rotational movement of the camera and the translational movement of the camera.
8. A method for generating a three-dimensional map of an environment, comprising: Receiving, by a processing circuit from a distance sensor, two or more consecutive scans of the environment performed by the distance sensor at different times, wherein the two or more consecutive scans represent information derived from signals reflected from an object in the environment; Receiving, by the processing circuit, two or more consecutive camera images of the environment captured by a camera, wherein each of the two or more consecutive camera images of the object is captured by the camera at different positions within the environment; Constructing, by the processing circuit, a spatial cost volume of the environment based on the two or more consecutive scans of the distance sensor and the two or more consecutive camera images; Determining, by the processing circuit, the pixel-level uncertainty of the spatial cost volume based on the two or more consecutive scans and the two or more consecutive camera images; Determining, by the processing circuit, a rough estimate of the depth of an object in the environment based on the spatial cost volume; Matching points in the two or more consecutive scans with points in the two or more consecutive camera images; Determining a fine estimate of the depth of the object based on the rough estimate of the depth of the object and further based on matching points on consecutive distance sensor images with consecutive camera images; and Generating, by the processing circuit, a three-dimensional map of the environment based on the fine depth of the object, the two or more consecutive scans, and the two or more consecutive camera images.
Citation Information
Patent Citations
Applying an annotation to an image based on keypoints
US10778916B2
Integrated radar and ADS-b
US20180246200A1
Digital active phased array radar
US20190113610A1
Mobile imaging platform calibration
CN105378506A
Automatic navigation method and device
CN106931961A