LIDAR SYSTEM AND METHOD FOR OPERATING A LIDAR SYSTEM - Patent application
The lidar system adjusts the transmit field-of-view independently of the receive field-of-view, enhancing object detection accuracy and reducing data processing by focusing on specific regions of interest, addressing the issue of fixed resolution grids in lidar systems.
Patent Information
- Application Number
- JP2025521240
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2023-10-12
- Publication Date
- 2025-11-05
AI Technical Summary
Lidar systems with fixed resolution grids face reduced quality in object detection when objects do not fully fill the field of view, leading to integration of non-reflected pulses and compromised estimated point clouds.
A lidar system with adjustable transmit field-of-view (Tx FoV) that moves transmit optics along a lateral axis without altering the receive field-of-view (Rx FoV), allowing for dynamic adjustment of Tx FoV to enhance object detection and reduce unnecessary data processing.
Improves object detection accuracy and reduces unnecessary data processing by focusing transmit optics on specific regions of interest, maintaining high detection probability and resolution without altering the receive field-of-view.
Smart Images

Figure 2025536279000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION Embodiments of the present invention relate to a lidar system.Embodiments of the present invention relate to a method of operating a lidar system. [Background technology]
[0002] A vehicle may include a sensor system that monitors the external environment for obstacle detection and avoidance. The sensor system may include multiple sensor assemblies for monitoring nearby objects near the vehicle and distant objects. Each sensor assembly may include at least one of a camera, a radar (radio detection and ranging) sensor, a lidar (light detection and ranging) sensor, and a microphone. A lidar sensor includes one or more emitters that transmit light pulses from the vehicle to the outside and one or more detectors that receive and analyze the reflected light pulses. A lidar sensor may include one or more optical elements for concentrating and directing the transmitted and received light from within a field-of-view (FOV) outside the vehicle. The sensor system may determine the position of an object in the external environment based on data acquired from the sensors. The vehicle may control one or more vehicle systems, including a powertrain, braking systems, and steering systems, based on the position of the object.
[0003] Lidar systems using a fixed resolution grid integrate a fixed number of pulses or detector readings to form a superpixel. However, if the object does not fully fill the field of view of the superpixel, pulses or detector readings that are not reflected by the object will be integrated, resulting in a reduced quality of the estimated point cloud. Lidar systems can also be used in other applications such as aircraft, ships, and / or mapping systems. Summary of the Invention [Problem to be solved by the invention]
[0004] In one embodiment, the lidar system includes a series of emitters, each configured to transmit light pulses from the vehicle along a transmit axis to form a transmit field-of-view (Tx FoV). At least one detector is configured to receive at least a portion of the light pulses reflected from objects within a receive field-of-view (Rx FoV) along a receive axis. The transmit optic is configured to be movable along a transverse axis and intersect each transmit axis without intersect- ing the receive axis to adjust the Tx FoV without adjusting the Rx FoV.
[0005] In another embodiment, a method for adjusting a transmission field-of-view (Tx FoV) is provided. Optical pulses are transmitted from a vehicle along at least one transmission axis to form a transmit field-of-view (Tx FoV). At least a portion of the optical pulses reflected from objects within a reception field-of-view (Rx FoV) are received along a reception axis. The transmit optical device is moved along a lateral axis to intersect each transmission axis without intersecting the reception axis to adjust only the Tx FoV without adjusting the Rx FoV.
[0006] In yet another embodiment, a non-transitory computer-readable medium includes stored instructions that, when executed by one or more computing devices, perform the following tasks: transmit optical pulses from a vehicle to form a transmission field-of-view (Tx FoV); receive at least a portion of the optical pulses reflected from objects within a reception field-of-view (Rx FoV); and move the transmit optical device along a horizontal axis to adjust only the Tx FoV without adjusting the Rx FoV.
[0007] The present disclosure relates to embodied systems and methods for operating a LIDAR system, the method including at least one of: each photodetector performing operations to facilitate measurements associated with an optical signal reflected from an object external to the LIDAR system; a processor receiving result values from the photodetectors indicating times at which photons were detected at or near a target wavelength; the processor combining different sets of the result values to generate a plurality of superpixels; the processor using the generated superpixels to obtain spatiotemporal coherence metrics; the processor selecting a subset of optical pulses or a group of result values based on the spatiotemporal coherence metrics; and / or the processor detecting a distance between the LIDAR system and the object based on the selected subset of optical pulses or group of result values.
[0008] The embodied system may include a processor and a non-volatile computer-readable storage medium having stored thereon programming instructions configured to cause the processor to execute a method for operating the lidar system. The methods described above may also be embodied by a computer program product including a memory and programming instructions configured to cause a processor to perform the operations.
[0009] The present disclosure relates to embodied systems and methods for operating a LIDAR system, the method including: a processor arranging pixels into a grid, where the pixels contain result values generated from a processed waveform generated by a photodetector of the LIDAR system; identifying a first region of interest within the grid based on correlations between distance values associated with the pixels and / or correlations between intensity values associated with the pixels; combining the result values associated with pixels located within the first region of interest to generate feature values; and generating superpixels whose values are set to the generated feature values.
[0010] The embodied system may include a processor and a non-volatile computer-readable storage medium having stored thereon programming instructions configured to cause the processor to perform a method of operating the lidar system. The methods described above may also be embodied by a computer program product including a memory and programming instructions configured to cause a processor to perform the operations. [Means for solving the problem]
[0011] A LiDAR system according to one embodiment of the present invention includes an array of transmitters configured to transmit light pulses from a vehicle along a transmission axis to form a Tx FoV (transmission field-of-view); at least one detector configured to receive at least a portion of the light pulses that reflect off objects in an Rx FoV (reception field-of-view) along a receive axis; and transmit optics mounted for movement along a lateral axis and configured to intersect each of the transmission axes but not the receive axis to adjust the Tx FoV without adjusting the Rx FoV.
[0012] According to one embodiment of the present invention, the Tx FoV and the Rx FoV overlap, and the adjusted Tx FoV is located within the area of the Tx FoV.
[0013] According to one embodiment of the present invention, the system further includes a collimator mounted adjacent to the series of emitters and configured to focus and direct the light pulses along each transmit axis to collectively form a transmit beam.
[0014] According to one embodiment of the present invention, the transmit optical device is arranged adjacent to the collimator and configured to focus the transmit beam in a region of the Tx FoV to form the adjusted Tx FoV, and the transmit optical device includes a cylindrical lens.
[0015] According to one embodiment of the present invention, the series of emitters includes a linear array of emitters arranged parallel to the horizontal axis, the linear array of emitters including a proximal emitter and a distal emitter arranged opposite the proximal emitter. According to one embodiment of the present invention, the system further includes an actuator coupled to the transmit optical device and configured to move the transmit optical device through a range between a rest position, in which the optical device does not intersect any transmit axis of the linear array of emitters, and a distal position, in which the distal emitter intersects the transmit axis.
[0016] According to one embodiment of the present invention, the system further includes a controller configured to move the transmit optical device along the horizontal axis, the controller determining from the received light pulses whether the object is an unknown object; and further configured to move the transmit optical device along the horizontal axis between a proximal position and a distal position while light pulses are transmitted via the transmit optical device.
[0017] According to one embodiment of the present invention, the controller is further configured to receive sweep data indicative of the light pulses reflecting from the unknown object while moving the transmit optical device; determine a position of the unknown object based on the sweep data; and move the transmit optical device to a position along the horizontal axis such that the adjusted Tx FoV is aligned with the position of the unknown object.
[0018] A LiDAR system according to one embodiment of the present invention includes a processor; and a non-transitory computer-readable storage medium including programming instructions configured to cause the processor to implement a method for operating a LiDAR system, the programming instructions including instructions to receive from a photodetector a time result value indicating a time at which the photodetector detects a photon at or near a target wavelength; combine different sets of the result values to generate superpixels; use the superpixels to obtain a first spatiotemporal coherence metric; select a subset of light pulses or a group of result values based on the first spatiotemporal coherence metric; and detect a distance between the LiDAR system and the object based on the selected subset of light pulses or the selected group of result values.
[0019] According to one embodiment of the invention, the first spatiotemporal coherence metric comprises a metric specifying a change in distribution between detection of two pulses or groups of two pulses, respectively, by the plurality of photodetectors, and the subset of light pulses or group of result values is selected based on the largest value of the metric.
[0020] According to one embodiment of the invention, the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the differences between consecutive timestamps sorted from lowest to highest or lowest to highest, and a selected subset of the light pulses or group of result values comprises light pulses or result values associated with a relatively low measured variance.
[0021] According to one embodiment of the present invention, the first spatiotemporal coherence metric comprises a score for each pulse of the optical signal indicating the reliability or validity of object detection, and the pulses are selected to be included in the subset if the score exceeds a value.
[0022] A LiDAR system according to one embodiment of the present invention includes a processor; and a non-transitory computer-readable storage medium including programming instructions configured to cause the processor to implement a method for operating a LiDAR system, the programming instructions including instructions to arrange a plurality of pixels in a grid, the plurality of pixels including result values generated from a processed waveform generated by a photodetector of the LiDAR system; identify a first region of interest based on at least one of a correlation between range values associated with the plurality of pixels and a correlation between intensity values associated with the plurality of pixels; combine result values associated with pixels located within the first region of interest to generate at least one first feature value; and generate a first superpixel having a value set to the at least one first feature value.
[0023] According to one embodiment of the present invention, the programming instructions further include instructions to obtain a kernel size and use the kernel size to identify the area of interest of the grid.
[0024] According to one embodiment of the present invention, the kernel size is obtained by finding the location of at least one of the plurality of pixels that is a closest neighbor to the pixel of interest of the grid in at least a range aspect; and defining the kernel size based on the location of the closest neighbor within the grid.
[0025] According to one embodiment of the present invention, the kernel size is obtained by obtaining a reference kernel size; identifying a region within the grid using the reference kernel size; identifying a central pixel of the region; calculating a score for each pixel of the region using the result value associated therewith, the score indicating the degree of correlation between the result value associated with the pixel and the central pixel; selecting a pixel from the plurality of pixels based on the score; and defining the kernel size based on the position of the selected pixel within the grid.
[0026] A method for operating a LiDAR system according to one embodiment of the present invention includes receiving, by a processor, result values from a plurality of photodetectors indicating times at which the plurality of photodetectors detect photons at or near a target wavelength, the result values being based on operations performed by each of the plurality of photodetectors to facilitate measurements associated with light signals reflected from an object external to the LiDAR system; combining, by the processor, other sets of the result values to generate superpixels; using, by the processor, the superpixels to obtain a first spatiotemporal coherence metric; selecting, by the processor, a subset of light pulses or a group of result values based on the first spatiotemporal coherence metric; and detecting, by the processor, a distance between the LiDAR system and the object based on the selected subset of light pulses or the selected group of result values.
[0027] According to one embodiment of the present invention, the first spatiotemporal coherence metric comprises at least one of a distribution comparison metric, a time-of-flight statistical metric, and a detection confidence score, the first spatiotemporal coherence metric comprising a metric specifying a change in distribution between detection of two pulses or groups of two pulses by the plurality of optical detectors, respectively, and the subset of optical pulses or the group of result values is selected based on the largest value of the metric.
[0028] According to one embodiment of the invention, the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the differences between consecutive timestamps sorted from lowest to highest or lowest to highest, and a selected subset of the light pulses or group of result values comprises light pulses or result values associated with a relatively low measured variance. [Brief explanation of the drawings]
[0029] [Figure 1]FIG. 1 is a front perspective view of an exemplary vehicle with a self-driving system (SDS) including a lidar sensor with an adjustable transmit field of view (Tx FoV) according to aspects of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram illustrating communication between an SDS and other systems and devices according to aspects of the present disclosure. [Figure 3] 1 is a diagram illustrating an example structure of a lidar sensor of an SDS, according to an aspect of the present disclosure. [Figure 4] FIG. 1 illustrates a plan view of a lidar sensor according to an aspect of the present disclosure. [Figure 5] FIG. 5 is a cross-sectional view of the lidar sensor of FIG. 4 taken along line VV, in accordance with an aspect of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of a lidar sensor providing a Tx FoV, according to an aspect of the present disclosure. [Figure 7] FIG. 10 is a schematic diagram of another lidar sensor illustrating the transmit optics adjusted to a first position to adjust the Tx FoV to a first region relative to the overall Tx FoV, in accordance with aspects of the present disclosure. [Figure 8] FIG. 8 is yet another schematic diagram of the lidar sensor of FIG. 7 illustrating the transmit optics adjusted to a second position to adjust the Tx FoV to a second region relative to the overall Tx FoV, in accordance with aspects of the present disclosure. [Figure 9] FIG. 8 is yet another schematic diagram of the lidar sensor of FIG. 7 illustrating the transmit optics adjusted to a third region relative to the overall Tx FoV. [Figure 10] 8 is a diagram illustrating the overall Tx FoV and receive field of view (Rx FoV) of FIG. 7 according to an aspect of the present disclosure. [Figure 11] 1 is a diagram illustrating the relationship between adjusted Tx FoV and Rx FoV according to an aspect of the present disclosure. [Figure 12] FIG. 10 is a flow diagram illustrating a method for adjusting a Tx FoV according to an aspect of the present disclosure. [Figure 13] 4 is a diagram illustrating a technique for combining results from the photodetectors of the lidar system shown in FIG. 3. [Figure 14a] We provide a description of a convolutional oversampling technique for combining non-overlapping result sets from the photodetectors of the lidar system shown in Figure 3. The combination is performed by functions F1 through F5. [Figure 14b] We provide a description of a convolutional oversampling technique for combining non-overlapping result sets from the photodetectors of the lidar system shown in Figure 3. The combination is performed by functions F1 through F5. [Figure 14c] We provide a description of a convolutional oversampling technique for combining non-overlapping result sets from the photodetectors of the lidar system shown in Figure 3. The combination is performed by functions F1 through F5. [Figure 14d] We provide a description of a convolutional oversampling technique for combining non-overlapping result sets from the photodetectors of the lidar system shown in Figure 3. The combination is performed by functions F1 through F5. [Figure 14e] We provide a description of a convolutional oversampling technique for combining non-overlapping result sets from the photodetectors of the lidar system shown in Figure 3. The combination is performed by functions F1 through F5. [Figure 15] 10 is a diagram illustrating an example in which only two of five pulses P1 to P5 are reflected from an object and integrated. [Figure 16] FIG. 1 provides a flow diagram of a method illustrating variable resolution refinement in a Geiger-mode lidar. [Figure 17] 1 is a diagram providing an example useful for understanding binning. [Figure 18] 1 is a diagram providing an example of a histogram. [Figure 19] 1 is a diagram illustrating a target object; [Figure 20] 1 provides a flow diagram of an exemplary method for spatial processing of lidar data according to the present solution. [Figure 21a] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21b] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21c] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21d] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21e] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21f] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 21g] 10A-10C provide an illustration of yet another technique for combining results generated from lidar waveforms. [Figure 22] 10A-10C provide examples showing modified or adjusted kernel sizes and / or regions of interest (ROIs). [Figure 23] 10A-10C provide examples showing modified or adjusted kernel sizes and / or regions of interest (ROIs). [Figure 24] 10 is a diagram providing an example of an ROI having a position adjusted or modified with a grid. [Figure 25] 10 provides a flow diagram of yet another exemplary method for spatial processing of lidar data in accordance with the present solution. [Figure 26] 1 is a diagram providing an example of a system. [Figure 27] 1 is a diagram providing a more detailed illustration of an autonomous vehicle. [Figure 28] FIG. 1 is a diagram providing an example block diagram of a vehicle trajectory planning process. [Figure 29] 1 is a diagram providing an example of a computer system. DETAILED DESCRIPTION OF THE INVENTION
[0030] Where necessary, detailed embodiments are disclosed herein; however, it should be understood that the disclosed embodiments are merely exemplary and may be embodied in various or alternative forms. The drawings are not necessarily made to scale, and some features may be exaggerated or reduced in size to clearly show the details of particular components. Therefore, specific structural and functional details disclosed herein should not be construed as limiting, but should merely be considered as representative standards for teaching those of ordinary skill in the art and guiding them to variously embody the present disclosure.
[0031] A lidar system can have a fixed resolution. This resolution is defined by attributes of the hardware and software system and is unrelated to the reflective performance of the lidar system. In practice, this means that for a particular target object with a certain size, the lidar system has a constant detection probability as a function of a particular reflectivity and distance, while also having a fixed resolution. However, this may not necessarily be a desirable characteristic in a computer vision pipeline, as described below. Such a pipeline may desire to maintain a constant detection probability for a particular range, and may sacrifice other characteristics of the range-sensing system (e.g., resolution) to achieve this. For example, if it is possible to maintain a constant (high) detection probability even when fewer measurement points are generated for objects with low reflectivity, even if this may result in reduced velocity estimation performance or lower data association, the perception system can be confident that an object is present despite the low reflectivity.
[0032] This specification describes embodiments of systems, apparatus, devices, methods, and / or computer program products, and combinations and subcombinations thereof, that provide variable-resolution refinement of Geiger-mode lidar systems to address detection probability issues in existing lidar systems. These features of the invention can improve the operation of lidar systems and enhance the performance of object detection and vehicle control using lidar data.
[0033] The method generally includes: each photodetector performing an operation to facilitate a measurement associated with an optical signal reflected from an object external to the LIDAR system; a processor receiving result values from the photodetectors indicating the time at which photons were detected at or near the target wavelength; the processor combining different sets of the result values to generate a plurality of superpixels; the processor using the generated superpixels to obtain spatiotemporal coherence metrics; the processor selecting a subset of optical pulses or a group of result values based on the spatiotemporal coherence metrics; the processor detecting a distance between the LIDAR system and the object based on the selected subset of optical pulses or group of result values; and / or the processor using the distance to control operation of the vehicle.
[0034] The spatiotemporal coherence metric can be obtained by considering superpixels for a fixed or variable number of optical signal pulses. In either case, the spatiotemporal coherence metric may include, but is not limited to, a distribution comparison metric, a time-of-flight (ToF) statistical metric, and / or a detection confidence score. In the case of a distribution comparison metric, the spatiotemporal coherence metric may include an metric indicating a distribution change that occurs between the detection of two pulses or two groups of pulses by multiple optical detectors. A subset of optical pulses or a group of result values is selected based on the metric having the largest value. In the case of a time-of-flight statistical metric, the spatiotemporal coherence metric may include a measured variance value for the difference between consecutive timestamps when the timestamps for each pulse are sorted from smallest to largest or from largest to smallest. The selected subset of optical pulses or a group of result values includes optical pulses or result values associated with relatively low measured variance values. In the case of a detection confidence score, the spatiotemporal coherence metric may include a score for each pulse of the optical signal that indicates the confidence or effectiveness of object detection. If the score exceeds a particular value, the pulse is selected for inclusion in the subset. Similarly, if the score associated with a particular pulse exceeds a particular value, the resulting value may be selected for inclusion in the group.
[0035] As used in this document, the singular forms "a," "an," and "the" include the plural unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used in this document have the same meaning as commonly understood by one of ordinary skill in the art. As used in this document, the term "including" means "including, but not limited to." As used in this document, the term "vehicle" refers to any form of transportation capable of carrying one or more human passengers and / or cargo and powered by any form of energy. The term "vehicle" includes, but is not limited to, automobiles, trucks, vans, trains, autonomous vehicles, airplanes, and unmanned aerial vehicles. An "autonomous vehicle (AV)" is a vehicle equipped with a processor and programming instructions and drivetrain that allows the processor to control the drivetrain without the intervention of a human driver. An AV may be fully autonomous, meaning it does not require a human driver in most or all driving conditions and functions, or it may be semi-autonomous, meaning it requires a human driver in certain conditions or for certain operations, or a human may override the autonomous systems and control the vehicle.
[0036] Rotating optical sensors, such as rotating lidar sensors, may include complex physical and electrical structures. Rotating lidar sensors can scan a wide, 360-degree field of view (FoV) around the vehicle. The area from which the lidar sensor's emitters transmit light is called the transmit field of view (Tx FoV), and the area from which the lidar sensor's detectors receive light is called the receive field of view (Rx FoV). Typically, the Tx FoV and Rx FoV overlap. Rotating lidar sensors may include a linear array of emitters to provide an extended Tx FoV over a wide vertical area. In such rotating lidar sensors, the detection range and Tx FoV have an inverse relationship. The larger the Tx FoV, the more widely the optical output is dispersed, and less light reaches small targets. In certain situations, autonomous driving systems (SDSs) may be interested in maximizing the Tx FoV, while in other situations they may be interested in maximizing the detection range within a more limited Tx FoV, such as when identifying unknown objects at long distances. Certain objects, such as tire debris, can be difficult to identify because they have low reflectivity and irregular shapes.
[0037] According to some aspects, the SDS adjusts only the Tx FoV without adjusting the Rx FoV to maintain a maximum Rx FoV under certain conditions, and adjusts the Tx FoV to maximize detection distance with a narrower Rx FoV under other conditions. The lidar sensor has a transmitter assembly including an adjustable transmit optics that is controlled to vertically move the transmit optics to adjust only the Tx FoV without adjusting the Rx FoV. The transmit optics changes the divergence angle of the transmitted beam to focus the Tx FoV on a smaller region of the Rx FoV, which reduces the size of the overall point cloud that must be analyzed and improves the performance of the lidar sensor. By reducing the Tx FoV, the number of photons irradiated from the emitter to the target increases in a smaller region of interest. In this case, the spatial resolution does not increase because the Rx FoV remains unchanged.
[0038] If a lidar sensor is to adjust both the Tx FoV and the Rx FoV, the adjustments must be synchronized so that the emitter and detector scan the same field of view. The advantage of adjusting only the Tx FoV without adjusting the Rx FoV is that the detector does not need to know the exact location where the Tx FoV is adjusted, as long as it is within the Rx FoV. Another advantage of adjusting only the Tx FoV without adjusting the Rx FoV is that the emitter and detector and associated detector lenses do not move, so no additional moving electronics are required.
[0039] 1 , a lidar sensor according to one or more embodiments is shown and generally designated by the reference numeral 100. The lidar sensor 100 is integrated with an autonomous driving system (SDS) 102 of a vehicle 104, such as an autonomous vehicle. The SDS 102 includes multiple sensors 106 that monitor the environment external to the vehicle 104. The lidar sensor 100 adjusts only its transmit field of view (Tx FoV) without adjusting its receive field of view (Rx FoV) to monitor a specific unidentified object 110, such as tire debris, in the environment external to the vehicle 104.
[0040] The SDS 102 includes multiple sensor assemblies, each including one or more sensors 106 that monitor a 360-degree field of view around the vehicle 104 at near and long range. The SDS 102 includes a top sensor assembly 112, two side sensor assemblies 114, two front sensor assemblies 116, and a rear sensor assembly 118. Each sensor assembly includes one or more of the following sensors 106: a camera, a lidar sensor, and a radar sensor.
[0041] The upper sensor assembly 112 is mounted on a loop around the vehicle 104 and includes multiple sensors 106, such as a lidar sensor and multiple cameras. The lidar sensor rotates around an axis to scan a 360-degree field of view (FoV) around the vehicle 104. The side sensor assemblies 114 are mounted on the sides of the vehicle 104, such as on the front fenders or inside the side mirrors of the vehicle 104, as shown in FIG. 1. Each side sensor assembly 114 includes multiple sensors 106, such as a lidar sensor and cameras, to monitor a near-field of view around the vehicle 104. The front sensor assemblies 116 are mounted on the front of the vehicle 104, such as below the headlights. Each front sensor assembly 116 includes multiple sensors 106, such as a lidar sensor, a radar sensor, and cameras, to monitor a long-field of view ahead of the vehicle 104. The rear sensor assemblies 118 are mounted, for example, at the rear top end of the vehicle 104, near the center high-mounted stop lamp (CHMSL). The rear sensor assembly 118 includes multiple sensors 106 , such as cameras and lidar sensors, for monitoring the rear view of the vehicle 104 .
[0042] 2 illustrates communication between the SDS 102 and other systems and devices in accordance with various aspects of the present invention. The SDS 102 includes a sensor system 200 and a controller 202. The controller 202 can communicate with other systems and devices directly or via a transceiver 204.
[0043] The sensor system 200 includes sensor assemblies, such as the upper sensor assembly 112 and the forward sensor assembly 116. The upper sensor assembly 112 includes one or more sensors, such as a lidar sensor 100, a radar sensor 208, and a camera 210. The camera 210 may be a visible light camera or an infrared camera. The sensor system 200 may include additional sensors, such as a microphone, a sonar sensor, a temperature sensor, a position sensor (e.g., a GPS), a position tracking sensor, a fuel sensor, a motion sensor (e.g., an inertial measurement unit (IMU)), a humidity sensor, an occupancy detection sensor, etc. The sensor system 200 provides sensor data 212 indicative of the environment external to the vehicle 104. The controller 202 analyzes the sensor data to identify and determine the location of objects external to the vehicle 104, such as traffic lights, distant vehicles, pedestrians, etc.
[0044] The SDS 102 also communicates with one or more vehicle systems 214, such as the engine, transmission, navigation system, and braking system, via a transceiver 204. The controller 202 can receive information from the vehicle systems 214 indicative of current operating conditions, such as vehicle speed, engine speed, status of turn signals, brake position, vehicle position, steering angle, ambient temperature, etc. The controller 202 can also control these vehicle systems 214 based on the sensor data 212; for example, the controller 202 can control the braking and steering systems for obstacle avoidance. The controller 202 can communicate directly with the vehicle systems 214 or indirectly with the vehicle systems 214 via a vehicle communication bus, such as a CAN bus 216.
[0045] The SDS 102 may also communicate with external objects 218, such as external vehicles and structures, to share external environmental information or collect additional external environmental information. The SDS 102 may include a vehicle-to-everything (V2X) transceiver 220 coupled to the controller 202 for communicating with the objects 218. For example, the SDS 102 may use the V2X transceiver 220 to perform vehicle-to-vehicle (V2V) communication directly with remote vehicles, vehicle-to-infrastructure (V2I) communication with structures (e.g., signs, buildings, traffic lights), or vehicle-to-motorcycle (V2M) communication with motorcycles.
[0046] The SDS 102 can communicate with a remote computing device 222 over a communications network 224 using one or more of the transceivers 204, 220, and can, for example, provide messages or visual information indicating the position of the object 218 relative to the vehicle 104 based on the sensor data 212. The remote computing device 222 can include one or more servers for handling one or more processes of the techniques described herein. The remote computing device 222 can also exchange data with a database 226 over the network 224.
[0047] The SDS 102 includes a user interface 228 that provides information to a user of the vehicle 104. The controller 202 can control the user interface 228 to provide messages or visual information indicating the position of the object 218 relative to the vehicle 104 based on the sensor data 212.
[0048] Although the controller 202 is described as a single controller, it may include multiple controllers or may be embodied in the form of software code within one or more other controllers. The controller 202 includes a processing unit or processor 230, which may include a microprocessor, ASIC, IC, memory (e.g., FLASH, ROM, RAM, EPROM, EEPROM, etc.), and software code that interact with each other to perform a series of operations. Such hardware and / or software may be grouped together into an assembly to perform a particular function. Any one of the controllers or devices described herein may include computer-executable instructions compiled or interpreted from a computer program created using various programming languages and technologies. The controller 202 also includes a memory 232 or non-volatile computer-readable storage medium capable of executing instructions of a software program. The memory 232 may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. Generally, the processor 230 receives and executes instructions from the memory 232, a computer-readable medium, etc. The controller 202 may also include predetermined data or "look-up tables" stored in memory.
[0049] FIG. 3 shows an example structure of a lidar sensor 300, such as the lidar sensor 100 of the upper sensor assembly 112. The lidar sensor 300 includes a base 302 mounted to the vehicle 104. The base 302 includes a motor 304 with a shaft 306 extending along an axis AA. The lidar sensor 300 also includes a housing 308 secured to the shaft 306 and rotatably mounted to the base 302 about the axis AA. The housing 308 includes an opening 310 and a cover 312 secured to the opening 310. The cover 312 is formed of an optically transparent material, such as glass. While FIG. 3 shows one cover 312, the lidar sensor 300 may include multiple covers 312 or a single cover 312 covering the entire exterior surface of the housing 308. The lidar sensor 300 may have a housing 308 that can rotate 360 degrees around the central axis or hub 306 of the motor 304. The housing 308 may include emitter / receiver openings 310 made of an optically transparent material. While only one opening is shown in FIG. 3 , the solution of the present invention is not limited in this respect, and multiple openings for emitting and receiving light may be provided in other situations. In any case, the lidar sensor 300 may emit light through one or more openings 310 and receive reflected light again through one or more openings 310 as the housing 308 rotates around the internal components. In an alternative situation, the outer shell of the housing 308 may be in the form of a fixed dome made at least partially of an optically transparent material, with a rotatable component disposed within the housing 308.
[0050] Within the rotating shell or fixed dome are emitters 316 constructed and arranged to generate and emit light pulses through the aperture 310 or transparent dome of the housing 308 via one or more laser emitting chips or other light emitting devices. Emitters 316 may include any number of individual emitters (e.g., 8 emitters, 64 emitters, or 128 emitters). The emitters can emit light at approximately the same intensity or at different intensities. The lidar sensor 300 also includes a detector 318, which includes an array of photodetectors. The photodetectors are positioned and configured to receive light reflected back into the system. Upon receiving the reflected light, the photodetectors generate a result (or electrical pulse) indicative of the measured intensity of the optical signal reflected from an object external to the lidar sensor. In Geiger mode applications, the photodetectors are activated when they detect a single photon near or at a target wavelength. The time at which the photodetectors are activated is recorded as a timestamp. The emitter 316 and detector 318 rotate with the rotating shell or within a fixed dome of the housing 308. One or more optics structures 322 may be positioned in front of the emitter 316 and / or detector 318 and act as lenses or wave plates to focus and direct light passing through the optics structures 322. The lidar sensor 300 includes one or more emitters 316 for transmitting light pulses 320 that pass through the cover 312 in a transmit field of view (Tx FoV, see FIG. 1 ) far away from the vehicle 104. The light pulses 320 strike one or more objects in the Rx FoV and then return to the lidar sensor 300 as reflected light pulses 328. The lidar sensor 300 also includes one or more detectors 318 for receiving the reflected light pulses 328 that pass through the cover 312. The detectors 318 may also receive light from an external light source, such as the sun. The lidar sensor 300 rotates about an axis AA to scan an area within its field of view (FoV). The emitter 316 and detector 318 can be stationary and mounted to the base 302, or movable and mounted to the housing 308. The emitter 316 is a light emitter, and the detector 318 is a light detector.
[0051] The emitters 316 may include laser emitting chips or other light emitting devices and may include any number of individual emitters (e.g., 8 emitters, 64 emitters, or 128 emitters). The emitters may be arranged in a linear array or laser bar configuration, as shown in FIG. 3. The emitters 316 may transmit light pulses 320 of approximately the same or varying intensities and may have various waveforms (e.g., sine waves, square waves, sawtooth waves, etc.). The lidar sensor 300 may include one or more optical structures 322 for focusing and guiding light passing through the cover 312. One or more optical structure 322 may be positioned in front of a mirror (not shown) to focus and guide the light passing through the optical structure. As shown in FIG. 3, a single optical structure 322 is positioned in front of the mirror and coupled to a rotating component of the system, so that the optical structure 322 rotates with the mirror. Optionally, the optical structure 322 may include multiple structures (e.g., lenses and / or wave plates). Alternatively, multiple optical device structures 322 may be arranged on or integrally formed with the shell portion of housing 308 .
[0052] Detector 318 may include a photodetector or photodetector array positioned to receive reflected light pulse 328. Detector 318 may be arranged in a linear array, as shown in FIG. 3. According to one aspect of the invention, detector 318 includes a plurality of pixels, each pixel including a Geiger-mode avalanche photodiode for detecting a reflection of the light pulse in each of a plurality of detection frames. In other embodiments, detector 318 may include a passive imager.
[0053] The lidar sensor 300 includes a controller 330, including a processor 332 and memory 334, that controls various components, such as the motor 304, emitter 316, and detector 318. The controller 330 also analyzes data collected from the detector 318 to determine characteristics of the received light and generate information about the environment external to the vehicle 104. For example, the controller 330 may generate a three-dimensional point cloud based on the data collected from the detector 318. The controller 330 may be integrated with other controllers, such as the controller 202 of the SDS 102. The lidar sensor 300 also includes a power unit 336, which receives power from a vehicle battery 338 and provides electrical power to the motor 304, emitter 316, detector 318, and controller 330. The lidar sensor 300 includes an analyzer 330A that includes elements such as the processor 332 and non-volatile computer-readable memory 334 containing programming instructions. The programming instructions configure the system to receive and analyze data from the light detector 318 and generate environmental data. Optionally, analyzer 330A may be integral with lidar sensor 300, as shown in the drawings, or may be externally located and coupled to the lidar sensor via a wired or wireless network. Analyzer 330A may include controller 330.
[0054] In a Geiger-mode lidar system scenario, the photodetector 318 operates when it detects a single photon at or near a target wavelength. The times the photodetector operates (and the associated emitter firing times) are accumulated in a histogram. A peak search operation is then performed on this histogram to obtain the depth of objects that reflected photons of the target wavelength in the spatial region covered by the photodetector. That is, a series of individual laser emissions across the field of view are obtained and tallied, using a peak-finding algorithm to determine the depth of the target object as a function of the time-of-flight of the observed photons (which may include, at a minimum, other signals).
[0055] Conventional systems do not require Geiger-mode lidar systems to use the same number of laser shots per pixel to recover histogram peaks. Indeed, the physical consequence of successively accumulating laser shots is that the frustum described by the rotating sensor simply expands along the relevant axis (i.e., the system's likelihood of summing reflections from various surfaces increases). Furthermore, there is also no restriction preventing the integration of return signals from adjacent photodetectors; this simply adjusts the frustum imaged by a particular lidar point. In fact, for a sensor with no spacing, each point represents the average depth of the nearest surfaces from each ray within a particular frustum. Incorporating these insights, the present invention modifies existing Geiger-mode lidar systems to use a variable-sized aggregation window depending on reflectivity or its proxy (e.g., return count, return signal noise, or observations from other related sensor systems, such as cameras).
[0056] The lidar sensor 300 uses the resulting values from the photodetectors 318 to generate measured 3D points by aggregating the photodetector results. Because a single photodetector generally cannot provide a sufficient depth measurement, the results from multiple photodetectors are combined into superpixels. One exemplary technique for generating superpixels is shown in FIG. 13.
[0057] 4 and 5 illustrate an exemplary lidar sensor 400. Similar to the lidar sensor 300 of FIG. 3, the lidar sensor 400 includes a housing 408 that includes an opening 410 and a cover 412 secured to the opening 410. The lidar sensor 400 includes one or more emitters 416 that transmit light pulses through the cover 412 and one or more detectors 418 that receive reflected light pulses that pass through the cover 412. In accordance with aspects of the invention, the emitters 416 and detectors 418 are each arranged in a linear array. The lidar sensor 400 includes a transmitting assembly 424 that includes the emitters 416 and a receiving assembly 426 that includes the detectors 418.
[0058] The transmitter assembly 424 includes a circuit board assembly 428 that controls the emitter 416. The circuit board assembly 428 includes a controller 430 that includes a processor 432 and a memory 434 mounted on a circuit board 435.
[0059] The transmit assembly 424 also includes multiple optical devices, including a collimator 436 and a transmit optics 438. The collimator 436 focuses and directs the optical pulses from each emitter 416 along a transmit (Tx) axis 440, as shown in FIG. 7, to form a single overall Tx beam. The transmit optics 438 are disposed between the collimator 436 and the cover 412 to concentrate the Tx beam into a smaller region of the Tx FoV. The transmit optics 438 may be a converging lens, such as a cylindrical lens, that focuses the optical pulses onto a single axis. The transmit assembly 424 also includes an actuator 442, such as a linear actuator, coupled to the transmit optics 438 and controlled by the controller 430. The actuator 442 moves the transmit optics 438 along a horizontal axis 444 that is perpendicular to the transmit axis 440. The actuator 442 linearly moves the transmit optics 438 from a rest position 446, which does not intersect any Tx axis, to a fully extended position 448, which intersects the Tx axis of the furthest-located emitter in the linear array of emitters 416. In accordance with aspects of the present invention, the actuator 442 can adjust the transmit optics 438 according to the rotational speed of the lidar sensor 400. For example, in one embodiment, when the lidar sensor 400 rotates at 10 Hz, or 600 revolutions per minute (rpm), the actuator 442 can adjust the transmit optics 438 from the rest position 446 to the distal position 448 over 100 milliseconds (ms). The actuator 442 can be a linear actuator, such as a voice coil. In accordance with aspects of the present invention, the stroke or linear adjustment of the actuator 442 is determined by the length of the linear array of emitters 416.
[0060] The receiver assembly 426 includes detectors 418 mounted on a circuit board 450. A controller 430 is coupled to the circuit board 450 to receive data from the detectors 418. The controller 430 analyzes the data collected by the detectors 418 and generates information about the surrounding environment of the lidar sensor 400. The receiver assembly 426 also includes one or more detector optics 452. The detector optics 452 may include a collimator to focus and direct received light pulses along a receive (Rx) axis 454 to each detector 418.
[0061] 4, the transmit assembly 424 is offset from the receive assembly 426. As the transmit optics 438 moves along the horizontal axis 444, the transmit optics 438 intersects the Tx axis 440 but does not intersect the Rx axis 454. As shown in FIG. 4, the Tx FoV and Rx FoV overlap. However, because the transmit optics 438 does not intersect the Rx axis 454, the lidar sensor 400 can adjust only the Tx FoV to track the object 510 without adjusting the Rx FoV.
[0062] 6 illustrates an example lidar sensor 600. Similar to lidar sensor 400, lidar sensor 600 transmits light pulses that generally form a Tx beam 660 within the Tx FoV. Unlike lidar sensor 400, lidar sensor 600 does not include transmit optics 438 for adjusting the Tx FoV.
[0063] 7-9 illustrate yet another exemplary lidar sensor 700. Similar to lidar sensor 400, lidar sensor 700 includes an array of emitters 716 that emit light pulses that collectively form a Tx beam 760 within the Tx FoV. Also similar to lidar sensor 400, lidar sensor 700 includes transmit optics 738 for forming an adjusted Tx FoV (Tx FoVADJ). Lidar sensor 700 includes a linear array of emitters 716 that includes a distal emitter 762, a central emitter 764, and a proximal emitter 766. FIGS. 7-9 illustrate a comparison between the Tx FoV and the adjusted Tx FoV (Tx FoVADJ) by moving the transmit optics 738 along a horizontal axis 744.
[0064] Figure 7 shows the transmit optics 738 adjusted to a distal position 748 to intersect the T x-axis of the distal emitter 762, forming a Tx FoV ADJ in an upper region 768 of the overall Tx FoV. Figure 8 shows the transmit optics 738 adjusted to a central position 770 to intersect the T x-axis of the central emitter 764, forming a Tx FoV ADJ in a central region 772 of the overall Tx FoV. Figure 9 shows the transmit optics 738 adjusted to a proximal position 774 to intersect the T x-axis of the proximal emitter 766, forming a Tx FoV ADJ in a lower region 776 of the overall Tx FoV. The controller 430 can adjust the position of the transmit optics 738 to generate an adjusted Tx FoV in order to track an object 710, such as a tire or tire debris.
[0065] FIGS. 10 and 11 show a comparison of the Tx FoV and the Rx FoV. Referring to FIG. 10, the Tx FoV and the Rx FoV are shown adjacent to each other so that the two fields of view have the same size, but as shown in FIG. 11, they overlap in the vehicle's 104 external environment. FIG. 11 also shows the adjusted Tx FoV after adjustment by the transmit optics 438. As the transmit optics 438 moves along the horizontal axis 444, the adjusted Tx FoV moves to overlap a different region of the Rx FoV. For example, referring to FIGS. 6 through 9, when the transmit optics 738 is in the distal position 748 (see FIG. 7), the lidar sensor 700 generates a Tx FoV ADJ in the upper region 768 of the Rx FoV. When the transmit optics 738 is adjusted to the central position 770 (see FIG. 8), the lidar sensor 700 generates a Tx FoV ADJ in the central region 772 of the Rx FoV. When the transmit optics 738 are adjusted to a proximal position 774 (see FIG. 9), the lidar sensor 700 generates a Tx FoVADJ in the lower region 776 of the Rx FoV. When the transmit optics 738 are adjusted to a rest position (or initial position) (see FIG. 4) that does not intersect any of the Tx axes, the Tx FoV returns to its original full range and overlaps with the Rx FoV, as shown on the right side of FIG. 11.
[0066] As the adjusted Tx FoV moves between different regions, the Rx FoV remains unchanged. This results in a wider FoV and a longer detection distance. Generally, the wider the FoV, the less capable the optical device is of scanning a small target area. That is, a wider FoV reduces the amount of light available to scan an object, so the width of the FoV must be reduced to obtain a clear image. However, by adjusting only the Tx FoV without adjusting the Rx FoV, a greater amount of light can be focused on a small target area to identify unknown objects 710, such as tires or tire debris, without reducing the width of the FoV.
[0067] 12, a flow diagram illustrating a method for adjusting a Tx FoV is shown and generally designated by the reference numeral 600. Method 600 is embodied in software code executed by controller 430 according to one or more embodiments. Although the flow diagram is shown as a series of sequential steps, one or more steps may be omitted or performed in other manners.
[0068] In step 602, the controller 430 controls the lidar sensor 400 to scan a 360-degree field of view around the vehicle 104 in the entire Tx FoV. The transmit optics 438 is located at an initial position 446 (see FIG. 4 ) and does not adjust the Tx FoV. The controller 430 analyzes the data from the emitter 416 to observe environmental changes. In step 604, the sensor 400 determines whether an unknown object 710, such as a tire or tire debris, has been detected within the Rx FoV. If the object is not detected, the controller 430 returns to step 602. If the controller 430 detects an unknown object 710 in step 604, the controller 430 proceeds to step 606.
[0069] In step 606, the controller 430 sweeps the Tx FoV and controls the lidar sensor 400 to perform an additional scan or series of scans. During step 606, the controller 430 controls the actuator 442 to move the transmit optics 438 at a predetermined speed within a determined range, such as between a proximal position 774 and a distal position 748, to sweep the Tx FoV. In one embodiment, the controller 430 can move the transmit optics 438 through a total range of motion of 10 mm or at 0.1 m / s in 100 ms.
[0070] In step 608, controller 430 analyzes the sweep data to determine the location of unknown object 710. If controller 430 determines the location of unknown object 710, it proceeds to step 610. If controller 430 cannot determine the location of unknown object 710, it returns to step 604.
[0071] After determining the location of the unknown object 710 in step 610, the controller 430 controls the LIDAR sensor 400 to focus on the unknown object 710 and perform an additional scan or series of scans while moving the transmit optics 438 to track the unknown object 710. During this step, the transmit optics 438 moves to a position corresponding to the region of the field of view (FoV) in which the unknown object 710 is located.
[0072] In step 612, controller 430 analyzes the focused scan data to identify unknown object 710. If controller 430 is unable to identify unknown object 710, it returns to step 610. If controller 430 identifies unknown object 710, it proceeds to step 614, where it returns transmit optics 438 to its initial position, and then returns to step 602.
[0073] By focusing the Tx FoV on the unknown object 710, the lidar sensor 700 can project more light onto the region of interest within the overall Tx FoV and collect more reflected light, thereby quickly identifying the unknown object 710. This approach improves the responsiveness of the SDS 102 to identify and react to the unknown object 710 compared to other lidar systems that do not adjust the Tx FoV, such as the lidar sensor 600 shown in FIG. 6. The method of adjusting the Tx FoV can be implemented using one or more controllers, such as the controller 430 or the computer system 1300 shown in FIG. 29.
[0074] The lidar sensor 300 uses the results output by the photodetectors 318 to aggregate the photodetector results to generate a measured three-dimensional point. Because a single photodetector generally cannot generate a sufficient depth measurement, the results from multiple photodetectors are combined into a superpixel. One exemplary technique for generating a superpixel is shown in Figures 3 and 13.
[0075] In FIGS. 3 and 13, the photodetector array includes photodetectors arranged in a grid configuration. The results p1, p2, ..., pk obtained from the photodetectors can be represented by a grid 550 defined by a plurality of cells, where each cell 552 corresponds to a respective photodetector, where k is an integer equal to the total number of photodetectors in the array. The cells 552 of the grid 550 may be arranged in the same pattern as the photodetectors (e.g., a 256 x 256 grid pattern). Each result is also referred to as a pixel of the lidar image. The photodetector pixels p1, p2, ..., pk may be simply aggregated in supercells to generate a series of 3D points. The size of the supercell is W x W, where W is an integer. In FIG. 13, each supercell is 2 cells x 6 cells. The 3D points associated with each supercell 204 are derived by combining each of the six pixels to obtain superpixels SP1, SP2, ..., SPy. The first superpixel SP1 may be defined by the following equation:
[0076]
number
[0077] Other superpixels SP2,...,Spy may also be defined by similar formulas. The pixel aggregation scheme may vary depending on the design of the particular lidar sensor and may vary depending on the application. For example, a simple summation or convolution scheme may be used for pixel aggregation.
[0078] An illustration useful for understanding the convolutional approach is provided in FIG. 13. The convolutional approach can use one or more convolution filters (or kernels) 552 applied to a lidar image 550 to calculate features F1, F2, F3, F4, ..., F12. When multiple convolutional kernels are used, each kernel extracts a different feature in the lidar image. The kernel size is 2 x 6, the image size is 12 x 12, and the stride is 6. Thus, the features generated by the kernels 552 are defined by the following equations (1)-(4).
[0079]
number
[0080] F1, F2, F3, and F4 represent features generated by the processor (or computation kernel) 552. Such features are also called superpixels. p1, p2, ..., p 144 Each of the vectors represents a result output from a photodetector of a lidar system (e.g., the lidar sensor 300 of FIG. 3 ). Such a result is also referred to as a pixel of a lidar image. As shown in Equations (1)-(4), a convolutional oversampling operation generally involves combining pixel values to generate superpixels. The output of the processor (or operation kernel) 552 is arranged into a grid 554 of features (or superpixels) shown in FIG. 14E. The grid 554 is also referred to as a feature image. Features may include, but are not limited to, depth, intensity, noise, confidence, or other features associated with the point cloud. The solution of the present invention is not limited to the specific scheme of FIG. 14 , and other convolutional oversampling techniques may be used.
[0081] The lidar sensor 300 may include a Geiger-mode lidar system. In a Geiger-mode lidar system, a photodetector operates by detecting single photons near or from a target wavelength, and the detection operation times (and associated emitter emission times) are accumulated in a histogram. A peak-search operation is then performed on this histogram to obtain the depth of objects that reflected photons of the target wavelength in the spatial domain of the photodetector. That is, a series of individual laser emissions across the field of view (FoV) are acquired and then aggregated using a peak-searching algorithm to determine depth as a function of the time-of-flight (ToF) of the observed photons (or to include additional signals).
[0082] The lidar sensor 300 does not need to use the same number of laser emissions per pixel to find histogram peaks. In fact, the physical consequence of successively accumulating laser emissions is that the frustum defined by the rotating sensor is simply expanded along the relevant axis. That is, the probability that the return signals from two or more surfaces will merge into a superpixel increases. Also, the lidar sensor 300 is not prevented from integrating the return signals from adjacent photodetectors; this only adjusts the frustum imaged at a particular lidar point. In the case of a non-spaced sensor, each point actually represents the average depth of the nearest surface for each ray of the frustum. Combining this insight, the lidar sensor 300 is configured to use an aggregation window of variable size as a function of reflectivity or its proxy (e.g., number of returns, return signal noise, observations from other associated sensor systems such as cameras, etc.).
[0083] The lidar sensor 300 is configured to obtain integrated measurements from a single target or surface. This is done using a spatiotemporal coherence metric that exploits both the temporal and spatial diversity of the range measurements obtained from the lidar sensor 300. That is, multiple measurements are obtained across the spatial axis (integrating different pixels into superpixels) and the temporal axis (integrating different pulses). If measurements are obtained from the same surface, the statistical properties of the measurements should be similar, i.e., consistent in the temporal and spatial domains.
[0084] For example, a lidar system may be configured to combine pixel px into superpixel Spy using a photodetector and a processor employing a 2x6 arithmetic kernel. Table 1 below shows 12 time-of-flight (ToF) measurements or timestamps obtained for five pulses emitted from the lidar system. The example in Table 1 illustrates a situation where one superpixel combines the measurements of five pulses and 12 pixels. However, the target object may not be large enough to cover all of the given pulses. It should be understood that the lidar system rotates while transmitting pulses. As the lidar system rotates, the horizontal extension of the object may be sufficient to reflect across all five pulses.
[0085] [Table 1]
[0086] In Table 1, each cell contains the time-of-flight measurement or timestamp of a detection event reported by a superpixel associated with one of the five emitted pulses. As shown in FIG. 15, each superpixel reports one detection event 540, with a total of 12 detection events 540 for each pulse P1, P2, P3, P4, and P5. For pulse P1, pixel p1 has a timestamp of 684, pixel p2 has a timestamp of 559, pixel p3 has a timestamp of 629, pixel p4 has a timestamp of 192, pixel p5 has a timestamp of 835, pixel p6 has a timestamp of 763, pixel p7 has a timestamp of 707, pixel p8 has a timestamp of 359, pixel p9 has a timestamp of 9, pixel p10 has a timestamp of 723, pixel p11 has a timestamp of 277, and pixel p12 has a timestamp of 754. The scene actually imaged by pulse P1, as shown in FIG. 15, contains only noise (no target object is present in the superpixel's field of view). Similarly, each of the other pulses P2, P3, P4, and P5 also has a timestamp of 12, as shown in Table 1 and FIG. 15. The first three pulses P1, P2, P3 capture only noise (no target object in the field of view), and the last two pulses P4, P5 capture measurements of the target object.
[0087] In a fixed resolution system, a fixed number of pulses are integrated. However, if the target object is smaller than the integration time range, the overall detection quality may be degraded. In the example of Table 1 and Figure 15 above, if all measurements from five pulses P1, P2, P3, P4, and P5 are integrated to estimate the range of the target object, the measurements from the first three pulses P1, P2, and P3 do not contain range information for the target object, which negatively impacts quality.
[0088] The essence of the solution of the present invention is to provide an embodied system and method for selecting a subset of optical pulses to integrate the results of the photodetectors to perform distance estimation of a target object. Such selection is based on a spatiotemporal statistical coherence metric. The spatiotemporal statistical coherence can be measured using a distribution comparison metric, time-of-flight (ToF) temporal statistics, and / or solution differences or advantages.
[0089] Distribution Comparison Metrics An example of a distribution comparison metric is the Kullback-Leibler (KL) divergence. For each pulse, the system aggregates the pixel detection values within the superpixel using time bins with a size of 100 to obtain a new probability distribution (normalized histogram). In other words, 12 measurements are assigned to bins with a width of 100. As previously explained, the timestamps associated with pulse P1 in Table 1 are [684, 559, 629, 192, 835, 763, 707, 359, 9, 723, 277, 754]. The timestamps can be assigned to bins as follows: 684 is assigned to bin b600-700, and 559 is assigned to bin b500-600. The histogram H of the timestamps is determined by the following expression: there is one (9) detected value in bin b0-100, one (192) in bin b100-200, one (277) in bin b200-300, one (359) in bin b300-400, zero in bin b400-500, one (559) in bin b500-600, two (629,684) in bin b600-700, four (763,707,723,754) in bin b700-800, and one (835) in bin b800-900. Therefore, the histogram H is defined as follows: H = [1,1,1,1,0,1,2,4,l].
[0090] This histogram H may be converted into a probability distribution (or probability mass function, PMF) over the total number of detections (12), as shown in the following expression: PMF1 = [1 / 12,1 / 12,1 / 12,1 / 12,0 / 12,1 / 12,2 / 12,4 / 12,1 / 12] = [0.0833,0.0833,0.0833,0.0833,0.0,0.0833,0.1667,0.3333,0.0833].
[0091] Performing the same procedure (bin assignment, histogram generation and transformation) we obtain the following additional PMFs for pulses P2, P3, P4 and P5: PMF2:[0.1667,0.0833,0.0,0.1667,0.1667,0.1667,0.0833,0.0833,0.0833] PMF3:[0.1,0.1,0.0,0.0,0.0,0.1,0.2,0.3,0.2] PMF4:[.3333,0.6667,0.0,0.0,0.0,0.0,0.0,0.0,0.0] PMF5:[0.1667,0.8333,0.0,0.0,0.0,0.0,0.0,0.0,0.0]
[0092] The KL divergence metric is then used to detect distribution changes. The KL metric generally provides a method for comparison between distributions. Differences in the distributions may be exploited as a signal of a distribution change, which indicates the onset of another integration range. The KL divergence metric is defined by the following formula:
[0093]
number
[0094] where Q is the current distribution and P is the additional distribution. The KL divergence can be calculated as follows: KL[P=PMF2,Q=PMF1] = 0.18 KL[P=PMF3,Q=PMF2] = 0.65 KL[P=PMF4,Q=PMF3] = 1.67 KL[P=PMF5,Q=PMF4] = 0.07
[0095] The KL divergence KL[P=PMF4,Q=PMF3] is larger than the other KL divergence values KL[P=PMF2,Q=PMF1], KL[P=PMF3,Q=PMF2], and KL[P=PMF5,Q=PMF4]. This indicates that there is a change in the detection distribution between the two pulses P3 and P4. Consequently, pulse integration is suspended until pulse P3, and integration begins (or resumes) from pulse P4.
[0096] Time of Flight (Tof) Temporal Statistics: The ToF temporal statistics may include the ToF time variance, etc. The ToF time variance can be obtained by calculating the variance of the aligned ToF difference values. First, the ToF measurements or timestamps for each superpixel are sorted from smallest to largest according to the following expression: P1:[9,192,277,359,559,629,684,707,723,754,763,835] P2:[70,87,174,314,396,472,486,551,599,600,705,804] P3:[72,115,537,600,677,709,755,777,845,849,916,976] P4:[7,99,99,99,100,100,100,100,101,101,101,102] P5:[98,99,100,100,100,100,100,100,101,101,101,102]
[0097] Next, calculate the set of differential timestamp values D for each pulse. The set of differential timestamp values Dp1 for the first pulse P1 can be found as follows, where each difference value is the difference between two consecutive timestamps: Dp1:[192-9,277-192,359-277,559-359,629-559,684-629,707-684,723-707,754-723,763-754,835-763] = [183,85,82,200,70,55,23,16,31,9,72]
[0098] Similar calculations are performed on the remaining pulses P2, P3, P4, and P5 to obtain the following sets of differential timestamp values Dp2, Dp3, Dp4, and Dp5: Dp2 = [17,87,140,82,76,14,65,48,1,105,99] Dp3 = [43,422,63,77,32,46,22,68,4,67,60] Dp4 = [2,0,0,1,0,0,0,1,0,0,l] Dp5 = [1,1,0,0,0,0,0,1,0,0,l]
[0099] Next, the system measures the variance for each set of difference values Dp1, Dp2, Dp3, Dp4, and Dp5. Each variance is calculated by finding the mean, then summing the squared differences from the mean, divided by the size of the set. For example, the variance VDP1 of set Dp1 can be calculated as follows: (i) calculate the mean of the set of distinct timestamp values; (ii) subtract the mean from each difference timestamp value in the set; (iii) calculate the square of the value obtained in step (ii); (iv) add all the values obtained in (iii); and (v) divide the value obtained in (iv) by the total number of difference timestamp values in the set. For example, for set Dp1, p1 Variance V DP1 is determined as follows:
[0100] (i)Mean = (183+85+82+200+70+55+23+16+31+9+72) / 11=75.09 (ii)(183-75.09=107.91),(85-75.09-9.91),(82-75.09=6.91),(200-75.09=124.91),(70-75.09=-5.09),(55-75. 09=-20.09),(23-75.09=-52.09),(16-75.09=-59.09),(31-75.09=-44.09),(9-75.09=-66.09),(72-75.09=-3.09) (iii)(107.91 2 =11,644.5681),(9.91 2 =98.2081),(6.912 =47.7481),124.91 2 =15602.5081),(-5.09 2 =25.9081),(-20.09 2 =403.6081),(-52.09 2 =2713.3681),(-59.09 2 =3491.6281),(-44.09 2 =1943.9281),(-66.09 2 =4367.8881),(-3.09 2 =9.5481) (iv)11644.5681+98.2081+47.7481+15602.5081+25.9081+403.6081+2713.3681+3491.6281+1943.9281+4367.8881+9.5481=40,348.9091 (v) 40,348.9091 / 11 = 3668.08
[0101] The next variances for the other sets Dp2, Dp3, Dp4, Dp5 can be calculated in a similar manner. VDp2 = 1684.74 VDp3 = 11990.15 VDp4 = 0.43 VDp5 = 0.23
[0102] As the variance values reveal, the temporal variance associated with pulses (P1-P3) is much larger than that of pulses (P4, P5), which may be considered an indication that pulses (P4, P5) are integrating measurements of other target objects with those captured by pulses (P1-P3).
[0103] The second method (temporal statistics) is important because it does not require the generation of histograms as in the first method (KL divergence), making it even less computationally expensive (though even KL divergence can use very coarse histograms). However, the second method cannot distinguish between two target measurements taken at two different distances that show similar temporal variance (measurements from both distances are all concentrated locally). In this case, a combined measurement of the mean and variance can be considered.
[0104] Instead of using a criterion to stop integrating measurements, it is also possible to use spatiotemporal coherence to remove noisy measurements. For example, the third pulse P3 of 10 pulses may have a statistically inconsistent measurement due to interference. Instead of stopping integration at two pulses, the system can remove the inconsistent measurement from the third pulse and then continue integrating.
[0105] In addition to the above-mentioned stopping criteria, various other "stopping criteria" can be used. Some examples of new stopping criteria are: accumulating a fixed number of laser shots until the confidence of the detected peak (or other statistical detection probability metric) exceeds a critical value; accumulating a fixed number of laser shots until the confidence of the detected peak begins to decrease by integrating additional pulses; accumulating laser shots in a sliding window until one of the above two criteria is met; and / or accumulating a fixed number of laser shots based on the size of the aggregation window and the reflectivity estimates of neighboring points, causing the lidar sensor to emit a quadtree on the 2D range image plane.
[0106] It is not necessary to use the same aggregation function or values for each other axis. Thus, the system can emit points corresponding to any square or non-square solid angle of the entire sensor frustum, as the design chooses. In particular, sliding-window aggregation schemes do not need to correspond exactly to "image plane pixels," since they can operate at the much smaller unit of individual laser firings. Representing observations as points in a range imaging system is a significant simplification; a better first-order approximation can be oriented planes or surface elements.
[0107] There is also no mandatory requirement to use reflectivity or similar metrics to guide the aggregation. Instead, other characteristics of intermediate measurements can be used. For example, recovered depth itself can be used to guide ongoing aggregation, which means transferring the function of reducing data on surfaces close to the vehicle to the lidar storage device itself. This implementation is an advantageous solution, as it helps maintain a consistent point density in world space (range sensors typically suffer from "data overload" in areas close to the sensor in order to obtain sufficient point density at longer distances).
[0108] The solution of the present invention allows for the design of software-efficient trade-offs between range imaging system characteristics such as reflectivity performance, detection probability, detection range, and resolution. While in existing systems, such axes are typically fixed at design time and determined as a function of hardware, the present invention moves such decisions into the software domain. Furthermore, use of the present invention as part of a lidar system allows for the generation of constant density point clouds for any volume in world space (a highly unique approach), or variable density point clouds or surface clouds as a function of other imaging characteristics, such as estimated surface reflectivity.
[0109] The spatiotemporal coherence metric has two particular advantages: it provides robustness against noise and interference; and it provides a robust integration stopping criterion to support variable resolution structures.
[0110] Autonomous vehicles (AVs) use sensors for situational awareness. Sensors that are part of an AV's autonomous driving system (SDS) may include cameras, lidar devices, inertial measurement units (IMUs), etc. These sensors are used to capture and analyze the scene around the AV. Scene analysis is performed to detect objects, including stationary objects (e.g., fixed structures) and dynamic objects (e.g., pedestrians and other vehicles). Data from the sensors may also be used to detect conditions such as road markings, lane turns, traffic lights, and traffic signs. Sometimes, camera footage can be combined with a point cloud obtained from the AV's lidar device to gain additional insight into the scene and situation around the AV.
[0111] A lidar device operating on an autonomous vehicle (AV) includes a transceiver device, which may include a transmitting assembly and a receiving assembly. The transmitter transmits optical signals, and the receiver can receive and process the received optical signals.
[0112] In some existing implementations, LIDAR devices use a fixed pixel size (angular resolution) to integrate a fixed amount of raw data per point. However, in some applications, it is preferable to use a more intelligent data integration method tailored to the characteristics of the target to improve detection probability and data quality (accuracy and precision of distance and intensity). As shown in Figure 19, target objects (e.g., vehicles) are larger than the basic pixel size, but using a larger pixel size risks merging data in areas where distance and intensity vary greatly, which can result in blurred or distorted data.
[0113] First, the choice of input data can be considered, e.g., one must decide what data stage to use for spatial processing.
[0114] Option 1: Pristine data: Robust, computationally expensive, requires firmware modifications.
[0115] Option 2: Waveform data: Robust and requires firmware modifications.
[0116] Option 3: Lidar data: Computationally less expensive and can be implemented in firmware or software as a post-processing step.
[0117] Lidar data can be a good choice because it has a relatively low initial cost and can be implemented in both firmware and software.
[0118] According to some aspects, the following spatial processing methods can be implemented for each of the above options. For Options 1 and 2 presented above, there are two approaches used for data processing. The first approach for Option 1 associated with the original data is as follows: The original data is completely processed using fixed superpixel sizes to obtain the distance and signal intensity of each superpixel, and the original data is then reprocessed using variable superpixel sizes based on the correlation of distance and intensity with neighboring pixels. The second approach is as follows: The total intensity of each pixel is calculated using fixed superpixel sizes, and the original data is processed using variable superpixel sizes based on the correlation of the total intensity with neighboring pixels. The signal intensity has two independent fields (signal intensity and noise intensity) and one dependent field (total intensity = signal + noise).
[0119] A first approach to Option 2 related to waveform data is as follows: fully process the waveform to obtain the distance and signal strength of each superpixel using a fixed superpixel size, then recreate and process the waveform with a variable superpixel size based on the distance and strength correlation with neighboring pixels. A second approach is as follows: fully process the histogram with a fixed superpixel size to obtain approximate distance and signal strength (before waveform analysis), then integrate the waveform using a variable superpixel size based on the distance and strength correlation with neighboring pixels using the approximate distance and strength.
[0120] For option 3 related to lidar data, the present invention embodies a new approach that may generally involve integrating neighboring pixels in a lidar data frame based on range and intensity correlations; and recalculating new range, signal strength, noise strength, and confidence based on the data fields reported in each superpixel.
[0121] More specifically, the new approach may involve arranging pixels in a grid (where the pixels contain waveform processing results generated by a photodetector in a lidar system) and identifying a region of interest (ROI) in the grid based on correlations between the pixels. Such correlations may include, but are not limited to, correlations between distance values and / or correlations between intensity values associated with the pixels. In some situations, the ROI can be identified by obtaining a kernel size and using the kernel size to define the region of interest in the grid. This kernel size may be variable. The kernel size may be obtained by finding the pixel closest to the point of interest (POI) in the grid in terms of distance and defining the kernel size based on the location of this closest pixel. Alternatively, a reference kernel size may be obtained, and the reference kernel size may be used to identify a region in the grid, identify a center pixel of the region, calculate a score for each pixel in the region indicating the degree of correlation between the result value of that pixel and the result value of the center pixel, select the pixel based on the score, and then define the kernel size based on the location of the selected pixel. The score may be a function of distance, intensity, and / or noise.
[0122] The size and / or location of a region of interest (ROI) within the grid may be adjusted to maximize the likelihood that the ROI includes more pixels associated with the object. Such adjustment can be achieved by identifying a point of interest (POI) within the ROI, identifying the pixel closest to the POI in terms of minimum distance, using the centroid of the closest pixel to obtain the likelihood that the POI is associated with a corner or edge point on the object's surface, and adjusting the size and / or location of the ROI based on this likelihood. The POI may be the center pixel of the ROI.
[0123] One or more pixels within the ROI may be selectively excluded from aggregation with other pixels within the ROI. Such exclusion may be based on how far the pixel is from the POI and / or surface in one or more dimensions. Such dimensions may include, but are not limited to, distance, intensity, noise, and confidence. The resulting values associated with the remaining pixels (i.e., selected pixels) located within the ROI are combined together to generate feature values. Superpixels populated with the feature values are generated.
[0124] The operations of the novel approach described above may be performed iteratively to generate additional superpixels. The size and / or shape of the ROI used in a first iteration to generate a first superpixel may be the same as or different from the size and / or shape of the ROI used in another iteration to generate a second superpixel. Superpixels may be used in some scenarios to control the operation of autonomous vehicles or other robotic devices (e.g., articulated arms or electrosurgical tools).
[0125] As shown in FIGS. 2 and 13, the lidar sensor 300 uses the output waveforms of the photodetectors 226 to generate results p1, p2, ..., px. Each result p1, p2, ..., px has a value associated with it. The values may include, but are not limited to, a distance value, an intensity value, a noise value, a confidence value, and / or a trial value. The results may be aggregated to generate superpixels. One exemplary simple technique for generating superpixels will be discussed below in connection with FIG. 13.
[0126] In a Geiger-mode lidar system, the sensor includes an avalanche detector (or photodiode) configured to generate an electrical pulse of a specific amplitude in response to photon absorption of the same or similar wavelength as the emitted optical signal. A histogram may then be constructed through many measurements, and the location of the object surface may be estimated from the peak of the histogram. As used herein, the term "test" refers to each measurement attempt. A measurement attempt involves transmitting a pulse and recording the detection time. A measurement is associated with, but not necessarily coincides with, a pulse. Detection results from various detectors can also be aggregated to obtain multiple measurements from a single pulse. The output of each detector is a single measurement. However, the accuracy of the histogram is primarily limited by the width of the bin. Therefore, the present invention relates to post-processing operations that update range, intensity, noise, and / or confidence values. This post-processing operation will be discussed in detail in connection with FIG. 20.
[0127] 13 and 21, the results p1, p2, ..., px can be represented as a grid 550 defined by a plurality of cells 552. Each result is also referred to as a pixel of the lidar image. The pixels p1, p2, ..., px can be simply aggregated in supercell units, thereby generating a series of 3D points. The size of a supercell is Q x Z, where Q and Z are integers. In FIG. 13, each supercell is 2 cells x 6 cells. The 3D points associated with each supercell 554 are derived by combining the six pixels to obtain superpixels SP1, SP2, ..., SPy. The first superpixel SP1 can be defined by the following equation (6):
[0128]
number
[0129] It should be understood that each of the other superpixels SP2,...,Spy can also be defined by a similar formula. The pixel counting scheme is specific to the design of a particular lidar system and may vary depending on the application. For example, simple counting may be used for pixel counting.
[0130] The approach of Figure 13 has certain drawbacks. Because this technique uses a fixed pixel size (angular resolution) and a fixed number of pixels per superpixel, it runs the risk of combining pixels from areas where distance and intensity vary greatly, which can cause distortions such as blurring in LIDAR images composed of superpixels. The solution of the present invention overcomes these drawbacks by implementing a new approach to generating superpixels. This new approach will be discussed in conjunction with Figure 20.
[0131] 20 and 21 provide a flow diagram of an example method 650 for generating and / or using superpixels. Method 650 can be performed in whole or in part by a processor (e.g., processor 222 of FIG. 3) of a lidar system (e.g., lidar sensor 200 of FIG. 3).
[0132] Method 650 begins at step 652 and continues to step 654, where a light pulse is transmitted from a lidar system. The transmitted light may reflect off an object and return to the lidar system. The reflected light may be received by a photodetector at step 656. The waveform output from the photodetector is then processed at step 658 to generate results. Each result has a value associated with it. Such values may include, but are not limited to, a distance value, an intensity value, a noise value, a confidence value, a calculated value, and a trial value. At step 660, the results are arranged into a grid of cells. The example grid 500 shown in FIG. 21 a shows results p1, p2, ..., p144 each assigned to its cell 502.
[0133] This grid is used to generate superpixels according to the solution of the present invention. This new approach can use at least one filter (or kernel) that computes features as it passes through the grid. If multiple computing kernels are used, each kernel extracts a different feature from the grid.
[0134] Thus, in step 662 of FIG. 20, the processor obtains a kernel size and a stride. The kernel size may be a predefined fixed value or a variable value. The kernel size may be defined as Q cells by Z cells. Here, Q and Z are integers and may be the same or different. When the intrinsic angular resolutions of the azimuth angle and the elevation angle are the same, Q and Z are selected to have the same value. For example, as shown in FIG. 21A, they may be selected as 3x3 cells (3x3), or as 5x5 cells or 7x7 cells (not shown). The solution of the present invention is not limited to these exemplary details. Even values may be used for Q and Z. However, in some applications, odd values may be desirable so that the central pixel can be used as the POI. A larger kernel size generates better features but increases computational costs. The stride S may be a predefined fixed value or a variable value, where S is an integer (e.g., 1 or 6).
[0135] Instead of restricting the search area with a fixed QxZ kernel, it may be more robust to find the nearest pixel. The nearest pixel can be identified in Cartesian space and / or pixel space. Pixel space is similar to spherical coordinates (range, azimuth, and elevation). The nearest pixel can be found by minimizing the angle between the POI and the neighboring pixel, including points within a distance tolerance. Therefore, in variable kernel size and / or stride scenarios, the kernel size and / or stride may be dynamically determined or derived based on the nearest pixel approach. The nearest pixel approach method, for example, obtains a reference kernel size, uses this kernel size to identify points of interest (POI) (where the POI is the central pixel of the area defined by the kernel size), calculates a score (A) for each pixel using the associated values of each pixel (e.g., results p1, p2, ..., p144), selects a certain number (e.g., 12) of nearest pixels of the POI based on the scores, and can define a new kernel size based on the selected nearest pixels and / or select a stride based on the score and / or the new kernel size. The score (A) indicates the degree of relationship between pixels. Each score (A) may be defined as the following equation (2) or (3):
[0136]
number
[0137] where R is distance, I is intensity, N is noise, C is confidence, T is test value, K is the number of calculations, and w1, ..., w6 are weights. The solution of the present invention is not limited to formulas (2) and (3). The score (A) may be a function of any combination of one or more of the values listed above (e.g., R, I, N, C, and / or T). The solution of the present invention is not limited to a specific nearest pixel approach, and other nearest pixel approaches may also be used. In the resulting kernel size, Q and Z may be the same or different from each other.
[0138] Once the kernel size and stride have been determined, method 650 proceeds to step 664 to identify an ROI in the grid. An exemplary ROI 804 is shown in FIG. 21a. The ROI 804 is comprised of cells located within the grid in an area bounded by a kernel-sized search window. While the ROI is shown as having equal length and width and a square shape, the present solution is not limited thereto. The ROI may also have shapes with different lengths and widths, such as a linear shape represented by the dotted line 850 in FIG. 21a.
[0139] Next, in step 666, a point of interest (POI) is identified within the ROI. A POI includes, but is not limited to, the center pixel of the ROI. For example, as shown in FIG. 21B, POI 806 is composed of pixel p14, which is the center pixel of ROI 804. However, the solution of the present invention is not limited to this. For example, if the reliability of the center pixel exceeds a reliability threshold, this indicates sufficient certainty or accuracy of the reported data field, and selecting another POI is pointless. However, if the reliability of the center pixel is lower than the reliability threshold, the reported distance is likely to be inaccurate, which could cause the spatial processing to fail completely. In this case, in certain situations, the distance value of the center pixel POI can be replaced with a representative distance value from the ROI. This approach involves first finding the nearest neighbor in pixel space whose reliability is greater than the reliability threshold and replacing the distance of the center pixel POI with this value. Other POI fields are not modified. Second, the system performs spatial processing as usual, tracking the quadrant location of the relevant pixels contained within the ROI, and only retains the final result if the correlated pixel is present in at least N quadrants of the ROI (most conservatively, all four). Otherwise, it ignores the point and moves on to the next stride pixel.
[0140] The size and / or location of the ROI may be selectively adjusted, as shown in steps 668-670. By moving the location of the ROI within the grid, the system can maximize the likelihood that the kernel will include more pixels that belong to the same target object as the POI (e.g., pixel p14 in FIG. 21a) or that are well-correlated with the POI.
[0141] The center of the pixel closest to the POI can indicate where the POI is located on the object's surface. If the center is offset to the top, bottom, left, or right of the POI, the POI is likely to be an edge point of the object. Conversely, if the center is offset to a corner of the POI or ROI, the POI is likely to be a corner point. Such information may be used to adjust the kernel size in one or more directions (e.g., expand / increase or contract / decrease). For example, if the POI is determined to be an edge point of the object, the kernel size and / or the ROI position within the grid are modified so that the POI is located at the edge, not at the center, of the ROI. If the POI is determined to be a corner point at the bottom left corner of the target object, the kernel size and / or the ROI position within the grid are modified so that the POI is located at the bottom left corner, not at the center, of the ROI. Similarly, if the POI is determined to be the upper left corner point of the target object, the kernel size and / or ROI position are changed so that the POI is located at the upper left corner of the ROI. The solution of the present invention is not limited thereto.
[0142] The kernel size may be adjusted to expand the ROI in both the Q and Z directions. For example, as shown in Figures 21A and 21C, the kernel size is expanded from 3 cells x 3 cells (3x3P) to 4 cells x 4 cells (4x4). As a result, the ROI 804 is expanded in both the Q and Z directions to form the ROI 804'. The solution of the present invention is not limited thereto.
[0143] Additionally or alternatively, the kernel size may be adjusted only in the Q direction, as shown in FIG. 22, or only in the Z direction, as shown in FIG. 23. Additionally or alternatively, the position of the ROI within the grid may also be changed. For example, as shown in FIG. 24, a 3×3 ROI is moved from a first position 850 (i.e., a position shifted one cell in each of the Q and Z directions) to a second position 852. The solution of the present invention is not limited to this. The position of the ROI may be shifted in one direction or both directions by a number of cells selected depending on the given application.
[0144] Referring again to FIG. 20 , the method continues from step 650 to step 672, where a distance tolerance, intensity tolerance, noise tolerance, and / or confidence threshold is obtained from the data store. These values may be preconfigured. In step 674, one or more pixels may be excluded from counting in the ROI according to the distance tolerance, intensity tolerance, noise tolerance, and / or confidence threshold. The tolerances are used to exclude points that are too far away from the POI in one or more dimensions. For example, if the distance value of the POI is 10 and the distance tolerance is ±1, the system determines whether other pixels in the ROI have distance values between 9 and 11. If so, the pixel is included in the count. If not, the pixel is excluded from the count. The solution of the present invention is not limited to this example. If the LIDAR system reports multiple return values per pixel, all return values for all pixels in the ROI must be checked for (in)consistency.
[0145] In some scenarios, a pixel may be considered a matching pixel if one of the following is true: (i) its associated distance, intensity, and / or noise value is within the tolerance; (ii) its associated confidence value is equal to or greater than the confidence threshold. A pixel may be considered a non-matching pixel if one of the following is true: (i) its associated distance, intensity, and / or noise value is outside the tolerance; or (ii) its associated confidence value is lower than the confidence threshold. For example, as shown in FIG. 21D, pixels p3, p37, and p40 of ROI 804' are considered non-matching pixels 808 for aggregation. The solution of the present invention is not limited thereto.
[0146] In other scenarios, surface normals may be estimated and distance tolerances applied based on the surface rather than the POI. For example, if a pixel faces the road, the surface normal will point upward. In other scenarios, surface normals may be estimated and distance tolerances applied based on the surface rather than the POI. For example, if a pixel faces the road, the surface normal will point upward. In this case, pixels within the integration window may be included if their distance values are sufficiently close to the road surface, rather than if they are close to the POI. Also, or alternatively, neighboring pixels may be included in the set of matching pixels if their signal strength confidence intervals overlap. Although Geiger-mode lidar intensities are noisy, confidence intervals can be calculated using binomial statistics based on the number of calculations and tests in the interval containing the returned signal. Alternatively or additionally, neighboring pixels may be included in the set of matching pixels if their noise strength confidence intervals overlap. In this case, the noise intensity confidence interval is a function of the number of noise calculations and the number of noise tests, which can be derived from the total number of counts and tests minus the number of counts and tests in the interval that contains the returned signal.
[0147] Upon completion of step 674, the method continues to step 676, where the remaining pixels in the ROI 804' are combined by a kernel to generate superpixels. For example, as shown in FIG. 21E, kernel 810 applies a function to the remaining pixels p1, p2, p4, p13, p14, p15, p16, p25, p26, p27, p28, p38, and p39 to calculate feature F1. Feature F1 may be defined by the following equation (4):
[0148]
number
[0149] Feature F1 is considered a superpixel (i.e., SP1 = F1). The features (or superpixels) may include, but are not limited to, distance, intensity, noise, and / or confidence. The manner in which pixels are counted will vary depending on the design of the particular lidar system and may vary depending on the application. For example, a simple sum or average value may be used for pixel counting. In this regard, the distance value of feature F1 may be, but is not limited to, the average distance of the remaining pixels within the ROI. The intensity value of feature F1 may be an intensity value derived from the sum of the signal counts and test numbers of the remaining pixels within the ROI, or the average intensity of the remaining matching pixels (if counts and test values are not provided). The noise value of feature F1 may be a noise value derived from the sum of the noise counts and test numbers of the remaining pixels within the ROI, or the average noise of the remaining pixels. The confidence value of feature F1 may be, but is not limited to, a confidence value derived from the updated noise value and the sum of the signal counts and test numbers within the kernel.
[0150] The process of steps 662 to 676 is repeated to generate another superpixel by the stride. For example, the next ROI is identified by moving the kernel search window by the stride, and the next superpixel is generated according to the process described above. For example, as shown in FIG. 21F, if the stride is 4, the kernel search window is moved four cells to the right. As a result, the next superpixel is set to feature F2 defined by the following mathematical formula (5).
[0151]
number
[0152] Other features F3,...,F12 are generated in a similar manner. These features define the feature map 512 shown in FIG. 21G. Other superpixels are established corresponding to these features (e.g., SP2 = F2, SP3 = F3,..., SP12 = F12). These superpixels can be used to control vehicle operation or deploy personnel to the scene, as shown in optional step 680 of FIG. 20. Thereafter, in step 682, method step 650 ends or other operations are performed.
[0153] The method step 650 described above provides various advantages over existing systems and methods. For example, the embodied system and method step 650 provides: (i) improved distance accuracy and precision; (ii) improved distance and intensity accuracy; (iii) improved detectability and quality of painted lines, such as road lane markings; (iv) removal of high-intensity artifacts and increased effective dynamic range; (v) improved probability of detection of black or dark objects; and (vi) increased probability of detection of targets at long distances. In relation to item (v), the resulting (or pixel) data may be contaminated (e.g., may contain erroneous distance values and / or relatively low confidence values below a critical value), which may hinder the detection of dark objects. In scenarios where the confidence value is low, the solution of the present invention can increase the confidence value above the critical value, thereby enabling dark objects to be detected with greater confidence.
[0154] Method step 650 can be implemented in the form of intelligent oversampling in the detection and waveform analysis steps at the Geiger-mode avalanche photodiode (GmAPD) data level. Method step 650 can also be implemented in the signal detection step by using a double-pass process in which data from a first detection step is fed back to a second detection step; or by calculating a prior based on the original data prior to histogram generation to determine which pixels to combine into a single histogram. In the latter case, a total flux can be calculated for each GmAPD pixel, which is a value that combines distance, signal strength, and noise.
[0155] 25, a flow diagram of yet another method 900 for operating a lidar system is provided. Method 900 can be performed in whole or in part by a processor (e.g., processor 332 of FIG. 3) of the lidar system (e.g., lidar sensor 300 of FIG. 3).
[0156] Method 900 begins at step 902, and at step 904, pixels (e.g., pixels p1,..., p144 in FIG. 21) are arranged in a grid (e.g., grid 800 in FIG. 21). The pixels include values generated by processing a waveform generated from a photodetector of a LIDAR system (e.g., photodetector 318 in FIG. 3). At step 906, a processor performs an operation of identifying an ROI in the grid based on correlations between the pixels, including, but not limited to, correlations between distance values and / or correlations between intensity values associated with the pixels.
[0157] In some scenarios, the ROI can be identified as follows: obtain a kernel size; use this kernel size to define a region of interest within a grid. The kernel size may be variable. The kernel size may be obtained by the following process: find the closest pixel on the side of the minimum distance to the POI within the grid; define the kernel size based on the location of the closest pixel. Alternatively, the kernel size can be obtained as follows: obtain a reference kernel size; use this kernel size to identify a region within the grid; identify a center pixel of the region; calculate a score (e.g., score A mentioned above) using the associated result value for each pixel within the region; this score indicates the degree of correlation of the result value between the pixel and the center pixel; select a pixel based on this score; and define the kernel size based on the location of the selected pixel. The score may be a function of distance, intensity, and / or noise.
[0158] In step 908, the size and / or location of the ROI within the grid is selectively adjusted to maximize the likelihood of including more pixels associated with the object. Such adjustment may be performed in the following manner: identify a POI within the ROI; identify the nearest pixels of the POI on the side of minimum distance; use the centers of these nearest pixels to calculate the likelihood that the POI is associated with an edge point or corner point of the object surface; and adjust the size and / or location of the ROI according to that likelihood. The POI may be the center pixel of the ROI.
[0159] At step 910, one or more pixels within the ROI may be selectively excluded from aggregation with other pixels, where the exclusion is determined based on how far the particular pixel is from the POI and / or road surface in one or more dimensions, including but not limited to distance, intensity, noise, and confidence.
[0160] In step 912, the processor combines the result values associated with pixels located within the ROI to generate a feature value (e.g., feature value F1 of FIG. 21g). In step 914, a superpixel is generated having this feature value as its value. The operations of steps 906-914 can be repeated to generate other superpixels, as shown in step 916 of FIG. 21g. The ROI used to generate the superpixel in the first iteration may be different in size and / or shape from the ROI used to generate other superpixels in the second iteration. Step 918 is then performed to end method 900, or other operations may be performed, such as returning to step 902.
[0161] The lidar systems described above can be used in a variety of applications. While the solutions of the present invention are described in the context of autonomous vehicles, they are not limited to autonomous vehicle applications. The solutions of the present invention can also be used in robotic applications (e.g., motion control of articulated robotic arms) and / or system performance enhancement applications.
[0162] FIG. 26 illustrates an exemplary system 1000 according to one aspect of the present invention. The system 1000 includes a vehicle 1002 that travels semi-autonomously or fully autonomously along a roadway. The vehicle 1002 is also referred to as an AV 1002 within this document. An AV 1002 includes, but is not limited to, a ground vehicle (as shown in FIG. 26), an aircraft, or a water vehicle. As noted above, unless expressly stated otherwise, the disclosure herein is not necessarily limited to embodiments of autonomous vehicles, and in some embodiments may also include semi-autonomous vehicles.
[0163] AV 1002 is generally configured to sense objects in its vicinity, including but not limited to vehicles 1003, bicyclists 1014 (e.g., riders of bicycles, electric scooters, motorcycles, or the like), pedestrians 1016, etc.
[0164] As shown in Figure 26, an autonomous vehicle (AV) 1002 may include a sensor system 1018, an in-vehicle computing device 1022, a communication interface 1020, and a user interface 1024. The autonomous vehicle system may also include certain components contained within the vehicle, as shown in Figure 27, which may be controlled by the in-vehicle computing device 1022 using various communication signals and / or commands, such as acceleration signals or commands, deceleration signals or commands, steering signals or commands, braking signals or commands, etc.
[0165] The sensor system 1018 may include one or more sensors coupled to or included with the AV 1002. For example, such sensors may include, without limitation, a LiDAR system, a radar system, a laser detection and ranging (LADAR) system, a sonar (sound navigation and ranging) system, one or more cameras (e.g., visible light spectrum cameras, infrared cameras, etc.), temperature sensors, position sensors (e.g., global positioning systems (GPS)), position sensing sensors, fuel sensors, motion sensors (e.g., inertial measurement units (IMUs)), humidity sensors, occupancy sensors, or similar sensors. Sensor data may include information describing the location of objects in the environment surrounding the AV 1002, information about the environment itself, information about the movement of the AV 1002, information about the path of the vehicle, or similar information. As the AV 1002 moves along the ground, at least some of the sensors may collect data about the ground.
[0166] The AV 1002 can also transmit sensor data collected by the sensor system to a remote computing device 1010 (e.g., a cloud processing system) via a communications network 1008. The remote computing device 1010 can be configured with one or more servers to perform one or more processes of the techniques described herein. The remote computing device 1010 can also be configured to send and receive data / instructions to and from the AV 1002 via the network 1008, as well as to and from a server and / or a data store 1012. The data store 1012 can include, but is not limited to, a database.
[0167] Network 1008 may include one or more wired or wireless networks. For example, network 1008 may include a cellular network (e.g., a Long Term Evolution (LTE) network, a Code Division Multiple Access (CDMA) network, a 3G network, a 4G network, a 5G network, or other next generation networks). The network may also include a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic based network, a cloud computing network, or similar networks and / or combinations of such networks.
[0168] The AV 1002 can search, receive, display, and edit information generated from local applications or transmitted over the network 1008 from a data store 1012. The data store 1012 may be configured to store and provide raw data, indexed data, structured data, roadmap data 1060, program instructions, or other configuration information.
[0169] The communication interface 1020 may be configured to enable communication between the AV 1002 and external systems (e.g., external devices, sensors, other vehicles, servers, data stores, databases, etc.). The communication interface 1020 may use any currently or future known protocol, security scheme, encoding, format, packaging, etc., such as, without limitation, Wi-Fi, infrared links, Bluetooth, etc. The user interface system 1024 may include a keyboard, a touchscreen display device, a microphone, a speaker, etc., as part of the peripheral devices embodied in the AV 1002. The vehicle may also receive status, description, or other information about devices or objects in its surrounding environment via communication links (e.g., vehicle-to-vehicle (V2V), vehicle-to-object (V2O), or other vehicle-to-external (V2X) communication links) via the communication interface 1020. The term "V2X" refers to communication with certain objects that a vehicle may encounter or affect in its surrounding environment.
[0170] FIG. 27 illustrates an exemplary system architecture 1100 for a vehicle in accordance with the present teachings. Vehicles 1002 and / or 1003 of FIG. 26 may have the same or similar system architecture as that shown in FIG. 27. Accordingly, the following description of system architecture 1100 is sufficient to understand vehicles 1002 and 1003 of FIG. 26. However, other types of vehicles are within the scope of the technology described herein, which may include more or fewer components than those described in connection with FIG. 27. As a non-limiting example, an air vehicle may not include brake or gear controls, but may include an altitude sensor. As another non-limiting example, a water vehicle may include a depth sensor. Those of ordinary skill in the art will understand that various propulsion systems, sensors, and controllers may be included depending on the type of vehicle.
[0171] 27, a vehicle system architecture 1100 includes an engine or motor 1102 and various sensors 1104-1118 for measuring various vehicle parameters. In the case of a gasoline or hybrid vehicle having a fuel-based engine, the sensors may include, for example, an engine temperature sensor 1104, a battery voltage sensor 1106, an engine rotational speed (RPM) sensor 1108, and a throttle position sensor 1110. If the vehicle is an electric or hybrid vehicle, the vehicle may have an electric motor and therefore may include sensors such as a battery monitoring system 1112 for measuring battery current, voltage, and / or temperature, motor current 1114 and voltage 1116 sensors, and a motor position sensor 1118, such as a resolver or encoder.
[0172] Operating parameter sensors common to both types of vehicles include, for example: position sensors 1136 such as accelerometers, gyroscopes, and / or inertial measurement units, speed sensors 1138, and odometer sensors 1140. The vehicle may also include a clock 1142 that the system uses to determine the vehicle's time during operation. The clock 1142 may be built into the in-vehicle computing device or may be a separate device, and there may be multiple clocks.
[0173] The vehicle may also include a variety of sensors that operate to gather information about the environment in which the vehicle is traveling. Such sensors may include, for example, a location sensor 1160, such as a global positioning system (GPS) device, object detection sensors, such as one or more cameras 1162, a lidar system 1164, radar, and / or sonar system 1166. The sensors may also include environmental sensors 1168, such as a precipitation sensor and / or an outside air temperature sensor. The object detection sensors enable the vehicle to detect objects within a certain distance range in any direction of the vehicle, while the environmental sensors collect data about environmental conditions within the vehicle's area of movement.
[0174] During operation, information from the sensors is transmitted to an in-vehicle computing device 1120. The in-vehicle computing device 1120 may be embodied using the computer system of FIG. 29. The in-vehicle computing device 1120 may analyze data collected by the sensors and selectively control the operation of the vehicle based on the analysis. For example, the in-vehicle computing device 1120 may control braking via a brake controller 1122, direction via a steering controller 1124, speed and acceleration via a throttle controller 1126 in the case of a gasoline-powered vehicle, speed and acceleration via a motor speed controller 1128 such as a current level controller in the case of an electric vehicle, a differential gear controller 1130 in the case of a vehicle with a transmission, and / or other controls. The auxiliary device controller 1134 may be configured to control one or more auxiliary devices, such as a test system, auxiliary sensors, or mobile devices carried by the vehicle.
[0175] Geographic location information may be transmitted from the location sensor 1160 to the in-vehicle computing device 1120, which can access a map of the environment corresponding to the location information to identify fixed environmental elements such as roads, buildings, stop lights, and / or stop / go signals. Images captured from the camera 1162 and / or object detection information collected from a sensor, such as a lidar system 1164, are transmitted from the sensor to the in-vehicle computing device 1120. The object detection information and / or captured images are processed by the in-vehicle computing device 1120 to detect objects around the vehicle. Object detection based on the sensor data and / or captured images may be performed using any known or later-known technology that can be used in the embodiments disclosed herein.
[0176] Lidar information is transmitted from the lidar system 1164 to the in-vehicle computing device 1120. Also, captured images are transmitted from the camera 1162 to the in-vehicle computing device 1120. The lidar information and / or captured images are processed by the in-vehicle computing device 1120 to detect objects near the vehicle. The object detection scheme performed by the in-vehicle computing device 1120 includes the functionality described in detail herein.
[0177] The system architecture 1100 may also include an on-board display device 1154 capable of generating and outputting an interface to a vehicle occupant that displays sensor data, vehicle status information, or output generated by the processes described herein. The display device may include an audio speaker to provide such information in audio form, or may be a separate device.
[0178] The in-vehicle computing device 1120 may include or be communicatively coupled to a routing controller 1132 that generates a navigation route from a start location of the autonomous vehicle to a destination location. The routing controller 1132 may access a map data store to identify possible routes and road segments that the vehicle can travel from the start location to the destination location. The routing controller 1132 may assign scores to possible routes to identify a preferred route for reaching the destination. For example, the routing controller 1132 may generate a navigation route based on minimizing Euclidean distance or other cost function over the route, and may access traffic and / or estimation information that may affect the time it takes to travel a particular route. Depending on the implementation, the routing controller 1132 may generate one or more routes using various routing methods, such as Dijkstra's algorithm, Bellman-Ford algorithm, or other algorithms. The routing controller 1132 may also use traffic information to generate a navigation route that reflects expected conditions for that route (e.g., the current day of the week or time of day, etc.), so that a route during rush hour may be different from a route during the middle of the night. The routing controller 1132 may also generate one or more navigation routes to a destination and may transmit two or more routes to the user for selection by the user from among various possible routes.
[0179] In various embodiments, the in-vehicle computing device 1120 can determine perception information about the AV's surrounding environment. Based on sensor data provided by one or more sensors and acquired position information, the in-vehicle computing device 1120 can determine perception information about the AV's surrounding environment. Perception information can indicate what a typical driver can perceive in the vehicle's surrounding environment. Perception data can include information related to one or more objects in the AV's surrounding environment. For example, the in-vehicle computing device 1120 can process sensor data (e.g., lidar or radar data, camera images, etc.) to identify objects and / or features in the AV's environment. Objects can include traffic signals, road boundaries, other vehicles, pedestrians, obstacles, etc. The in-vehicle computing device 1120 can use object recognition algorithms, video tracking algorithms, and computer vision algorithms (e.g., repeatedly tracking an object frame by frame over multiple time intervals) that may be known now or in the future to determine perception.
[0180] In some embodiments, the in-vehicle computing device 1120 may also determine, for one or more identified objects in the environment, the current state of the object. The state information may include, without limitation, for each object: current position, current speed and / or acceleration, current direction, current attitude, current form, size or occupancy, type (e.g., vehicle, pedestrian, bicycle, stationary object or obstacle), and / or other state information.
[0181] The in-vehicle computing device 1120 may perform one or more prediction and / or forecasting tasks. For example, the in-vehicle computing device 1120 may predict the future position, trajectory, and / or motion of one or more objects. For example, the in-vehicle computing device 1120 may predict the future position, trajectory, and / or motion of an object based, at least in part, on perception information (e.g., state data for each object, including estimated form and pose, as described below), location information, sensor data, and / or other data describing the past and / or current state of the object, AV, surrounding environment, and / or the relationships between them. For example, if the object is a vehicle and the current driving environment includes an intersection, the in-vehicle computing device 1120 may predict whether the object is likely to proceed straight or turn. If perception data indicates that the intersection does not have a traffic light, the in-vehicle computing device 1120 may also predict the likelihood that the vehicle will have to come to a complete stop before entering the intersection.
[0182] In various embodiments, the in-vehicle computing device 1120 may determine a motion plan for the autonomous vehicle. For example, the in-vehicle computing device 1120 may determine a motion plan for the autonomous vehicle based on perception data and / or prediction data. Specifically, based on predictions of future positions of neighboring objects and other perception data, the in-vehicle computing device 1120 may determine a motion plan that optimally drives the autonomous vehicle, taking into account the objects at those future positions.
[0183] In some embodiments, the in-vehicle computing device 1120 can receive predictive information and determine how to handle objects and / or actors in the AV's environment. For example, for a particular actor (e.g., a vehicle with a particular speed, direction, and turning angle), the in-vehicle computing device 1120 determines whether to overtake, yield, stop, and / or pass based on traffic conditions, map data, the state of the autonomous vehicle, etc. The in-vehicle computing device 1120 also plans not only the path but also driving parameters (e.g., distance, speed, and / or turning angle) so that the AV can travel along a given path. That is, for a given object, the in-vehicle computing device 1120 determines what to do with the object and how to do it. For example, for a given object, the in-vehicle computing device 1120 can determine to pass the object and whether to pass the left or right side of the object (including moving parameters such as speed). The in-vehicle computing device 1120 can also evaluate the collision risk between the sensed object and the AV. If the collision risk exceeds an acceptable threshold, it can be determined whether the collision can be avoided if the autonomous vehicle follows a defined vehicle trajectory or performs one or more dynamically generated emergency avoidance maneuvers within a predefined time period (e.g., N milliseconds). If the collision can be avoided, the in-vehicle computing device 1120 can execute one or more control commands to perform deliberate avoidance maneuvers (e.g., slowly reducing speed, accelerating, changing lanes, or making a sharp turn). Conversely, if the collision cannot be avoided, the in-vehicle computing device 1120 can execute one or more control commands to perform emergency maneuvers (e.g., braking and / or changing direction).
[0184] As described above, plans and control data for the autonomous vehicle's movements are generated for execution. The in-vehicle computing device 1120 can control, for example, braking via a brake control, direction control via a steering control, speed and acceleration control via a throttle control in the case of a gasoline-powered vehicle, speed and acceleration control via a motor speed control such as a current level controller in the case of an electric vehicle, a differential gear control in the case of a vehicle with a transmission, and / or other controls.
[0185] 28 provides a block diagram useful for understanding how AV operation or movement is accomplished in accordance with the inventive solution. All operations performed in blocks 1202-1212 may be performed by an on-board computing device (e.g., on-board computing device 1022 in FIG. 26 and / or 1120 in FIG. 27) of the vehicle (e.g., AV 1002 in FIG. 26).
[0186] In block 1202, the position of an AV (e.g., AV 1002 in FIG. 26) is sensed. This sensing may be performed based on sensor data output from a position sensor of the AV (e.g., position sensor 1160 in FIG. 27). This sensor data may include, but is not limited to, GPS data. The sensed position of the AV is transmitted to block 1206.
[0187] In block 1204, an object (e.g., vehicle 1003 in FIG. 26) is detected near (e.g., less than 100 meters from) the AV (e.g., 1002 in FIG. 26). This detection is performed based on sensor data 1216 output from the AV's camera (e.g., camera 1162 in FIG. 27) and / or lidar system (e.g., lidar system 1164 in FIG. 27). For example, image processing is performed to detect instances of a particular class of object (e.g., vehicle, bicyclist, or pedestrian) from the image. This image processing / object detection may be performed by any image processing / object detection algorithm known now or in the future. Lidar sensor data may include, but is not limited to, superpixels generated by the methods described in methods 500, 600, 650, and 900.
[0188] Additionally, block 1204 determines a predicted trajectory for the object. Block 1204 predicts the object's trajectory based on the object's class, cube geometry, cube heading, and / or the contents of map 1218 (e.g., sidewalk location, lane location, lane direction, driving rules, etc.). How the cube geometry and heading are determined will become clearer in the following description. At this point, it should be noted that the cube geometry and / or heading may be determined using various types of sensor data (e.g., 2D images, 3D lidar point clouds) and vector map 1218 (e.g., lane geometry). Techniques for predicting the object's trajectory based on cube geometry and heading may include, for example, predicting that the object will move along a straight line path in the same direction as the cube's heading. Predicted object trajectories include, but are not limited to, the following trajectories: a trajectory defined by the object's actual speed (e.g., 1 mph) and actual driving direction (e.g., west); a trajectory defined by the object's actual speed (e.g., 1 mph) and other possible directions of movement (e.g., south, southwest, or X degrees AV (e.g., 40°) from the object's actual direction of movement); a trajectory defined by other possible speeds (e.g., 2-10 mph) and the object's actual direction of movement (e.g., west); and / or a trajectory defined by other possible speeds (e.g., 2-10 mph) and other possible directions of movement (e.g., south, southwest, or X degrees AV (e.g., 40°) from the object's actual direction of movement). The possible speeds and / or possible driving directions may be predefined values for the same class and / or subclass of the object. Note again that the cube defines the entire range and direction of the object. The heading defines the direction in which the object's front is pointing and therefore provides information on the object's actual and / or possible running direction.
[0189] Information 1220 specifying the predicted trajectory of the object and the geometry / orientation of the cube is provided to block 1206. In some scenarios, the classification of the object is also communicated to block 1206. Block 1206 generates a vehicle trajectory using information from blocks 1202 and 1204. Techniques for determining a vehicle trajectory using the cube may include, for example, determining a trajectory along which the AV can pass the object when the following conditions are met: the object is located in front of the AV, the cube's orientation is aligned with the AV's direction of travel, and the cube's length is greater than a critical value. The solution of the present invention is not limited to this particular scenario. Vehicle trajectory 1208 may be determined based on the position information from block 1202, the object detection information from block 1204, and / or map information 1214 pre-stored in the vehicle's data store. Map information 1214 may include, but is not limited to, all or a portion of road map 1060 of FIG. 26 . The vehicle trajectory 1208 may represent a smooth path without any abrupt changes that may cause inconvenience to passengers. For example, a vehicle trajectory may be defined as a path traveled along a particular lane of a road that the object is not expected to travel within a given time. The vehicle trajectory 1208 is provided to a subsequent block 1210.
[0190] In block 1210, steering angle and velocity commands are generated based on the vehicle trajectory 1208. The steering angle and velocity commands are provided to block 1210 for vehicle dynamics control, i.e., the steering angle and velocity commands cause the AV to follow the vehicle trajectory 1208.
[0191] Various embodiments may be implemented using one or more computer systems similar to, for example, computer system 1300 shown in Figure 29. Computer system 1300 may be any computer capable of performing the functions described herein.
[0192] 29, computer system 1300 may be any computer capable of performing the functions described herein. Computer system 1300 also includes a user input / output interface 1302 and user input / output devices 1303, such as buttons, a monitor, a keyboard, a pointing device, etc.
[0193] Computer system 1300 includes one or more processors (also referred to as central processing units or CPUs), such as processor 1304. Processor 1304 is coupled to a communications infrastructure or bus 1306. Processor 1304 may be a graphics processing unit (GPU), a specialized electronic circuit with a parallel architecture designed to process mathematically complex applications for parallel processing of mathematically complex data common in computer graphics applications, images, video, and the like.
[0194] Computer system 1300 also includes main memory 1308 (e.g., one or more levels of random access memory (RAM), including cache, for storing control logic (i.e., computer software) and / or data). Computer system 1300 may also include one or more secondary storage devices or secondary memory 1310 (e.g., hard disk drive 1312) and / or portable storage device 1314 (which may interface with portable storage unit 1318). Portable storage device 1314 and portable storage unit 1318 may be floppy disk drives, magnetic tape drives, compact disk drives, optical storage devices, tape backup devices, and / or other storage devices / drives.
[0195] Secondary memory 1310 may include other means, devices, or accesses that allow computer system 1300 to access computer programs and / or other instructions and / or data, such as interfaces 1320 and portable storage units 1322 (e.g., program cartridges and cartridge interfaces (similar to those found in video game devices), portable memory chips (such as EPROMs or PROMs) and associated sockets, memory sticks and USB ports, memory cards and associated memory card slots, and / or other portable storage units and associated interfaces).
[0196] Computer system 1300 also includes a network or communication interface 1324 that can communicate and interact with any combination of remote devices, remote networks, remote entities, etc. (individually or collectively referred to by reference numeral 1328). For example, communication interface 1324 enables computer system 1300 to communicate with remote devices 1328 over communication path 1326, which may be wired and / or wireless. This communication path 1326 may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 1300 over communication path 1326.
[0197] In one embodiment, a physical, non-transitory device or article of manufacture, also referred to herein as a computer program product or program storage device, includes a physical, non-transitory computer-usable or readable medium having control logic (software) stored thereon, including, but not limited to, physical articles of manufacture embodying computer system 1300, main memory 1308, secondary memory 1310, portable storage units 1318 and 1322, or combinations thereof. Such control logic, when executed by one or more data processing devices (e.g., computer system 1300), causes the data processing devices to operate as described herein.
[0198] Based on the teachings herein, one skilled in the relevant art will clearly understand how to make and use embodiments herein using data processing devices, computer systems and / or computer architectures other than those shown in Figure 29. In particular, embodiments may operate in conjunction with software, hardware and / or operating system implementations other than those described herein.
[0199] Relevant terms in this specification include the following: The term "electronic device" or "computing device" refers to a device that includes a processor and memory. Each device may have its own processor and / or memory, or the processor and / or memory may be shared with other devices, such as in a virtual machine or container configuration. The memory may contain or receive programming instructions that, when executed by a processor, cause the electronic device to perform one or more operations in accordance with the programming instructions.
[0200] The terms "memory," "memory device," "data store," "data storage facility," and the like refer to non-transitory devices in which computer-readable data, programming instructions, or both are stored. Unless expressly stated otherwise, the terms "memory," "memory device," "data store," "data storage facility," and the like are intended to include embodiments comprised of a single device, embodiments in which multiple memory devices together or collectively store a single set of data or instructions, and individual sectors within such devices. A computer program product is a memory device in which programming instructions are stored.
[0201] The terms "processor" and "processing device" refer to hardware components of an electronic device configured to execute programming instructions. Unless explicitly stated otherwise, the singular term "processor" or "processing device" is intended to include embodiments consisting of a single processing device as well as embodiments in which multiple processing devices, which may be components of one device or separate devices, work together or in concert to perform a process.
[0202] The term "object," when referring to an object sensed by a vehicle perception system or simulated by a simulation system, is intended to include both stationary objects and moving or potentially moving actors, except when the terms "actor" or "stationary object" are explicitly used.
[0203] The term "trajectory," when used in the context of autonomous vehicle motion planning, refers to the plan that a vehicle motion planning system generates and that a vehicle motion control system follows to control the vehicle's motion. The trajectory includes the vehicle's planned position and attitude at various points in time over a time horizon, as well as the vehicle's planned steering angle and steering angle rate over the same time horizon. The autonomous vehicle motion control system uses the trajectory to send commands to the vehicle's steering controller, braking controller, throttle controller, and / or other motion control subsystems to cause the vehicle to move along the planned path.
[0204] A "trajectory" of an actor, which a vehicle's perception or prediction system can generate, refers to the path that the actor is predicted to follow over a time horizon, and the actor's predicted speed and / or position at various points along that path.
[0205] As used herein, the terms "street," "lane," "road," and "intersection" are illustratively described with respect to a vehicle traveling on one or more roads, but embodiments are intended to include lanes and intersections in other locations, such as parking lots. Also, in the case of an autonomous vehicle designed for indoor use, such as an automated pickup device in a warehouse, a "street" may refer to an aisle in the warehouse, and a "lane" may refer to a portion of the aisle. In the case of an autonomous vehicle that is a drone or other aircraft, a "road" or "lane" may refer to an airway or a portion thereof. In the case of an autonomous vehicle that is a ship, a "road" or "lane" may refer to a waterway or a portion thereof.
[0206] In this specification, when terms such as "first" and "second" modify a noun, such use is merely to distinguish one item from another and does not require sequential order unless otherwise specified. Also, terms indicating relative positions, such as "vertical" and "horizontal," "front" and "rear," are intended to indicate relative relationships to one another and are not necessarily absolute, referring to only one possible position depending on the orientation of the device associated with the term.
[0207] It should be recognized that the Detailed Description section, and not other sections of the specification, is intended to be used to interpret the claims. Other sections may present one or more exemplary embodiments contemplated by the inventors, but are not intended to limit the invention or the appended claims in any manner.
[0208] While this specification describes embodiments for exemplary fields and applications, it should be understood that this specification is not limited to the disclosed examples. Other embodiments and variations thereon are possible and fall within the spirit and scope of the invention. For example, without limiting the generality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or components shown and described in this disclosure. Also, embodiments (even if not explicitly described) have meaningful potential for use in fields and applications beyond the examples described herein.
[0209] Embodiments have been described herein using functional building blocks to illustrate the implementation of certain functions and relationships. The boundaries of such functional building blocks have been arbitrarily defined herein for the convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or equivalent functions and relationships) are properly performed. Furthermore, alternative embodiments may perform the functional blocks, steps, operations, methods, etc. in an order different from that described herein. Features of different embodiments disclosed herein may be freely combined. For example, one or more features of a method embodiment may be combined with any of the system or product embodiments. Similarly, features of a system or product embodiment may be combined with any of the method embodiments disclosed herein.
[0210] As used herein, the phrases "one embodiment," "an embodiment," "an example embodiment," or similar phrases indicate that a described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with one embodiment, a person of ordinary skill in the relevant art will understand that such feature, structure, or characteristic may also be included in other embodiments even if not explicitly mentioned or described herein. Some embodiments may also be described using the terms "coupled" and "connected," along with their derivatives. These terms are not necessarily used synonymously. For example, some embodiments may use the terms "coupled" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements cooperate or interact with each other even when they are not in direct contact with each other. Features of different embodiments disclosed herein may be freely combined. For example, one or more features of a method embodiment may be combined with any of the system or product embodiments. Similarly, features of a system or product embodiment may be combined with any of the method embodiments disclosed herein.
Claims
1. As a LiDAR system, an array of transmitters configured to transmit light pulses from the vehicle along a transmission axis to form a Tx transmission field-of-view (FoV); at least one detector configured to receive at least a portion of the light pulses that reflect off objects within an Rx reception field-of-view (FoV) along a receive axis; a transmit optical device mounted for movement along a horizontal axis and configured to intersect the respective transmit axis but not the receive axis to adjust the Tx FoV without adjusting the Rx FoV; LiDAR system.
2. the Tx FoV and the Rx FoV overlap, and the adjusted Tx FoV is located within an area of the Tx FoV. The LiDAR system of claim 1 .
3. a collimator mounted adjacent to the series of emitters and configured to focus and direct the light pulses along each transmit axis to collectively form a transmit beam; The LiDAR system of claim 1 .
4. the transmit optics is arranged adjacent to the collimator and configured to focus the transmit beam on a region of the Tx FoV to form the adjusted Tx FoV. The LiDAR system of claim 3.
5. the transmitting optical device includes a cylindrical lens; The LiDAR system of claim 4.
6. the series of emitters includes a linear array of emitters arranged parallel to the horizontal axis, the linear array of emitters including a proximal emitter and a distal emitter arranged opposite the proximal emitter; The LiDAR system of claim 1 .
7. further comprising an actuator coupled to the transmit optical device and configured to move the transmit optical device through a range between a rest position, in which the optical device does not intersect any transmit axis of the linear array of emitters, and a distal position for intersecting the transmit axis of the distal emitter. The LiDAR system of claim 6.
8. a controller configured to move the transmit optics along the transverse axis; The controller determining from the received light pulses whether the object is an unknown object; and further configured to move the transmit optical device along the transverse axis between a proximal position and a distal position while light pulses are transmitted via the transmit optical device. The LiDAR system of claim 1 .
9. The controller receiving sweep data indicative of the light pulses reflecting from the unknown object while moving the transmitting optical device; determining a position of the unknown object based on the sweep data; and further configured to move the transmit optics to a position along the horizontal axis such that the adjusted Tx FoV is aligned with the position of the unknown object.
10. The LiDAR system of claim 9.
10. As a LiDAR system, processor; a non-transitory computer-readable storage medium including programming instructions configured to cause the processor to implement a method for operating a LiDAR system; The programming instructions are: receiving a result value from the photodetector indicating a time when the photodetector detects a photon at or near the target wavelength; combining other sets of the result values to generate superpixels; using the superpixels to obtain a first spatiotemporal coherence metric; selecting a subset of light pulses or a group of result values based on the first spatiotemporal coherence metric; and and detecting a distance between the LiDAR system and the object based on the selected subset of light pulses or the selected group of result values. LiDAR system.
11. the first spatiotemporal coherence metric includes a metric specifying a change in distribution between detection of two pulses or groups of two pulses, respectively, by the plurality of photodetectors, and the subset of optical pulses or the group of result values is selected based on the largest value of the metric.
11. The LIDAR system of claim 10.
12. the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the differences between successive timestamps sorted from lowest to highest or highest to lowest, and a selected subset of the light pulses or group of result values comprises light pulses or result values associated with a relatively low measured variance. The LiDAR system of claim 10.
13. the first spatiotemporal coherence metric comprises a score for each pulse of the optical signal indicative of a confidence or validity of object detection, the pulse being selected for inclusion in the subset if the score exceeds a value. The LiDAR system of claim 10.
14. As a LiDAR system, processor; a non-transitory computer-readable storage medium including programming instructions configured to cause the processor to implement a method for operating a LiDAR system; The programming instructions are: arranging a plurality of pixels in a grid, the plurality of pixels including resultant values generated from processed waveforms produced by photodetectors of the LiDAR system; identifying a first region of interest based on at least one of a correlation between range values associated with the plurality of pixels and a correlation between intensity values associated with the plurality of pixels; combining result values associated with pixels located within the first region of interest to generate at least one first feature value; and a command to generate a first superpixel having a value set to the at least one first feature value; LiDAR system.
15. the programming instructions further include instructions to obtain a kernel size and use the kernel size to identify the area of interest of the grid.
15. The LiDAR system of claim 14.
16. The kernel size is locating at least one of the plurality of pixels that is a nearest neighbor of the pixel of interest of the grid in at least a range dimension; and obtained by defining the kernel size based on the location of the nearest neighbor within the grid.
16. The LiDAR system of claim 15.
17. The kernel size is Get the reference kernel size; identifying regions within the grid using the reference kernel size; identifying a central pixel of said region; calculating a score for each pixel of the region, the score indicating the degree of correlation between the result value associated with the pixel and the central pixel, using the result value associated therewith; selecting a pixel from the plurality of pixels based on the score; and obtained by defining the kernel size based on the position of the selected pixel within the grid; 16. The LiDAR system of claim 15.
18. 1. A method for operating a LiDAR system, comprising: receiving, by a processor, result values from a plurality of photodetectors indicative of times at which the plurality of photodetectors detect photons at or near a target wavelength, the result values being based on operations performed by each of the plurality of photodetectors to facilitate measurements associated with optical signals reflected from objects external to the LiDAR system; combining, by the processor, other sets of the result values to generate superpixels; using, by the processor, the superpixels to obtain a first spatiotemporal coherence metric; selecting a subset of light pulses or a group of result values based on the first spatiotemporal coherence metric; and detecting, by the processor, a distance between the LiDAR system and the object based on the selected subset of light pulses or the selected group of result values. method.
19. the first spatiotemporal coherence metric comprises at least one of a distribution comparison metric, a time-of-flight statistical metric, and a detection confidence score; the first spatiotemporal coherence metric includes a metric specifying a change in distribution between detection of two pulses or groups of two pulses, respectively, by the plurality of photodetectors, and the subset of optical pulses or the group of result values is selected based on the largest value of the metric.
20. The method of claim 18.
20. the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the differences between successive timestamps sorted from lowest to highest or highest to lowest, and a selected subset of the light pulses or group of result values comprises light pulses or result values associated with a relatively low measured variance.
20. The method of claim 19.