Lidar system, method for operating a lidar system
By adjusting the transmitting field of view in the lidar system without affecting the received field of view, using the processor to select a subset of light pulses and generate superpixels, the problem of point cloud quality degradation when the object in the prior art is not large enough to fill the superpixel field of view is solved, and higher detection accuracy and object recognition capabilities are achieved.
Patent Information
- Application Number
- CN202380072895.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-23
AI Technical Summary
Existing lidar systems integrate pulses that are not reflected from the object cause point cloud quality degradation when the object is not large enough to fill the superpixel field of view.
By adjusting the transmitting field of view without adjusting the received field of view, subset selection and superpixel generation of optical pulses are performed using adjustable transmitting optics and processors to improve detection accuracy of the lidar system.
It is realized that the detection accuracy and object recognition capabilities of the lidar system are improved without reducing the quality of the received field of view, especially when identifying unknown objects.
Smart Images

Figure CN120035772A_ABST
Abstract
Description
Technical Field
[0001] Embodiments relate to a lidar system. Embodiments relate to a method for operating a lidar system. Background Art
[0002] The vehicle may include a sensor system to monitor its external environment for obstacle detection and avoidance. The sensor system may include multiple sensor components for monitoring objects approaching the vehicle in the near field and distant objects in the far field. Each sensor component may include one or more sensors, such as cameras, radio detection and ranging (radar) sensors, light detection and ranging (lidar) sensors, and microphones. A laser radar (LiDAR) sensor includes one or more transmitters for emitting light pulses out of the vehicle, and one or more detectors for receiving and analyzing reflected light pulses. The laser radar sensor may include one or more optical elements to focus and guide the emitted light and the received light within the field of view outside the vehicle. The sensor system may determine the position of an object in the external environment based on data from the sensor. The vehicle may control one or more vehicle systems, such as a powertrain, a braking system, and a steering system, based on the position of the object.
[0003] LiDAR systems that use a fixed resolution grid integrate a fixed number of pulses or detector readings to form a superpixel. However, if the object is not large enough to fill the field of view of the superpixel, integrating pulses or detector readings that are not reflected from the object will degrade the quality of the estimated point cloud. LiDAR systems can also be used in other applications, such as in aircraft, ships, and / or mapping systems. Summary of the invention Technical issues
[0004] In one embodiment, a lidar system is provided with a series of transmitters, each transmitter configured to transmit light pulses out of a vehicle along a transmit axis to form a transmit field of view (Tx FoV). At least one detector is configured to receive at least a portion of the light pulses reflected from an object within a receive field of view (Rx FoV) along the receive axis. The transmit optics are mounted for translation along a lateral axis and are configured to intersect each transmit axis without intersecting the receive axis to adjust the Tx FoV without adjusting the Rx FoV.
[0005] In another embodiment, a method for adjusting a transmit field of view is provided. Light pulses are transmitted from a vehicle along at least one transmit axis to form a transmit field of view (Tx FoV). At least a portion of the light pulses reflected from an object within a receive field of view (Rx FoV) along the receive axis are received. The transmit optics are translated along a transverse axis to intersect each transmit axis without intersecting the receive axis to adjust the Tx FoV without adjusting the Rx FoV.
[0006] In yet another embodiment, a non-transitory computer-readable medium has instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations including: transmitting light pulses from a vehicle to form a transmit field of view (Tx FoV); receiving at least a portion of the light pulses reflected from an object within a receive field of view (Rx FoV); and translating the transmit optics along a lateral axis to adjust the Tx FoV without adjusting the Rx FoV.
[0007] The present disclosure relates to implementing systems and methods for operating a laser radar system. The method includes: performing operations by each photodetector to facilitate measurements associated with light signals reflected from an object external to the laser radar system; receiving, by a processor, a result value from the photodetector, the result value indicating a time when the photodetector detected a photon at or near a target wavelength; combining, by a processor, different sets of result values to generate a plurality of superpixels; using, by a processor, the plurality of superpixels to obtain a spatiotemporal coherence metric; selecting, by a processor, a subset of light pulses or a set of result values based on the spatiotemporal coherence metric; and / or detecting, by a processor, a distance between the laser radar system and an object based on the selected subset of light pulses or the selected set of result values.
[0008] The implementation system may include: a processor; and a non-transitory computer-readable storage medium, which includes programming instructions configured to cause the processor to implement a method for operating a lidar system. The above method can also be implemented by a computer program product including a memory and programming instructions, which are configured to cause the processor to perform operations.
[0009] The present disclosure relates to an implementation system and method for operating a laser radar system. The method includes: arranging pixels in a grid by a processor (wherein the pixels include result values generated from processing waveforms generated by a photodetector of the laser radar system); identifying, by the processor, a first region of interest in the grid based on correlations between range values associated with the pixels and / or correlations between intensity values associated with the pixels; combining, by the processor, the result values associated with pixels located within the first region of interest to generate (one or more) feature values; and generating, by the processor, a superpixel having (one or more) values set as (one or more) feature values.
[0010] The implementation system may include: a processor; and a non-transitory computer-readable storage medium, which includes programming instructions configured to cause the processor to implement a method for operating a lidar system. The above method can also be implemented by a computer program product including a memory and programming instructions, which are configured to cause the processor to perform operations. Solutions to the problem
[0011] A laser radar system according to an embodiment of the present invention includes: a series of transmitters, each transmitter is configured to transmit light pulses from a vehicle along a transmission axis to form a transmission field of view (Tx FoV); at least one detector, configured to receive at least a portion of the light pulses reflected from an object within a receiving field of view (Rx FoV) along the receiving axis; and a transmitting optical device, which is mounted for translation along a transverse axis and is configured to intersect each transmission axis without intersecting the receiving axis to adjust the Tx FoV without adjusting the Rx FoV.
[0012] According to an embodiment of the present invention, the Tx FoV and the Rx FoV overlap, and wherein the adjusted Tx FoV is located within the area of the Tx FoV.
[0013] According to an embodiment of the present invention, a collimator is also included, which is mounted adjacent to the series of emitters and is configured to focus and direct the light pulses along each emission axis to collectively form an emission light beam.
[0014] According to an embodiment of the invention, the emission optics is arranged adjacent to the collimator and is configured to focus the emission beam onto the area of the Tx FoV to form an adjusted Tx FoV. The emission optics comprises a cylindrical lens.
[0015] According to an embodiment of the present invention, the series of emitters includes a linear array of emitters arranged parallel to the transverse axis, the linear array of emitters including a proximal emitter and a distal emitter arranged opposite to the proximal emitter. According to an embodiment of the present invention, it also includes: an actuator connected to the emission optical device and configured to translate the emission optical device through a range between a static position and a distal position, in which the emission optical device does not intersect with any emission axis of the linear array of emitters and in which the emission optical device intersects with the emission axis of the distal emitter.
[0016] According to an embodiment of the present invention, it also includes a controller configured to translate the emitting optical device along a horizontal axis, wherein the controller is further configured to: determine from the received light pulse that the object is an unknown object; and translate the emitting optical device along the horizontal axis between a proximal position and a distal position while emitting light pulses through the emitting optical device.
[0017] According to an embodiment of the present invention, the controller is further configured to: receive scanning data indicating light pulses reflected from an unknown object when translating the transmitting optical device; determine the position of the unknown object based on the scanning data; and translate the transmitting optical device along the horizontal axis to a certain position so that the adjusted Tx FoV is aligned with the position of the unknown object.
[0018] According to an embodiment of the present invention, a laser radar system includes: a processor; a non-temporary computer-readable storage medium, including programming instructions configured to cause the processor to implement a method for operating a laser radar system, wherein the programming instructions include instructions to: receive a result value from a photodetector, the result value indicating the time when the photodetector detects a photon at or near a target wavelength; combine different sets of result values to generate a superpixel; use the superpixel to obtain a first spatiotemporal coherence metric; select a subset of light pulses or a set of result values based on the first spatiotemporal coherence metric; and detect the distance between the laser radar system and an object based on the selected subset of light pulses or the selected set of result values.
[0019] According to an embodiment of the invention, the first spatiotemporal coherence metric comprises metrics respectively specifying a change in the distribution between detections of two pulses or two groups of pulses by the plurality of photodetectors, and a subset of light pulses or a result value group is selected based on a maximum metric among the metrics.
[0020] According to an embodiment of the present invention, the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the difference between consecutive timestamps that have been sorted from lowest value to highest value or from highest value to lowest value, and the selected subset of optical pulses or group of result values comprises optical pulses or result values associated with relatively low measured variances.
[0021] According to an embodiment of the invention, the first spatiotemporal coherence measure comprises a score for each pulse of the optical signal indicating the confidence or validity of the object detection, and the pulse is selected for inclusion in the subset when the score exceeds a certain value.
[0022] According to an embodiment of the present invention, a laser radar system includes: a processor; a non-temporary computer-readable storage medium, which includes programming instructions configured to cause the processor to implement a method for operating a laser radar system, wherein the programming instructions include instructions to: arrange a plurality of pixels in a grid, the plurality of pixels including result values generated from processing a waveform generated by a photodetector of the laser radar system; identify a first region of interest in the grid based on at least one of a correlation between range values associated with the plurality of pixels and a correlation between intensity values associated with the plurality of pixels; combine the result values associated with pixels located within the first region of interest to generate at least one first eigenvalue; and generate a first superpixel, the first superpixel having a value set to the at least one first eigenvalue.
[0023] According to an embodiment of the present invention, the programming instructions also include instructions for obtaining a kernel size and using the kernel size to identify a region of interest in the grid.
[0024] According to an embodiment of the present invention, the kernel size is obtained by the following steps: locating a pixel among the plurality of pixels that is a closest neighbor of the pixel of interest in the grid at least in terms of range; and defining the kernel size based on the position of the closest neighbor in the grid.
[0025] According to an embodiment of the present invention, the kernel size is obtained by the following steps: obtaining a reference kernel size; using the reference kernel size to identify a region in a grid; identifying a central pixel of the region; using a result value associated with each of the pixels in the region to calculate a score for the pixel, the score indicating a correlation between the result values associated with the pixel and the central pixel; selecting a pixel from the plurality of pixels based on the score; and limiting the kernel size based on the position of the selected pixel in the grid.
[0026] A method for operating a lidar system according to an embodiment of the present invention includes: receiving, by a processor, result values from a plurality of photodetectors, the result values indicating a time at which the plurality of photodetectors detect photons at or near a target wavelength, the result values being based on operations performed by each of the plurality of photodetectors to facilitate measurements associated with light signals reflected from an object outside the lidar system; combining, by a processor, different sets of result values to generate a superpixel; using, by a processor, the superpixel to obtain a first spatiotemporal coherence metric; selecting, by the processor, a subset of light pulses or a set of result values based on the first spatiotemporal coherence metric; and detecting, by the processor, a distance between the lidar system and an object based on the selected subset of light pulses or the selected set of result values.
[0027] According to an embodiment of the present invention, the first spatiotemporal coherence metric includes at least one of a distribution comparison metric, a flight time statistical metric, and a detection confidence score, wherein the first spatiotemporal coherence metric includes a metric that respectively specifies the distribution change between the detection of two pulses or two groups of pulses by the multiple photodetectors, and selects a subset of light pulses or a group of result values based on the maximum one of the metrics.
[0028] According to an embodiment of the present invention, the first spatiotemporal coherence metric comprises, for each pulse, a measured variance of the difference between consecutive timestamps that have been sorted from lowest value to highest value or from highest value to lowest value, and the selected subset of optical pulses or group of result values comprises optical pulses or result values associated with relatively low measured variances. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a front perspective view of an exemplary vehicle with a automated driving system (SDS) including a lidar sensor with an adjustable transmit field of view (Tx FoV) according to aspects of the present disclosure.
[0030] Figure 2 is a schematic diagram illustrating communication between an SDS and other systems and devices according to aspects of the present disclosure.
[0031] Figure 3 is an exemplary architecture of a lidar sensor of an SDS according to aspects of the present disclosure.
[0032] Figure 4 is a top view of a lidar sensor according to aspects of the present disclosure.
[0033] Figure 5 is taken along section line VV according to various aspects of the present disclosure Figure 4 A cross-sectional view of a lidar sensor.
[0034] Figure 6 is a schematic diagram of a lidar sensor providing Tx FoV according to aspects of the present disclosure.
[0035] Figure 7 is a schematic diagram of another lidar sensor according to aspects of the present disclosure, illustrating a transmit optic adjusted to a first position to adjust the Tx FoV to a first region relative to the total Tx FoV.
[0036] Figure 8 According to various aspects of the present disclosure Figure 7 Another schematic diagram of a lidar sensor of the type shown having a transmit optical device adjusted to a second position to adjust the Tx FoV to a second area relative to the total Tx FoV.
[0037] Fig. 9 yes Figure 7 Another schematic diagram of a lidar sensor of , illustrating a transmit optical device that is adjusted to a third position to adjust the Tx FoV to a third area relative to the total Tx FoV.
[0038] Fig.10 It shows that according to various aspects of the present disclosure Figure 7 The total Tx FoV and receive field of view (Rx FoV).
[0039] Fig.11 The relationship between the adjusted Tx FoV and Rx FoV according to aspects of the present disclosure is shown.
[0040] Fig.12 is a flow chart illustrating a method for adjusting Tx FoV according to aspects of the present disclosure.
[0041] Fig.13 is used to combine Figure 3 Illustration of the technique showing the lidar waveform results for the lidar system.
[0042] Figures 14a to 14e (collectively referred to as "FIG. 14") provides an illustration of the method for combining Figure 3 Illustration of the convolution oversampling technique for a non-overlapping set of results for the photodetectors of a LiDAR system shown. The combining is accomplished by functions F1 to F5.
[0043] Fig.15 An example of an object is shown which reflects only two of the five pulses P1 to P5 which are integrated in this example.
[0044] Fig.16 A flow chart of an illustrative method for variable resolution refinement in a Geiger-mode lidar is provided.
[0045] Fig.17 A diagram is provided to help understand binning.
[0046] Fig.18 A graphical representation of the histogram is provided.
[0047] Fig.19 is a graphical representation of the target object.
[0048] Fig. 20 A flow chart of an illustrative method for spatial processing of lidar data according to the present solution is provided.
[0049] Figures 21a to 21g(collectively referred to as "FIG. 21") provides an illustration showing another technique for combining results generated from lidar waveforms.
[0050] Figures 22 to 23 Diagrams showing modified or otherwise adjusted kernel sizes and / or regions of interest (ROIs), respectively, are provided.
[0051] Fig.24 An illustration of a ROI with an adjusted or otherwise modified position in a grid is provided.
[0052] Fig.25 A flow chart of another illustrative method for spatial processing of lidar data according to the present solution is provided.
[0053] Fig.26 A graphical representation of the system is provided.
[0054] Fig. 27 A more detailed illustration of an autonomous vehicle is provided.
[0055] Fig.28 A block diagram of an illustrative vehicle trajectory planning process is provided.
[0056] Fig.29 A diagram of a computer system is provided. DETAILED DESCRIPTION
[0057] As required, detailed embodiments are disclosed herein; however, it should be understood that the disclosed embodiments are merely exemplary and may be embodied in various and alternative forms. The drawings are not necessarily drawn to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but merely as a representative basis for teaching those skilled in the art to embody the present disclosure in various ways.
[0058] A lidar system can have a fixed resolution. The resolution is defined as a property of the hardware and software systems and is independent of the reflectivity performance of the lidar system. In practice, this means that for a given target of fixed size, the lidar system has a certain detection probability as a function of reflectivity and range, and a fixed resolution. However, this is not necessarily what is desired in a downstream computer vision pipeline. Such a pipeline may prefer to have a constant detection probability over a given range, and be willing to sacrifice other properties of the range sensing system (such as resolution) to obtain it. For example, for objects of lower reflectivity, having a constant (high) detection probability but having fewer points on the target would allow the perception system to be confident that the object is present even if it is of low reflectivity, potentially at the expense of poorer velocity estimates or poorer data association.
[0060] This document describes system, device, apparatus, method and / or computer program product embodiments, and / or any combination and sub-combination of the foregoing, for variable resolution refinement of Geiger mode lidar resolution to address the detection probability problem of conventional lidar systems. This feature of the present solution provides improved lidar system operation, object detection using lidar data, and / or vehicle control.
[0061] The method generally includes: performing operations by each photodetector to facilitate measurements associated with light signals reflected from an object external to the lidar system; receiving, by a processor, a result value from the photodetector, the result value indicating the time when the photodetector detected a photon at or near a target wavelength; combining, by the processor, different sets of result values to generate a plurality of superpixels; using, by the processor, the plurality of superpixels to obtain a spatiotemporal coherence metric; selecting, by the processor, a subset of light pulses or a group of result values based on the spatiotemporal coherence metric; detecting, by the processor, a distance between the lidar system and the object based on the selected subset of light pulses or the selected group of result values; and / or causing, by the processor, the distance to be used to control operation of a vehicle.
[0062] Spatiotemporal coherence metrics can be obtained by considering superpixels for a fixed or variable number of pulses of the light signal. In both cases, the spatiotemporal coherence metrics may include, but are not limited to, distribution comparison metrics, flight time statistics metrics, and / or detection confidence scores. In the case of distribution comparison metrics, the spatiotemporal coherence metrics may include metrics that specify the distribution changes between the detections of two pulses or two groups of pulses by the multiple photodetectors, respectively. A subset of light pulses or a result value group is selected based on the largest one of the metrics. In the case of flight time statistics metrics, the spatiotemporal coherence metrics may include a measurement variance for each pulse of the difference between consecutive timestamps, which has been sorted from the lowest value to the highest value or from the highest value to the lowest value. The selected subset of light pulses or result value group includes light pulses or result values associated with relatively low measurement variances. In the case of detection confidence scores, the spatiotemporal coherence metrics may include a score for each pulse of the light signal indicating the confidence or validity of object detection. When the score exceeds a certain value, the pulse is selected to be included in the subset. Similarly, when the score associated with the corresponding pulse exceeds a certain value, the resulting value may be selected for inclusion in the group.
[0064] As used in this document, the singular forms "a", "an" and "the" include plural references unless the context clearly indicates otherwise. Unless otherwise specified, all technical and scientific terms used in this document have the same meaning as those generally understood by those of ordinary skill in the art. As used in this document, the term "including" means "including but not limited to". In this document, the term "vehicle" refers to any mobile form of transportation that can carry one or more human occupants and / or cargo and is powered by any form of energy. The term "vehicle" includes but is not limited to cars, trucks, vans, trains, autonomous vehicles, airplanes, drones, etc. "Autonomous driving vehicle" (or "AV") is a vehicle with a processor, programming instructions, and powertrain components that can be controlled by the processor without the need for a human operator. An autonomous driving vehicle can be fully autonomous, where it does not require an operator for most or all driving conditions and functions, or it can be semi-autonomous, where an operator may be required under certain conditions or for certain operations, or the operator can override the autonomous driving system of the vehicle and can control the vehicle.
[0066] Rotating optical sensors, such as rotating lidar sensors, can include complex physical and electrical architectures. Rotating lidar sensors can scan a wide 360-degree field of view (FoV) around a vehicle. The area where the lidar sensor's transmitter emits light is called the emission (Tx) FoV, and the area where the lidar sensor's detector receives light is called the receiving (Rx) FoV. Typically, the Tx FoV and Rx FoV overlap. A rotating lidar sensor can include a linear array of transmitters to provide a Tx FoV that extends over a wide vertical area. In such a rotating lidar sensor, the range and Tx FoV are inversely related. The larger the Tx FoV, the more the optical power is extended, and less light is emitted to small targets. In some cases, a self-driving system (SDS) may be more interested in maximizing the Tx FoV, whereas in some cases, the SDS may be more interested in maximizing the range on a more limited Tx FoV, for example, when identifying unknown objects in the far field. Some objects may be difficult to identify, such as tire debris, because they are not reflective and have irregular shapes.
[0067] According to some aspects, the SDS adjusts the Tx FoV without adjusting the Rx FoV to maximize the Rx FoV under certain conditions and to maximize the range over a smaller Rx FoV under other conditions. The lidar sensor includes a transmitter assembly having an adjustable transmit optics that is controlled to translate vertically to adjust the TxFoV without adjusting the Rx FoV. The transmit optics changes the divergence of the emitted light beam to focus the Tx FoV within a smaller area of the Rx FoV, which increases the processing power of the lidar sensor by reducing the overall size of the point cloud to be analyzed. By reducing the Tx FoV, the number of photons emitted by the transmitter onto the target increases for a smaller area of interest. In this case, since the Rx FoV does not change, the spatial resolution is not increased.
[0068] If the LiDAR sensor is to adjust the Tx FoV and Rx FoV, the adjustment needs to be synchronized to ensure that the emitter and detector scan the same area of the FoV. One benefit of adjusting the Tx FoV without adjusting the Rx FoV is that the detector does not need to know the exact position to which the Tx FoV is adjusted, as long as it remains within the Rx FoV. Another benefit of adjusting the Tx FoV without adjusting the Rx FoV is that this can be done without any additional moving electronics because the emitter, detector, and associated detector lens do not translate.
[0069] Reference Figure 1 , a lidar sensor is shown in accordance with one or more embodiments and is generally designated by reference numeral 100. The lidar sensor 100 is integrated with a self-driving system (SDS) 102 of a vehicle 104, such as an autonomous vehicle. The SDS 102 includes a plurality of sensors 106 to monitor an environment external to the vehicle 104. The lidar sensor 100 adjusts a transmit field of view (Tx FoV) without adjusting a receive field of view (Rx FoV) to monitor certain unknown objects 110 within the environment external to the vehicle 104, such as tire debris.
[0070] The SDS 102 includes a plurality of sensor assemblies, each of which includes one or more sensors 106 to monitor a 360-degree FoV in the near field and the far field around the vehicle 104. According to aspects of the present disclosure, the SDS 102 includes a top sensor assembly 112, two side sensor assemblies 114, two front sensor assemblies 116, and a rear sensor assembly 118. Each sensor assembly includes one or more sensors 106, such as a camera, a lidar sensor, and a radar sensor.
[0071] The top sensor assembly 112 is mounted to the top of the vehicle 104 and includes a plurality of sensors 106, such as a lidar sensor and a plurality of cameras. The lidar sensor rotates about an axis to scan a 360 degree FoV around the vehicle 104. The side sensor assembly 114 is mounted to the side of the vehicle 104, such as a camera. Figure 1 The front fenders 114 are shown, or mounted in side mirrors. Each side sensor assembly 114 includes a plurality of sensors 106, such as lidar sensors and cameras, to monitor the FoV of the adjacent vehicle 104 in the near field. The front sensor assembly 116 is mounted to the front of the vehicle 104, for example, below the headlights. Each front sensor assembly 116 includes a plurality of sensors 106, such as lidar sensors, radar sensors, and cameras, to monitor the FoV in front of the vehicle 104 in the far field. The rear sensor assembly 118 is mounted to the upper rear of the vehicle 104, for example, adjacent to the center high mounted stop light (CHMSL). The rear sensor assembly 118 includes a plurality of sensors 106, such as cameras and lidar sensors, for monitoring the FoV behind the vehicle 104.
[0072] Figure 2 The communication between the SDS 102 and other systems and devices according to various aspects of the present disclosure is shown. The SDS 102 includes a sensor system 200 and a controller 202. The controller 202 can communicate with other systems and devices directly or through a transceiver 204.
[0073] The sensor system 200 includes sensor components, such as the top sensor component 112 and the front sensor component 116. The top sensor component 112 includes one or more sensors, such as the lidar sensor 100, the radar sensor 208, and the camera 210. According to aspects of the present disclosure, the camera 210 can be a visible spectrum camera, an infrared camera, etc. The sensor system 200 can include additional sensors, such as microphones, sound navigation and ranging (SONAR) sensors, temperature sensors, position sensors (e.g., global positioning system (GPS)), etc.), position sensors, fuel sensors, motion sensors (e.g., inertial measurement unit (IMU)), etc.), humidity sensors, occupancy sensors, etc. The sensor system 200 provides sensor data 212 indicating the external environment of the vehicle 104. The controller 202 analyzes the sensor data to identify and determine the location of external objects relative to the vehicle 104, such as the location of traffic lights, remote vehicles, pedestrians, etc.
[0074] The SDS 102 also communicates with one or more vehicle systems 214, such as an engine, a transmission, a navigation system, a braking system, etc., through the transceiver 204. The controller 202 can receive information indicating current operating conditions from the vehicle system 214, such as vehicle speed, engine speed, turn signal state, brake position, vehicle position, steering angle, and ambient temperature. The controller 202 can also control one or more of the vehicle systems 214 based on the sensor data 212. For example, the controller 202 can control the braking system and the steering system to avoid obstacles. The controller 202 can communicate directly with the vehicle system 214 or indirectly with the vehicle system 214 through a vehicle communication bus (such as a CAN bus 216).
[0075] The SDS 102 may also communicate with external objects 218 (e.g., remote vehicles and structures) to share external environmental information and / or collect additional external environmental information. The SDS 102 may include a vehicle-to-everything (V2X) transceiver 220 connected to the controller 202 to communicate with the object 218. For example, the SDS 102 may use the V2X transceiver 220 to communicate directly with a remote vehicle via vehicle-to-vehicle (V2V) communication, directly with a structure (e.g., a sign, building, or traffic light) via vehicle-to-infrastructure (V2I) communication, and directly with a motorcycle via vehicle-to-motorcycle (V2M) communication.
[0076] The SDS 102 can use one or more of the transceivers 204, 220 to communicate with a remote computing device 222 over a communication network 224, for example, to provide a message or visual indication of the location of the object 218 relative to the vehicle 104 based on the sensor data 212. The remote computing device 222 may include one or more servers to process one or more processes of the technology described herein. The remote computing device 222 may also communicate data with a database 226 over the network 224.
[0077] SDS 102 includes user interface 228 to provide information to a user of vehicle 104 Controller 202 may control user interface 228 to provide a message or visual indicating the location of object 218 relative to vehicle 104 based on sensor data 212 .
[0078] Although the controller 202 is described as a single controller, it may include multiple controllers or may be implemented as software code within one or more other controllers. The controller 202 includes a processing unit or processor 230, which may include any number of microprocessors, ASICs, ICs, memories (e.g., FLASH, ROM, RAM, EPROM, and / or EEPROM) and software codes to perform a series of operations in cooperation with each other. Such hardware and / or software may be grouped together in a component to perform certain functions. Any one or more of the controllers or devices described herein include computer executable instructions that may be compiled or interpreted from a computer program created using various programming languages and / or techniques. The controller 202 also includes a memory 232 or a non-temporary computer-readable storage medium that is capable of executing instructions of a software program. The memory 232 may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. Typically, the processor 230 receives instructions, for example, from a memory 232, a computer-readable medium, etc., and executes these instructions. According to aspects of the present disclosure, the controller 202 also includes predetermined data or a "look-up table" stored in memory.
[0079] Figure 3 An exemplary architecture of a lidar sensor 300, such as the lidar sensor 100 of the top sensor assembly 112, according to aspects of the present disclosure is shown. The lidar sensor 300 includes a base 302 mounted to the vehicle 104. The base 302 includes a motor 304 having an axis 306 extending along an axis AA. The lidar sensor 300 also includes a housing 308 secured to the axis 306 and mounted for rotation about the axis AA relative to the base 302. The housing 308 includes an opening 310 and a cover 312 secured within the opening 310. The cover 312 is formed of a light-transmissive material such as glass. Although Figure 3 308, the lidar sensor 300 may include multiple covers 312, or one cover 312 that spans the entire outer surface of the housing 308. The lidar sensor 300 includes a housing 308 that can rotate 360° about a central axis (such as the shaft 306 or hub of the motor 304). The housing 308 may include a transmitter / receiver opening 310 made of a light-transmissive material. Although Figure 3310, but the present solution is not limited thereto. In other cases, multiple openings for transmitting and / or receiving light may be provided. Either way, as the housing 308 rotates about the internal components, the lidar sensor 300 may transmit light through one or more of the openings 310 and receive reflected light back toward one or more of the openings 310. In an alternative case, the housing of the housing 308 may be a fixed dome at least partially made of a light-transmitting material, with the rotatable component being inside the housing 308.
[0080] Inside the rotating shell or fixed dome is a transmitter 316, which is configured and positioned to generate and emit light pulses through the opening 310 or through the transparent dome of the shell 308 via one or more laser emitter chips or other light emitting devices. The transmitter 316 may include any number of individual transmitters (for example, 8 transmitters, 64 transmitters, or 128 transmitters). The transmitter may emit light of substantially the same intensity or varying intensities. The lidar sensor 300 also includes a detector 318 containing an array of photodetectors. The photodetector is positioned and configured to receive light reflected back into the system. Upon receiving the reflected light, the photodetector generates a result (or electrical pulse) indicating the measured intensity of the light signal reflected from an object outside the lidar sensor. In Geiger mode applications, the photodetector is excited when a single photon at or near the target wavelength is detected. The time at which the photodetector is excited is recorded as a timestamp. The transmitter 316 and the detector 318 rotate with the rotating shell, or they rotate within the fixed dome of the shell 308. One or more optical element structures 322 may be positioned in front of the emitter 316 and / or detector 318 to act as one or more lenses or wave plates that focus and direct light passing through the optical element structures 322. The lidar sensor 300 includes one or more emitters 316 for transmitting light pulses 320 through the cover 312 and away from the vehicle 104 to the Tx FoV ( Figure 1 ). Light pulses 320 are incident on one or more objects within the Rx FoV and are reflected back toward the lidar sensor 300 as reflected light pulses 328. The lidar sensor 300 also includes one or more detectors 318 for receiving reflected light pulses 328 that pass through the cover 312. The detectors 318 also receive light from external light sources (e.g., the sun). The lidar sensor 300 rotates about axis AA to scan the area within its FoV. The emitter 316 and detector 318 can be fixed, for example, mounted to the base 302, or dynamic and mounted to the housing 308. The emitter 316 is a light emitter and the detector 318 can be a light detector.
[0081] The emitters 316 may include laser emitter chips or other light emitting devices, and may include any number of individual emitters (e.g., 8 emitters, 64 emitters, or 128 emitters). The emitters may be arranged in a linear array or laser bar, such as Figure 3 As shown. The emitter 316 can emit light pulses 320 of substantially the same intensity or different intensities, and can be in various waveforms, such as sine, square wave, and sawtooth. The lidar sensor 300 may include one or more optical elements 322 to focus and direct the light passing through the cover 312. One or more optical element structures 322 can be positioned in front of the mirror (not shown) to focus and direct the light passing through the optical element structure. Figure 3 As shown, a single optical element structure 322 is positioned in front of the mirror and connected to the rotating element of the system so that the optical element structure 322 rotates with the mirror. Alternatively or additionally, the optical element structure 322 may include a plurality of such structures (e.g., lenses and / or wave plates). Optionally, the plurality of optical element structures 322 may be arranged in an array on or integral with the shell portion of the housing 308.
[0083] Detector 318 may include a photodetector or an array of photodetectors positioned to receive reflected light pulses 328. Detector 318 may be arranged in a linear array, such as Figure 3 According to various aspects of the present disclosure, the detector 318 includes a plurality of pixels, wherein each pixel includes a Geiger-mode avalanche photodiode for detecting reflections of the light pulse during each of a plurality of detection frames. In other embodiments, the detector 318 includes a passive imager.
[0084] The lidar sensor 300 includes a controller 330 having a processor 332 and a memory 334 to control various components such as the motor 304, the transmitter 316, and the detector 318. The controller 330 also analyzes the data collected by the detector 318 to measure the characteristics of the received light and generate information about the environment outside the vehicle 104. For example, the controller 330 can generate a three-dimensional point cloud based on the data collected by the detector 318. The controller 330 can be integrated with another controller (such as the controller 202 of the SDS102). The lidar sensor 300 also includes a power supply unit 336, which receives power from the vehicle battery 338 and supplies power to the motor 304, the transmitter 316, the detector 318, and the controller 330. The lidar sensor 300 includes an analyzer 330A having elements such as a processor 332 and a non-transitory computer-readable memory 334 containing programming instructions. The programming instructions are configured to enable the system to receive data collected by the light detector 318, analyze the received data to measure characteristics of the received light, and generate information that the connected system can use to make decisions about operating in the environment in which the data was collected. Optionally, the analyzer 330A can be integrated with the lidar sensor 300 as shown, or part or all of it can be external to the lidar sensor and communicatively connected to the lidar sensor via a wired or wireless communication network or link. The analyzer 330A can include a controller 330.
[0085] In the case of a Geiger-mode lidar system, the photodetector 318 fires when a single photon at or near the target wavelength is detected. The times of the photodetector fires (and associated illuminator fires) are accumulated into a histogram. A peak-finding operation is then run on the histogram to obtain the depth of objects that reflect photons at the target wavelength for the spatial region covered by the photodetector(s). That is: a series of individual echoes from the laser fires are obtained over a certain field of view; and an aggregation and peak-finding algorithm is used to determine the depth of the target object as a function of the observed photon's time of flight (which may include at least other signals).
[0087] In conventional systems, there is nothing that forces a Geiger-mode lidar system to use the same number of laser excitations on each pixel to recover the peak of the histogram. In fact, the physical consequence of continuing to accumulate laser excitations is purely that the frustum described by the rotating sensor expands along the axis in question as one continues to do so (that is, the system increases the probability of merging returns from more than one surface). Furthermore, there is nothing that prevents the system from merging returns from adjacent photodetectors - again, this merely adjusts the frustum imaged by a particular lidar point. In fact, in any gapless sensor, each point is effectively the average depth of the nearest surface image along each ray of some frustum. Putting these insights together, a conventional Geiger-mode lidar system is modified according to the present solution to use an aggregation window of variable size as a function of reflectivity or its surrogate, such as echo count, echo noise, observations from related other sensor systems (such as cameras), etc.
[0088] The lidar sensor 300 uses the results output from the photodetector 318 to produce a measured 3D point by accumulating the results from the photodetector(s). A single photodetector is typically not sufficient to produce a depth measurement, so instead, the results from multiple photodetectors are combined in a superpixel. Fig.13 An illustrative technique for generating superpixels is shown.
[0090] Figure 4 and Figure 5 An exemplary lidar sensor 400 is shown. Figure 3 Similar to the lidar sensor 300 of the embodiment of the present invention, the lidar sensor 400 includes a housing 408 having an opening 410 and a cover 412 fixed in the opening 410. The lidar sensor 400 includes: one or more transmitters 416 for transmitting light pulses through the cover 412; and one or more detectors 418 for receiving reflected light pulses passing through the cover 412. According to various aspects of the present disclosure, the transmitters 416 and the detectors 418 are arranged in linear arrays, respectively. The lidar sensor 300 includes a transmitter assembly 424 and a receiver assembly 426, the transmitter assembly including the transmitter 416, and the receiver assembly including the detector 418.
[0091] The transmitter assembly 424 includes a circuit board assembly 428 for controlling the transmitter 416. The circuit board assembly 428 includes a controller 430 having a processor 432 and a memory 434 mounted to a circuit board 435.
[0092] The emitter assembly 424 also includes a plurality of optical components, including a collimator 436 and a transmit optic 438. The collimator 436 focuses and directs the light pulses from each emitter 416 along a transmit (Tx) axis 440 to collectively form a Tx beam, such as Figure 7 As shown. The emission optics 438 are arranged between the collimator 436 and the cover 412 to focus the Tx beam toward a smaller area of the Tx FoV. The emission optics 438 can be a converging lens, such as a cylindrical lens, that focuses the light pulses onto a single axis. The emitter assembly 424 also includes an actuator 442, such as a linear actuator, which is connected to the emission optics 438 and controlled by the controller 430. The actuator 442 translates the emission optics 438 along a transverse axis 444 arranged perpendicular to the Tx axis 440. The actuator 442 provides linear adjustment of the emission optics 438 from a rest position 446 to a fully extended position 448, in which the emission optics 438 does not intersect any Tx axis, and in which the emission optics 438 intersects the Tx axis of the farthest emitter of the linear array of emitters 416. According to aspects of the present disclosure, the actuator 442 can adjust the transmitting optical device 438 based on the rotation speed of the lidar sensor 400. For example, in one embodiment, the lidar sensor 400 rotates at 10 Hz or 600 revolutions per minute (RPM), and the actuator 442 adjusts the transmitting optical device 438 from the static position 446 to the distal position 448 within 100 milliseconds (ms). The actuator 442 can be a linear actuator, such as a voice coil. According to aspects of the present disclosure, the travel or linear adjustment of the actuator 442 is based on the length of the linear array of the emitter 416.
[0093] The receiver assembly 426 includes a detector 418 mounted to a circuit board 450. A controller 430 is connected to the circuit board 450 to receive data from the detector 418. The controller 430 analyzes the data collected by the detector 418 and generates information about the environment surrounding the lidar sensor 400. The receiver assembly 426 also includes one or more detector optics 452. The detector optics 452 may include a collimator to focus and direct received light pulses to each detector 418 along a receive (Rx) axis 454.
[0094] Reference Figure 4 , the transmitter assembly 424 is offset from the receiver assembly 426. When the transmit optical device 438 is translated along the horizontal axis 444, the transmit optical device 438 intersects the Tx axis 440, but does not intersect the Rx axis 454. The Tx FoV and Rx FoV overlap, as shown in FIG. Figure 4However, because transmit optics 438 does not intersect Rx axis 454, lidar sensor 400 can adjust the Tx FoV to track object 510 without adjusting the Rx FoV.
[0095] Figure 6 An exemplary lidar sensor 600 is shown. Similar to lidar sensor 400, lidar sensor 600 transmits light pulses that collectively form a Tx beam 660 within the Tx FoV. Unlike lidar sensor 400, lidar sensor 600 does not include transmit optics 438 for adjusting the Tx FoV.
[0096] Figures 7 to 9 Another exemplary lidar sensor 700 is shown. Similar to the lidar sensor 400, the lidar sensor 700 includes a series of emitters 716 that emit light pulses that collectively form a Tx beam 760 within the TxFoV. Also, similar to the lidar sensor 400, the lidar sensor 700 includes a transmit optics 738 to form an adjusted Tx FoV (Tx FoVADJ). The lidar sensor 700 includes a series of emitters 716 arranged in a linear array, including a far-end emitter 762, a center emitter 764, and a near-end emitter 766. Figures 7 to 9 A comparison between Tx FoV and Tx FoVADJ is shown as transmit optics 738 is translated along horizontal axis 744 .
[0097] Figure 7 Transmit optics 738 are shown adjusted to a distal position 748 to intersect the Tx axis of the distal transmitter 762 and generate a Tx FoVADJ at an upper region 768 of the full or unadjusted Tx FoV. Figure 8 Transmit optics 738 are shown adjusted to a neutral position 770 to intersect the Tx axis of center transmitter 764 and generate a Tx FoVADJ at a center region 772 of the total Tx FoV. Fig. 9 Transmit optics 738 are shown adjusted to a proximal position 774 to intersect the Tx axis of proximal transmitter 766 and generate a Tx FoVADJ at a lower region 776 of the total Tx FoV. Controller 430 may adjust the position of transmit optics 738 so that the adjusted Tx FoV tracks object 710, such as a tire or tire fragment.
[0098] Figure 10 to Figure 11 A comparison of Tx FoV and Rx FoV is shown. Fig.10, the Tx FoV and the Rx FoV are oriented adjacent to each other to illustrate that the two fields of view are the same size, but they overlap in the environment outside the vehicle 104, such as Fig.11 shown. Fig.11 Also shown is the adjusted Tx FoV after being adjusted by the transmit optics 438. As the transmit optics 438 translates along the horizontal axis 444, the adjusted Tx FoV shifts to overlap different areas of the Rx FoV. For example, and returning to reference Figures 6 to 9 , when the launch optical device 738 is at the distal position 748 ( Figure 7 ), the lidar sensor 700 generates a Tx FoVADJ at the upper region 768 of the Rx FoV. Subsequently, when the transmit optical device 738 is adjusted to the middle position 770 ( Figure 8 ), the lidar sensor 700 generates a Tx FoV at the center region 772 of the Rx FoV. When the transmit optics 738 is adjusted to the near end position 774 ( Fig. 9 ), the lidar sensor 700 generates a Tx FoVADJ at the lower region 776 of the Rx FoV. Once the transmit optics 738 are adjusted to a stationary position ( Figure 4 ), the Tx FoV returns to its full range and overlaps with the Rx FoV, as shown in Fig.11 as shown on the right.
[0099] As the adjusted Tx FoV shifts between different areas, the Rx FoV remains constant. This allows for a longer range in a wider FoV. In general, the wider the FoV, the less ability the optics have to scan a small target area. In other words, a wider FoV means less light is available to scan the object, so to get a clear image the width is reduced. However, by adjusting the Tx FoV without adjusting the Rx FoV, more light can be focused on a small target area to identify unknown objects such as tires or tire fragments 710 without minimizing the width of the FoV.
[0100] Reference Fig.12 , a flow chart depicting a method for adjusting Tx FoV is shown according to one or more embodiments, and the method is generally indicated by reference numeral 600. According to one or more embodiments, the method 600 is implemented using software code executed by the controller 430. Although the flow chart is shown as having multiple sequential steps, one or more steps may be omitted and / or performed in another manner without departing from the scope and intent of the present disclosure.
[0101] At step 602, the controller 430 controls the lidar sensor 400 to scan a 360 degree field of view about the vehicle 104 at the full Tx FoV. The transmitting optics 438 is located at the stationary position 446 ( Figure 4 ) and does not adjust the Tx FoV. Controller 430 analyzes the data from transmitter 416 to observe any environmental changes. At step 604, sensor 400 determines whether an unknown object 710, such as a tire or tire debris, is detected outside the vehicle and within the Rx FoV. If no such object is detected, controller 430 returns to step 602. If controller 430 detects an unknown object 710 at step 604, it proceeds to step 606.
[0102] At step 606, the controller 430 controls the lidar sensor 400 to perform another scan or series of scans while scanning the Tx FoV. During step 606, the controller 430 scans the Tx FoV by controlling the actuator 442 to translate the transmit optics 438 at a predetermined rate through a predetermined range, such as between the proximal position 774 and the distal position 748. In one embodiment, the controller 430 controls the transmit optics 438 to translate through its entire range of 10 mm in 100 ms, or at 0.1 m / s.
[0103] At step 608, the controller 430 analyzes the scan data to determine the location of the unknown object 710. If the controller 430 determines the location of the unknown object 710, it proceeds to step 610. If the controller 430 does not determine the location of the unknown object 710, it returns to step 604.
[0104] At step 610, after determining the location of the unknown object 710, the controller 430 controls the lidar sensor 400 to track the unknown object 710 by performing another scan or a series of scans with the transmit optics 438 focused on the unknown object 710. During this step, the transmit optics 438 is translated to a position corresponding to the area within the FoV where the unknown object 710 is located.
[0105] At step 612, the controller 430 analyzes the focused scan data to identify the unknown object 710. If the controller 430 cannot identify the unknown object 710, it returns to step 610. Once the controller 430 identifies the unknown object 710, it proceeds to step 614 and returns the emission optics 438 to the rest position, and then returns to step 602.
[0106] By focusing the Tx FoV on the unknown object 710, the lidar sensor 700 can quickly identify the unknown object 710 by projecting more light onto the unknown object, thereby collecting more reflected light from the area of interest within the total Tx FoV. Compared with other lidar systems that do not adjust the Tx FoV, such as Figure 6 This approach improves the responsiveness of SDS 102 in identifying and responding to unknown objects 710 compared to the lidar sensor 600 shown in FIG. 1 . The method for adjusting the Tx FoV may be implemented using one or more controllers, such as controller 430 or Fig.29 Computer system 1300 is shown.
[0108] The lidar sensor 300 uses the results output from the photodetector 318 to produce a measured 3D point by accumulating the results from the photodetector(s). A single photodetector is typically not sufficient to produce a depth measurement, so instead, the results from multiple photodetectors are combined in a superpixel. Figure 3 and Fig.13 An illustrative technique for generating superpixels is shown.
[0110] exist Figure 3 and Fig.13 In, the photodetector array includes photodetectors arranged in a grid pattern. The results p1, p2, ..., px from the photodetectors can be represented in a grid 550 defined by a plurality of cells, where each cell 552 is associated with a corresponding one of the photodetectors and x is an integer equal to the total number of photodetectors in the array. The cells 552 of the grid 550 can be arranged in the same pattern as the photodetectors, for example, a 256 x 256 grid pattern. Each result is also referred to herein as a pixel of the lidar image. The pixels p1, p2, ..., px from the photodetectors can be naturally accumulated in a super-cell manner to produce a set of 3D points. The super-cell has a size of W x W, where W is an integer. In Fig.13 , each supercell is 2 cells x 6 cells. The 3D point associated with each supercell 204 is derived by combining the corresponding six pixels with each other to obtain superpixels SP1, SP2, ..., SPy. The first superpixel SP1 can be defined by the following mathematical equation.
[0111] SP1=f(P1, P2, P3, P4, P5, P6, Px+1, Px+2, Px+3, Px+5, Px+6)
[0112] Other super pixel SP 2 , ..., SP ywill be defined by similar mathematical equations, as should be understood. The mechanism for accumulating pixels is specific to the individual lidar sensor being designed and may vary depending on the application. For example, simple addition or convolution methods may be used for pixel aggregation.
[0113] Fig.13 A diagram that helps to understand the convolution method is provided in FIG. The convolution method can employ at least one convolution filter (or kernel) 552 that operates on the lidar image 550 and calculates the feature F 1 、F 2 、F 3 、F 4 , ..., F 12 . In the case of using multiple computing kernels, each computing kernel extracts different features from the lidar image. The computing kernel has a size of 2x6. The image has a size of 12x12. The stride is 6. Therefore, the features generated by the computing kernel 552 are defined by the following mathematical equations (1)-(4).
[0114] F1=f(P1, P2, P3, P4, P5, P6, P13, P14, P15, P16, P17, P18) (1)
[0115] F2=f(P7, P8, P9, P10, P11, P12, P19, P20, P21, P22, P23, P24) (2)
[0116] F3=f(P25, P26, P27, P28, P29, P30, P37, P38, P39, P40, P41, P42) (3)
[0117] F4=f(P31, P32, P33, P34, P35, P36, P43, P44, P45, P46, P47, P48) (4) ...
[0119] F 1 、F 2 、F 3 and F 4 represents features generated by processor (or computing core) 552. These features are also referred to herein as superpixels. 1 、p 2 , ..., p 144 Respectively represent the data from the LiDAR system (for example, Figure 3The convolution operation generally involves combining pixel values to generate superpixels, as demonstrated by mathematical equations (1)-(4). The output of processor (or computational kernel) 552 is arranged in a grid 554 of features (or superpixels), such as Fig.14e As shown. Grid 554 is also referred to herein as a feature image. Features may include, but are not limited to, depth, intensity, noise, confidence, or other features associated with a point cloud. The present solution is not limited to the details of FIG. 14 , and other convolution oversampling techniques may also be used.
[0120] The lidar sensor 300 may include a Geiger-mode lidar system. In a Geiger-mode lidar system, photodetectors are excited when single photons at or near a target wavelength are detected, and the times at which they are excited (and the associated illuminator excited) are accumulated into a histogram. A peak finding operation is then run on the histogram to obtain the depth of an object that reflects photons at the target wavelength for the spatial region covered by the photodetectors. That is: a series of individual echoes from the laser excitation are obtained over a certain field of view (FoV); and an aggregation and peak finding algorithm is used to determine the depth as a function of the time of flight (ToF) of the observed photons (which may include at least other signals).
[0122] There is nothing that forces the lidar sensor 300 to use the same number of laser excitations on each pixel to recover the peak of the histogram. In fact, the physical consequence of continuously accumulating laser excitations is purely that the frustum described by the rotating sensor expands along the axis in question. That is, the probability of merging returns from more than one surface into a superpixel is increased. In addition, there is nothing that prevents the lidar sensor 300 from merging returns from adjacent photodetectors - again, this only adjusts the frustum imaged by a particular lidar point. In any gapless sensor, each point is actually the average depth of the nearest surface along the ray imaged for each ray of a certain frustum. Putting these insights together, the lidar sensor 300 is configured to use a variable size aggregation window as a function of reflectivity or its surrogate (such as echo count, echo noise, observations from related other sensor systems (such as cameras), etc.).
[0124] The lidar sensor 300 is configured to ensure that integrated measurements are obtained from one target or one surface. This is achieved using a spatiotemporal coherence metric that exploits the joint spatiotemporal diversity in the range measurements obtained by the lidar sensor 300. In other words, multiple measurements are obtained in the spatial axis (integrating multiple pixels into superpixels) and the temporal axis (integrating multiple pulses). If the measurements are obtained from the same surface, the statistical properties of the measurements should be similar, i.e., coherent in the temporal and spatial domains.
[0126] For example, a lidar system includes a photodetector and a processor configured to combine pixels px into superpixels SPy using a processor (or computing kernel) having a size of 2x6. Table 1 below shows twelve ToF measurements or timestamps obtained for five pulses emitted from the lidar system. The example of Table 1 shows the measurement results of a single superpixel integrating five pulses and twelve pixel measurements. However, the object is not large enough to span all given pulses. It should be noted that the lidar system here rotates as the pulses are sent. Therefore, when the lidar system rotates, the span of the object across the horizontal direction may be large enough to intercept and reflect all five pulses.
[0127] Table 1
[0128] In Table 1, each cell includes a ToF measurement or timestamp associated with a detection event reported by a superpixel associated with a particular one of the five transmit pulses. Fig.15 As shown, each superpixel reports a single detection event 540, so that for each pulse P1, P2, P3, P4, P5 there are a total of twelve detection events 540. For pulse P1, the timestamps are: 684 for pixel P1, 559 for pixel P2, 629 for pixel P3, 192 for pixel P4, 835 for pixel P5, 763 for pixel P6, 707 for pixel P7, 359 for pixel P8, 9 for pixel P9, 723 for pixel P10, 277 for pixel P11, and 754 for pixel P12. The actual situation imaged by pulse P1 is just noise (no target object is present in the superpixel field of view), as shown in FIG. Fig.15 Similarly, for each of the other pulses P2, P3, P4, and P5, in Tables 1 and Fig.15 There are twelve detection timestamps in . The first three pulses P1, P2, P3 capture only noise measurements (no target in the field of view), while the last two pulses P4, P5 capture target measurements.
[0130] In a system with fixed resolution, a fixed number of pulses are integrated. However, if the target object is smaller than the span of the integration time, the quality of the overall detection will deteriorate. Fig.15 In the example provided in , if all measurements from five pulses P1, P2, P3, P4, P5 are integrated to estimate the range of a target object, the measurements of the first three pulses P1, P2, P3 will negatively affect the quality of the estimate because they do not contain information about the range of the target object.
[0132] The essence of the present solution is to provide an implementation system and method for selecting a subset of light pulses, where the photodetector results for the subset of light pulses should be integrated together to perform a distance estimation for a target object. This selection is based on a spatiotemporal statistical coherence metric. There are several ways of measuring spatiotemporal statistical coherence, for example by using distribution comparison metrics, time-of-flight (ToF) temporal statistics, and / or differences / advantages of the solutions.
[0133] Distribution comparison metrics:
[0134] The distribution comparison metric is the Kullback-Leibler (KL) divergence. For each pulse, the system obtains a new probability distribution (normalized histogram) by integrating the detections of the pixels in the superpixel with a time bin step of 100. In other words, the twelve measurements are assigned to bins of width 100. As mentioned above, the timestamps of the measurements associated with pulse P1 in Table 1 are [684, 559, 629, 192, 835, 763, 707, 359, 9, 723, 277, 754]. The timestamps can be assigned to bins as follows: 684 is placed in bin b600-700; 559 is placed in bin b600-600, and so on. The histogram H of the timestamps is determined by the following expression: 1 detection (9) for b0-100; 1 detection (192) for b100-200; 1 detection (277) for b200-300; 1 detection (359) for b300-400; 0 detection for b400-500; 1 detection (559) for b500-600; 2 detections (629, 684) for b600-700; 4 detections (763, 707, 723, 754) for b700-800; and 835 for b800-900. Therefore, the histogram H can be defined by the following expression:
[0136] H=[1, 1, 1, 1, 0, 1, 2, 4, l].
[0137] The histogram H can be transformed into a probability distribution (or probability mass function (PMF)) by dividing the counts by the total number of counts (12), as shown in the following expression:
[0139] PMF1=[1 / 12, 1 / 12, 1 / 12, 1 / 12, 0 / 12, 1 / 12, 2 / 12, 4 / 12, 1 / 12]=
[0140] [0.0833, 0.0833, 0.0833, 0.0833, 0.0, 0.0833, 0.1667, 0.3333, 0.0833].
[0141] Following the same procedure (binning, histogram creation, and histogram transformation), the following additional PMFs were obtained for pulses P2, P3, P4, P5.
[0142] PMF2: [0.1667, 0.0833, 0.0, 0.1667, 0.1667, 0.1667, 0.0833, 0.0833, 0.0833]
[0143] PMF3: [0.1, 0.1, 0.0, 0.0, 0.0, 0.1, 0.2, 0.3, 0.2]
[0144] PMF4: [0.3333, 0.6667, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]
[0145] PMF5: [0.1667, 0.8333, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]
[0147] Next, the KL divergence metric is used to detect changes in distributions. The KL metric generally provides a way to compare distributions. Differences in distributions can be used as a signal of a distribution change, which indicates the beginning of a different integration span. The KL divergence metric is defined by the following mathematical formula (5).
[0148]
[0149] Where Q is the current distribution and P is the distribution to be added. The KL divergence is calculated as follows.
[0150] KL[P=PMF2, Q=PMF1]=0.18
[0151] KL[P=PMF3, Q=PMF2]=0.65
[0152] KL[P=PMF4, Q=PMF3]=1.67
[0153] KL[P=PMF5, Q=PMF4]=0.07
[0155] The value of the KL divergence KL[P=PMF4, Q=PMF3] is higher than the other KL divergences KL[P=PMF2, Q=PMF1], KL[P=PMF3, Q=PMF2], KL[P=PMF5, Q=PMF4]. This indicates a distribution change between the detection of the two pulses P3 and P4. As a result, the integration of pulses stops until pulse P3, and the pulse integration starts (or restarts) at pulse P4.
[0156] ToF time statistics:
[0157] ToF time statistics may include but are not limited to ToF time expansion. ToF time expansion can be obtained by calculating the variance of the sorted ToF differences. First, the ToF measurement results or timestamps of each superpixel are sorted from the smallest value to the largest value, as shown in the following expression.
[0158] P1: [9, 192, 277, 359, 559, 629, 684, 707, 723, 754, 763, 835]
[0159] P2: [70, 87, 174, 314, 396, 472, 486, 551, 599, 600, 705, 804]
[0160] P3: [72, 115, 537, 600, 677, 709, 755, 777, 845, 849, 916, 976]
[0161] P4: [7, 99, 99, 99, 100, 100, 100, 100, 101, 101, 101, 102]
[0162] P5: [98, 99, 100, 100, 100, 100, 100, 100, 101, 101, 101, 102]
[0164] Next, a calculation is performed to calculate a set D of difference timestamp values for each pulse. The set Dp1 of difference timestamp values for the first pulse P1 is calculated as follows, where each different timestamp value is the difference between two consecutive timestamps.
[0165] Dp1: [192-9, 277-192, 359-277, 559-359, 629-559, 684-629, 707-684, 723- 707, 754-723, 763-754, 835-763] = [183, 85, 82, 200, 70, 55, 23, 16, 31, 9, 72]
[0167] Similar calculations are performed to obtain the following set Dp3, Dp4, Dp5 of difference timestamp values for the other four pulses P2, P3, P4, P5.
[0168] Dp2=[17, 87, 140, 82, 76, 14, 65, 48, 1, 105, 99]
[0169] Dp3=[43, 422, 63, 77, 32, 46, 22, 68, 4, 67, 60]
[0170] Dp4=[2,0,0,1,0,0,0,1,0,0,l]
[0171] Dp5=[1,1,0,0,0,0,0,1,0,0,l]
[0173] Next, the system measures the variance of the difference timestamp values in each set Dp1, Dp2, Dp3, Dp4, and Dp5. Each variance is calculated by the following steps: (i) calculating the average of a set of different timestamp values; (ii) subtracting the average from each difference timestamp value in the set; (iii) taking the square of the value obtained in step (ii); (iv) adding all the values obtained in (iii); and (v) dividing the value obtained in (iv) by the total number of difference timestamp values in the set. For example, the set D p1 Variance V Dp1 Determine as follows.
[0174] (i)Mean=(183+85+82+200+70+55+23+16+31+9+72) / 11=75.09
[0175] (ii) (183-75.09=107.91), (85-75.09-9.91), (82-75.09=6.91), (200-75.09=124.91), (70-75.09=-5.09), (55-75.09=-20.09), (23-75.09=-52.09), (16-75.09=-59.09), (31-75.09=-44.09), (9-75.09=-66.09), (72-75.09=-3.09)
[0176] (iii)(107.91 2 =11,644.5681),(9.91 2 =98.2081), (6.91 2 =47.7481), 124.91 2 =15602.5081), (-5.09 2 =25.9081), (-20.09 2 =403.6081), (-52.09 2 =2713.3681), (-59.09 2 =3491.6281), (-44.09 2 =1943.9281), (-66.09 2 =4367.8881), (-3.09 2 =9.5481)
[0177] (iv)11644.5681+98.2081+47.7481+15602.5081+25.9081+403.6081+2713.3681+3491.6281+1943.9281+4367.8881+9.5481=40,348.9091
[0178] (v)40,348.9091 / 11=3668.08
[0179] The following variances of other sets Dp2, Dp3, Dp4, Dp5 can be calculated in a similar manner.
[0180] VDp2=1684.74
[0181] VDp3=11990.15
[0182] VDp4=0.43
[0183] VDp5=0.23
[0185] It is clear from these variances that the time spread measured for pulses P1-P3 is very high compared to the time spread measured for pulses P4 and P5. This can be considered an indication that pulses P4 and P5 are integrating measurements of different target objects captured by pulses P1-P3.
[0187] It is important to point out that the second method is computationally cheaper since it does not require the creation of a histogram as in the first method (although a very coarse histogram can be used for the KL divergence). However, the second method cannot distinguish between measurements obtained from two targets at two different ranges, which may produce similar temporal spreads (local measurements of the two ranges). In this case, a joint mean and variance measurement can be considered.
[0188] Spatiotemporal coherence can be used to reject noisy measurements rather than using it as a stopping criterion for measurement integration. For example, the third pulse P3 of ten pulses may have incoherent statistical measurements due to interference. Instead of stopping integration at two pulses, the system can discard the incoherent measurements obtained by the third pulse and resume integration.
[0189] In addition to the stopping criteria proposed above, various "stopping criteria" are possible. Some examples of new stopping criteria: accumulating multiples of a fixed number of laser shots until the confidence of the detected peak (or other statistical measure of detection probability) exceeds a given threshold; accumulating multiples of a fixed number of laser shots until the confidence of the detected peak begins to decrease by integrating more pulses; accumulating laser shots in a sliding window manner until either of the above two criteria is met; and / or accumulating multiples of a fixed number of laser shots constrained by the neighbor estimate aggregation window size and reflectivity estimate, so that the lidar sensor emits a quadtree in the 2D range image plane.
[0190] Note that there is no requirement to use the same aggregation function or values on different axes. Thus, the system can (if chosen at design time) emit points corresponding to square or non-square solid angles of the entire sensor frustum. In particular, sliding window aggregation schemes have no practical correspondence to "image plane pixels" since they can operate on much smaller units excited by a single laser. In any range-finding imaging system, representing observations as points is a gross simplification, and a better first-order approximation may be oriented planes or surface elements.
[0192] Nor is it required to use reflectivity etc. to guide aggregation. Other properties of the intermediate measurements could be used instead. For example, depth recovery itself could be used to guide continued aggregation - moving the decimation of surfaces close to the vehicle into the lidar accumulator itself. Such an implementation would be useful to establish a constant point density in world space (ranging sensors typically suffer from "data overload" close to the sensor making it difficult to obtain sufficient point density far from the sensor).
[0193] The present solution can be designed to effectively compromise among the software properties of a depth imaging system, such as reflectivity performance, probability of detection, range, and resolution. These axes are typically fixed at design time and as a function of the hardware, but the present solution can move these decisions into the software domain. Furthermore, use of the present solution allows the generation of point clouds with constant density for any volume of world space as part of a LiDAR system (which is highly unconventional), or allows the generation of point clouds or surface clouds of varying density as a function of other imaging properties such as estimated surface reflectivity.
[0195] The spatiotemporal coherence metric has two advantages in particular: it provides immunity to noise and interference, and it provides a robust integration stopping criterion that supports variable resolution architectures.
[0196] AVs use sensors for situational awareness. Sensors may be part of a self-driving system (SDS) in an AV, which may include camera(s), lidar (light detection and ranging) devices, inertial measurement units (IMUs), and / or the like. Sensors are used to capture and analyze the scene around the AV. Scene analysis is performed to detect objects including static objects (such as fixed structures) and dynamic objects (such as pedestrians and other vehicles). Data from sensors may also be used to detect conditions such as road markings, lane curvature, traffic lights, and traffic signs. Sometimes, a scene representation (such as point cloud(s) obtained from the lidar device(s) of the AV) may be combined with images from a camera to gain further understanding of the scene or situation around the AV.
[0197] A laser radar device operating on an AV includes a transceiver device, which includes a transmitter component and a receiver component. The transmitter can transmit light signals, and the receiver can receive and process the received light signals.
[0198] In some current implementations, LiDAR devices may use a fixed pixel size (angular resolution) where each point integrates a fixed amount of raw data. In some applications, it is desirable to use a more intelligent data integration method that adapts to the characteristics of the target to improve detection probability and data quality (range and intensity accuracy and precision). Fig.19 As shown, target objects (such as vehicles) are larger than the default pixel size, but using a larger pixel size creates the risk of combining data from areas with large variations in range and intensity, which can cause blurring and other distortions in the data.
[0199] Initially, one can consider the choice of input data, for example, deciding which data level to use for spatial processing.
[0200] Option 1: Raw Data: Robustness, high computational cost, and requires firmware modifications.
[0201] Option 2: Waveform: Robust and requires firmware modification.
[0202] Option 3: LiDAR data: Computationally cheap and can be implemented in firmware or software as a post-processing step.
[0203] LiDAR data presents a good option because it has a relatively low upfront cost and can be implemented in both firmware and software.
[0205] According to some aspects, the following methods can be implemented for each of the above options to perform spatial processing. For options 1-2 listed above, there are two methods of using data for spatial processing. Regarding option 1 involving raw data, the first method includes: fully processing the raw data with a fixed superpixel size to obtain the range and signal strength of each superpixel; and reprocessing the raw data with a variable superpixel size based on the range and intensity correlation with neighboring pixels. The second method includes: using a fixed superpixel size to calculate the total intensity of each pixel; and processing the raw data with a variable superpixel size based on the total intensity correlation with neighboring pixels. The signal intensity has two independent fields (signal intensity and noise intensity) and one correlated field (total intensity = signal plus noise).
[0206] With respect to Option 2 involving waveforms, a first approach includes fully processing a waveform with a fixed superpixel size to obtain the range and signal strength of each superpixel, and regenerating and processing a waveform with a variable superpixel size based on range and strength correlations with neighboring pixels. A second approach includes fully processing a histogram with a fixed superpixel size to obtain approximate range and signal strength (prior to waveform analysis), and using the approximate range and strength to integrate the waveform with a variable superpixel size based on range and strength correlations.
[0207] With respect to Option 3 involving LiDAR data, the present solution implements a new approach that may generally include: aggregating neighboring pixels in a LiDAR data frame based on range and intensity correlation; and recalculating new ranges, signal strengths, noise strengths, and confidence levels based on the data fields reported for each superpixel.
[0208] More specifically, the new method may include: arranging pixels in a grid (wherein the pixels include result values generated from processing waveforms produced by a photodetector of a laser radar system); and identifying a region of interest (ROI) in the grid based on correlations between pixels. Correlations may include, but are not limited to, correlations between range values associated with pixels and / or correlations between intensity values associated with pixels. In some cases, the ROI may be identified by the following steps: obtaining a kernel size; and defining a region of interest in the grid using the kernel size. The kernel size may be variable. The kernel size may be obtained by the following steps: locating a pixel that is a closest neighbor of a POI in the grid, at least in terms of range; and defining the kernel size based on the position of the closest neighbor in the grid. Alternatively, the kernel size may be obtained by the following steps: obtaining a reference kernel size; identifying a region in the grid using the reference kernel size; identifying a center pixel of the region; calculating a score for each pixel using the result values associated with each pixel in the region (wherein the score indicates a correlation between the result values associated with the pixel and the center pixel); selecting a pixel based on the score; and defining the kernel size based on the position of the selected pixel in the grid. The score may be a function of range, intensity, and / or noise.
[0209] The size and / or position of the ROI in the grid may optionally be adjusted to maximize the likelihood that the ROI contains a greater number of pixels associated with the object. This adjustment may be achieved by: identifying a point of interest (POI) in the ROI; identifying a pixel that is the closest neighboring pixel to the POI, at least in terms of range; using the centroid of the closest neighboring pixels to obtain a likelihood that the POI is associated with an edge point or a corner point on the surface of the object; and adjusting the size and / or position of the ROI based on the likelihood that the POI is associated with an edge point or a corner point on the surface of the object. The POI may be a central pixel of the ROI.
[0210] One or more pixels in the ROI may optionally be disqualified from aggregation with other pixels in the ROI. The disqualification may be based on how far the pixel is from the POI and / or surface in one or more dimensions. The dimensions may include, but are not limited to, range, intensity, noise, and confidence. The result values associated with the remaining (or qualified) pixels within the ROI are combined with each other to produce an eigenvalue. A superpixel is generated, the superpixel having a value set as the eigenvalue.
[0211] The above operations of the new method can be repeated iteratively to generate other superpixels. The ROI used to generate a first superpixel in a first iteration can have a size and / or shape that is the same or different from the size and / or shape of the ROI used to generate a second superpixel in another iteration. In some cases, superpixels can be used to control the operation of autonomous vehicles and / or other robotic devices (e.g., articulated arms or electrosurgical instruments).
[0212] like Figure 2 and Fig.13 The illustrated lidar sensor 300 uses the waveform output from the photodetector 226 to generate results p1, p2, px. Each result p1, p2, ..., px has a value associated with it. These values include, but are not limited to, range values, intensity values, noise values, confidence values, and / or test values. The results can be aggregated to generate superpixels. Fig.13 An illustrative primitive technique for generating superpixels is discussed.
[0213] In a Geiger-mode lidar system, the sensor includes an avalanche detector or (photodiode) that is configured to generate an electrical pulse of a given amplitude in response to the absorption of photons of the same or similar wavelength as the emitted light signal. Subsequently, a histogram can be combined over many tests, and the position of the object surface can be estimated from the peak of the histogram. As used herein, the term "test" refers to each measurement attempt. A measurement attempt consists of sending a pulse and recording the detection time. The test is associated with a measurement result, but not necessarily with a pulse. By grouping detections from multiple detectors, there can be multiple tests from a single pulse. Each detector output is a measurement result. However, the accuracy of the histogram is essentially limited by the width of the bin. Therefore, the present solution involves post-processing operations for updating range, intensity, noise and / or confidence values. Fig. 20 Post-processing operations are discussed in detail.
[0214] Now refer to Fig.13 21, the results p1, p2, ..., px can be represented in a grid 550 defined by a plurality of cells 552. Each result is also referred to herein as a pixel of the lidar image. The pixels p1, p2, ..., px can be naturally aggregated in a super-cell to super-cell manner to produce a set of 3D points. The super-cell has a size of Q x Z, where Q and Z are integers. Fig.13 , each supercell is 2 cells x 6 cells. The 3D point associated with each supercell 554 is derived by combining the corresponding six pixels with each other to obtain superpixels SP1, SP2, ..., SPy. The first superpixel SP1 can be defined by the following mathematical equation (1).
[0215] SP1=f(P1, P2, P3, P4, P5, P6, Px+1, Px+2, Px+3, Px+5, Px+6)(1)
[0216] As should be understood, each of the other superpixels SP2, ..., SPy will be defined by similar mathematical equations. The mechanism for aggregating pixels is specific to the individual LiDAR system being designed and may vary depending on the application. For example, simple addition may be employed for pixel aggregation.
[0218] Fig.13 There are certain disadvantages to the method of using superpixels. Since this technique uses a fixed pixel size (angular resolution), where each superpixel integrates a fixed number of pixels, there is a risk of combining pixels from areas with large range and intensity variations, which can cause blurring and other distortions in the resulting LiDAR image including the superpixels. The present solution addresses these disadvantages by implementing a new method for generating superpixels. Fig. 20 Discuss this new approach.
[0219] Fig. 20 21 provides a flow chart of an illustrative method 650 for generating and / or using superpixels. The method 650 may be performed in whole or in part by a lidar system (e.g., Figure 3 A processor (e.g., Figure 3 processor 222) to execute.
[0221] Method 650 begins at step 652 and proceeds to step 654, where a light pulse is emitted from the laser radar system. The light may be reflected from the object and returned to the laser radar system. As shown in step 656, the reflected light may be received by a photodetector. Subsequently, as shown in step 658, the waveform output from the photodetector is processed to generate results. Each result has a value associated with it. These values may include, but are not limited to, range values, intensity values, noise values, confidence values, count values, and test values. In step 660, the results are arranged in a grid of cells. Fig.21a An illustrative grid 500 is shown in FIG. 5 , wherein results p1 , p2 , . . . , p144 are respectively assigned to cells 502 thereof.
[0223] The grid is then used to generate superpixels according to the novel approach of the present solution. This novel approach may employ at least one filter (or kernel) that operates on the grid and computes features. In the case of employing multiple computer cores, each computing core extracts different features from the grid.
[0225] Therefore, in Fig. 20In step 662, the processor obtains the kernel size and the stride. The kernel size may be a predefined fixed value or a variable value. The kernel size may be defined as Q units multiplied by Z units. Both Q and Z are integers that may be the same or different from each other. When the inherent angular resolution is equal in azimuth and elevation, the kernel size is selected so that Q and Z are equal to each other. For example, the kernel size is selected as Fig.21a Three cells times three cells (3x3) as shown, five cells times five cells (5x5) (not shown), or seven cells times seven cells (7x7) (not shown). The present solution is not limited to the details of this example. Even numbers can be used for Q and Z. However, in some applications, odd numbers may be desirable for Q and Z because they allow for the center pixel of a POI to be used as described below. The larger the kernel size, the better the resulting features, but the higher the computational cost. The span S can be a predefined fixed value or a variable value, where S is an integer (e.g., 1 or 6).
[0227] It may be more robust to use a fixed Q x Z kernel to locate the nearest neighboring pixels rather than defining a search space. The nearest neighboring pixels may be located in Cartesian space and / or pixel space. Pixel space is similar to spherical coordinates (range, azimuth, elevation). The nearest neighbor can be found by including points less than the range tolerance and minimizing the angle between the POI and the neighbor. Therefore, in the case of variable kernel size and / or span, the kernel size and / or span can be dynamically determined or otherwise obtained based on the nearest neighbor method. The nearest neighbor method may include, for example, the following steps: obtaining a reference kernel size; identifying a POI using the kernel size (where the POI is the center pixel of the area defined by the kernel size); calculating a score A for each pixel using a value associated therewith (e.g., results p1, p2, ..., p144); selecting a given number (e.g., twelve) of nearest neighbor pixels of the POI based on the score; defining a new kernel size based on the selected nearest neighbor pixels; and / or selecting a span based on the score and / or the new kernel size. The score A indicates the degree to which pixels are related to each other. Each score A can be defined by the following mathematical equation (2) or (3).
[0228] A=f(R,I,N,C,T,K) (2)
[0229] A=f(w1*R, w2*I, W3*N, w4*C, w5*T, w6*K) (3)
[0230] Wherein, R represents range, I represents intensity, N represents noise, C represents confidence, T represents test, K represents count quantity, and w1, ..., w6 represent weights respectively. The present solution is not limited to mathematical equations (2) and (3). The score A can be a function of any combination of one or more of the listed values (i.e., R, I, N, C and / or T). The present solution is not limited to this specific nearest neighbor method. Other nearest neighbor methods can be used here. The resulting kernel size can have Q and Z values that are the same or different from each other.
[0231] Once the kernel size and stride have been obtained, method step 650 continues to step 664 where a ROI is identified in the grid. Fig.21a An illustrative ROI 804 is shown in FIG. ROI 804 includes cells within the region of a grid bounded by a kernel search window having a kernel size. The ROI is shown as having a square shape because its length and width are the same. The present solution is not limited in this respect. The ROI may have other shapes with different lengths and widths, such as a grid with different lengths and widths. Fig.21a The linear shape shown by the dashed line 850.
[0233] Next, in step 666, a POI in the ROI is identified. The POI may include, but is not limited to, the center pixel of the ROI. Figure 21b As shown, POI 806 includes pixel p14 because it is the center pixel of ROI 804. The present solution is not limited in this respect. For example, if the confidence of the center pixel is above the confidence threshold, it means that the reported data field has sufficient certainty or accuracy, and the value of selecting an alternative POI is zero. However, if the confidence of the center pixel is below the confidence threshold, the probability that the reported range is correct will decrease, which will cause the spatial processing to fail completely. In this case, the center pixel POI range can be replaced by a representative range from the ROI in some cases. The method first includes: finding the closest neighbor in pixel space with a confidence greater than the confidence threshold; and replacing the range of the center pixel POI with this range. No other POI fields are modified. Second, the system performs spatial processing as usual; tracks the quadrant positions of the relevant pixels included in the ROI; and keeps the final result if and only if there are relevant pixels in at least N quadrants of the ROI (most conservatively, all four quadrants). Otherwise, ignore the point and continue to the next pixel in your span.
[0235] The size and / or position of the ROI may optionally be adjusted as shown in blocks 668 to 670. By shifting the position of the ROI in the grid, the system may cause the kernel to contain a greater number of pixels belonging to the same target object (e.g., Fig.21aThe likelihood of pixels p14) or a larger number of pixels being well associated with the POI is maximized.
[0237] The centroid of the nearest neighbor pixels relative to the POI can indicate where the POI is located on the surface of the object. If the centroid is offset upward, downward, left, or right relative to the POI, the POI may be an edge point of the object. Conversely, if the centroid is biased toward the corner of the POI or ROI, the POI may be a corner point. This information can be used to adjust (e.g., expand / increase or shrink / decrease) the kernel size in one or more directions. For example, if the POI is considered to be an edge point of the object, the kernel size and / or ROI position in the grid are changed so that the POI is located at the edge of the ROI instead of the center of the ROI. If the POI is considered to be a corner point at the bottom left of the target object, the kernel size and / or ROI position in the grid are changed so that the POI is located at the bottom left corner of the ROI instead of the center of the ROI. Similarly, if the POI is considered to be a corner point at the top left of the target object, the kernel size and / or ROI position in the grid are changed so that the POI is located at the top left corner of the ROI instead of the center of the ROI. The present solution is not limited in this respect.
[0239] The kernel size can be adjusted to expand the ROI in both the Q and Z directions. For example, the kernel size is expanded from three cells by three cells (3x3) to four cells by four cells (4x4), as shown in Fig.21a and Fig.21c In practice, ROI 804 is extended in both the Q and Z directions to form ROI 804'. The present solution is not limited in this respect. The kernel size may additionally or alternatively be limited only in the direction of Fig. 22 Q direction as shown or only along Fig.23 The position or location of the ROI in the grid may alternatively or additionally be changed. For example, Fig.24 As shown, the 3x3 ROI is shifted from a first position 850 (i.e., one unit along the Q direction and one unit along the Z direction) to a second position 852. The present solution is not limited in this respect. The position of the ROI can be shifted in one or both directions by any number of units selected according to a given application.
[0241] Return to reference Fig. 20, method step 650 continues to step 672, wherein a range tolerance, an intensity tolerance, a noise tolerance, and / or a confidence threshold are obtained from a database. These values may be preconfigured values. In step 674, one or more pixels in the ROI may be disqualified from aggregation based on the range tolerance, the intensity tolerance, the noise tolerance, and / or the confidence threshold. These tolerances are used to exclude points that are too far from the POI in one or more dimensions from aggregation. For example, if the range value of the POI is 10 and the range tolerance is ±1, the system determines whether the range value of another pixel in the ROI is between 9 and 11. If so, the other pixel is eligible for aggregation. Otherwise, the pixel is disqualified from aggregation. The present solution is not limited to the details of this example. If the lidar system reports multiple echoes per pixel, all echoes in all pixels of the ROI must be checked for (no) eligibility or (not) suitability for aggregation.
[0243] In some cases, a pixel may be considered a qualified pixel when (i) its associated range, intensity, and / or noise value falls within the tolerance and / or (ii) its associated confidence value is equal to or greater than a confidence threshold. A pixel may be considered a non-qualified pixel when (i) its associated range, intensity, and / or noise value falls outside the tolerance and / or (ii) its associated confidence value is less than a confidence threshold. For example, Fig.21d As shown, pixels p3, p37, and p40 of ROI 804' are considered ineligible (or unsuitable) pixels 808 for aggregation. The present solution is not limited in this respect.
[0245] In some cases, the surface normal can be estimated and the range tolerance can be applied relative to the surface rather than relative to the POI. For example, if the pixel is looking at the road, the surface normal will point upward. If the range is close enough to the road surface, rather than if the range is close enough to the POI, the pixel within the integration window can be included. Additionally or alternatively, when there is an overlap in the signal strength confidence interval, the adjacent pixels can be included in the set of qualified pixels. The Geiger mode lidar intensity is noisy, but under the conditions of multiple counts and multiple trials within a given span containing the echo signal, the confidence interval can be calculated using binomial statistics. When there is an overlap in the noise intensity confidence interval, the adjacent pixels can be alternatively or additionally included in the set of qualified pixels. In this case, the noise intensity confidence interval is a function of the number of noise counts and noise trials, which can be derived from the total counts and trials minus the number of counts and trials in the span containing the echo signal.
[0247] After completing step 674, the method continues to step 676, where the remaining pixels in ROI 804' are combined by the kernel to produce superpixels. Fig.21e As shown, kernel 810 applies the function to the remaining p1, p2, p4, p13, p14, p15, p16, p25, p26, p27, p28, p38, p39 to obtain feature F1. Feature F1 can be defined by the following mathematical equation (4).
[0248] F1=f(P1, P2, P4, P13, P14, P15, P16, P25, P26, P27, P28, P38, P39) (4)
[0249] Feature F1 is considered to be a superpixel (i.e., SP1=F1). Features (or superpixels) may include, but are not limited to, range, intensity, noise, and / or confidence. The mechanism for aggregating pixels is specific to the individual lidar systems being designed and may vary depending on the application. For example, simple addition or averaging may be employed for pixel aggregation. In this regard, the range value of feature F1 may include, but is not limited to, the average range of the remaining pixels within the ROI. The intensity value of feature F1 may include, but is not limited to: an intensity value derived from the sum of the signal counts and trials of the remaining pixels within the ROI; or the average intensity of the remaining qualified pixels within the ROI (when counts and trials are not available). The noise value of feature F1 may include, but is not limited to: a noise value derived from the sum of the trials and noise counts of the remaining pixels within the ROI; or the average noise of the remaining pixels within the ROI. The confidence value of feature F1 may include, but is not limited to, a confidence value derived from the updated noise value and the sum of the signal counts and trials in the kernel.
[0251] Repeat the process of steps 662 to 676 to generate other superpixels based on the span. For example, the next ROI is identified by shifting the kernel search window according to the span and generating the next superpixel according to the above process. Fig.21f As shown, the stride is 4, so the kernel search window is shifted four cells to the right. Therefore, the next superpixel is set to feature F2 defined by the following mathematical equation (5).
[0252] F2=f(P5, P6, P7, P8, P17, P18, P20, P29, P30, P31, P32, P41, P42, P43, P44) (5)
[0253] Other features F3, ..., F12 are generated in a similar manner. Figure 21gThe feature map 512 is shown. The other superpixels are set to these features (i.e., SP2=F2, SP3=F3, ..., SP12=F12). The superpixels can then be used to control the operation of the vehicle and / or dispatch personnel to the scene, as shown in optional step 680. Then, step 682 is performed, where method step 650 ends or other operations are performed.
[0255] The above-described method step 650 provides several advantages over existing systems and methods. For example, implementing the system and method step 650 provides: (i) improved ranging precision and accuracy; (ii) improved range and intensity accuracy; (iii) improved detectability and quality of painted lines (e.g., lane markings on a road); (iv) elimination of high-intensity artifacts and increased effective dynamic range; (v) improved detection probability of black or dark targets; and (vi) increased detection probability of targets at longer ranges. With respect to item (v), it should be noted that the result (or pixel) data may be corrupted (e.g., containing incorrect range values and / or relatively low confidence values that fall below a threshold), which prevents detection of dark objects. In the case of a low confidence value, the present solution can cause an increase in the confidence value such that it now exceeds the threshold, thereby enabling detection of dark objects with increased confidence.
[0257] Method step 650 can be implemented at the Geiger-mode avalanche photodiode (GmAPD) data level as a form of intelligent oversampling in the detection and waveform analysis steps. Method step 650 can also be implemented at the signal detection stage by using a two-pass process where data from a first detection attempt is fed back into a second detection attempt, or using a prior calculated from the raw data before histogramming to determine which pixels to integrate into a single histogram. In the latter case, the total flux can be calculated for each GmAPD pixel since it combines range, signal strength, and noise.
[0258] Now refer to Fig.25 , a flowchart of another method 900 for operating a laser radar system is provided. The method 900 can be performed in whole or in part by a laser radar system (e.g., Figure 3 A processor (e.g., Figure 3 processor 332) to execute.
[0260] Method 900 begins at step 902 and proceeds to step 904, where pixels (e.g., pixels p1, ..., p144 of FIG. 21) are arranged in a grid (e.g., grid 800 of FIG. 21). The pixels include pixels processed by a photodetector (e.g., Figure 3The result value generated by the waveform generated by the photodetector 318 of the image sensor 310 is shown in FIG. 906. In step 906, the processor performs operations to identify the ROI in the grid based on the correlation between the pixels. The correlation may include, but is not limited to, the correlation between the range values associated with the pixels and / or the correlation between the intensity values associated with the pixels.
[0262] In some cases, the ROI may be identified by the following steps: obtaining a kernel size; and using the kernel size to define a region of interest in a grid. The kernel size may be variable. The kernel size may be obtained by the following steps: locating those pixels that are closest neighbors of the POI in the grid in terms of at least range; and defining the kernel size based on the position of the closest neighbor in the grid. Alternatively, the kernel size may be obtained by the following steps: obtaining a reference kernel size; using the reference kernel size to identify a region in the grid; identifying a center pixel of the region; using the result value associated with each of the pixels in the region to calculate a score for the pixel, the score (e.g., score A described above) indicating a correlation between the result values associated with the pixel and the center pixel; selecting a pixel based on the score; and defining the kernel size based on the position of the selected pixel in the grid. The score may be a function of range, intensity, and / or noise.
[0264] In step 908, the size and / or position of the ROI in the grid is optionally adjusted to maximize the likelihood that the ROI contains a greater number of pixels associated with the object. This adjustment may be achieved by the following steps: identifying a POI in the ROI; identifying pixels that are nearest neighbor pixels of the POI at least in terms of range; using the centroid of the nearest neighbor pixels to obtain a likelihood that the POI is associated with an edge point or a corner point on the surface of the object; and adjusting the size and / or position of the ROI based on the likelihood that the POI is associated with an edge point or a corner point on the surface of the object. The POI may be a central pixel of the ROI.
[0266] In step 910, one or more pixels in the ROI may be optionally disqualified from aggregation with other pixels in the ROI. The disqualification may be based on how far the pixel is from the POI and / or road surface in one or more dimensions. The dimensions may include, but are not limited to, range, intensity, noise, and confidence.
[0268] In step 912, the processor combines the result values associated with the pixels located within the ROI to generate a feature value (eg, Figure 21gThe eigenvalue F1 of FIG. 1 is set as the eigenvalue F1 of FIG. 1 . In step 914, a superpixel is generated, and the superpixel has a value set as the eigenvalue. The operations of block steps 906 to 914 can be iteratively repeated to generate other superpixels, as shown in block step 916. The ROI used to generate the first superpixel in the first iteration can have a size and / or shape different from the size and / or shape of the ROI used to generate the second superpixel in another iteration. Subsequently, step 918 is performed, wherein the method 900 ends or performs other operations (e.g., returns to step 902).
[0270] The above-described LiDAR system can be used in a variety of applications. The present solution will now be described in the context of an autonomous vehicle. However, the present solution is not limited to autonomous vehicle applications. The present solution can be used in other applications, such as robotic applications (e.g., controlling the movement of an articulated arm) and / or system performance applications.
[0272] Fig.26 1 shows an example system 1000 according to aspects of the present disclosure. The system 1000 includes a vehicle 1002 that travels along a road in a semi-autonomous or autonomous manner. The vehicle 1002 is also referred to as AV 1002 in this document. AV 1002 may include, but is not limited to, a land vehicle (e.g., Fig.26 As described above, except where otherwise noted, the present disclosure is not necessarily limited to AV embodiments, and in some embodiments, it may include non-autonomous driving vehicles.
[0274] AV 1002 is generally configured to detect objects in its vicinity. Objects may include, but are not limited to, vehicles 1003, riders 1014 (such as riders of bicycles, electric scooters, motorcycles, etc.), and / or pedestrians 1016.
[0276] like Fig.26 As shown, AV 1002 may include sensor system 1018, onboard computing device 1022, communication interface 1020, and user interface 1024. The autonomous vehicle system may also include certain components contained in the vehicle (e.g., Fig. 27 ), these components can be controlled by the onboard computing device 1022 using various communication signals and / or commands, such as acceleration signals or commands, deceleration signals or commands, steering signals or commands, braking signals or commands, etc.
[0278] The sensor system 1018 may include one or more sensors coupled to and / or included within the AV 1002. For example, such sensors may include, but are not limited to, a lidar system, a radio detection and ranging (radar) system, a laser detection and ranging (lidar) system, a sound navigation and ranging (sonar) system, one or more cameras (e.g., visible spectrum cameras, infrared cameras, etc.), temperature sensors, positioning sensors (e.g., global positioning systems (GPS), etc.), position sensors, fuel sensors, motion sensors (e.g., inertial measurement units (IMU), etc.), humidity sensors, occupancy sensors, etc. The sensor data may include information describing the location of objects within the surrounding environment of the AV 1002, information about the environment itself, information about the motion of the AV 1002, information about the route of the vehicle, etc. As the AV 1002 travels over a surface, at least some of the sensors may collect data related to the surface.
[0280] AV 1002 may also transmit sensor data collected by the sensor system to a remote computing device 1010 (e.g., a cloud processing system) over a communication network 1008. The remote computing device 1010 may be configured with one or more servers to perform one or more processes of the techniques described in this document. The remote computing device 1010 may also be configured to transmit data / instructions to / from AV 1002, to / from (one or more) servers, and / or (one or more) data storages 1012 over the network 1008. The (one or more) data storages 1012 may include, but are not limited to, (one or more) databases.
[0282] The network 1008 may include one or more wired or wireless networks. For example, the network 1008 may include a cellular network (e.g., a long term evolution (LTE) network, a code division multiple access (CDMA) network, a 3G network, a 4G network, a 5G network, another type of next generation network, etc.). The network may also include a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cloud computing network, etc. and / or a combination of these or other types of networks.
[0284] AV 1002 may retrieve, receive, display and edit information generated from local applications or delivered from data storage 1012 via network 1008. Data storage 1012 may be configured to store and serve raw data, indexed data, structured data, road map data 1060, program instructions or other configurations known in the art.
[0286] The communication interface 1020 may be configured to allow communication between the AV 1002 and external systems, such as external devices, sensors, other vehicles, servers, data storage, databases, etc. The communication interface 1020 may use any now or hereafter known protocol, protection scheme, encoding, format, packaging, etc., such as but not limited to Wi-Fi, infrared link, Bluetooth, etc. The user interface system 1024 may be part of a peripheral device implemented within the AV 1002, including, for example, a keyboard, a touch screen display device, a microphone and speakers, etc. The vehicle may also receive status information, descriptive information, or other information about devices or objects in its environment over a communication link (such as a vehicle-to-vehicle, vehicle-to-object, or other V2X communication link) through the communication interface 1020. The term "V2X" refers to communication between a vehicle and any object that the vehicle may encounter or affect in its environment.
[0287] Fig. 27 An example system architecture 1100 for a vehicle is shown according to aspects of the present disclosure. Fig.26 The vehicles 1002 and / or 1003 may have Fig. 27 Therefore, the following discussion of system architecture 1100 is sufficient to understand Fig.26 (one or more) vehicles 1002, 1003. However, other types of vehicles are considered within the scope of the technology described in this document and may include more or less such as a combination of Fig. 27 As a non-limiting example, an aircraft may not include brakes or gear controls, but may include a depth sensor. In another non-limiting example, a water-based vehicle may include a depth sensor. Those skilled in the art will appreciate that other propulsion systems, sensors, and controls may be included, as known, based on the type of vehicle.
[0289] like Fig. 27As shown, the system architecture 1100 of a vehicle includes an engine or motor 1102 and various sensors 1104 to 1118 for measuring various parameters of the vehicle. In a gasoline-powered or hybrid vehicle having a fuel-powered engine, the sensors may include, for example, an engine temperature sensor 1104, a battery voltage sensor 1106, an engine revolutions per minute (RPM) sensor 1108, and a throttle position sensor 1110. If the vehicle is an electric or hybrid vehicle, the vehicle may have an electric motor and therefore include sensors such as a battery monitoring system 1112 (to measure the current, voltage, and / or temperature of the battery), motor current 1114 and voltage 1116 sensors, and a motor position sensor 1118, such as a resolver and encoder.
[0291] Operating parameter sensors common to both types of vehicles include, for example: a position sensor 1136, such as an accelerometer, gyroscope, and / or inertial measurement unit; a speed sensor 1138; and an odometer sensor 1140. The vehicle may also have a clock 1142, which the system uses to determine the vehicle time during operation. The clock 1142 may be encoded into the vehicle onboard computing device, it may be a separate device, or multiple clocks may be available.
[0293] The vehicle may also include various sensors that operate to collect information about the environment in which the vehicle is traveling. These sensors may include, for example: a position sensor 1160 (such as a global positioning system (GPS) device); an object detection sensor, such as one or more cameras 1162; a lidar system 1164; and / or a radar and / or sonar system 1166. The sensors may also include environmental sensors 1168, such as a precipitation sensor and / or an ambient temperature sensor. The object detection sensors may enable the vehicle to detect objects within a given distance range of the vehicle in any direction, while the environmental sensors collect data about the environmental conditions within the vehicle's area of travel.
[0295] During operation, information is transmitted from the sensors to the vehicle onboard computing device 1120. The vehicle onboard computing device 1120 may use Fig.29The vehicle onboard computing device 1120 analyzes the data captured by the sensors and optionally controls the operation of the vehicle based on the analysis results. For example, the vehicle onboard computing device 1120 can control: braking via a brake controller 1122; direction via a steering controller 1124; speed and acceleration via a throttle controller 1126 (in a gasoline-powered vehicle) or a motor speed controller 1128 (such as a current level controller in an electric vehicle); a differential gear controller 1130 (in a vehicle with a transmission); and / or other controllers. The auxiliary device controller 1134 can be configured to control one or more auxiliary devices, such as a test system, an auxiliary sensor, a mobile device transported by the vehicle, etc.
[0297] The geographic location information may be transmitted from the location sensor 1160 to the vehicle onboard computing device 1120, which may then access a map of the environment corresponding to the location information to determine known fixed features of the environment, such as streets, buildings, stop signs, and / or stop / go signals. Captured images from the camera 1162 and / or object detection information captured from sensors such as a lidar system 1164 are transmitted from these sensors to the vehicle onboard computing device 1120. The object detection information and / or captured images are processed by the vehicle onboard computing device 1120 to detect objects approaching the vehicle. Any known or to-be-known techniques for object detection based on sensor data and / or captured images may be used in the embodiments disclosed in this document.
[0299] The lidar information is transmitted from the lidar system 1164 to the vehicle onboard computing device 1120. Additionally, captured images are transmitted from the camera(s) 1162 to the vehicle onboard computing device 1120. The lidar information and / or captured images are processed by the vehicle onboard computing device 1120 to detect objects approaching the vehicle. The manner in which object detection is performed by the vehicle onboard computing device 1120 includes such capabilities as detailed in this disclosure.
[0300] In addition, the system architecture 1100 may include an onboard display device 1154 that may generate and output an interface on which sensor data, vehicle state information, or output generated by the processes described in this document is displayed to passengers of the vehicle. The display device may include an audio speaker that presents such information in an audio format, or a separate device may be an audio speaker that presents such information in an audio format.
[0302] The vehicle onboard computing device 1120 may include and / or communicate with a route controller 1132 that generates a navigation route from the starting location of the autonomous vehicle to the destination location. The route controller 1132 may access a map data memory to identify possible routes and road segments that the vehicle can travel to reach the destination location from the starting location. The route controller 1132 may score possible routes and identify a preferred route to the destination. For example, the route controller 1132 may generate a navigation route that minimizes the Euclidean distance or other cost function of the route during travel, and may further access traffic information and / or estimates that may affect the amount of time it will take to travel on a particular route. According to an embodiment, the route controller 1132 may generate one or more routes using various route methods (such as the Dijkstra algorithm, the Bellman-Ford algorithm, or other algorithms). The route controller 1132 may also use traffic information to generate a navigation route that reflects the expected conditions of the route (e.g., the current day of the week or the current time, etc.), so that the route generated for traveling during peak hours may be different from the route generated for traveling late at night. The route controller 1132 may also generate more than one navigation route to a destination, and send more than one of these navigation routes to the user for the user to select from among various possible routes.
[0304] In various embodiments, the vehicle onboard computing device 1120 may determine perception information of the surrounding environment of the AV. Based on sensor data provided by one or more sensors and the obtained position information, the vehicle onboard computing device 1120 may determine perception information of the surrounding environment of the AV. The perception information may represent what an average driver would perceive in the surrounding environment of the vehicle. The perception data may include information related to one or more objects in the environment of the AV. For example, the vehicle onboard computing device 1120 may process sensor data (e.g., lidar or radar data, camera images, etc.) to identify objects and / or features in the environment of the AV. Objects may include traffic signals, road boundaries, other vehicles, pedestrians, and / or obstacles, etc. The vehicle onboard computing device 1120 may determine perception using any now or later known object recognition algorithm, video tracking algorithm, and computer vision algorithm (e.g., repeatedly tracking an object frame by frame over a number of time periods).
[0306] In some embodiments, the vehicle onboard computing device 1120 may also determine the current state of the object for one or more identified objects in the environment. The state information may include, but is not limited to, for each object: current location; current speed and / or acceleration, current orientation; current pose; current shape, size, or footprint; type (e.g., vehicle, pedestrian, bicycle, stationary object, or obstacle); and / or other state information.
[0308] The vehicle onboard computing device 1120 may perform one or more prediction and / or forecasting operations. For example, the vehicle onboard computing device 1120 may predict the future position, trajectory, and / or motion of one or more objects. For example, the vehicle onboard computing device 1120 may predict the future position, trajectory, and / or motion of an object based at least in part on sensory information (e.g., state data for each object, including an estimated shape and pose determined as discussed below), positional information, sensor data, and / or any other data describing the past and / or current state of the object, the AV, the surrounding environment, and / or their relationship. For example, if the object is a vehicle and the current driving environment includes an intersection, the vehicle onboard computing device 1120 may predict whether the object will likely move straight ahead or turn. If the sensory data indicates that the intersection does not have a traffic light, the vehicle onboard computing device 1120 may also predict whether the vehicle may have to stop completely before entering the intersection.
[0309] In various embodiments, the onboard computing device 1120 may determine a motion plan for the autonomous vehicle. For example, the onboard computing device 1120 may determine a motion plan for the autonomous vehicle based on the perception data and / or the prediction data. Specifically, given predictions and other perception data regarding future locations of neighboring objects, the onboard computing device 1120 may determine a motion plan for the autonomous vehicle AV that optimally navigates the autonomous vehicle relative to the objects at their future locations.
[0311] In some embodiments, the vehicle onboard computing device 1120 can receive predictions and make decisions about how to handle objects and / or actors in the AV's environment. For example, for a particular actor (e.g., a vehicle with a given speed, direction, turning angle, etc.), the vehicle onboard computing device 1120 decides whether to overtake, yield, stop, and / or pass based on, for example, traffic conditions, map data, the state of the autonomous vehicle, etc. In addition, the vehicle onboard computing device 1120 also plans the AV's driving path on a given route and driving parameters (e.g., distance, speed, and / or turning angle). That is, for a given object, the vehicle onboard computing device 1120 decides what to do with the object and determines how to do it. For example, for a given object, the vehicle onboard computing device 1120 can decide to pass the object and can determine whether to pass on the left or right side of the object (including motion parameters such as speed). The vehicle onboard computing device 1120 can also assess the risk of collision between the detected object and the AV. If the risk exceeds an acceptable threshold, it may be determined whether the collision can be avoided if the autonomous vehicle follows a defined vehicle trajectory and / or implements one or more dynamically generated emergency maneuvers executed in a predefined time period (e.g., N milliseconds). If the collision can be avoided, the vehicle onboard computing device 1120 may execute one or more control instructions to perform a prudent maneuver (e.g., slow down, accelerate, change lanes, or turn appropriately). Conversely, if the collision cannot be avoided, the vehicle onboard computing device 1120 may execute one or more control instructions for performing an emergency maneuver (e.g., braking and / or changing direction of travel).
[0313] As described above, planning and control data for the motion of the autonomous vehicle are generated for execution. The vehicle onboard computing device 1120 may control braking, for example, via a brake controller; direction, via a steering controller; speed and acceleration, via a throttle controller (in a gasoline-powered vehicle) or a motor speed controller (such as a current level controller in an electric vehicle); a differential gear controller (in a vehicle with a transmission); and / or other controllers.
[0314] Fig.28 A block diagram is provided for understanding how to implement motion or movement of an AV according to the present solution. All operations performed in blocks 1202 and 1212 may be performed by a vehicle (e.g., Fig.26 AV 1002) of an onboard computing device (e.g., Fig.26 The onboard computing device 1022 and / or Fig. 27 The onboard computing device 1120) is executed.
[0316] In block 1202, an AV (e.g., Fig.26 The position of the AV 1002 can be determined based on the position sensor of the AV (e.g., Fig. 27 The detection is performed using sensor data output by a position sensor 1160 of the AV. The sensor data may include, but is not limited to, GPS data. Subsequently, the detected position of the AV is passed to block 1206.
[0318] In block 1204, an object (e.g., Fig.26 The detection is based on the camera of the AV (e.g., Fig. 27 1162) and / or a lidar system of the AV (e.g., Fig. 27 The laser radar system 1164 may be configured to detect the sensor data 1216 output by the laser radar system 1164. For example, image processing may be performed to detect instances of a particular class of objects (e.g., vehicles, cyclists, or pedestrians) in the image. The image processing / object detection may be implemented according to any known or later known image processing / object detection algorithm. The laser radar sensor data may include, but is not limited to, superpixels generated according to the above methods 500, 600, 650, and 900.
[0319] Additionally, a predicted trajectory of the object is determined in block 1204. In block 1204, based on the object category, The object's trajectory is predicted based on the cube geometry(s), cube orientation(s), and / or the contents of map 1218 (e.g., sidewalk locations, lane locations, lane direction of travel, driving rules, etc.). The manner in which the cube geometry(s) and orientation(s) are determined will become apparent as the discussion proceeds. At this point, it should be noted that various types of sensor data (e.g., 2D images, 3D lidar point clouds) and vector map 1218 (e.g., lane geometry) are used to determine the cube geometry(s) and / or orientation(s). Techniques for predicting an object's trajectory based on cube geometry and orientation may include, for example, predicting that an object is moving on a linear path in the same direction as the orientation direction. Predicted object trajectories may include, but are not limited to, the following trajectories: a trajectory defined by the object's actual speed (e.g., 1 mile per hour) and actual direction of travel (e.g., west); a trajectory defined by the object's actual speed (e.g., 1 mile per hour) and another possible direction of travel of the object (e.g., south, southwest, or in a direction toward the AV at X (e.g., 40°) degrees from the object's actual direction of travel); a trajectory defined by another possible speed of the object (e.g., 2 to 10 miles per hour) and the object's actual direction of travel (e.g., west); and / or a trajectory defined by another possible speed of the object (e.g., 2 to 10 miles per hour) and another possible direction of travel of the object (e.g., south, southwest, or in a direction toward the AV at X (e.g., 40°) degrees from the object's actual direction of travel). For objects in the same category and / or subcategory as the object, possible travel speed(s) and / or possible travel directions(s) may be predefined. It should be noted again that the cube defines the entire extent of the object and the orientation of the object. The orientation defines the direction in which the front of the object is pointing, and thus provides an indication of the actual and / or possible direction of travel of the object.
[0321] Information 1220 specifying a predicted trajectory of the object, the cube geometry / or orientation(s) is provided to box 1206. In some cases, a classification of the object is also passed to box 1206. In box 1206, a vehicle trajectory is generated using information from boxes 1202 and 1204. Techniques for determining a vehicle trajectory using a cube may include, for example, determining a trajectory of the AV that will pass the object when the object is in front of the AV, the cube has an orientation aligned with the direction of movement of the AV, and the cube has a length greater than a threshold. The present solution is not limited to the details of this scenario. The vehicle trajectory 1208 can be determined based on position information from box 1202, object detection information from box 1204, and / or map information 1214 (which is pre-stored in a data memory of the vehicle). Map information 1214 may include, but is not limited to Fig.26 The vehicle trajectory 1208 may represent a smooth path without sudden changes that would otherwise cause discomfort to passengers. For example, the vehicle trajectory is defined by a travel path along a given lane of a road that the predicted object has not traveled for a given amount of time. The vehicle trajectory 1208 is then provided to block 1210.
[0323] In block 1210, steering angle and speed commands are generated based on the vehicle trajectory 1208. The steering angle and speed commands are provided to block 1210 for vehicle dynamics control, i.e., the steering angle and speed commands cause the AV to follow the vehicle trajectory 1208.
[0325] For example, one or more computer systems such as Fig.29 The illustrated computer system 1300 implements various embodiments. The computer system 1300 may be any computer capable of performing the functions described in this document.
[0326] like Fig.29 As shown, the computer system 1300 can be any computer capable of performing the functions described herein. The computer system 1300 also includes (one or more) user input / output interfaces 1302 and (one or more) user input / output devices 1303, such as buttons, monitors, keyboards, pointing devices, etc.
[0327] Computer system 1300 includes one or more processors (also referred to as central processing units or CPUs), such as processor 1304. Processor 1304 is connected to a communications infrastructure or bus 1306. Processor 1304 may be a graphics processing unit (GPU), for example, a specialized electronic circuit designed to process mathematically intensive applications having a parallel structure for processing large blocks of data in parallel, such as mathematically intensive data common to computer graphics applications, images, video, etc.
[0328] The computer system 1300 also includes a main memory 1308, such as a random access memory (RAM), which includes one or more levels of cache and stored control logic (i.e., computer software) and / or data. The computer system 1300 may also include one or more secondary storage devices or secondary memories 1310, such as a hard disk drive 1312; and / or a removable storage device 1314 that may interact with a removable storage unit 1318. The removable storage device 1314 and the removable storage unit 1318 may be a floppy disk drive, a tape drive, an optical disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0329] Secondary memory 1310 may include other devices, means, or other methods for allowing computer system 1300 to access computer programs and / or other instructions and / or data, such as an interface 1320 and a removable storage unit 1322, such as a program cartridge and cartridge interface (such as found in video game devices), a removable memory chip (such as an EPROM or PROM) and an associated socket, a memory stick and USB port, a memory card and an associated memory card slot, and / or any other removable storage unit and associated interface.
[0330] The computer system 1300 may also include a network or communication interface 1324 to communicate and interact with any combination of remote devices, remote networks, remote entities, etc. (individually and collectively referenced by reference numeral 1328). For example, the communication interface 1324 may allow the computer system 1300 to communicate with the remote device 1328 via a communication path 1326, which may be wired and / or wireless and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to or from the computer system 1300 via the communication path 1326.
[0331] In an embodiment, a tangible non-transitory device or article of manufacture including a tangible non-transitory computer usable or readable medium having control logic (software) stored thereon is also referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1300, primary memory 1308, secondary memory 1310, and removable storage units 1318 and 1322, as well as tangible articles of manufacture embodying any combination of the foregoing. When executed by one or more data processing devices (such as computer system 1300), such control logic causes such data processing devices to operate as described herein.
[0332] Based on the teachings contained in this disclosure, how to use Fig.29 It will be apparent to those skilled in the relevant art(s) to make and use embodiments of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in the drawings. In particular, embodiments may operate using software, hardware, and / or operating system implementations other than those described in this document.
[0334] Terms related to this disclosure include:
[0335] An "electronic device" or "computing device" refers to a device that includes a processor and memory. Each device may have its own processor and / or memory, or the processor and / or memory may be shared with other devices, as in a virtual machine or container arrangement. The memory will contain or receive programming instructions that, when executed by the processor, cause the electronic device to perform one or more operations in accordance with the programming instructions.
[0337] The terms "memory," "memory device," "data storage facility," and the like refer to a non-transitory device on which computer-readable data, programming instructions, or both are stored, respectively. Unless specifically stated otherwise, the terms "memory," "memory device," "data storage facility," and the like are intended to include single device embodiments, embodiments in which multiple memory devices together or collectively store a set of data or instructions, and individual sectors within such devices. A computer program product is a memory device on which programming instructions are stored.
[0339] The terms "processor" and "processing device" refer to a hardware component of an electronic device that is configured to execute programmed instructions. Unless specifically stated otherwise, the singular term "processor" or "processing device" is intended to include single processing device embodiments and embodiments in which multiple processing devices together or collectively perform processing, which multiple processing devices may be components of a single device or components of separate devices.
[0341] When referring to objects detected by the vehicle perception system or simulated by the simulation system, the term "object" is intended to include stationary objects and moving (or potentially moving) actors unless otherwise specified by use of the terms "actor" or "stationary object."
[0343] When used in the context of motion planning for an autonomous vehicle, the term "trajectory" refers to a plan that the vehicle's motion planning system will generate and that the vehicle's motion control system will follow when controlling the vehicle's motion. The trajectory includes the planned position and orientation of the vehicle at multiple points in time over a time range, as well as the planned steering wheel angle and angular rate of the vehicle over the same time range. The autonomous vehicle's motion control system will consume the trajectory and send commands to the vehicle's steering controller, brake controller, throttle controller, and / or other motion control subsystems to move the vehicle along the planned path.
[0345] A "trajectory" of an agent that a vehicle perception or prediction system may generate refers to a predicted path that the agent will follow over a time horizon, as well as the agent's predicted speed and / or the agent's position along the path at various points along the time horizon.
[0346] In this document, the terms "street," "lane," "road," and "intersection" are illustrated by the example of a vehicle traveling on one or more roads. However, embodiments are intended to include lanes and intersections in other locations, such as parking areas. In addition, for autonomous vehicles designed for use indoors (such as automated picking devices in warehouses), a street may be a corridor of a warehouse and a lane may be a portion of a corridor. If the autonomous vehicle is a drone or other aircraft, the term "street" or "road" may refer to a route and a lane may be a portion of a route. If the autonomous vehicle is a ship, the term "street" or "road" may refer to a waterway and a lane may be a portion of a waterway.
[0348] In this document, when terms such as "first" and "second" are used to modify nouns, unless otherwise specified, such usage is intended only to distinguish one item from another, and is not intended to require a sequential order. In addition, relative position terms, such as "vertical" and "horizontal", or "front" and "rear", when used, are intended to be relative to each other, not necessarily absolute, and refer only to one possible position of the device associated with these terms according to the orientation of the device.
[0350] It should be understood that the detailed description section, but not any other section, is intended to be used to interpret the claims. The other sections may set forth one or more but not all exemplary embodiments contemplated by the inventors and, therefore, are not intended to limit the present disclosure or the appended claims in any way.
[0352] Although the present disclosure describes exemplary embodiments of exemplary fields and applications, it should be understood that the present disclosure is not limited to the disclosed examples. Other embodiments and modifications thereof are possible and are within the scope and spirit of the present disclosure. For example, and without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware and / or entities shown in the figures and / or described in this document. In addition, the embodiments (whether explicitly described or not) have significant practicality for fields and applications outside the examples described in this document.
[0353] In this document, embodiments have been described with the aid of functional building blocks that illustrate implementation methods of specific functions and their relationships. For ease of description, the boundaries of these functional building blocks have been arbitrarily defined in this document. Alternative boundaries can be defined as long as the specified functions and relationships (or their equivalents) are properly performed. Moreover, alternative embodiments can perform functional blocks, steps, operations, methods, etc. in an order different from that described in this document. Features from different embodiments disclosed herein can be freely combined. For example, one or more features from a method embodiment can be combined with any one of a system or product embodiment. Similarly, features from a system or product embodiment can be combined with any one of a method embodiment disclosed herein.
[0354] References to "one embodiment", "embodiment", "example embodiment" or similar phrases in this document indicate that the described embodiment may include specific features, structures or characteristics, but each embodiment may not necessarily include specific features, structures or characteristics. In addition, such phrases do not necessarily refer to the same embodiment. In addition, when describing specific features, structures or characteristics in conjunction with an embodiment, it is within the knowledge of a technician in the relevant field to combine such features, structures or characteristics into other embodiments (whether or not explicitly mentioned or described in this document). In addition, some embodiments may be described using the expressions "connection" and "connection" and their derivatives. These terms are not necessarily intended to be synonyms of each other. For example, some embodiments may be described using the terms "connection" and / or "connection" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "connection" can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. Features from different embodiments disclosed herein can be freely combined. For example, one or more features from a method embodiment can be combined with any one of a system or product embodiment. Similarly, features from a system or product embodiment can be combined with any one of the method embodiments disclosed herein.
Claims
1. A laser radar system, include: a series of transmitters, each transmitter configured to transmit light pulses out of the vehicle along a transmit axis to form a transmit field of view (Tx FoV); at least one detector configured to receive at least a portion of the light pulses reflected along a receive axis by an object within a receive field of view (Rx FoV); as well as A transmit optic is mounted for translation along a lateral axis and is configured to intersect each transmit axis without intersecting the receive axis to adjust the Tx FoV without adjusting the Rx FoV.
2. The laser radar system according to claim 1, in, The Tx FoV and the Rx FoV overlap, and wherein the adjusted Tx FoV is located within the area of the Tx FoV.
3. The lidar system of claim 1 further comprising a collimator mounted adjacent to the series of emitters and configured to focus and direct the light pulses along each transmit axis to collectively form a transmit light beam.
4. The laser radar system according to claim 3, in, The transmit optics are arranged adjacent to the collimator and are configured to focus the transmit beam onto an area of the Tx FoV to form an adjusted TxFoV.
5. The laser radar system according to claim 4, in, The emission optics include a cylindrical lens.
6. The laser radar system according to claim 1, in, The series of transmitters includes a linear array of transmitters arranged parallel to the transverse axis, the linear array of transmitters including a proximal transmitter and a distal transmitter arranged opposite the proximal transmitter.
7. The laser radar system according to claim 6, further comprising: include: An actuator is connected to the emission optics and is configured to translate the emission optics through a range between a rest position and a distal position, wherein at the rest position, the emission optics does not intersect any emission axis of the linear array of emitters and at the distal position, the emission optics intersects the emission axis of the distal emitter.
8. The laser radar system of claim 1 , further comprising a controller configured to translate the transmit optics along the lateral axis, in, The controller is further configured to: determining from the received light pulses that the object is an unknown object; and The transmit optics are translated along the transverse axis between a proximal position and a distal position while transmitting light pulses through the transmit optics.
9. The laser radar system according to claim 9, in, The controller is further configured to: receiving scan data indicative of light pulses reflected from the unknown object while translating the transmitting optics; determining a location of the unknown object based on the scan data; and The transmit optics are translated along the lateral axis to a position such that the adjusted Tx FoV is aligned with the position of the unknown object.
10. A laser radar system, include: processor; A non-transitory computer-readable storage medium comprising programming instructions configured to cause the processor to implement a method for operating a lidar system, wherein the programming instructions include instructions for: receiving a result value from a photodetector, the result value indicating a time when the photodetector detected a photon at or near a target wavelength; combining different sets of said result values to generate superpixels; using the superpixels to obtain a first spatiotemporal coherence metric; selecting a subset or a result value group of light pulses based on said first spatiotemporal coherence measure; and The distance between the lidar system and the object is detected based on the selected subset of light pulses or the selected group of result values.
11. The laser radar system according to claim 10, in, The first spatiotemporal coherence metric comprises a metric specifying a variation in distribution between detections of two pulses or two groups of pulses by the plurality of photodetectors, respectively, and a subset of the light pulses or a resulting value group is selected based on a maximum one of the metrics.
12. The laser radar system according to claim 10, in, The first spatiotemporal coherence metric comprises, for each pulse, a measured variance of differences between consecutive timestamps that have been ordered from lowest to highest value or from highest to lowest value, and the selected subset of light pulses or group of result values comprises light pulses or result values associated with relatively low measured variances.
13. The laser radar system according to claim 10, in, The first spatiotemporal coherence metric comprises a score for each pulse of the optical signal, the score being indicative of confidence or validity of object detection, and the pulse is selected for inclusion in the subset when the score exceeds a certain value.
14. A laser radar system, include: processor; A non-transitory computer-readable storage medium comprising programming instructions configured to cause the processor to implement a method for operating a lidar system, wherein the programming instructions include instructions for: arranging a plurality of pixels in a grid, the plurality of pixels comprising result values generated from processing a waveform produced by a photodetector of the lidar system; identifying a first region of interest in the grid based on at least one of a correlation between range values associated with the plurality of pixels and a correlation between intensity values associated with the plurality of pixels; combining result values associated with pixels located within the first region of interest to produce at least one first feature value; and A first superpixel is generated, the first superpixel having a value set to the at least one first eigenvalue.
15. The laser radar system according to claim 14, in, The programming instructions also include instructions for obtaining a kernel size and using the kernel size to identify a region of interest in the grid.
16. The laser radar system according to claim 15, in, The kernel size is obtained by following the steps below: locating those pixels of the plurality of pixels that are closest neighbors, at least in terms of range, to a pixel of interest in the grid; as well as The kernel size is defined based on the location of the nearest neighbor in the grid.
17. The laser radar system according to claim 15, in, The kernel size is obtained by following the steps below: Get the reference kernel size; identifying regions in the grid using the reference kernel size; identifying a center pixel of the region; calculating a score for each of the pixels in the region using the result value associated with the pixel, the score indicating a correlation between the result values associated with the pixel and the center pixel; selecting a pixel from the plurality of pixels based on the score; as well as The kernel size is defined based on the position of the selected pixel in the grid.
18. A method for operating a lidar system, include: receiving, by a processor, result values from a plurality of photodetectors, the result values indicating times when the plurality of photodetectors detected photons at or near a target wavelength, the result values being based on operations performed by each of the plurality of photodetectors to facilitate measurements associated with light signals reflected from an object external to the lidar system; combining, by the processor, different sets of the result values to generate a superpixel; using, by the processor, the superpixel to obtain a first spatiotemporal correlation metric; selecting, by the processor, a subset or a group of result values of light pulses based on the first spatiotemporal coherence metric; as well as The distance between the lidar system and the object is detected by the processor based on the selected subset of light pulses or the selected group of result values.
19. The method according to claim 18, in, The first spatiotemporal coherence metric comprises at least one of a distribution comparison metric, a time-of-flight statistics metric, and a detection confidence score, The first spatiotemporal coherence metric comprises metrics respectively specifying distribution changes between detections of two pulses or two groups of pulses by the plurality of photodetectors, and a subset of light pulses or a result value group is selected based on a maximum one of the metrics.
20. The method according to claim 19, in, The first spatiotemporal coherence metric comprises, for each pulse, a measured variance of differences between consecutive timestamps that have been ordered from lowest to highest value or from highest to lowest value, and the selected subset of light pulses or group of result values comprises light pulses or result values associated with relatively low measured variances.