Doppler-assisted object mapping for autonomous vehicle applications
By employing Doppler-assisted speed sensing technology and point cloud mapping algorithms, the efficiency and accuracy issues of object recognition and tracking in autonomous driving have been resolved, thereby improving the path selection and safety of autonomous driving systems.
Patent Information
- Application Number
- CN202111319937.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-09
- Filing Date
- 2021-11-09
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-11-09
AI Technical Summary
Existing autonomous driving technologies struggle to efficiently identify and track dynamic objects in the environment, especially when the Doppler effect is not fully utilized, leading to insufficient path selection and safety.
Using Doppler-assisted velocity sensing technology, coherent LiDAR is used to detect frequency changes caused by the motion of reflective surfaces. Combined with point cloud mapping algorithms, objects in the environment can be quickly identified and tracked.
It improves the efficiency and accuracy of object recognition and tracking, and enhances the path selection and safety of autonomous driving systems.
Smart Images

Figure CN114527478B_ABST
Abstract
Description
Technical Field
[0001] This specification generally relates to autonomous vehicles. More specifically, it relates to improving autonomous driving systems and components by using speed sensing data to assist in the detection, identification, and tracking of objects encountered in autonomous driving environments. Background Technology
[0002] Autonomous (fully and partially autonomous) vehicles operate by sensing their external environment using various electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., audio and humidity) sensors. Some autonomous vehicles map driving paths through the environment based on the sensed data. Driving paths can be determined based on Global Positioning System (GPS) data and road map data. While GPS and road map data can provide information about static aspects of the environment (buildings, street layouts, road closures, etc.), dynamic information (such as information about other vehicles, pedestrians, streetlights, etc.) is obtained from the simultaneously collected sensor data. The accuracy and safety of the driving path and rate mechanisms selected by autonomous vehicles depend on the timely and accurate identification of the various objects present in the driving environment, and on the ability of the driving algorithm to process information about the environment and provide instructions to the vehicle control and drivetrain systems. Summary of the Invention
[0003] In one embodiment, a method is disclosed, comprising: obtaining by a computing device a plurality of sensing data frames of an environment surrounding an autonomous vehicle (AV), each of the plurality of sensing data frames including a plurality of points, wherein each of the plurality of points corresponds to a reflection of a signal emitted by a sensing system of the AV from a surface of an object in the environment, and includes a distance to the corresponding reflecting surface, and wherein one or more of the plurality of points include velocity data of the corresponding reflecting surface. The disclosed method further comprises evaluating by the computing device an assumption that a first set of points from a first set of sensing data frames corresponds to a second set of points from a second set of sensing data frames, wherein evaluating the assumption includes mapping the first set of points to the second set of points based on velocity data from one or more of the first or second set of points; and determining an evaluation metric for the assumption based on the performed mapping. The disclosed method further comprises determining a driving path for the AV based on the evaluation metric.
[0004] In another embodiment, a system is disclosed comprising a memory storing instructions and a computing device for executing instructions from the memory to obtain a plurality of sensed data frames of the environment surrounding an autonomous vehicle (AV). Each of the plurality of sensed data frames includes a plurality of points, wherein each of the plurality of points corresponds to a reflection of a signal emitted by the AV's sensing system from a surface of an object in the environment, and includes a distance to the corresponding reflective surface, and wherein one or more of the plurality of points includes velocity data of the corresponding reflective surface. The computing device also evaluates a hypothesis that a first set of points from a first sensed data frame corresponds to a second set of points from a second sensed data frame, wherein, in order to evaluate the hypothesis, the computing device maps the first set of points to the second set of points based on velocity data from one or more of the first or second set of points; and determines an evaluation metric for the hypothesis based on the performed mapping. The computing device also causes a driving path for the AV to be determined based on the evaluation metric.
[0005] In another embodiment, a non-transitory computer-readable medium having instructions stored thereon is disclosed, which, when executed by a computing device, causes the computing device to obtain a plurality of sensed data frames of the environment surrounding an autonomous vehicle (AV), each of the plurality of sensed data frames including a plurality of points, wherein each of the plurality of points corresponds to a reflection of a signal emitted by the AV's sensing system from a surface of an object in the environment, and includes distances to the corresponding reflective surface, and wherein one or more of the plurality of points include velocity data of the corresponding reflective surface. The instructions also cause the computing device to evaluate a hypothesis that a first set of points from a first sensed data frame corresponds to a second set of points from a second sensed data frame, wherein, in order to evaluate the hypothesis, the computing device maps the first set of points to the second set of points based on velocity data from one or more of the first or second set of points; and determines an evaluation metric for the hypothesis based on the performed mapping. The instructions also cause the computing device to determine a driving path for the AV based on the evaluation metric. Attached Figure Description
[0006] This disclosure is illustrated by way of example rather than limitation, and can be more fully understood when considered in conjunction with the accompanying drawings, with reference to the following detailed description, in which:
[0007] Figure 1 This is a diagram illustrating components of an example autonomous vehicle using Doppler-assisted object recognition and tracking, according to some embodiments of this disclosure.
[0008] Figure 2A This is a schematic diagram illustrating Doppler-assisted object identification and tracking using point cloud mapping as part of a perception system for an autonomous vehicle, according to some embodiments of this disclosure.
[0009] Figure 2B The mapping from a first point cloud (corresponding to a first sensing data frame) to a second point cloud (corresponding to a second sensing data frame) according to some embodiments of the present disclosure is shown.
[0010] Figure 3 This is a schematic diagram of rolling shutter correction for accurate point cloud mapping as part of the perception system of an autonomous vehicle, according to some embodiments of the present disclosure.
[0011] Figure 4A This is a schematic diagram of a dual-sensor setup utilizing point cloud mapping as part of a perception system for an autonomous vehicle, according to some embodiments of this disclosure.
[0012] Figure 4B This is a schematic diagram illustrating the determination of the lateral velocity associated with the return point using a dual-LIDAR triangulation scheme according to some embodiments of this disclosure.
[0013] Figure 5 A flowchart is depicted illustrating an example method for Doppler-assisted object recognition and point cloud tracking for autonomous vehicle applications according to some embodiments of the present disclosure.
[0014] Figure 6 A flowchart depicts an example method for evaluating the assumption that a set of points in a first sensing frame corresponds to a set of points in a second sensing frame, according to some embodiments of this disclosure.
[0015] Figure 7 A block diagram of an example computer device for autonomous vehicle applications, capable of enabling Doppler-assisted object recognition and tracking, is depicted according to some embodiments of the present disclosure. Detailed Implementation
[0016] Autonomous vehicles can use Light Detection and Ranging (LIDAR) technology to detect the distance to various objects in the environment, and sometimes the speed of these objects. A LIDAR emits one or more laser signals (pulses) traveling toward an object and then detects the arrival signal reflected from that object. By determining the time delay between signal emission and the arrival of the reflected wave, a time-of-flight (ToF) LIDAR can determine the distance to the object. A typical LIDAR emits signals in multiple directions to obtain a wide field of view of the external environment. For example, a LIDAR device can cover the entire 360-degree field of view by scanning to collect a series of consecutive frames with timestamps. As a result, each sector in space is sensed with a time increment Δt, which is determined by the angular velocity of the LIDAR's scan rate. As used herein, a "frame" or "sensing data frame" can refer to the entire 360-degree field of view of the environment obtained by a scan of the LIDAR, or alternatively, to any smaller sector obtained by a partial scan or by a scan designed to cover a limited angle, such as 1 degree, 5 degrees, 10 degrees, or any other angle.
[0017] Each frame can include several return points (or simply "points") corresponding to reflections from various objects in the environment. Each point can be associated with a distance to the corresponding object, or more specifically, with a distance to a surface element of the object responsible for the corresponding return point. The set of points can be represented as a frame or otherwise associated with frames and is sometimes referred to as a "point cloud." A point cloud can include returns from multiple objects. Typically, the priori is unknown how many objects are within a given frame. A single object (such as another vehicle, road sign, pedestrian, etc.) can generate multiple return points. For example, a 10-degree frame can include returns from one or more road signs, returns from multiple vehicles located at different distances from the LiDAR device (which may be mounted on an autonomous vehicle) and moving at different speeds in different directions, returns from pedestrians crossing a roadway, returning from pedestrians walking along a sidewalk or standing on the side of the road, and returns from many other objects. Segmenting a given point cloud (which can be performed by the autonomous vehicle's perception system) into clusters corresponding to different objects is useful in autonomous driving applications. Nevertheless, points that are close together (e.g., spaced at small angular distances and corresponding to reflections from similar distances) can belong to different objects. For example, a traffic sign and a pedestrian standing near the sign can generate a nearby return point. Similarly, a car moving along a bicycle in an adjacent lane can generate a nearby return point.
[0018] Identifying points corresponding to different objects is typically performed using mappings of point clusters belonging to different frames (such as frames with consecutive timestamps). Specifically, it can be assumed that a selected first set of points (point clusters) of a frame identified by timestamp t and a selected second set of points of frame t+Δt belong to the same object (e.g., a car or a pedestrian). Using a suitable (e.g., best-fit) geometric transformation, the first set of points is then mapped onto the second set of points, and it is determined whether the obtained mapping is within acceptable accuracy to assess whether the hypothesis is validated or invalidated. In this case, different hypotheses can be selected and the process repeated. Such hypothesis selection and validation can be performed in parallel, evaluating multiple hypotheses simultaneously.
[0019] Various algorithms can be used to find optimal geometric transformations, such as the Iterative Closest Point (ICP) algorithm, which identifies the optimal transformation using a series of progressively improving convergence steps (iterations). Traditional ICP (or other mapping) algorithms are based on point mappings in a coordinate space using angular (or linear lateral) coordinates and longitudinal (or radial) range (distance) values, and suffer from several drawbacks. In particular, choosing a small or large time increment Δt has corresponding disadvantages. For example, a small Δt may weaken the algorithm's ability to invalidate erroneous assumptions based on a small number of frames. Conversely, a large Δt may reduce the sensing rate. Furthermore, using Time-of-Flight (ToF) LiDAR techniques, which can detect the gradual separation of objects over several (or more) consecutive time frames, objects located close to each other may require a considerable amount of time to distinguish.
[0020] Time-of-Flight (ToF) LiDAR is typically used for ranging. ToF can also determine the velocity (speed and direction of motion) of a returning point by rapidly transmitting two or more signals (as part of different sensing frames) and detecting the position of the reflecting surface as it moves between each additional frame. The intervals between consecutive signals can be short enough that the object does not significantly change its position relative to other objects in the environment between consecutive signals (frames), but long enough to allow the LiDAR to accurately detect changes in the object's position. However, ToF LiDAR devices typically cannot determine the velocity of an object based on a single sensing frame.
[0021] The aspects and embodiments of this disclosure enable methods for identifying objects in an autonomous vehicle environment using Doppler-assisted velocity sensing. Specifically, coherent LiDAR utilizes phase information encoded in the transmitted signal and carried to and from the target by the emitted electromagnetic waves, providing additional functionality not available in standard ToF LiDAR technology. Coherent LiDAR detects changes in the frequency (and accompanying phase) of the reflected wave caused by the motion of the reflecting surface, a phenomenon known as the Doppler effect. The frequency / phase of the reflected wave corresponds to the velocity component V of the reflecting surface parallel to the wave propagation direction. rSensitive, referred to in this paper as "radial" or "longitudinal" velocity. In addition to obtaining range information, coherent LiDAR allows the radial velocity to be correlated with the return points of the point cloud. As described in more detail below, this additional information contributes to more efficient object recognition and tracking. In particular, radial velocity data allows for more effective hypothesis formation. For example, the radial velocity V based on at least some points in the first point set of the first frame. r Based on the positions of corresponding points in the second point set of the second frame, the perception system can quickly discard some of the hypotheses that are inconsistent with the measured velocity (e.g., a hypothesis can be discarded if the second point set moves too far or too little from the first point set, depending on the measured velocity). Conversely, in some cases, hypotheses can be formed based on point clusters that have similar radial velocities or radial velocities that are different from each other but consistent with the radial velocities of objects performing a combination of linear motion and rotation (corresponding to point clusters).
[0022] The use of velocity information can also facilitate the verification of formed hypotheses. For example, a hypothesis formed based on a single sensing frame can be tested (evaluated) when data from a second sensing frame is collected: the position of the midpoint in the second frame can be compared with the position of the midpoint in the first frame, and the hypothesis can be confirmed or invalidated based on how consistent the movement of the corresponding object is with the velocity of the point in the first frame. Similarly, if a hypothesis is formed based on a mapping of points in the first and second frames, the radial velocity of subsequent (third, fourth, etc.) frames can then be used to verify the hypothesis by comparing the actual distance traveled by each point of the (hypothetical) object with the displacement predicted by the velocity measurement.
[0023] Figure 1 This is a diagram illustrating components of an example autonomous vehicle (AV) 100 using Doppler-assisted object recognition and tracking, according to some embodiments of this disclosure. Figure 1 The operation of an example autonomous vehicle is illustrated. Autonomous vehicles can include motor vehicles (cars, trucks, buses, motorcycles, all-terrain vehicles, recreational vehicles, any specialized agricultural or construction vehicles, etc.), aircraft (airplanes, helicopters, drones, etc.), naval vehicles (ships, boats, yachts, submarines, etc.), or any other self-propelled vehicle capable of operating in a self-driving mode (without human input or with reduced human input) (e.g., a sidewalk delivery robot vehicle).
[0024] The driving environment 110 may include any object (active or inactive) located outside the AV, such as roadways, buildings, trees, shrubs, sidewalks, bridges, mountains, other vehicles, pedestrians, etc. The driving environment 110 may be urban, suburban, rural, etc. In some embodiments, the driving environment 110 may be a non-road environment (e.g., aquaculture or agricultural land). In some embodiments, the driving environment may be an indoor environment, such as an industrial plant, shipping warehouse, hazardous area of a building, etc. In some embodiments, the driving environment 110 may be substantially flat, with various objects moving parallel to a surface (e.g., parallel to the surface of the Earth). In other embodiments, the driving environment may be three-dimensional and may include objects capable of moving in all three directions (e.g., balloons, leaves, etc.). Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement of a self-propelled vehicle can occur. For example, "driving environment" may include any possible flight environment for an aircraft or the marine environment for a naval vessel. Objects in the driving environment 110 may be located at any distance from the AV, from a few feet (or less) to several miles (or more).
[0025] Exemplary AV 100 may include a sensing system 120. The sensing system 120 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., acoustic) sensing subsystems and / or devices. As used throughout this disclosure, the terms “optical” and “light” should be understood to include any electromagnetic radiation (waves) that can be used in object sensing to facilitate autonomous driving, such as distance sensing, speed sensing, acceleration sensing, rotational motion sensing, etc. For example, “optical” sensing may utilize the range of light visible to the human eye (e.g., a wavelength range of 380 to 700 nm), the UV range (below 380 nm), the infrared range (above 700 nm), the radio frequency range (above 1 m), etc. In embodiments, “optical” and “light” may include any other suitable range of the electromagnetic spectrum.
[0026] Sensing system 120 may include radar unit 126, which may be any system that uses radio or microwave frequency signals to sense objects within the driving environment 110 of AV 100. The radar unit may be configured to sense both the spatial position of objects (including their spatial dimensions) and their velocity (e.g., using Doppler shift technology). Hereinafter, “velocity” refers to both how fast an object is moving (the object’s rate) and the direction of its motion. The term “angular velocity” refers to how fast an object is rotating about an axis and the direction of that axis of rotation. For example, a car turning left (right) has an axis of rotation pointing upwards (downwards), and the value of the angular velocity is equal to the rate of change of the angle of rotation (e.g., measured in radians per second).
[0027] Sensing system 120 may include one or more LIDAR sensors 122 (e.g., LIDAR rangefinders), which may be laser-based units capable of determining the distance to objects in driving environment 110 (e.g., using Time-of-Flight (ToF) technology). LIDAR sensors can utilize wavelengths of electromagnetic waves shorter than radio waves, and therefore can provide higher spatial resolution and sensitivity compared to radar units. LIDAR sensors may include coherent LIDAR sensors, such as frequency-modulated continuous wave (FMCW) LIDAR sensors. LIDAR sensors may use optical heterodyne detection for velocity determination. In some embodiments, the functionality of ToF and coherent LIDAR sensors is combined into a single (e.g., hybrid) unit capable of determining both the distance to a reflecting object and the radial velocity of the reflecting object. Such a hybrid unit may be configured to operate in an incoherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection) or simultaneously in both modes. In some implementations, multiple LIDAR sensor 122 units may be mounted on the AV, for example, at spatially separated locations, to provide additional information about the transverse component of the velocity of the reflecting object, as described in more detail below.
[0028] LIDAR sensor 122 may include one or more laser sources that generate and emit signals, and one or more detectors that receive signals reflected from an object. LIDAR sensor 122 may include a spectral filter to filter out parasitic electromagnetic waves having wavelengths (frequency) different from the emitted signal. In some embodiments, LIDAR sensor 122 may include a directional filter (e.g., aperture, diffraction grating, etc.) to filter out electromagnetic waves that may reach the detector in a direction different from the retroreflection direction of the emitted signal. LIDAR sensor 122 may use a variety of other optical components (lenses, mirrors, gratings, optical films, interferometers, spectrometers, local oscillators, etc.) to enhance the sensor's sensing capabilities.
[0029] In some embodiments, the LIDAR sensor 122 may, for example, scan a 360-degree field of view in the horizontal direction. In some embodiments, the LIDAR sensor 122 may be capable of spatial scanning along both the horizontal and vertical directions. In some embodiments, the field of view may reach 90 degrees in the vertical direction (e.g., at least a portion of the area above the horizon is scanned by the LIDAR signal). In some embodiments, the field of view may be a complete sphere (composed of two hemispheres). For the sake of brevity, when “LIDAR technology,” “LIDAR sensing,” “LIDAR data,” and “LIDAR” are generally referred to in this disclosure, such references should also be understood to include other sensing technologies that typically operate at near-infrared wavelengths, but may also include sensing technologies that operate at other wavelengths.
[0030] The sensing system 120 may also include one or more cameras 129 to capture images of the driving environment 110. The images may be two-dimensional projections (planar or non-planar, e.g., fisheye) of the driving environment 110 (or portions thereof) onto the projection plane of the camera. Some of the cameras 129 of the sensing system 120 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 110. The sensing system 120 may also include one or more sonars 128, which in some embodiments may be ultrasonic sonars.
[0031] The sensing data acquired by sensing system 120 can be processed by data processing system 130 of AV 100. For example, data processing system 130 may include perception system 132. Perception system 132 can be configured to detect and track objects in driving environment 110 and identify detected objects. For example, perception system 132 can analyze images captured by camera 129 and can be able to detect traffic light signals, road signs, lane layouts (e.g., boundaries of traffic lanes, topology of intersections, designation of parking spots, etc.), the presence of obstacles, etc. Perception system 132 can also receive LiDAR sensing data (coherent Doppler data and incoherent ToF data) to determine the distance to various objects in environment 110 and the velocity of such objects (radial velocity and lateral velocity in some embodiments, as described below). In some embodiments, perception system 132 can combine LiDAR data with data captured by camera 129. In one example, camera 129 can detect an image of a stone partially obstructing a traffic lane. Using data from camera 129, perception system 132 can be able to determine the angular size of the stone, but not the linear size of the stone. Using LIDAR data, the perception system 132 can determine the distance from the stone to the AV, and therefore, by combining the distance information with the angular size of the stone, the perception system 132 can also determine the linear dimension of the stone.
[0032] In another embodiment, using LiDAR data, perception system 132 can determine how far a detected object is from the AV, and can also determine the component of the object's velocity along the direction of motion of the AV. Furthermore, using a series of rapid images acquired by the camera, perception system 132 can also determine the lateral velocity of the detected object in a direction perpendicular to the direction of motion of the AV. In some embodiments, the lateral velocity can be determined solely from LiDAR data, for example, by identifying the edges of the object (using a horizontal scan) and further determining how fast the edges of the object move in the lateral direction. Perception system 132 may have a point cloud module (PCM) 133 to perform mapping of return points from different sensing frames in order to identify and track various objects in the driving environment 110. PCM 133 may be a velocity-assisted (Doppler-assisted) module that uses velocity data to augment range data for more efficient and reliable object detection and tracking, as described in more detail below.
[0033] The sensing system 132 can also receive information from a GPS transceiver (not shown) configured to obtain information about the AV's position relative to the Earth. The GPS data processing module 134 can use GPS data in conjunction with sensing data to help accurately determine the position of the AV relative to fixed objects in the driving environment 110 (such as roadways, lane boundaries, intersections, sidewalks, crosswalks, road signs, surrounding buildings, etc.), the positions of which can be provided by map information 135. In some embodiments, the data processing system 130 can receive non-electromagnetic data, such as sonar data (e.g., ultrasonic sensor data), temperature sensor data, pressure sensor data, meteorological data (e.g., wind speed and direction, precipitation data), etc.
[0034] The data processing system 130 may also include an environmental monitoring and prediction component 136, which can monitor how the driving environment 110 evolves over time, for example, by maintaining tracking of the position and velocity of moving objects (relative to the Earth). In some embodiments, the environmental monitoring and prediction component 136 may maintain tracking of changes in the appearance of the environment due to the motion of the AV relative to the environment. In some embodiments, the environmental monitoring and prediction component 136 may make predictions about how various moving objects in the driving environment 110 will be positioned within a predicted time frame. The predictions may be based on the current position and velocity of the moving objects and on the tracking dynamics of the moving objects during a certain (e.g., predetermined) time period. For example, based on stored data of object 1 indicating accelerated motion of object 1 during a previous 3-second time period, the environmental monitoring and prediction component 136 may infer that object 1 is resuming its motion from a stop sign or red traffic light signal. Therefore, given the lane layout and the presence of other vehicles, the environmental monitoring and prediction component 136 may predict where object 1 might be in the next 3 or 5 seconds of motion. As another example, based on stored data indicating the deceleration of object 2 during the previous 2-second time period, environmental monitoring and prediction component 136 can infer that object 2 is stopped at a stop sign or a red traffic light. Therefore, environmental monitoring and prediction component 136 can predict where object 2 might be in the next 1 or 3 seconds. Environmental monitoring and prediction component 136 can perform periodic checks on the accuracy of its predictions and modify the predictions based on new data obtained from sensing system 120.
[0035] Data generated by the perception system 132, the GPS data processing module 134, and the environmental monitoring and prediction component 136 can be used by autonomous driving systems, such as an AV vehicle control system (AVCS) 140. AVCS 140 may include one or more algorithms to control how the AV behaves in various driving situations and environments. For example, AVCS 140 may include a navigation system for determining a global driving route to a destination. AVCS 140 may also include a driving path selection system for selecting a specific path through the current driving environment, which may include selecting a driving lane, navigating traffic congestion, selecting a location for a U-turn, selecting a trajectory for parking operations, etc. AVCS 140 may also include an obstacle avoidance system for safely avoiding various obstacles (rocks, broken-down vehicles, jaywalking pedestrians, etc.) within the AV's driving environment. The obstacle avoidance system can be configured to assess the size and trajectory of obstacles (if the obstacle is moving) and select the optimal driving strategy (e.g., braking, steering, acceleration, etc.) for obstacle avoidance.
[0036] The AVCS140's algorithms and modules can generate instructions for various vehicle systems and components, such as powertrain and steering (150), vehicle electronics (160), signaling (170), and more. Figure 1 Other systems and components not explicitly shown. Powertrain and Steering 150 may include an engine (internal combustion engine, electric motor, etc.), transmission, gearbox, axles, wheels, steering mechanism, and other systems. Vehicle Electronics 160 may include an onboard computer, engine management, ignition system, communication system, carputers, telematic, in-vehicle entertainment system, and other systems and components. Signaling 170 may include high and low headlights, parking lights, turning and reversing lights, horn and alarm, interior lighting system, instrument panel notification system, passenger notification system, radio and wireless network transmission system, etc. Some of the commands output by AVCS 140 may be directly delivered to Powertrain and Steering 150 (or Signaling 170), while other commands output by AVCS 140 are first delivered to Vehicle Electronics 160, which generates commands for Powertrain and Steering 150 and / or Signaling 170.
[0037] In one example, AVCS 140 can determine that an obstacle identified by data processing system 130 needs to be avoided by slowing the vehicle down to a safe speed and then steering the vehicle around the obstacle. AVCS 140 can output commands to powertrain and steering 150 (directly or via vehicle electronics 160) to 1) reduce fuel flow to the engine to lower engine speed by modifying throttle settings, 2) downshift the drivetrain to a lower gear via the automatic transmission, 3) engage the braking unit to reduce (while coordinating with the engine and transmission) the vehicle speed until a safe speed is reached, and 4) use the power steering mechanism to perform steering maneuvers until the obstacle is safely avoided. Subsequently, AVCS 140 can output commands to powertrain and steering 150 to restore the vehicle's previous speed settings.
[0038] Figure 2A This is a schematic diagram of Doppler-assisted object identification and tracking 200 using point cloud mapping as part of a perception system for an autonomous vehicle, according to some embodiments of the present disclosure. Figure 2AThe image depicts an AV 202 (which could be AV 100 or any other AV) approaching an intersection. AV 202 has a sensor 206, which can be a LIDAR (such as a coherent LIDAR, FMCW LIDAR, hybrid coherent / ToF LIDAR, a combination of coherent and incoherent LIDAR, etc.) or any other device that allows sensing radial velocity information in addition to range (distance) information. Sensor 206 performs a scan of the driving environment of AV 202. Specifically, sensor 206 can sense multiple return points (“points”) for each sensing frame (sensing data frame). Sensing frames can be separated by a time increment Δt. The time increment refers to the time difference between signals emitted in the same direction (or returning from the same direction), since different directions can be detected by signals at slightly different times. More specifically, Δt can be the period of rotation of the sensor (e.g., a LiDAR transmitter); and when there are N points around a 360-degree horizontal field of view, the detection of any two adjacent sensing directions may have a time lead / lag of Δt / N.
[0039] Object 210 (e.g., a car, truck, bus, motorcycle, or any other object) is approaching the intersection and making a left turn, such as Figure 2A The two consecutive positions of the AV corresponding to two consecutive LIDAR frames captured at times t and t+Δt, for example, 204(1) and 204(2), are shown. Similarly, the positions of object 210 at the two frames t and t+Δt are shown as 212(1) and 212(2), respectively.
[0040] It should be understood that Figure 2A The displacements of AV 202 and object 210 between two consecutive frames shown are exaggerated for illustrative purposes; in reality, various objects may change their positions more insignificantly over the time increment Δt. For example, when object 210 completes... Figure 2A When the left-hand turn is depicted, several frames can be sensed by sensor 206.
[0041] Object 210 performs a combination of translational and rotational motions. For example, the translation vector of a reference point of object 210. And object 210 rotates around this reference point by an angle Δφ. (As explained below, and for the sake of the effectiveness of the described method, the choice of reference point is quite arbitrary). In a flat driving environment, it is sufficient to describe the rotational motion by a single component value Δφ, but in a non-flat 3D (especially flight or navigation) environment, the rotation angle can be a vector. Its three components describe the pitch angle, yaw angle, and roll angle, respectively. The angular velocity of object 210 characterizes the rate of rotation (turning) of object 210. Similarly, the linear velocity of the reference point determines the rate of translation of object 210. If object 210 is rigid, with respect to a certain reference point O (having coordinates) angular velocity and linear velocity This knowledge allows us to determine the velocity of any other point on the rigid body according to the following equation (referred to as the rigid body equation in this paper):
[0042]
[0043] The choice of reference point O can be arbitrary, because the same relationship exists for any other reference point O', i.e.
[0044]
[0045] in, This is the line reference velocity at reference point O'. Although the line reference velocity changes when the reference point changes, the angular velocity is independent of the choice of reference point. This independence provides additional flexibility when performing point cloud mapping.
[0046] Reference point O can be considered as the center of rotation of the object. The freedom to choose the reference point reflects the possibility of representing any displacement of a rigid object through a infinite number of possible combinations of rotations (about an arbitrarily chosen center of rotation, but to the same angle and around the same axis) and translations. (An exception is pure translation.) Therefore, in some embodiments, the reference point O can conveniently be chosen somewhere inside the object (although not necessarily near the object's geometric center or centroid). In other embodiments, the motion of the object can conveniently be represented as about an axis parallel to... The rotation of the axis (without translation in a plane perpendicular to the axis) and the translation along the axis. The choice of such a center of rotation (hereinafter referred to as the "pure rotation" setting) is unique (at most, arbitrary translation along the axis) and can be determined from conditions. It is determined in the middle, and it gives
[0047]
[0048] Where C is any number. (In two-dimensional motion, C is zero.) For small angular velocities (when the object primarily performs translational motion), as seen in the last expression, the center of rotation is located at a considerable distance. Therefore, for numerical accuracy, in some implementations, the distance from the object to its center of rotation can be limited; for example, once the center of rotation is determined to be farther than a predetermined distance (e.g., a number of the object's longest dimension), a pure rotation setting can be changed to a combined rotation-translation setting. Alternatively, a bias of a large distance to the center of rotation can be used, as described in more detail below.
[0049] like Figure 2A As shown, at position 212(1), object 210 can reflect several signals output by sensor 206 (represented by solid lines) and generate several return points (shown by black circles) for the first frame. Return points should be understood as data items generated by sensing system 132 based on measurements performed by sensor 206 (e.g., indexed by the angular direction of the output signal or in any other way), as part of sensing system 120. Each return point may include (or be associated with) the distance r to the actual physical reflection point and the radial velocity V. r It is equal to the direction along which the sensor 206 is directed (by the unit vector). (Description) speed Quantity: In some implementations, only a few of the return points may include radial velocity values. For example, while a ToF range measurement can be performed for each return point, only a few of these points (e.g., one in every five, one in every ten, etc.) can be detected by coherent LiDAR and include velocity data. Radial velocity V r The velocity V is measured within the reference frame of AV202. Therefore, since AV202 is normally moving, the measured velocity V... r It may differ from the velocity of the corresponding physical reflection point relative to the ground, which can then be determined by adding the velocity of object 210 (in vector form) measured in frame AV 202 to the velocity of AV 202 relative to the ground (which can be known independently, for example, from rate meter / odometer data, map / GPS data, etc.).
[0050] At position 212(2), object 210 can similarly reflect the new set of signals output by sensor 206 (represented by dashed lines) and generate several return points for the second frame. One or more mapping algorithms implemented by PCM 133 can determine the geometric transformation that maps the point cloud of the first frame to the point cloud of the second frame. Figure 2B A mapping 250 from a first point cloud 262 (e.g., corresponding to a first frame) to a second point cloud 264 (e.g., corresponding to a second frame) according to some embodiments of this disclosure is shown. The mapping shown is equivalent to a geometric transformation of the point cloud associated with a rigid object (e.g., object 260). Mapping 250 is achieved by identifying translation vectors. and rotation angle To determine the duration of a given time increment Δt, the linear velocity of the object was also identified. and angular velocity (For example, the average velocity over time interval Δt). Mapping 250 can use the Iterative Closest Point (ICP) algorithm, which iteratively corrects the transformation and minimizes an error metric (e.g., mean squared error or some other predetermined metric) based on a comparison (and vice versa) of the transformed first point cloud 262 with the second point cloud 264. In some implementations, other mapping algorithms can be used, such as the Kabsch algorithm, Procrustes superposition, etc. Although only two sensing frames (with corresponding point clouds) are depicted for simplicity, similar mappings can be generated between various consecutive sensing frames (e.g., between the second and third frames, between the third and fourth frames, etc.) for both object recognition and tracking.
[0051] Refer again Figure 2A As object 210 moves from position 212(1) to position 212(2), the return point in the second frame corresponds to the reflective surface of object 210, which may be different from the surface that caused the signal reflection in the first frame. For example, as Figure 2A As depicted, when a portion of a previously occluded rotating object 210 enters the field of view of sensor 206, additional return points can be detected. Conversely, some of the previously exposed return points may not exist (because the corresponding physical reflective surfaces disappear from the field of view), etc. To address this dynamic aspect of the point cloud, the algorithm executed by PCM 133 can draw a bounding box, which can be a projection of a 3D bounding box onto the field of view. The projection of the bounding box evolves dynamically as the bounding box rotates relative to the field of view. After setting the bounding box around the point cloud (e.g., as a hypothetical formed portion), PCM 133 can map the actual (currently visible) and currently occluded points to anticipate when such occluded points may enter the field of view, enabling faster and more efficient point cloud mapping.
[0052] Using radial velocity data to exclude or validate hypotheses can be performed before, after, or concurrently with point cluster mapping (e.g., ICP mapping). For example, to reduce the computational cost of mapping, hypotheses inconsistent with the radial velocity data can be discarded before mapping. In some implementations, mapping can be performed on the contours of two point clouds. In some implementations, the PCM 133 can retain those hypotheses whose point clouds have passed velocity validation and are suitable for successful mapping, within a set alignment accuracy.
[0053] In one example implementation, point cloud mapping aided by radial velocity data can be performed as follows. The perception system 132 can identify a first point cloud 262 (e.g., a source point cloud) acquired at time t and a second point cloud 264 (e.g., a target point cloud) acquired at time t+Δt. The PCM 133 can make assumptions about associating points in the first point cloud 262 with points in the second point cloud 264 using various mapping methods (such as ICP mapping). For example, one of the assumptions formed might be that each point in the first point cloud is matched with its closest point in the second point cloud. In some implementations, such an assumption might be just one of many, because the motion of the underlying object might be such that (especially over longer periods) points are not necessarily mapped to the closest points. For example, point A′ in the second point cloud 264 might be closest to point B, but is correctly mapped to point A in the first point cloud 262.
[0054] After forming one or more hypotheses, PCM 133 can identify pairs of points in two clouds that map to each other (enumerated using index j): R j (1)→R j (2). Furthermore, PCM 133 can identify a set of fitting parameters {β} = β1, β2…, which will parameterize the motion of the object (e.g., a rigid body) corresponding to the assumed mapping. Parameters can include translational velocities, rotational velocities, centers of rotation, etc. The number of parameters can depend on the mapped motion. For example, two-dimensional planar motion (e.g., the motion of a vehicle on a flat surface) can be characterized by a single angular velocity value (the rate of rotation about a vertical axis), while three-dimensional motion (e.g., the motion of a flying object) can be characterized by three components of angular velocity. The table below provides examples of parameters that can be used to describe mappings of point clouds.
[0055]
[0056]
[0057] Here, the plane x′y′ is perpendicular to the angular velocity. Direction; V OΩ The translational velocity is along The directional component.
[0058] Based on the actual coordinates R of the point at time t (first frame) j (1) and use fitting parameters, for example, V O With Ω, PCM133 can predict the future coordinates of the first point cloud at time t+Δt (e.g., the time of the second frame).
[0059]
[0060] For ease of annotation, vectors are represented by bold letters instead of arrows. Similarly, in some implementations, PCM 133 can also perform reverse "prediction" of the past positions (at the time of the first frame) of the coordinates of points in the second point cloud:
[0061]
[0062] After making such a prediction, PCM 133 can compare the predicted coordinates of the first point cloud at a future time instance. With the second cloud R j (2) The degree of approximation of the actual coordinates of the points (e.g., obtained by sensing system 120). The accuracy of the prediction can be characterized by the forward-looking residual.
[0063]
[0064] It is the sum of the residuals at each point j. Residual It can include several contributions. For example, This can include penalties for errors in the predicted coordinates:
[0065]
[0066] Among them, R j|| R represents the radial component of the radius vector (radial distance) to point j, while R j⊥ Represents the components of vectors with the same radius in the transverse plane. Depending on whether the motion is two-dimensional or three-dimensional, R... j⊥ Each component can have one or two components. In some implementations, weights a and b can be different from each other to account for the fact that radial and lateral distances may be known with different degrees of accuracy. In some implementations, weights a and b can be considered equal to each other.
[0067] residual It may also include the radial velocity V j|| The penalty for prediction error, which can be known from coherent LiDAR sensors:
[0068]
[0069] Among them, R j|| (1)+V j|| (1) Δt is based on the previously measured radial distance R j|| (1) and expected increment V j||(1) Δt) The predicted radial distance at time t+Δt. The weight c can be different from the weights a and b. In some implementations, the relative values of a and b and c can depend on the relative accuracy with which the radial distance, azimuth distance, and radial velocity can be known. Therefore, the total residual It can be represented as a sum.
[0070]
[0071] In some implementations, a backward-looking residual can be defined similarly. For example, a backward-looking residual...
[0072]
[0073] Similarly, a penalty for errors in coordinate prediction could be included.
[0074]
[0075] Furthermore, it may also include the radial velocity V. j|| The penalty for prediction error, which can be known from the LiDAR sensing system:
[0076]
[0077] Among them, R j|| (2)-V j|| (2) Δt is based on the radial distance R measured subsequently. j|| (2) and reduction -V j|| (2) The “predicted” radial distance of Δt at a past time t. Total retrospective residuals. It can be the sum The total residual error of the mapping can be the sum S({β}) = S > ({β})+S < ({β}), whose optimization (e.g., minimization) can determine the optimal value of the fitted parameters, for example, {β} = V O Ω, R O In some implementations, as a definition of forward-looking residuals... and retrospective residuals The speed error can be replaced by a combination error utilizing the average speed value, for example...
[0078] S j (vel) = c[R] j|| (1)-R j|| (2)+Δt(V j|| (1)+V j|| (2)) / 2]2 .
[0079] Because the residual error S is the fitting parameter β k Since S is a nonlinear function, various iterative methods can be used to optimize it, such as gradient descent, Gauss-Newton method, Levenberg-Marquardt method, etc. For example, in gradient descent, the parameter β... k In the gradient Iteration in the defined direction (β) k →β k +Δβ k ), which is obtained through the Jacobian matrix. In this representation, the system of linear equations (in matrix form) in each iteration of the gradient descent method is as follows:
[0080] λΔβ=J T ΔS
[0081] The vector Δβ = (Δβ1, Δβ2, ...) for the fitting parameter increment is determined based on the residual error ΔS = (ΔS1, ΔS2, ...) that appears in the corresponding iteration; an adjustable parameter λ can be selected to obtain the maximum accuracy, the fastest convergence, or based on other considerations.
[0082] In the Gauss-Newton method, a system of linear equations can be used:
[0083] (J T J)Δβ=J T ΔS
[0084] Determine the local minimum in each iteration.
[0085] In the Levenberg-Marquart method (damped Gauss-Newton method), a system of equations can be used to interpolate between the equations of the gradient descent method and the Gauss-Newton method:
[0086] (J T J+λ1)Δβ=J T ΔS
[0087] Determine the increment in each iteration.
[0088] Iteration can continue until sufficient convergence is achieved, for example, when further improvement from additional iterations falls below a predetermined target level. The determined optimal fitting parameters {β} are then determined. min It can be approximated by the minimum residual error S used for the specific assumptions under consideration. min =S({β}) min In some implementations, if the obtained residual error S minWithin the target range, PCM 133 can accept this hypothesis as the current working hypothesis. In other embodiments, PCM 133 can form multiple hypotheses and perform the above mapping process for each of the formed hypotheses, wherein the hypothesis with the smallest residual error is accepted as the current working hypothesis. Different hypotheses may include different associations between points in the first point cloud and points in the second point cloud. In some embodiments, the various hypotheses may involve different numbers of points. For example, some points may be included in some hypotheses but excluded from others. In some embodiments, if the hypotheses involve different numbers of points, the correspondingly determined smallest residual error S can be used before comparison with other hypotheses. min Normalization (e.g., dividing by the number of points).
[0089] In some implementations, the PCM 133 may retain multiple hypotheses for subsequent validation, rather than selecting a single working hypothesis. For example, the retained hypotheses may be validated using third point clouds from the third frame t+2Δt, the fourth frame t+3Δt, and so on. Mapping can be performed between any (or each) consecutive frame pairs until the final hypothesis is confirmed. In some implementations, the PCM 133 may also identify (e.g., using the shape of the point cloud) the type of an object (e.g., car, truck, bicycle, road sign, etc.). In some implementations, after an object is identified, the PCM 133 may continue mapping the point clouds of subsequent frames for object tracking.
[0090] In some implementations, the residual error S({β}) may include some, but not all, of the contributions described above. For example, in some implementations, the coordinate error S(coord) may be included in the residual error, while the velocity error S(vel) may initially be excluded. Instead, S(vel) can be used for hypothesis verification. Specifically, after a hypothesis is formed and mapping is performed using one or more of the possible hypotheses selected as feasible point cloud mappings, S(vel) can be computed using the optimal set of parameters determined for the respective hypothesis. Hypotheses that result in a residual error S(vel) exceeding a certain predetermined threshold can be discarded. Hypotheses with a residual error S(vel) below the threshold can be retained (e.g., for subsequent verification using additional frames). In some implementations, the hypothesis with the lowest residual error S(vel) can be accepted (and used for object tracking / subsequent verification using additional frames). In some implementations, as an alternative to using radial velocity to compute residual errors (such as S(vel)), radial velocity data can be used to filter out hypotheses inconsistent with radial data. Specifically, the error can be computed for each mapping hypothesis 1→2.
[0091]
[0092] And the error exceeds a preset threshold (e.g., determined empirically), E > E T Those assumptions can be discarded. In some implementations, velocity data can be used for the initial elimination of infeasible assumptions and for evaluating the remaining assumptions (e.g., using S(vel) residuals, as described above).
[0093] In various implementations, various methods can be used to address the ambiguity between selecting the rotation center and identifying the translational velocity (both in a plane perpendicular to the angular axis). In one implementation, the rotation center can be chosen randomly, for example, near the center of the first (or second) point cloud. In some implementations, the performed mapping can be biased towards a smaller translational velocity, or optionally (depending on preference settings), a smaller rotation radius. For example, residual prospective error can include contributions:
[0094]
[0095] Here, d and g represent the biases that are unfavorable for high translational velocities and large distances to the rotation center (from both point clouds), respectively, while N is the hypothetical number of points. Higher values of d favor pure rotation settings, while higher values of g favor rotation centers closer to the center of the point cloud. The specific values of d and g can be chosen empirically, for example, by maximizing the accuracy and efficiency of the resulting point cloud mapping. In some implementations, one (or both) of the coefficients d and g can be very small or zero. A small but still non-zero coefficient g can prevent R... O It becomes too large, and therefore can help avoid R O Large but Ω small and product R O In the case where Ω is neither small nor large, the error of Ω may be caused by a large value of R. O enlarge.
[0096] Figure 3 This is a schematic diagram of rolling shutter correction for accurate point cloud mapping as part of a perception system for an autonomous vehicle, according to some embodiments of this disclosure. In some embodiments, the time increment between the detection of points in the first cloud and corresponding points in the second cloud can be different from the time interval Δt between different sensing frames. For example, although the LIDAR sensor 206 can detect the same spatial orientation with a period Δt (e.g., the LIDAR transmitter can use angular velocity...), the time increment between the detection of points in the first cloud and corresponding points in the second cloud can be different. (Rotation), but within one cycle of the LIDAR transmitter's rotation (rolling shutter), different spatial orientations can receive the LIDAR sensor's attention at different times. Therefore, the motion of an object generating a return point (reflection) can cause a time shift between consecutive frames different from Δt. For example, the j-th point 301 of the first point cloud can be identified (during cloud point mapping, e.g., as part of the formed hypothesis) as a point 302 mapped onto the second point cloud. Although a certain reference orientation is detected at time t (first frame) and t+Δt (second frame), the j-th point 301 of the first point cloud can be at a time t different from the angular coordinates of the j-th point 301 of the first point cloud. The time of angular lag (or lead, depending on the relative position with respect to the reference direction). The location was detected. Similarly, the j-th point 302 of the second point cloud can be located at a time t+Δt, determined by the angular coordinates of the j-th point 302 of the second point cloud. A definite amount of time The location was detected. As a result, the time difference between the detections of the two points can be Δt + δt. j For example, correction by specific points The magnitude and sign of the correction can be determined by the rate and direction of motion of the j-th point. For additional accuracy in point cloud mapping, such corrections can be considered during the mapping process. For example, the predicted position of a point in the first point cloud at the time of the second data frame can be determined using the modified time.
[0097]
[0098] The residual error in the radial velocity can be calculated using the following equation:
[0099]
[0100] Similar modifications can be made to other parameters encountered during the mapping process (e.g., backtracking). Correcting δt j It can be a positive number, a negative number, or zero, depending on the specific point in the cloud.
[0101] Figure 4A This is a schematic diagram of a dual LiDAR setup 400 utilizing point cloud mapping as part of a perception system for an autonomous vehicle, according to some embodiments of this disclosure. Figure 4AThe diagram depicts an AV 402 with multiple LiDAR sensors (two shown for illustration), such as a first sensor 406 and a second sensor 407. These can be any type of coherent (or a combination of coherent and incoherent) LiDAR device capable of sensing the distance to and radial velocity of reflective surfaces of objects in the driving environment. Sensors 406 and 407 can perform a scan of the driving environment and generate return points corresponding to various objects. Each sensor can output a signal with a phase characteristic (e.g., chirp or any other phase or frequency modulation characteristic) unique to the sensor, such that the return signals from the sensors do not interfere with each other. Sensors 406 and 407 can be located at a distance from each other (baseline distance) to improve lateral velocity resolution. In some embodiments, the baseline distance can be as large as possible (e.g., limited by the length or width of the AV 402). In some embodiments, because the lateral velocity resolution is greatest in the direction perpendicular to the baseline and least in the direction parallel to the baseline, more than two sensors can be used in a non-collinear (e.g., triangular) arrangement. For example, the third sensor may be located near the centerline of AV 402 (e.g., near the front or rear of the vehicle).
[0102] In some implementations, the processing logic of the sensing system (e.g., sensing system 120) can synchronize the sensing frames of sensors 406 and 407, so that the sensing signals are output at the same time instances (e.g., at t, t+Δt, t+2Δt, t+3Δt, etc.). In other implementations, the sensing frames can be staggered (e.g., to reduce potential interference or to improve instantaneous resolution), such that one sensor outputs a signal at times t, t+Δt, t+2Δt, t+3Δt, while another sensor outputs a sensing signal at times t+Δt / 2, t+3Δt / 2, t+5Δt / 2, and so on. Each sensor can detect its corresponding point cloud, and due to the different positioning and timing of the sensing frames, these point clouds may differ from the point clouds of other sensors even at the same time. The processing logic of the sensing system (e.g., sensing system 132) can be tailored to the first sensor cloud. Each point identifies the nearest second sensor cloud The processing logic identifies the points and associates these two points with the same reflective portion of object 410. In some implementations, the processing logic may approximate the reflective portion as being located at the midpoint. Place. Figure 4A The diagram shows a cluster of points 420 corresponding to object 410 (for simplicity, only one cluster of points is shown, for example, detected by the first sensor 406).
[0103] Because the two sensors are in different positions, the radial velocity V1 measured by the first sensor 406 along the first radial direction 408 for the return point 422 of the point cluster 420 may differ from the radial velocity V2 measured along the second radial direction 409 for the same (or nearby) return point of the cluster sensed by the second sensor 407. The difference V1-V2 indicates the lateral velocity of the return point 422. When V1 = V2, the lateral velocity of the return point 422 is zero (within the measurement accuracy), while V1 > V2 (or conversely, V1 < V2) indicates that the return point 422 has moved to the left (or right). Figure 4B This is a schematic diagram illustrating the use of a dual LiDAR triangulation scheme 400 to determine the lateral velocity associated with a return point according to some embodiments of this disclosure. The x-axis shown coincides with the baseline between the first sensor 406 and the second sensor 407 (which are a distance b apart from each other). The y-axis is perpendicular to the x-axis, and the origin of the coordinate system is located at the midpoint between the sensors. (The processing logic can use any other choice of coordinate axes, depending on the implementation.) As shown, the first sensor 406 can determine a first distance r1 to the return point 422, an angular direction α to the return point, and a radial velocity V1. Similarly, the second sensor 407 can determine a second distance r2 from the second sensor to the return point 422, an angular direction β to the return point, and a corresponding radial velocity V2. Assuming the velocity points in the xy plane (which is the most common case in driving applications), the velocity... The components can be determined based on the fact that the velocity projection in direction 408 is V1 and in direction 409 is V2 (as depicted by the dashed line). Specifically, according to
[0104] V1 = -V x sinα-V y cosα,
[0105] V2 = V x sinβ-V y cosβ,
[0106] The lateral velocity can be deduced as
[0107]
[0108] The approximation applies to small angles α and β, where cosα≈cosβ≈1 and sin(α+β)≈α+β≈L / r, where r is either distance r1 or r2 (r1 and r2 are nearly identical in terms of the accuracy used). Similarly, three or more LiDAR sensors can determine the velocity associated with the return point. All three components. Such a determination may be particularly advantageous in autonomous flight or navigation applications involving objects capable of moving in all three directions.
[0109] Lateral velocity data obtained using a multi-sensor setup can be used for hypothesis formation and subsequent validation. In one implementation, lateral velocity data can be used to enhance the evaluation of hypotheses based on radial velocity fields, such as those concerning... Figures 2A-2B As disclosed, based on lateral velocity data, the PCM 133 can separate objects with similar radial velocities but different lateral velocities (e.g., vehicles passing each other in opposite directions). When the lateral velocities of different objects are sufficiently different, feasible hypotheses can already be formed based on a single sensing frame. In some cases, the PCM 133's ability to use lateral velocity data can be limited. This may be based on the fact that the accuracy of lateral velocity determination may be lower than the accuracy of radial velocity measurement. Based on the above-derived data for V... x The equation, if the radial velocity has an accuracy δV, is given. r (meaning the measured value V) r Indicates actual speed in [V] r -δV r V r +δV r Within the interval), and assuming the distance r is exactly known or has high precision, then Figure 4B The triangulation-based lateral velocity measurement shown in the figure has accuracy:
[0110]
[0111] For example, if δV r =0.1m / s, and LIDAR sensors 406 and 407 are located at a distance of L = 1.2m, then the accuracy of the lateral velocity determination at a distance of r = 60m will be δV x = 5.0 m / s. For larger (smaller) distances, the accuracy will be higher (lower). This level of accuracy is quite satisfactory for distinguishing (based on a single sensing frame) vehicles moving in opposite directions or for distinguishing between cars and bicycles (pedestrians), regardless of their direction of movement.
[0112] In some implementations, lateral velocity data can be used for cloud point mapping (e.g., regarding...). Figures 2A-2B The described method (for hypothesis formation and / or verification) is similar to the method used with radial velocity. For example, lateral velocity can be used as a constraint on possible mappings, quantified with one or more additional residual errors, such as prospective lateral velocity errors.
[0113]
[0114] And similarly, for backtracking errors:
[0115]
[0116] The weight assigned to the lateral velocity error is described by a coefficient h. In some implementations, the weight h can be smaller than the weight c assigned to the radial velocity error, for example, inversely proportional to the ratio of the corresponding accuracy in the determination of the lateral and radial velocities, h / a ~ (δV) r / δV x ) 2 The symbol ~ indicates an order-of-magnitude estimate rather than an exact ontology.
[0117] In some implementations, various other evaluation metrics can be designed to assess errors in point cloud mapping. For example, while quadratic residuals have been described with respect to the above implementations, any one or more functions (e.g., monotonic functions) can be used to quantify errors in coordinate matching, radial velocity matching, lateral velocity matching, and biases used in identifying translational velocities and rotation centers.
[0118] In some implementations, a machine learning model can be used to determine evaluation metrics such as weights a (assigned to radial distance mismatch), b (assigned to lateral distance mismatch), c (assigned to radial velocity mismatch), h (assigned to lateral velocity mismatch), and biases such as d (for large translational velocities) and g (for large rotational radii), or other metrics used in the evaluation metrics. More specifically, the model can be trained using several point clouds and two or more sensing frames (corresponding point clouds with correction point associations used as training (target) mappings). Correction associations can be labeled by a human operator / developer, or pre-obtained using various point cloud mapping algorithms (such as ICP), or any combination thereof. The evaluation metrics can be determined during the training of the machine learning model (e.g., a neural network).
[0119] After identifying one or more objects in the sensing frame and the data processing system 132 provides information about the identified objects and their motion to the AVCS 140, the AVCS 140 can determine the driving path of the AV and, based on the motion of the identified objects, provide corresponding instructions (e.g., speed, acceleration, braking, steering instructions, etc.) to the powertrain and steering 150 and the vehicle electronics 160.
[0120] Figure 5A flowchart is depicted of an example method 500 for Doppler-assisted object recognition and point cloud tracking for autonomous vehicle applications according to some embodiments of this disclosure. Method 500, and method 600 described below, and / or each of their respective functions, routines, subroutines, or operations can be executed by a computing device having one or more processing units (CPUs) and a storage device communicatively coupled to the CPUs. In some embodiments, methods 500 and 600 can be executed by a single processing thread. Alternatively, methods 500 and 600 can be executed by two or more processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In the illustrative example, the processing threads implementing method 500 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing methods 500 and 600 can execute asynchronously with respect to each other. Various operations of methods 500 and 600 can be performed in comparison to… Figure 5 and Figure 6 The order in which operations are executed differs from the order shown. Some operations of a method can be executed concurrently with other operations. Some operations can be optional.
[0121] Method 500 can be applied to autonomous vehicles (such as...) Figure 1The system and components of the autonomous vehicle (100). Method 500 can use sensing data obtained by sensing system 120 and data processing system 130 to identify and track objects present in the driving environment. Information about the identified objects can be provided to the autonomous vehicle control system 140. Method 500 may include, at block 510, obtaining multiple sensing data frames of the environment surrounding the AV by a computing device. Each of the multiple sensing data frames may include multiple points. Points correspond to the reflection of signals emitted by the sensing system of the AV from the surface of an object in the environment. Each point may include various data, such as the timestamp of the sensing frame and the coordinates of the reflecting surface. The coordinates may include the distance to the corresponding reflecting surface and an angle (or any other value, rather than explicitly identifying the location of the reflecting surface) specifying the direction to the reflecting surface. At least some of the points may also include velocity data of the corresponding reflecting surface; the velocity data may include the radial velocity of the reflecting surface. Each point may also include the intensity of the reflected signal, the polarization of the reflected signal, etc. The radial distance may be determined from LIDAR data, while the angle may be known independently from synchronizer data, clock data, for example, based on the known rotation frequency of the transmitter of the sensing system in a plane of rotation (e.g., a horizontal plane). Velocity data can be acquired by a sensing system, which may include a coherent optical detection and ranging device (LIDAR) capable of using, for example, Doppler-assisted sensing, to detect radial velocity. In some embodiments, the coherent LIDAR may be a frequency-modulated continuous wave (FMCW) LIDAR, and the signal emitted by the sensor may include phase-modulated or frequency-modulated electromagnetic waves. The sensing system may also be able to simultaneously emit various other signals, such as pulse signals, which can be used for Time-of-Flight (ToF) distance measurements. In some embodiments, the sensor may include separate ToF LIDARs and coherent LIDARs, each emitting separate signals that are synchronized, mixed, and transmitted along the same optical path.
[0122] Various sensing frames can correspond to different periods (e.g., rotations) of the transmitter in the sensing system. For example, a first sensing frame can correspond to a first period, and a second sensing frame can correspond to a different period (e.g., earlier or later). The terms "first" and "second" should be understood only as identifiers and should not presuppose a specific order. In particular, there can be any number of intermediate frames between the first and second frames.
[0123] In some implementations, the AV sensing system may include multiple sensors. For example, the sensing system may include a first sensor (e.g., sensor 406) capable of sensing the velocity of objects in the environment. When measured by the first sensor, the velocity data of the corresponding reflective surface may include a first component of the velocity of the corresponding reflective surface along a first direction from the first sensor to the corresponding reflective surface. Furthermore, the sensing system may include a second sensor (e.g., sensor 407) capable of sensing the velocity of objects in the environment. When measured by the second sensor, the velocity data of the corresponding reflective surface may include a second component of the velocity of the corresponding reflective surface along a second direction from the second sensor to the corresponding reflective surface.
[0124] At block 520, method 500 may continue by evaluating, using a computing device (e.g., a device executing software instructions issued by PCM 133 as part of sensing system 132), the hypothesis that a first set of points from a first sensing data frame corresponds to a second set of points from a second sensing data frame. In some embodiments, the first and second set of points may be selected (e.g., by the sensing system) based on the spatial proximity of the selected points, based on the proximity of the velocity values associated with the points, using various segmentation algorithms, or through any other selection process. The first and second set of points may be identified as part of the hypothesis that such points correspond to a single object. In some embodiments, evaluating the hypothesis includes identifying the first set of points as having velocity data consistent with a rigid body performing a combination of translational and rotational motion.
[0125] As by Figure 5 As illustrated in the blowout section, at box 522, the evaluation hypothesis can include the hypothesis that the selected object (corresponding to the first set of points in the first frame) is moving in a certain way, for example, as can be specified by a set of fitting parameters. Fitting parameters can include the translational properties of rigid bodies (such as the translational velocity of the object). and / or the rotation of the rigid body (such as the angular velocity of the object) (such as rotation center).
[0126] At block 524, method 500 may continue to map the first point set to the second point set using a computing device. For example, mapping the first point set to the second point set may include using an iterative nearest-point algorithm. In some embodiments, mapping the first point set to the second point set may include determining the position of the first point set after a time increment corresponding to the time difference between the first and second sensing data frames based on fitting parameters (forward inference). Similarly, in some embodiments, mapping the first point set to the second point set may include inferring the position of the second point set before a time decrement corresponding to the time difference between the second and first sensing data frames based on fitting parameters (backward inference). In some embodiments, mapping the first point set to the second point set is performed based on velocity data (e.g., radial velocity data) of at least some points in the first and / or second point sets. In some embodiments, mapping the first point set to the second point set includes determining the lateral velocity of the corresponding reflective surface based on the velocity data. The lateral velocity can be determined using velocity data obtained from the first and second sensors, for example, as per the... Figure 4A and Figure 4B As described.
[0127] At box 526, the computing device may determine an evaluation metric for the hypothesis based on the performed mapping. The evaluation metric may be a single value (or multiple values) characterizing how closely the predicted motion of the hypothesized object (using velocity values associated with points in the first and / or second point sets) matches (aligns) with the first and second point sets. For example, the evaluation metric may depend on a mismatch (forward-looking metric) between the inferred future position of the first point set and the actual position of the second point set, and / or a mismatch (backward-looking metric) between the inferred past position of the second point set and the actual position of the first point set. In some embodiments, the evaluation metric may be determined according to method 600 described below. In those embodiments, where the evaluation of the hypothesis involves matching velocity data with the motion of the rigid body, the evaluation metric may be based at least in part on the difference between the velocity data of the first point set and the velocity distribution of the rigid body. For example, the evaluation metric may be a weighted sum of the squared differences between the rigid body's velocity and the actual velocity measured by sensors (e.g., a first sensor and (optionally) a second sensor). As another example, the evaluation metric can describe the mismatch between the coordinates and / or velocities of the first set of points and the coordinates and / or velocities of the second set of points (e.g., the sum of weighted squared errors).
[0128] At box 520, the hypothetical evaluation may include determining the values of the fitting parameters based on an evaluation metric. In some embodiments, the fitting parameters are determined based on velocity data. The velocity data may include the radial velocities of at least some points in a first set of points or a second set of points, or both. In some embodiments, the velocity data may include radial velocities detected by multiple sensors (e.g., at least two sensors located at different positions from each other).
[0129] At box 530, the computing device can (e.g., based on an evaluation metric that meets predetermined criteria) identify a hypothetical object that matches a first set of points (at time in the first frame) and a second set of points (at time in the second frame), and that the motion of the hypothetical object is consistent with velocity data associated with the corresponding points (either the first set or the second set of points, or both). The predetermined criteria may be that the evaluation metric has a specific relationship with a threshold (e.g., higher than, equal to, or lower than) a threshold.
[0130] Alternatively, the computing device can also identify objects in the environment corresponding to the first and second point sets as specific types of objects (cars, trucks, buses, motorcycles, pedestrians, etc.) based on the evaluated assumptions.
[0131] At box 540, method 500 can continue using a computing device to determine the driving path of the AV based on an evaluation metric. More specifically, after using the evaluation metric to validate the hypothesis, the computing device can generate a representation of the identified object. This representation can be a set of geometric descriptors of the identified object, such as a set of coordinates relative to the object's bounding box, or a set of parameters identifying the relative orientation of various elements of the object (e.g., ribs or faces) and another set of parameters identifying the object's position and orientation in space (e.g., relative to the ground, other objects, or map layout). Those skilled in the art will recognize that the number of possible representations of various objects is infinite. The computing device can provide the representation of the identified object to the AV's control system (e.g., AVCS140). The control system can then determine the AV's driving path based on the provided representation. Based on the position and motion (translation and rotation) of the identified object, the control system can determine whether the AV should accelerate, brake, turn, change lanes, stop, reverse, etc. (or perform any combination of these actions). The control system can then output corresponding commands to the powertrain and steering 150, vehicle electronics 160, signaling 170, etc., to ensure that the AV follows the determined driving path.
[0132] Figure 6A flowchart depicts an example method 600 for evaluating an assumption that a point set of a first sensing frame corresponds to a point set of a second sensing frame, according to some embodiments of this disclosure. Method 600 may be performed in conjunction with block 520 of method 500 for Doppler-assisted object recognition and point cloud tracking for autonomous vehicle applications. At block 610, a computing device performing method 600 (e.g., a device executing software instructions issued by PCM 133 as part of perception system 132) may select the assumption that a first point set from the first sensing data frame corresponds to a second point set from the second sensing data frame. At block 620, the method may continue to obtain estimates of one or more components of the translational velocity or rotational (angular) velocity of the object assumed to be associated with the first point set using the computing device. For example, the components of the translational velocity and / or rotational velocity may be determined by fitting the velocity data of the first frame (and, optionally, the second frame) using rigid body equations.
[0133] At block 630, method 600 may continue using a computing device to predict the position of the hypothetical object after a time increment corresponding to the time difference between the first and second sensing data frames, based on estimates of one or more components of the translational velocity or one or more components of the rotational velocity of the hypothetical object. At block 640, the computing device may compare the predicted position with the coordinates of a second set of points and determine an evaluation metric, such as the weighted squared error of the position of the second set of points relative to the corresponding position predicted based on the estimated components of the translational velocity and / or rotational velocity.
[0134] Figure 7 A block diagram depicts an example computer device 700 for autonomous vehicle applications, capable of enabling Doppler-assisted object recognition and tracking, according to some embodiments of the present disclosure. The example computer device 700 can be connected to other computer devices in a LAN, intranet, extranet, and / or the Internet. The computer device 700 can operate with server capabilities in a client-server network environment. The computer device 700 can be a personal computer (PC), set-top box (STB), server, network router, switch, or bridge, or any device capable of executing a set of instructions (ordered or otherwise) specifying actions to be performed by the device. Furthermore, although only a single example computer device is shown, the term "computer" should also be considered to include any collection of computers that individually or jointly execute one or more sets of instructions to perform any one or more of the methods discussed herein.
[0135] Example computer device 700 may include processing device 702 (also called processor or CPU), main memory 704 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM)) etc.), static memory 706 (e.g., flash memory, static random access memory (SRAM) etc.), and auxiliary memory (e.g., data storage device 718), which can communicate with each other via bus 730.
[0136] Processing device 702 represents one or more general-purpose processing devices, such as microprocessors, central processing units, etc. More specifically, processing device 702 may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. Processing device 702 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. According to one or more aspects of this disclosure, processing device 702 may be configured to execute instructions that perform a method 500 for Doppler-assisted object recognition and point cloud tracking, and a method 600 for evaluating an assumption that a point set of a first sensing frame corresponds to a point set of a second sensing frame.
[0137] Example computer device 700 may also include a network interface device 708 that can be communicatively coupled to network 720. Example computer device 700 may also include a video display 710 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 712 (e.g., a keyboard), a cursor control device 714 (e.g., a mouse), and a sound signal generation device 716 (e.g., a speaker).
[0138] Data storage device 718 may include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 728 thereon storing one or more sets of executable instructions 722. According to one or more aspects of this disclosure, the executable instructions 722 may include executable instructions for performing a method 500 of using Doppler-assisted object recognition and point cloud tracking, and a method 600 of evaluating an assumption that a point set of a first sensing frame corresponds to a point set of a second sensing frame.
[0139] The executable instructions 722, during their execution by the example computer device 700, may also reside wholly or at least partially in the main memory 704 and / or the processing device 702, which also constitute a computer-readable storage medium. The executable instructions 722 may also be transmitted or received over a network via a network interface device 708.
[0140] Although the computer-readable storage medium 728 is Figure 7 As illustrated as a single medium, the term "computer-readable storage medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of VM operation instructions. The term "computer-readable storage medium" should also be understood to include any medium capable of storing or encoding a set of instructions for execution by a machine, which causes the machine to perform any one or more of the methods described herein. Therefore, the term "computer-readable storage medium" should be understood to include, but is not limited to, solid-state storage as well as optical and magnetic media.
[0141] Some parts of the detailed description above are presented in terms of algorithms and symbolic representations of operations on data bits within computer memory. These descriptions and representations of algorithms are means used by those skilled in the art of data processing to most efficiently communicate the essence of their work to others skilled in the art. Here, and generally speaking, an algorithm is considered a self-consistent sequence of steps leading to a desired result. These steps are those that require physical manipulation of physical quantities. Typically, though not always, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. It is sometimes convenient, primarily for customary usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.
[0142] However, it should be remembered that all these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, as is apparent from the following discussion, it should be understood that throughout the description, the use of terms such as “identify,” “determine,” “store,” “adjust,” “make,” “return,” “compare,” “create,” “stop,” “load,” “copy,” “throw,” “replace,” “execute,” etc., refers to the actions and processes of a computer system (or similar electronic computing device) that manipulate or transform data represented as physical (electronic) quantities in the registers and memory of the computer system into other data similarly represented as physical quantities in the memory or registers or other such information storage, transmission, or display devices of the computer system.
[0143] Examples of this disclosure also relate to apparatus for performing the methods described herein. This apparatus may be specifically constructed for a desired purpose, or it may be a general-purpose computer system selectively programmed by a computer program stored in a computer system. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk (including optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0144] The methods and displays presented herein are not inherently connected to any particular computer or other device. Various general-purpose systems can be used in conjunction with the teachings and procedures herein, or more specialized devices can be constructed to perform the required method steps, as may prove convenient. The necessary structures for various such systems will appear as set forth in the description below. Furthermore, the scope of this disclosure is not limited to any particular programming language. It should be understood that the teachings of this disclosure can be implemented using various programming languages.
[0145] It should be understood that the above description is intended to be illustrative and not restrictive. Many other examples of implementation will become apparent to those skilled in the art upon reading and understanding the above description. Although specific examples have been described in this disclosure, it should be recognized that the systems and methods of this disclosure are not limited to the examples described herein but can be practiced with modifications within the scope of the appended claims. Accordingly, the specification and drawings are considered illustrative and not restrictive. Therefore, the scope of this disclosure should be determined by reference to the appended claims and the full scope of their equivalents.
Claims
1. A method for autonomous vehicle AV, comprising: Multiple points are obtained by a computing device to image the environment surrounding the AV, wherein each of the multiple points is: Corresponding to the reflection of signals emitted by the sensing system of the AV from the surface of an object in the environment, and This includes the position data of the corresponding reflective surface and the velocity data of the corresponding reflective surface; Multiple hypotheses are generated, each of which maps a first set of points among the multiple points associated with a first time to a second set of points among the multiple points associated with a second time. For each of the plurality of hypotheses: The computing device uses the position data and the velocity data to predict the motion of the first set of points between the first time and the second time. Using the motion of the second point set and the predicted motion of the first point set, an evaluation metric for the hypothesis is obtained; and The driving path of the AV is determined based on the evaluation metrics obtained for the plurality of assumptions.
2. The method according to claim 1, wherein, Determining the driving path of the AV includes: Based on the evaluation metrics obtained for the multiple hypotheses, the objects corresponding to the environment of the first point set and the second point set are identified; and The system provides a representation of the identified object to the control system of the AV to determine the driving path of the AV.
3. The method according to claim 1, wherein, The sensing system of the AV includes a coherent optical detection and ranging device (LIDAR), and the signal emitted by the sensing system is a phase-modulated or frequency-modulated electromagnetic wave.
4. The method according to claim 3, wherein, The coherent LIDAR is a frequency-modulated continuous wave LIDAR.
5. The method according to claim 1, wherein, The sensing system of the AV includes a first sensor, and wherein the velocity data of the corresponding reflective surface includes a first component of the velocity of the corresponding reflective surface along a first direction from the first sensor to the corresponding reflective surface.
6. The method according to claim 5, wherein, The sensing system of the AV includes a second sensor located at a different location from the first sensor, and wherein the velocity data of the corresponding reflective surface includes a second component of the velocity of the corresponding reflective surface along a second direction from the second sensor to the corresponding reflective surface.
7. The method according to claim 6, wherein, Mapping the first point set to the second point set includes determining the lateral velocity of the corresponding reflective surface based on the velocity data.
8. The method according to claim 1, wherein, The motion of the first set of points is constrained by a rigid body that performs a combination of translational and rotational motion.
9. The method according to claim 8, wherein, The evaluation metric is based, in at least part, on at least one of the following: A first metric representing the difference between one or more locations associated with the motion of the first set of points and one or more locations associated with the motion of the second set of points; and A second metric representing the difference between one or more velocities associated with the motion of the first set of points and one or more velocities associated with the motion of the second set of points.
10. The method according to claim 1, wherein, Mapping the first set of points to the second set of points involves using an iterative nearest-point algorithm.
11. The method according to claim 1, wherein, Mapping the first point set to the second point set includes: Associating multiple fitting parameters with the first point set, wherein the multiple fitting parameters characterize the motion of the rigid body; and The prediction of the motion of the first point set includes: inferring the positional change of the first point set associated with the time difference between the first time and the second time based on the plurality of fitting parameters.
12. The method of claim 11, further comprising: The values of the plurality of fitting parameters are determined at least in part based on minimizing the mismatch between the predicted motion of the first set of points and the position of the second set of points.
13. The method according to claim 12, wherein, At least some of the plurality of fitting parameters are determined based on the velocity data, wherein the velocity data includes radial velocities of at least some of the first set of points or the second set of points.
14. The method according to claim 11, wherein, The second time is earlier than the first time, and the motion of the first set of points is a time-reverse motion.
15. The method according to claim 11, wherein, The plurality of fitting parameters include at least one of the rotation of the rigid body or the translation of the rigid body between the first time and the second time.
16. A system for autonomous vehicle AV, comprising: Memory, which stores instructions; and The computing device executes the instructions from the memory to: Obtain multiple points for imaging the environment surrounding the AV, wherein each of the multiple points is: The reflection of signals emitted by the sensing system of the AV, corresponding to the surface of an object in the environment, and This includes the position data of the corresponding reflective surface and the velocity data of the corresponding reflective surface; Multiple hypotheses are generated, each of which maps a first set of points among the multiple points associated with a first time to a second set of points among the multiple points associated with a second time. For each of the plurality of hypotheses: The computing device uses the position data and the velocity data to predict the motion of the first set of points between the first time and the second time. Using the motion of the second point set and the predicted motion of the first point set, an evaluation metric for the hypothesis is obtained; and The driving path of the AV is determined based on the evaluation metrics obtained for the plurality of assumptions.
17. The system according to claim 16, wherein, The sensing system of the AV includes a coherent optical detection and ranging device (LIDAR), and the signal emitted by the sensing system is a phase-modulated or frequency-modulated electromagnetic wave.
18. The system according to claim 16, wherein, The sensing system of the AV includes a first sensor, and wherein the velocity data of the corresponding reflective surface includes a first component of the velocity of the corresponding reflective surface along a first direction from the first sensor to the corresponding reflective surface.
19. The system according to claim 18, wherein, The sensing system of the AV includes a second sensor located at a different location from the first sensor, and wherein the velocity data of the corresponding reflective surface includes a second component of the velocity of the corresponding reflective surface along a second direction from the second sensor to the corresponding reflective surface.
20. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a computing device, causes the computing device to: Multiple points were obtained to image the environment surrounding the autonomous vehicle (AV), among which, Each of the plurality of points: Corresponding to the reflection of signals emitted by the sensing system of the AV from the surface of an object in the environment, and This includes the position data of the corresponding reflective surface and the velocity data of the corresponding reflective surface; Multiple hypotheses are generated, each of which maps a first set of points among the multiple points associated with a first time to a second set of points among the multiple points associated with a second time. For each of the plurality of hypotheses: The computing device uses the position data and the velocity data to predict the motion of the first set of points between the first time and the second time. Using the motion of the second point set and the predicted motion of the first point set, an evaluation metric for the hypothesis is obtained; as well as The driving path of the AV is determined based on the evaluation metrics obtained for the plurality of assumptions.
21. The computer-readable medium of claim 20, wherein, The sensing system of the AV includes a coherent optical detection and ranging device (LIDAR), and the signal emitted by the sensing system is a phase-modulated or frequency-modulated electromagnetic wave.
22. The computer-readable medium of claim 20, wherein, The sensing system of the AV includes a first sensor, and wherein the velocity data of the corresponding reflective surface includes a first component of the velocity of the corresponding reflective surface along a first direction from the first sensor to the corresponding reflective surface.
23. The computer-readable medium according to claim 22, wherein, The sensing system of the AV includes a second sensor located at a different location from the first sensor, and wherein the velocity data of the corresponding reflective surface includes a second component of the velocity of the corresponding reflective surface along a second direction from the second sensor to the corresponding reflective surface.
Citation Information
Patent Citations
Control of Autonomous Vehicle Based on Environmental Object Classification Determined Using Phase Coherent LIDAR Data
US20190317219A1
Surround vehicle tracking and motion prediction
WO2019136479A1