Object tracking to support autonomous vehicle navigation

The optical object tracking system in autonomous vehicles improves object detection and navigation by using image-based tracking with predicted trajectories, addressing the challenges of tracking and navigating around objects in complex environments.

DE102020134834B4Active Publication Date: 2026-01-29MOTIONAL AD LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102020134834
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2020-12-23
Publication Date
2026-01-29
Estimated Expiration
2040-12-23

AI Technical Summary

Technical Problem

Autonomous vehicles face challenges in reliably detecting, continuously tracking, and anticipating the future actions of objects in their surroundings, particularly in complex environments.

Method used

An optical object tracking system that utilizes image-based tracking to determine the position of objects using optical sensors, incorporating location information from satellite navigation and refining positions with predicted trajectories, and combining detected and predicted positions to enhance accuracy, applicable to various types of moving and stationary objects.

Benefits of technology

Enhances the ability of autonomous vehicles to accurately track and navigate around objects, reducing the impact of temporary errors and improving navigation efficiency in diverse environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Procedures, encompassing the following: Obtained - using a processing circuit - from tracking data associated with an object, where the tracking data includes an initial position of the object at an initial time; Predicting – using the processing circuit – a second position of the object at a second time point, which lies after the first time point, based on the tracking data; Capturing – using a sensor from an autonomous vehicle – an image containing the object at the second time point; Detecting – using the processing circuit – the object at a third position using the image; and according to a provision that the third position lies within a threshold distance from the second position: Determining a fourth position of the object at the second time point based on the predicted second position and the detected third position of the object; and Navigating the autonomous vehicle - using a control circuit - according to the fourth position of the object; wherein the threshold distance is based at least partially on a difference between a feature of the object visible in the captured image and the feature of the object in one or more previously captured images used to generate the tracking data, and where the threshold distance remains constant in response to the fact that predicted changes in the orientation of the object relative to the autonomous vehicle are identified as the cause of the difference in the feature.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF INVENTION

[0001] This description concerns an optical object tracking system that supports the navigation of an autonomous vehicle. BACKGROUND

[0002] Autonomous vehicles can be used to transport people and / or cargo (e.g., packages, objects, or other items) from one place to another. For example, an autonomous vehicle can navigate to a person's location, wait for the person to board, and then navigate to a specific destination (e.g., a location chosen by the person). To navigate their environment, these autonomous vehicles are equipped with various types of sensors to detect objects in the surroundings.

[0003] DE 10 2016 114 168 A1 relates to a method for detecting an object in the vicinity of a motor vehicle using a sequence of images of the vicinity provided by a camera of the motor vehicle, comprising the steps of: recognizing a first object feature in a first image of the sequence, wherein the first object feature describes at least a part of the object in the vicinity; estimating a position of the object in the vicinity based on a predetermined motion model that describes a movement of the object in the vicinity; determining a predictor feature in a second image following the first image in the sequence based on the first object feature and the estimated position; determining a second object feature in the second image; and assigning the second object feature to the predictor feature in the second image if a predetermined assignment criterion is met.and confirming the second object feature as originating from the object, if the second object feature is assigned to the predictive feature. DE 10 2016 203 472 A1 describes a method for detecting an object in the vicinity of a vehicle. The method includes determining a first position and a first radial velocity of a first object point at a first time. Furthermore, the method includes determining a second position of a second object point at a second time. The method also includes determining a predicted position of the first object point at the second time, based on the first position and the first radial velocity. Finally, the method includes detecting an object in the vicinity of the vehicle, based on the predicted position of the first object point and on the second position of the second object point. SUMMARY

[0004] The use of sensors to detect objects near autonomous vehicles is known; however, improvements are desirable in the vehicle's capabilities for detecting, continuously tracking, and anticipating the future actions of these objects. The subject matter described in this document focuses on a computer system and techniques for detecting objects in the surrounding environment of an autonomous vehicle. The computer system is generally designed to receive input from one or more sensors of the vehicle, to detect one or more objects in the vehicle's surrounding environment based on the received input, and to operate the vehicle based on the object detection.

[0005] In particular, the autonomous vehicle is designed to reliably identify and track objects located in its vicinity. An exemplary tracking procedure involves: obtaining – using a processing circuit – tracking data associated with an object, where the tracking data includes an initial position of the object at a first time point; predicting – using the processing circuit – a second position of the object at a second time point, after the first, based on the tracking data; capturing – using a sensor of the autonomous vehicle – an image containing the object at the second time point; and detecting – using the processing circuit – the object at a third position using the image.According to a determination that the third position lies within a threshold distance from the second position: determining a fourth position of the object at the second time based on the predicted second position and the detected third position of the object; and navigating the autonomous vehicle - using a control circuit - according to the fourth position of the object.

[0006] These and other aspects, features, and implementations can be expressed as methods, devices, systems, components, program products, means or steps for performing a function, and in other ways.

[0007] These and other aspects, features and implementations will become apparent from the following descriptions, including the patent claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] They show: Fig. 1. An example of an autonomous vehicle with autonomous capabilities; Fig. 2. An example of a “cloud” computing environment; Fig. 3 a computer system; Fig. 4 an exemplary architecture for an autonomous vehicle; Fig. 5. An example of inputs and outputs that can be used by a perception module; Fig. 6. An example of a LiDAR system; Fig. 7 the LiDAR system in operation; Fig. 8. The functionality of the LiDAR system with additional details; Fig. 9 a block diagram of the relationships between the inputs and outputs of a planning module; Fig. 10 a directed graph used in railway planning; Fig. 11 a block diagram of the inputs and outputs of a control module; Fig. 12 a block diagram of the inputs, outputs and components of a controller; Fig. 13 a block diagram of a visual system for object tracking; Fig. 14A to 14C Block diagrams illustrating the processing of webs in an image plane and webs in a ground plane according to some embodiments; Fig. 15 a flowchart of the procedure, which refers to a block of the Fig. 14A is described in more general terms. Fig. Figures 16A to 16B are exemplary representations of images captured by an object detection system positioned on board an autonomous vehicle as the autonomous vehicle approaches a traffic intersection; Fig. 17A to 17C a top view of the in the Fig. The intersection shown in 16A to 16B, together with an autonomous vehicle and its optical sensor; Fig. 18 a flowchart of a process for assigning newly detected objects to previously detected obsolete paths, more generally with reference to a block of the Fig. 14B is described. DETAILED DESCRIPTION

[0009] For explanatory purposes, numerous specific details are presented in the following description to facilitate a thorough understanding of the present invention. However, it is obvious that the present invention can be implemented without these specific details. In other instances, known structures and devices are represented in the form of block diagrams to avoid unnecessarily complicating the understanding of the present invention.

[0010] The drawings illustrate specific arrangements or sequences of schematic elements, such as those representing devices, modules, instruction blocks, and data elements, for ease of description. However, it should be obvious to those skilled in the art that the specific sequence or arrangement of schematic elements in the drawings is not intended to imply that a particular sequence or order of processing, or a separation of processes, is required. Furthermore, the inclusion of a schematic element in a drawing is not intended to imply that this element is required in all embodiments, or that the features represented by this element cannot be incorporated into other elements or combined with other elements in some embodiments.

[0011] Furthermore, when connecting elements, such as solid or dashed lines or arrows, are used in the drawings to represent a connection, relationship, or association between or among two or more other schematic elements, the absence of such connecting elements should not be interpreted as meaning that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings to avoid ambiguity in the disclosure. Moreover, for the sake of simplicity, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents communication of signals, data, or instructions, those skilled in the art should understand that such an element can represent one or more signal paths (e.g.,a bus) represents the ability to influence communication.

[0012] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. The following detailed description presents numerous specific details to provide a thorough understanding of the various described embodiments. However, it will be obvious to those skilled in the art that the various described embodiments can be implemented without these specific details. In other cases, generally known methods, procedures, components, circuits, and networks have not been described in detail in order to avoid unnecessarily complicating the understanding of certain aspects of the embodiments.

[0013] Several features are described below, each of which can be used independently or in any combination with other features. However, a single feature may not solve any of the problems discussed above, or may only solve one of them. Some of the problems discussed above may not be completely solved by any of the features described here. Although headings are given, information relating to a particular heading, but not found in the section with that heading, may also be found elsewhere in this description. Embodiments are described here according to the following outline: 1. General Overview 2. Hardware Overview 3. Architecture of the autonomous vehicle 4. Inputs from the autonomous vehicle 5. Planning the autonomous vehicle 6. Control of the autonomous vehicle 7. Computing system for object detection using columns 8. Example point clouds and columns 9. Exemplary process for object detection and vehicle operation based on object detection General overview

[0014] Autonomous vehicles operating in complex environments (e.g., urban areas) present a significant technological challenge. To navigate these environments, autonomous vehicles use sensors such as optical sensors, LiDAR, and / or radar to detect various types of objects, including vehicles, pedestrians, and bicycles, in real time. One approach to object detection utilizes image analysis to track objects near an autonomous vehicle.

[0015] In particular, the system and techniques described here implement an image-based tracking system that can determine the position of objects surrounding an autonomous vehicle based on the objects' positions in images captured by one or more optical sensors mounted on or in the immediate vicinity of the autonomous vehicle. The position of the optical sensor at the time of image capture can be known from location information provided by a satellite-based navigation receiver of the autonomous vehicle. If the position of the optical sensor at the time of capture is known, the position of an object within an image provides a bearing line along which the object was located at the time the image was captured.Prior and / or subsequent detection of the object in additional images allows for further refinement of the object's position along the bearing line. In this way, location information for the object can be incorporated into an active trajectory that the autonomous vehicle can reference when navigating around the object.

[0016] In addition to identifying object positions based on imagery, an object's position can be refined by referencing previously acquired detection data from previously acquired images. Analyzing previously acquired tracking data allows for the determination of a predicted object position. This predicted position can be helpful in several ways. First, it can help identify which new object detections should be associated with which active trajectory. For example, a vehicle traveling at high speed relative to the autonomous vehicle may be a considerable distance from its last detected position.By predicting the position of a rapidly moving object based on its previously detected velocity, new trajectories can be linked to the predicted object locations to reconcile new and old tracking data. Secondly, periodic inaccuracies in the tracking data, caused by unexpected vibrations of the autonomous vehicle and / or temporary instability of the image acquisition sensor, can produce image data that differs significantly from previously acquired tracking data. In some embodiments, the detected position and the predicted position can be combined to generate a weighted average of the two positions. This can help reduce the severity of temporary errors resulting from image acquisition problems.In some embodiments, metadata associated with the image can include inputs—for example, from a gyroscope or other motion detection device—that serve to indicate stabilization problems of the optical sensor. Thirdly, the predicted position can also be used to correlate new tracking data with tracking data associated with an older or outdated trajectory. In this way, historical data associated with an object can be re-identified to more accurately predict the object's behavior.

[0017] In addition to the advantages described above, this system has the benefit of being applicable to many different types of objects—including moving objects such as cars and trucks, as well as stationary objects such as fire hydrants, lampposts, buildings, and the like. The described embodiments are also not necessarily limited to objects on the ground and could also be applied to vehicles that can move in the air or on water. Furthermore, it should be noted that the described system can be designed to have a stateless data flow pipeline, which makes it easier to run processes in parallel. This is because all state information associated with the tracking data remains, at least partially, with the associated object's path. Hardware overview

[0018] The Fig. Figure 1 shows an example of an autonomous vehicle 100 with autonomous capability.

[0019] As used here, the term “autonomous capability” refers to a function, feature or device that enables a vehicle to operate partially or completely without human intervention in real time – including, without limitation, fully autonomous vehicles, highly autonomous vehicles and conditionally autonomous vehicles.

[0020] As used here, an autonomous vehicle (AV) is a vehicle that possesses an autonomous capability.

[0021] As used here, the term "vehicle" encompasses means of transport for goods or people. Examples include automobiles, buses, trains, airplanes, drones, trucks, boats, ships, underwater vehicles, airships, etc. A driverless car is an example of a vehicle.

[0022] As used here, "trajectory" refers to a path or route for navigating an AV from a first spatiotemporal location to a second spatiotemporal location. In one embodiment, the first spatiotemporal location is referred to as the starting or initial location, and the second spatiotemporal location is referred to as the destination, final location, target, target position, or destination point. In some examples, a trajectory consists of one or more segments (e.g., road segments), and each segment consists of one or more blocks (e.g., portions of a lane or an intersection). In one embodiment, the spatiotemporal locations correspond to real-world locations. For example, the spatiotemporal locations are pickup or drop-off points for collecting or delivering people or goods.

[0023] As used here, the term "sensor(s)" encompasses one or more hardware components that detect information about the sensor's surrounding environment. Some of the hardware components may include sensing components (e.g., image sensors, biometric sensors), transmitting and / or receiving components (e.g., laser or radio frequency wave transmitters and receivers), electronic components such as analog-to-digital converters, a data storage device (such as RAM and / or non-volatile memory), software or firmware components, and data processing components such as an ASIC (application-specific integrated circuit), a microprocessor, and / or a microcontroller.

[0024] As used here, a “scene description” is a data structure (e.g., a list) or data stream that contains one or more classified or labeled objects detected by one or more sensors on the AV vehicle or provided by a source outside the AV.

[0025] As used here, a "road" is a physical area that can be traversed by a vehicle and may correspond to a named thoroughfare (e.g., city street, highway, etc.) or an unnamed thoroughfare (e.g., a driveway to a house or office building, a section of a parking lot, a section of an empty lot, a dirt track in a rural area, etc.). Because some vehicles (e.g., four-wheel-drive pickup trucks, sport utility vehicles, etc.) can traverse a variety of physical areas not specifically designed for vehicular traffic, a "road" may be a physical area that has not been formally defined as a thoroughfare by a municipality or other governmental or administrative body.

[0026] As used here, a "lane" is a portion of a road that can be traversed by a vehicle and can correspond to most or all of the space between lane markings, or only a portion (e.g., less than 50%) of the space between the lane markings. For example, a road with widely spaced lane markings could accommodate two or more vehicles between the markings, allowing one vehicle to overtake another without crossing the lane markings, and could therefore be interpreted as having one lane narrower than the space between the lane markings, or as having two lanes between the lane markings. A lane can also be interpreted without lane markings. For example, a lane can be defined based on physical features of an environment, e.g.,Rocks and trees along a main road in a rural area are defined.

[0027] “One or more” includes a function that is executed by one element, a function that is executed by more than one element, e.g. in a distributed manner where multiple functions are executed by one element, multiple functions are executed by multiple elements, or any combination of the above.

[0028] It is also understood that, although the terms "first," "second," etc., are used here in some cases to describe different elements, these elements should not be restricted by these terms. These terms are used merely to distinguish one element from another. For example, a first contact could be called a second contact, and analogously, a second contact could be called a first contact, without altering the scope of the various described embodiments. The first contact and the second contact are both contacts, but—unless otherwise specified—they are not the same contact.

[0029] The terminology used in the description of the various embodiments described herein serves solely to describe particular embodiments and is not to be understood restrictively. As used in the description of the various described embodiments and the attached patent claims, the singular forms "a," "an," and "the," "a," and "a" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It is also understood that the term "and / or," as used here, refers to and encompasses any and all possible combinations of one or more of the associated listed elements. Furthermore, it is understood that the terms "contains" and "includes" are to be understood as including or encompassing all possible combinations of one or more of the associated listed elements.“Includes”, “comprises” and / or “comprehensive” when used in this description indicate the presence of the features, integers, steps, operations, elements and / or components set forth, but do not exclude the presence or addition of one or more further features, integers, steps, operations, elements, components and / or groups thereof.

[0030] As used here, the term "if" or "in case" is interpreted—depending on the context—optionally to mean "when" or "at" or "in response to determining" or "in response to detection." Similarly, the phrase "when determined" or "when [a specified condition or event] is detected" is optionally interpreted—depending on the context—to mean "in determining" or "in response to determining" or "in detecting [the specified condition or event]" or "in response to detecting [the specified condition or event]."

[0031] As used here, an AV system refers to the AV together with the arrangement of hardware, software, stored data, and real-time generated data that supports the operation of the AV. In one embodiment, the AV system is integrated into the AV. In another embodiment, the AV system is distributed across multiple locations. For example, part of the AV system's software is implemented in a cloud computing environment, which is described below with reference to the Fig. The cloud computing environment described in section 2 is similar to 200.

[0032] This document describes, in general terms, technologies applicable to vehicles with one or more autonomous capabilities, including fully autonomous vehicles, highly autonomous vehicles, and conditionally autonomous vehicles, such as those designated as Level 5, 4, and 3 vehicles (for more details on the classification of vehicle autonomy levels, see SAE International Standard J3016: "Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems"). The technologies described in this document are also applicable to semi-autonomous vehicles and driver-assisted vehicles, such as those designated as Level 2 and 1 vehicles (see SAE International Standard J3016: "Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems").In one embodiment, one or more of the vehicle systems of levels 1, 2, 3, 4, and 5 can automate certain vehicle functions (e.g., steering, braking, and map use) under specific operating conditions based on the processing of sensor inputs. Vehicles of all levels—from fully autonomous vehicles to human-operated vehicles—can benefit from the technologies described in this document.

[0033] Referring to the Fig. 1 an AV system 120 controls the AV 100 along a trajectory 198 through an environment 190 to a destination 199 (sometimes referred to as end location), avoiding objects (e.g. natural obstacles 191, vehicles 193, pedestrians 192, cyclists and other obstacles) and observing road traffic rules (e.g. operating rules or driving preferences).

[0034] In one embodiment, the AV system 120 includes the devices 101, which are instrumented to receive operating instructions from the computing processors 146 and to act upon them. In one embodiment, the computing processors 146 are described below with reference to the Fig. The processor 304 described in Section 3 is similar. Examples of the devices 101 include a steering control device 102, brakes 103, a gearshift, an accelerator pedal or other mechanisms for acceleration control, windshield wipers, side door locks, window controls and turn signals.

[0035] In one embodiment, the AV system 120 comprises sensors 121 for measuring or deriving properties of the state or condition of the AV 100, such as the position, linear and angular velocity and acceleration, and the direction of travel of the AV (e.g., the orientation of the front end of the AV 100). Examples of the sensors 121 are GPS, inertial measurement units (IMUs) that measure both linear vehicle accelerations and angular rates, wheel speed sensors for measuring or estimating wheel slip ratios, wheel brake pressure or brake torque sensors, engine torque or wheel torque sensors, and steering angle and angular velocity sensors.

[0036] In one embodiment, the sensors 121 also include sensors for detecting or measuring properties of the AV's environment. For example, the sensors 121 include monocular or stereo video cameras 122 in the visible light spectrum, the infrared spectrum, or the thermal spectrum (or both), a LiDAR 123, radar, ultrasonic sensors, time-of-flight (TOF) depth sensors, velocity sensors, temperature sensors, humidity sensors, and precipitation sensors.

[0037] In one embodiment, the AV system 120 comprises a data storage unit 142 and a memory 144 for storing machine instructions assigned to computer processors 146 or data acquired by sensors 121. In one embodiment, the data storage unit 142 is similar to the ROM 308 or the storage device 310 described below in relation to the Fig. 3. In one embodiment, the memory 144 is similar to the main memory 306 described below. In one embodiment, the data storage unit 142 and the memory 144 store historical information, real-time information, and / or predictive information about the environment 190. In one embodiment, the stored information includes maps, driving performance, updates on traffic jams, or weather conditions. In one embodiment, data relating to the environment 190 are transmitted via a communication channel from a remotely located database 134 to the AV 100.

[0038] In one embodiment, the AV system 120 includes communication devices 140 for transmitting measured or derived properties of the states and conditions of other vehicles, such as positions, linear and angular velocities, linear and angular accelerations, and linear and angular orientations, to the AV 100. These devices include vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication devices and devices for wireless communication via point-to-point or ad-hoc networks, or both. In one embodiment, the communication devices 140 communicate via the electromagnetic spectrum (including radio communication and optical communication) or via other media (e.g., air and acoustic media).A combination of vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication (and in some implementations, one or more other communication types) is sometimes referred to as vehicle-to-everything (V2X) communication. V2X communication typically conforms to one or more communication standards for communicating with, between, or among autonomous vehicles.

[0039] In one embodiment, the communication devices 140 have communication interfaces. For example, the communication interfaces include wired, wireless, WiMAX, WiFi, Bluetooth, satellite, cellular, optical, near-field, infrared, or radio interfaces. The communication interfaces transmit data from a remotely located database 134 to the AV system 120. In one embodiment, the remotely located database 134 is, as in the Fig. 2 described, embedded in a cloud computing environment 200. The communication interfaces 140 transmit the data collected by the sensors 121 or other data relating to the operation of the AV 100 to the remotely located database 134. In one embodiment, the communication interfaces 140 transmit information relating to teleoperations to the AV 100. In another embodiment, the AV 100 communicates with other remote servers 136 (e.g., with "cloud" servers).

[0040] In one embodiment, the remotely located database 134 also stores and transmits digital data (e.g., data such as location information for highways and other roads). Such data is stored in the memory 144 of the AV 100 or transmitted from the remotely located database 134 to the AV 100 via a communication channel.

[0041] In one embodiment, the remotely located database 134 stores and transmits historical information about driving characteristics (e.g., speed and acceleration profiles) of vehicles that previously traveled along the trajectory 198 at similar times of day. In one implementation, such data can be stored in the memory 144 of the AV 100 or transmitted from the remotely located database 134 to the AV 100 via a communication channel.

[0042] The computing devices 146 arranged in the AV 100 algorithmically generate control actions based on both real-time sensor data and previous information, enabling the AV system 120 to perform its autonomous driving capabilities.

[0043] In one embodiment, the AV system 120 includes computer peripherals 132 coupled to the computing units 146 to provide information and warnings to a user (e.g., an inmate or a remote user) of the AV 100 and to receive input from it. In one embodiment, the peripherals 132 resemble the display 312, the input device 314, and the cursor control 316, which are subsequently referred to in the Fig. 3. The connection is wireless or wired. Any two or more of the interface devices can be integrated into a single device.

[0044] The Fig. Figure 2 illustrates an example of a "cloud" computing environment. Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services). In typical cloud computing systems, the machines used to deliver the services provided by the cloud are located in one or more large cloud data centers. Referring to the Fig. 2. Cloud computing environment 200 includes cloud data centers 204a, 204b, and 204c, which are interconnected via cloud 202. Data centers 204a, 204b, and 204c provide cloud computing services to computer systems 206a, 206b, 206c, 206d, 206e, and 206f, which are connected to cloud 202.

[0045] The cloud computing environment 200 includes one or more cloud data centers. Generally, a cloud data center refers to, for example, the one in the Fig. 2 Cloud data center 204a shown, referring to the physical arrangement of servers that make up a cloud, for example the one in the Fig. 2 Cloud 202 shown, or constitute a specific part of a cloud. In the cloud data center, servers are physically arranged in rooms, groups, rows, and racks, for example. A cloud data center has one or more zones, each containing one or more server rooms. Each room has one or more rows of servers, and each row contains one or more racks. Each rack contains one or more individual server nodes. In some implementations, servers are grouped into zones, rooms, racks, and / or rows based on the physical infrastructure requirements of the data center facility, which include power, energy, thermal, heat, and / or other requirements. In one embodiment, the server nodes resemble those shown in the Fig. The computer system described in section 3. Data center 204a has many computing systems distributed across many racks.

[0046] Cloud 202 comprises the cloud data centers 204a, 204b, and 204c, along with the network and interconnection resources (e.g., network devices, nodes, routers, switches, and network cables) that connect the cloud data centers 204a, 204b, and 204c and enable the computing systems 206a to 206f to access cloud computing services. In one embodiment, the network is a combination of one or more local area networks, wide area networks, or internetworks coupled via wired or wireless connections using terrestrial or satellite links. Data exchanged over the network is transmitted using any number of network layer protocols, such as Internet Protocol (IP), Multiprotocol Label Switching (MPLS), Asynchronous Transfer Mode (ATM), Frame Relay, etc.In embodiments where the network is a combination of several subnetworks, different network layer protocols are used for each of the underlying subnetworks. In some embodiments, the network represents one or more interconnected internet networks, such as the public internet.

[0047] The consumers of the computing systems 206a to 206f or the cloud computing services are connected to the cloud 202 via network connections and network adapters. In one embodiment, the computing systems 206a to 206f are implemented as various computing devices, for example, as servers, desktops, laptops, tablets, smartphones, Internet of Things (IoT) devices, autonomous vehicles (including cars, drones, shuttles, trains, buses, etc.), and consumer electronics. In another embodiment, the computing systems 206a to 206f are implemented in other systems or as part of other systems.

[0048] The Fig. Figure 3 illustrates a Computer System 300. In one implementation, the Computer System 300 is a specialized computing device. The specialized computing device is hardwired to perform the techniques or includes digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) permanently programmed to perform the techniques, or it may include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such specialized computing devices may also combine user-defined hardwired logic, ASICs, or FPGAs with user-specific programming to achieve the techniques.In various embodiments, the special computer devices are desktop computer systems, portable computer systems, handheld devices, network devices or other devices that contain hardwired and / or program logic for implementing the techniques.

[0049] In one embodiment, the computer system 300 includes a bus 302 or other communication mechanism for transmitting information and a hardware processor 304 coupled to a bus 302 for processing information. The hardware processor 304 is, for example, a general-purpose microprocessor. The computer system 300 also includes main memory 306, for example, random-access memory (RAM) or other dynamic storage device, which is coupled to the bus 302 for storing information and instructions executable by the processor 304. In one implementation, the main memory 306 is used during the execution of instructions executable by the processor 304 to store temporary variables or other intermediate information.If such instructions are stored in non-volatile storage media that the 304 processor can access, the 300 computer system becomes a special purpose machine, customized to perform the operations specified in the instructions.

[0050] In one embodiment, the computer system 300 further includes a read-only memory (ROM) 308 or other static storage device, which is coupled to the bus 302 for storing static information and instructions for the processor 304. A storage device 310, such as a magnetic disk, an optical disk, a solid-state drive, or a three-dimensional cross-point memory, is provided and coupled to the bus 302 for storing information and instructions.

[0051] In one embodiment, the computer system 300 is coupled via the bus 302 to a display 312, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, light-emitting diode (LED) display, or organic light-emitting diode (OLED) display, for displaying information to a computer user. An input device 314, which includes alphanumeric and other keys, is coupled to the bus 302 to transmit information and command selections to the processor 304. Another type of user input device is a cursor control 316, such as a mouse, trackball, touch-enabled display, or cursor direction keys, for transmitting directional information and command selections to the processor 304 and for controlling cursor movement on the display 312. This input device typically has two degrees of freedom in two axes, a first axis (e.g., an x-axis) and a second axis (e.g.,a y-axis), which allows the device to specify positions in a plane.

[0052] According to one embodiment, the techniques described herein are performed by the computer system 300 in response to the processor 304 executing one or more sequences of instructions contained in the main memory 306. Such instructions are read into the main memory 306 from another storage medium, such as the storage device 310. The execution of the instruction sequences contained in the main memory 306 causes the processor 304 to perform the process steps described herein. In alternative embodiments, a hard-wired circuit arrangement is used instead of, or in combination with, software instructions.

[0053] As used here, the term "storage media" refers to any non-volatile media that store data and / or instructions that cause a machine to operate in a specific way. Such storage media include non-volatile and / or volatile media. Non-volatile media include, for example, optical disks, magnetic disks, solid-state drives, or three-dimensional "crosspoint" memory, such as the Storage Device 310. Volatile media include dynamic memory, such as the Main Memory 306.Common forms of storage media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with hole patterns, a RAM, a PROM, an EPROM, a FLASH EPROM, NV-RAM or any other memory chip or any other memory cartridge.

[0054] Storage media differ from transmission media but can be used in conjunction with them. Transmission media are involved in the transfer of information between storage media. Examples of transmission media include coaxial cables, copper wire, and optical fibers, including the wires that make up the 302 bus. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data transmissions.

[0055] In one embodiment, various forms of media are involved in transmitting one or more sequences of one or more instructions to the processor 304 for execution. For example, the instructions are initially stored on a magnetic disk or solid-state drive of a remote computer. The remote computer loads the instructions into its dynamic memory and transmits them over a telephone line using a modem. A modem located locally to the computer system 300 receives the data over the telephone line and uses an infrared transmitter to convert the data into an infrared signal. An infrared detector receives the data carried in the infrared signal, and a suitable circuit arrangement places the data on the bus 302. The bus 302 carries the data to the main memory 306, from which the processor 304 retrieves and executes the instructions.Instructions received by the main memory 306 can optionally be stored in the storage device 310 either before or after execution by the processor 304.

[0056] The Computer System 300 also includes a communication interface 318, which is coupled to the bus 302. The communication interface 318 provides two-way data communication, coupled to a network connection 320, which is connected to a local area network 322. For example, the communication interface 318 could be an ISDN (Integrated Service Digital Network) card, a cable modem, a satellite modem, or a modem for providing a data communication connection to a corresponding type of telephone line. Another example is the communication interface 318 being a local area network (LAN) card for providing a data communication connection to a compatible LAN. Wireless connections are also implemented in some implementations.In any such implementation, the communication interface 318 sends and receives electrical, electromagnetic, or optical signals carrying digital data streams that represent various types of information.

[0057] The network connection 320 typically provides data communication to other data facilities via one or more networks. For example, the network connection 320 provides a connection via the local area network 322 to a host computer 324, a cloud data center, or equipment operated by an Internet Service Provider (ISP) 326. The ISP 326, in turn, provides data communication services via the worldwide packet data communication network, now commonly referred to as the "Internet" 328. Both the local area network 322 and the Internet 328 use electrical, electromagnetic, or optical signals to carry digital data streams. The signals through the various networks and the signals on the network connection 320 and via the communication interface 318, which carry the digital data to and from the computer system 300, are exemplary forms of transmission media.In one embodiment, the network 320 contains the Cloud 202 described above, or a part of the Cloud 202.

[0058] The computer system 300 sends messages and receives data, including program code, via the network(s), the network connection 320, and the communication interface 318. In one embodiment, the computer system 300 receives code for processing. Upon receipt, the received code is executed by the processor 304 and / or stored in the memory device 310 or other non-volatile memory for later execution. Autonomous vehicle architecture

[0059] Fig. Figure 4 shows an exemplary Architecture 400 for an autonomous vehicle (e.g., the one in Fig. 1 AV 100 shown). The architecture 400 includes a perception module 402 (sometimes called the perception circuit), a planning module 404 (sometimes called the planning circuit), a control module 406 (sometimes called the control circuit), a localization module 408 (sometimes called the localization circuit), and a database module 410 (sometimes called the database circuit). Each module plays a role in the operation of the AV 100. Modules 402, 404, 406, 408, and 410 can together form part of the system shown in the Fig. 1 AV system 120 shown. In some embodiments, any one of the modules 402, 404, 406, 408 and 410 is a combination of computer software (e.g., executable code stored on a computer-readable medium) and computer hardware (e.g., one or more microprocessors, microcontrollers, application-specific integrated circuits [ASICs]), hardware storage devices, other types of integrated circuits, other types of computer hardware, or a combination of some or all of these).

[0060] In operation, the planning module 404 receives data representing a destination 412 and determines data representing a trajectory 414 (sometimes referred to as a route) that the AV 100 can follow to reach the destination 412 (e.g., arrive there). In order for the planning module 404 to determine the data representing the trajectory 414, it receives data from the perception module 402, the localization module 408, and the database module 410.

[0061] The perception module 402 identifies nearby physical objects using one or more sensors 121, e.g., as in the Fig. 1 shown. The objects are classified (e.g. grouped into object types such as pedestrian, bicycle, automobile, traffic sign, etc.) and a scene description including the classified objects 416 is provided to the planning module 404.

[0062] The planning module 404 also receives data representing the AV position 418 from the localization module 408. The localization module 408 determines the AV position using data from sensors 121 and data from the database module 410 (e.g., a geographic datum) to calculate a position. For example, the localization module 408 uses data from a GNSS (Global Navigation Satellite System) sensor and geographic data to calculate the longitude and latitude of the AV.In one embodiment, the data used by the localization module 408 include highly accurate maps of the geometric properties of the roadway, maps describing the connection properties of the road network, maps describing the physical properties of the roadway (such as traffic speed, traffic volume, the number of lanes for vehicles and cyclists, the lane width, the traffic directions of the lanes or lane marking types and positions or combinations thereof), and maps describing the spatial location of road features such as pedestrian crossings, traffic signs or other various traffic signals.

[0063] The control module 406 receives the data representing the trajectory 414 and the data representing the AV position 418, and executes the control functions 420a to 420c (e.g., steering, throttling, braking, ignition) of the AV in such a way that the AV 100 travels the trajectory 414 to the destination 412. For example, if the trajectory 414 includes a left turn, the control module 406 executes the control functions 420a to 420c such that the steering angle causes the AV 100 to turn left, and the acceleration and braking cause the AV 100 to stop and wait for pedestrians and vehicles crossing before the turn is executed. Autonomous vehicle inputs

[0064] The Fig. Figure 5 shows an example for inputs 502a to 502d (e.g., those in Fig. 1 shown sensors 121) and outputs 504a to 504d (e.g. sensor data) from the perception module 402 ( Fig. 4) can be used. An input 502a is a LiDAR system (Light Detection and Ranging) (e.g., the one in the Fig. 1 LiDAR shown 123). LiDAR is a technology that uses light (e.g., pulses of light such as infrared light) to obtain data about physical objects in its line of sight. A LiDAR system produces LiDAR data as output 504a. LiDAR data are, for example, collections of 3D or 2D points (also known as point clouds) used to build a representation of the environment 190.

[0065] Another input 502b is a RADAR system. RADAR is a technology that uses radio waves to obtain data about nearby physical objects. RADARs can obtain data about objects that are not in the line of sight of a LiDAR system. A RADAR system 502b produces RADAR data as output 504b. Radar data includes, for example, one or more high-frequency electromagnetic signals used to build a representation of the environment 190.

[0066] Another input 502c is a camera system. A camera system uses one or more cameras (e.g., digital cameras that use a light sensor such as a charge-coupled device, CCD) to obtain information about nearby physical objects. A camera system produces camera data as output 504c. Camera data often takes the form of image data (e.g., data in an image data format such as RAW, JPEG, PNG, etc.). In some examples, the camera system has multiple independent cameras—e.g., for stereopsis (stereo vision)—which allows the camera system to perceive depth. The objects perceived by the camera system are described here as "nearby," but this refers to proximity relative to the AV (audio-visual device). In practice, the camera system may be designed to "see" distant objects, e.g., up to a kilometer or more in front of the AV.Accordingly, the camera system may have features such as sensors and lenses that are optimized for detecting distant objects.

[0067] Another input 502d is a traffic light detection system (TLD system). A TLD system uses one or more cameras to gather information about traffic lights, road signs, and other physical objects that provide visual navigation information. A TLD system produces TLD data as output 504d. TLD data often takes the form of image data (e.g., data in an image data format such as RAW, JPEG, PNG, etc.). A TLD system differs from a system that includes a camera in that a TLD system uses a camera with a wide field of view (e.g., using a wide-angle or fisheye lens) to gather information about as many physical objects as possible that provide visual navigation information, so that the AV 100 has access to all relevant navigation information provided by these objects.For example, the viewing angle of the TLD system can be approximately 120 degrees or more.

[0068] In some embodiments, outputs 504a to 504d are combined using sensor fusion technology. Thus, either the individual outputs 504a to 504d are provided to other systems of the AV 100 (e.g., to a planning module 404, as in the Fig. (as shown in Figure 4), or the combined output can be provided to the other systems—either as a single combined output or as multiple combined outputs of the same type (e.g., using the same combination technique, or by combining the same outputs, or both), or of different types (e.g., using different respective combination techniques, or by combining different respective outputs, or both). In some embodiments, an early fusion technique is used. An early fusion technique is characterized by combining outputs before one or more data processing steps are applied to the combined output. In some embodiments, a late fusion technique is used. A late fusion technique is characterized by combining outputs after one or more data processing steps have been applied to the individual outputs.

[0069] The Fig. Figure 6 shows an example of a LiDAR system 602 (e.g., the one in the Fig. 5. Input 502a shown). The LiDAR system 602 emits light 604a to 604c from a light emitter 606 (e.g., a laser transmitter). Light emitted by a LiDAR system is not usually in the visible spectrum; for example, infrared light is frequently used. Some of the emitted light 604b strikes a physical object 608 (e.g., a vehicle) and is reflected back to the LiDAR system 602. (Light emitted by a LiDAR system does not usually penetrate physical objects, e.g., solid physical objects.) The LiDAR system 602 also has one or more light detectors 610 that detect the reflected light. In one embodiment, one or more data processing systems associated with the LiDAR system generate an image 612, which represents the field of view 614 of the LiDAR system. The image 612 contains information that represents the boundaries 616 of a physical object 608.In this way, image 612 is used to determine the boundaries 616 of one or more physical objects near an AV.

[0070] Fig. Figure 7 shows the LiDAR system 602 in operation. In the scenario depicted in this figure, the AV 100 receives both the output 504c of the camera system in the form of an image 702, and the output 504a of the LiDAR system in the form of LiDAR data points 704. During operation, the data processing systems of the AV 100 compare the image 702 with the data points 704. In particular, a physical object 706 identified in the image 702 is also identified among the data points 704. In this way, the AV 100 perceives the boundaries of the physical object based on the contour and density of the data points 704.

[0071] The Fig. Figure 8 shows the functionality of the LiDAR system 602 in more detail. As described above, the AV 100 detects the boundary of a physical object based on properties of the data points detected by the LiDAR system 602. As shown in the Fig. As shown in Figure 8, a flat object, such as the ground 802, reflects light 804a to 804d emitted by a LiDAR system 602 in a consistent manner. In other words, since the LiDAR system 602 emits light using a consistent distance, the ground 802 reflects light back to the LiDAR system 602 at the same consistent distance. As the AV 100 travels over the ground 802, the LiDAR system 602 will continue to detect light reflected by the next valid ground point 806, provided the road is not blocked. However, if an object 808 blocks the road, light 804e to 804f emitted by the LiDAR system 602 will be reflected by points 810a to 810b in a manner inconsistent with the expected consistent way. From this information, the AV 100 can determine that object 808 exists. Railway planning

[0072] The Fig. Figure 9 shows a block diagram 900 of the relationships between the inputs and outputs of a planning module 404 (such as in the Fig. (shown in Figure 4). The output of a planning module 404 is generally a route 902 from a starting point 904 (e.g., source or starting point) to an endpoint 906 (e.g., destination or end point). The route 902 is usually defined by one or more segments. A segment is, for example, a distance traveled over at least part of a road, highway, expressway, driveway, or other physical area suitable for vehicular traffic. In some examples, such as when the AV 100 is an off-road vehicle like a four-wheel drive (4WD) or all-wheel drive (AWD) vehicle, an SUV, a light truck, or the like, the route 902 includes "off-road" segments, such as dirt roads or open terrain.

[0073] In addition to Route 902, a planning module also outputs lane-specific route planning data 908. This lane-specific route planning data 908 is used to traverse segments of Route 902 based on the segment's conditions at a given time. For example, if Route 902 includes a multi-lane highway, the lane-specific route planning data 908 includes trajectory planning data 910, which the AV 100 can use to select a lane among the multiple lanes, based on factors such as whether it is approaching an exit, whether there are other vehicles in one or more of the lanes, or other factors that vary over a few minutes or less. Similarly, in some implementations, the lane-specific route planning data 908 includes speed limits 912 specific to a segment of Route 902.For example, if the segment contains pedestrians or unexpected traffic, the speed limits 912 may restrict the AV 100 to a driving speed that is lower than an expected speed, e.g., a speed based on speed limit data for the segment.

[0074] In one embodiment, the inputs to the planning module 404 include database data 914 (e.g., from the one in the Fig. 4 database module 410 shown), current location data 916 (e.g., the one in the Fig. 4 AV position shown 418), destination data 918 (e.g. for the one in the Fig. 4 destination shown 412) and object data 920 (e.g. the classified objects 416, as shown by the one in the Fig. (4 shown in the perception module 402). In some embodiments, the database data 914 includes rules used in planning. Rules are specified using a formal language, e.g., using Boolean logic. In any given situation encountered by the AV 100, at least some of the rules for that situation will apply. A rule applies to a given situation if the rule has conditions that are satisfied based on the information available to the AV 100, e.g., information about the surrounding environment. Rules can have a priority. For example, a rule stating "if the road is a highway, drive in the leftmost lane" may have a lower priority than "if the exit is within a mile, drive in the rightmost lane."

[0075] The Fig. Figure 10 shows a directed graph 1000, which is used in railway planning, e.g. by the planning module 404 ( Fig. 4). Generally, a directed graph 1000 like the one in the Fig. The diagram shown is used to determine a path between any starting point 1002 and endpoint 1004. In reality, the distance separating the starting point 1002 and the endpoint 1004 can be relatively large (e.g., in two different urban areas) or relatively small (e.g., two intersections adjacent to a city block, or two lanes of a multi-lane road).

[0076] In one embodiment, the directed graph 1000 has nodes 1006a to 1006d, which represent different locations between the start point 1002 and the end point 1004 that could be traversed by an AV 100. In some examples, e.g., if the start point 1002 and the end point 1004 represent different urban areas, nodes 1006a to 1006d represent segments of roads. In other examples, e.g., if the start point 1002 and the end point 1004 represent different locations on the same road, nodes 1006a to 1006d represent different positions on that road. In this way, the directed graph 1000 contains information at varying levels of granularity. In one embodiment, a directed graph with high granularity is also a subgraph of another directed graph with a larger scale.For example, a directed graph in which the starting point 1002 and the endpoint 1004 are far apart (e.g., many miles apart) has most of its information at a low granularity and is based on stored data, but also includes some high-granular information for the part of the graph that represents physical locations in the AV 100's field of view.

[0077] Nodes 1006a to 1006d differ from objects 1008a and 1000b, which cannot overlap with a node. In one embodiment, at low granularity, objects 1008a and 1000b represent regions that cannot be traversed by a motor vehicle, e.g., areas without roads or paths. At high granularity, objects 1008a and 1000b represent physical objects within the AV 100's field of view, e.g., other motor vehicles, pedestrians, or other objects with which the AV 100 cannot share physical space. In one embodiment, some or all of objects 1008a to 1008b are static objects (e.g., an object that does not change its position, such as a street lamp or a power pole) or dynamic objects (e.g., an object that can change its position, such as a pedestrian or another motor vehicle).

[0078] Nodes 1006a to 1006d are connected by edges 1010a to 1010c. If two nodes 1006a and 1000b are connected by an edge 1010a, an AV 100 can travel between node 1006a and node 1006b without having to stop at an intermediate node before reaching node 1006b. (When referring to an AV 100 traveling between nodes, this means that the AV 100 travels between the two physical positions represented by the respective nodes.) Edges 1010a to 1010c are often bidirectional in the sense that an AV 100 can travel from a first node to a second node or from the second node to the first node. In one embodiment, the edges 1010a to 1010c are unidirectional in the sense that an AV 100 can travel from a first node to a second node, but the AV 100 cannot travel from the second node to the first node.Edges 1010a to 1010c are unidirectional if they represent, for example, one-way streets, single lanes of a city street, road or expressway, or other features that can only be traversed in one direction due to legal or physical restrictions.

[0079] In one embodiment, the planning module 404 uses the directed graph 1000 to identify a path 1012 consisting of nodes and edges between the start point 1002 and the end point 1004.

[0080] An edge 1010a to 1010c has an associated effort 1014a to 1014b. The effort 1014a to 1014b is a value representing the resources expended when AV 100 selects this edge. A typical resource is time. For example, if an edge 1010a represents a physical distance twice that of another edge 1010b, then the associated effort 1014a of the first edge 1010a may be twice the associated effort 1014b of the second edge 1010b. Other factors affecting time include expected traffic, the number of intersections, speed limits, etc. Another typical resource is fuel consumption. Two edges 1010a and 1010b can represent the same physical distance, but one edge 1010a may require more fuel than another edge 1010b - e.g. due to road conditions, the expected weather, etc.

[0081] If the planning module 404 identifies a path 1012 between the starting point 1002 and the endpoint 1004, the planning module 404 usually selects a path optimized for effort, e.g. the path with the lowest total effort when the individual effort of the edges is added. Control of the autonomous vehicle

[0082] The Fig. Figure 11 shows a block diagram 1100 of the inputs and outputs of a control module 406 (as e.g. in the Fig. 4 shown). A control module operates according to a controller 1102, which contains, for example, one or more processors (e.g., one or more computer processors such as microprocessors or microcontrollers, or both) similar to the processor 304, a short-term and / or long-term data storage (e.g., random access memory or flash memory, or both) similar to the main memory 306, the ROM 308, and the memory device 210, and instructions stored in memory that perform operations of the controller 1102 when the instructions are executed (e.g., by the one or more processors).

[0083] In one embodiment, the controller 1102 receives data representing a target output 1104. The target output 1104 typically includes a speed, e.g., a rotational speed, and a direction of travel. The target output 1104 can, for example, be based on data received from a planning module 404 (such as in the Fig. (shown in Figure 4). According to the target output 1104, the controller 1102 generates data that can be used as accelerator pedal input 1106 and steering input 1108. The accelerator pedal input 1106 represents the magnitude to which the throttle valve (e.g., acceleration control) of an AF 100 should be actuated—e.g., by pressing the steering pedal or by actuating another throttle valve control—in order to achieve the target output 1104. In some examples, the accelerator pedal input 1106 also includes data that can be used to actuate the brake (e.g., deceleration control) of the AF 100. The steering input 1108 represents a steering angle, e.g., the angle at which the steering control (e.g., steering wheel, steering angle adjuster, or other functionality for controlling the steering angle) of the AV should be positioned to achieve the target output 1104.

[0084] In one embodiment, the controller 1102 receives feedback that is used to adjust the inputs provided to the throttle and steering. For example, if the AV 100 encounters an obstacle 1110, such as a hill, the measured speed 1112 of the AV 100 is reduced below the target output speed. In another embodiment, each measured output 1114 is provided to the controller 1102 so that the necessary adjustments—for example, based on the difference 1113 between the measured speed and the target output—can be made. The measured output 1114 includes the measured position 1116, the measured speed 1118 (including rotational speed and direction of travel), the measured acceleration 1120, and other outputs measurable by sensors of the AV 100.

[0085] In one embodiment, information about the obstacle 1110 is detected in advance—e.g., by a sensor such as a camera sensor or a LiDAR sensor—and provided to a predictive feedback module 1122. The predictive feedback module 1122 then provides information to the controller 1102, which the controller 1102 can use for appropriate adjustments. For example, if the sensors of the AV 100 detect (“see”) a hill, this information can be used by the controller 1102 to prepare to actuate the throttle valve at the appropriate time to avoid a significant delay.

[0086] The Fig. Figure 12 shows a block diagram 1200 of the inputs, outputs, and components of the controller 1102. The controller 1102 has a speed profiler 1202 that influences the operation of an accelerator / brake controller 1204. For example, the speed profiler 1202 instructs the accelerator / brake controller 1204 to initiate acceleration or deceleration using the accelerator / brake 1206, depending on the feedback received by the controller 1102 and processed by the speed profiler 1202.

[0087] The controller 1102 also includes a lateral tracking controller 1208, which influences the operation of a steering controller 1210. For example, the lateral tracking controller 1208 instructs the steering controller 1210 to adjust the position of the steering angle actuator 1212, e.g., depending on feedback received by the controller 1102 and processed by the lateral tracking controller 1208.

[0088] The controller 1102 receives several inputs that are used to determine how to control the accelerator / brake 1206 and the steering angle actuator 1212. A planning module 404 provides information that the controller 1102 uses, for example, to select a direction of travel when the AV 100 starts operating and to determine which road segment to traverse when the AV 100 reaches an intersection. A localization module 408 provides the controller 1102 with information that describes, for example, the current location of the AV 100, so that the controller 1102 can determine whether the AV 100 is in an expected location based on how the accelerator / brake 1206 and the steering angle actuator 1212 are controlled. In one embodiment, the controller 1102 receives information from other inputs 1214, e.g. information received from databases, computer networks, etc. Architecture of object tracking

[0089] The Fig. Figure 13 shows a block diagram describing a visual object tracking system 1300 suitable for use with the AV 100. The object detector 1302 is used to capture one or more images of an area surrounding an autonomous vehicle and to perform analysis of these images. The images can be captured by a capture device, such as a high-speed camera or a video camera. The capture device will generally be designed to capture a wide field of view in order to minimize the number of capture devices required for tracking objects near the autonomous vehicle. In some embodiments, the capture device may be designed as a wide-field-of-view image capture device with a fixed focus distance, which is well suited for tracking objects near the autonomous vehicle.The object detector 1302 can have a processing circuit containing one or more processors 146 that execute instructions for analyzing the one or more images in order to identify and classify objects contained in the one or more images (see the descriptions of processors 146 and 304 in the text relating to the respective ). Fig. 1 and Fig. 3) In some embodiments, the processor(s) 146 may be configured to execute a classification routine, which may be stored, for example, in main memory 306 or in other storage media. The classification routine may be a correlation filter tracker, a deep tracker, and / or a Kalman filter that performs the identification and / or classification processes. The classification routine may be configured to distinguish between a stationary object, such as a lamppost, a fire hydrant, or a building, and a mobile object, such as a car, a truck, a motorcycle, or a pedestrian. Objects classified as stationary that are located far outside a planned path of movement of the autonomous vehicle (e.g.,Objects beyond a threshold distance from the planned trajectory can be safely ignored, while stationary objects within or near the planned trajectory can be tracked. Objects identified as moving, and objects classified as mobile (whether moving or stationary), can be forwarded as image detection data to a visual tracker 1304 for further analysis and tracking, provided the objects are within a threshold distance from the planned trajectory. Image detection data, when derived from multiple frames, can take many forms and include metrics such as object position, velocity, orientation, angular velocity, approach velocity, and the like. Different types of image detection data can influence the threshold distance for a given object.For example, the threshold distance can depend on the direction and speed of movement of the object and may change over time if the direction and speed of movement of the object change.

[0090] The visual tracker 1304 is designed to generate image plane trajectories from the image detection data. The visual tracker 1304 can include discrete processors 146 for executing instructions stored in main memory 306 that perform further visual tracking, or alternatively, these instructions can be executed by shared processors 146 allocated to the entire visual tracking system 1300. The visual tracker 1304 may require multiple image detections to generate image plane trajectories that accurately estimate the speed at which the detected objects traverse the image plane. In some embodiments, the image detections, together with the position data of the autonomous vehicle at the time of each image acquisition, can be collectively referred to as tracking data.The image plane paths generated by the visual tracker 1304, along with the associated tracking data, are forwarded to the tracking and fusion engine 1306.

[0091] The tracking and fusion engine 1306 has one or more processing circuits that use one or more processors 146 to execute instructions stored in the main memory 306 for converting image plane orbits into ground plane orbits. In some embodiments, the instructions may be stored elsewhere, such as in the ROM 308, in different partitions of the main memory 306, or in a separate memory module or storage medium (e.g., the memory device 310) associated with the object tracking system. Ground plane orbits describe a position of the detected object relative to the Earth's surface, not a position within the image plane.To convert image plane trajectories into ground plane trajectories, the position of the image plane trajectories within the image plane and the position of the autonomous vehicle at the time of each image acquisition can be used to establish a bearing line between the autonomous vehicle and the corresponding detected objects. Determining a precise distance between the autonomous vehicle and the detected objects from image-based data can be more challenging. In some embodiments, the distance information can be determined by correlating image cues within the image frame surrounding the detected object with known positions associated with previously acquired image data.For example, the position of an object resembling a car, located laterally offset from the autonomous vehicle between lane markings, can be determined with a reasonable degree of certainty by identifying the intersection of the bearing line with the lane defined by the lane markings visible in the imagery. The greater the lateral offset and the smaller the distance, the more accurate this method of determination, as the angle between the lane and the bearing line derived from the images will be larger. The proximity of the tracked object to visually distinct features, such as traffic signals and other recognizable objects within the frame, can also contribute to determining the distance between that specific object and the autonomous vehicle.In some embodiments, the distance information can be determined by a separate distance sensor such as a radar, LiDAR, laser rangefinder, or infrared distance sensor. In some embodiments, the distance sensor may only be directed towards the areas corresponding to the locations where objects are detected.

[0092] After using the bearing information and the information about the specified distance to convert the image plane trajectories into ground plane trajectories, the 1306 tracking and fusion system can add additional predictive information, such as that based on map data, to the ground plane trajectories to more accurately determine the likely behavior of the object based on its position relative to the map data. The map data might contain information about the direction a vehicle can turn at a particular intersection, thus narrowing down the number of likely directions a vehicle in a given lane will take. Behavioral prediction can take many forms and rely on many different factors, which are explained in more detail below.The tracking and fusion system 1306 can transmit the ground-level trajectories and prediction information to an autonomous navigation system 1308 to assist the autonomous navigation system 1306 in determining a path for the autonomous vehicle (e.g., to avoid detected objects). In some embodiments, the autonomous navigation system 1306 includes a control circuit responsible for controlling the navigation of the autonomous vehicle.

[0093] The Fig. Figure 14A shows a block diagram 1400, which provides a more detailed description of how the visual tracker 1304 converts image detection data into image plane trajectories. Specifically, the image detection data is transferred from the object detector 1302 to the trajectory generator 1402. The trajectory generator 1402 converts the image detection data into image plane trajectories. The image plane trajectory can contain at least one detected position of each object identified by the object detector 1302. In cases where the image detection data is generated from more than one image frame, additional metrics, such as velocity or frame traverse rate, can be determined from each movement of the object through the image frame that can be extracted from the multiple image frames. In some embodiments, 5 to 10 image frames can be analyzed simultaneously.With a system that considers every captured image, for example, 6-12 trajectory analyses per second could be performed by simultaneously analyzing 5 to 10 images from an optical sensor that captures 60 images per second. In some embodiments, the object detector 1302 and the visual tracker 1304 could be configured to sample at a slower rate of, for example, 5 to 10 images per second to reduce processing loads. In some embodiments, the sampling rate can vary based on the presence of detected objects positioned in such a way that a collision might be imminent.

[0094] The orbit generator 1402 then delivers generated image frame orbits, containing at least the detected positions of the objects detected within the image frame at a given time, to the matching engine 1408 (e.g., a correlation engine). The visual tracker 1304 also receives data of active orbits from the orbit data store 1404. The orbit data store 1404 is shown as a subcomponent of the tracking and fusion system 1306; however, in some embodiments, the orbit data store 1404 can be separate from the tracking and fusion system 1306. The orbit data store 1404 can contain both image plane orbits and ground plane orbits. In some embodiments, the orbit data can be organized in orbit files or objects that can contain both image plane and ground plane position information.The active orbit data transferred from the orbit data store 1404 to the visual tracker 1304 can be processed by the prediction engine 1406. The prediction engine 1406 can be configured to determine the probable position of each object contained in the active orbit data. Since the active orbit data may contain ground orbit data, the prediction engine 1406 can also benefit from any prediction information stored in the active orbit data derived from map data that constrains a possible movement of the object at a given time based on the latest tracking data contained in the active orbit data.

[0095] The prediction engine 1406 outputs an image plane orbit containing at least one predicted position of the objects. The matching engine 1408 compares the predicted positions of the objects, derived from data of active orbits, with the detected positions of the objects from the image orbits generated by the orbit generator 1402. In some embodiments, additional metrics such as direction and speed of movement, object shape, object size, object color, and the like can be included in the comparison. This comparison assists the matching engine 1408 in identifying which of the detected positions correlate with predicted positions generated from the data of active orbits provided by the orbit data store 1404.In some embodiments, all image plane trajectories that do not correlate with a predicted position of the active trajectories are forwarded to the Tracking and Fusion System 1306 as a new trajectory without historical tracking data. Those image trajectory files with detected positions that correlate with a predicted position can be combined with the data of the active trajectory. In some embodiments, the predicted and detected positions of the object can be averaged. How the predicted and detected positions vary can differ based on various factors, such as the degree of confidence in the tracking data associated with the active trajectory, as well as the degree of confidence in the quality of the object detection data.

[0096] The Fig. Figure 14B shows a block diagram 1410, which contains the same components as in the Fig. The operation described in section 14A differs, except that the path data store 1404 is designed to supply data from outdated paths to the matching engine 1408. Outdated paths are those paths that have not yet been deleted as being too old and / or unreliable, but which also do not meet a classification criterion for active paths. The criteria for active and outdated paths are generally distinguished based on factors such as the age of the tracking data, the object's position relative to the autonomous vehicle, and the tracking data confidence. For example, a path might be considered an active path if the tracking data has been updated within a threshold time period and / or if the path is assigned a tracking data confidence threshold based on the position and / or direction of movement of the object associated with the path.The "outdated path" criterion can also be based, at least in part, on a number of image frames acquired without an object being detected. In some embodiments, the tracking data confidence level can also be based on the temporal consistency of the data from which the tracking data is derived. The temporal consistency of the data is a measure of how closely the sensor data used to generate the secondary tracking data follow expected trends.

[0097] Since in some embodiments no predicted position is calculated, a match criterion for stale path data may be based primarily on the size and shape of the object. The match criterion for stale path data may also be applied only after an attempt has been made to find an active path that matches one of the newly generated image plane path(s). In some embodiments, position data associated with the most recent detection(s) may also be used as a factor in determining the probability that the stale path data are associated with the same object as the newly generated image plane path. For example, if the object would have had to accelerate at an improbable speed to reach the newly detected position, the correlation with the previously detected object might be considered to fail a match criterion.When combining data from outdated orbits with newly generated image plane orbits, the Matching Engine 1408 generally uses the detected position rather than attempting to average the detected position with the predicted position information associated with the old orbit data. It should be noted that in some embodiments, the Prediction Engine 1406 may be configured to receive the data from outdated orbits prior to the Matching Engine 1408, enabling the Prediction Engine 1406 to generate predicted position data for outdated orbit data that exhibit particularly high confidence levels.For example, older position data for a parked car that was detected when its lights were switched on might still have a high level of confidence even if the outdated track would normally be excluded from prediction due to its age, thus allowing data from outdated tracks to be used to refine a position of the parked car that is expected to move soon.

[0098] The Fig. Figure 14C shows a block diagram illustrating additional details regarding another specific implementation for processing image plane trajectories by the Tracking and Fusion System 1306. In particular, the Fig. 14C, how the tracking and fusion system 1306 did not use newly collected tracking data, as in the preceding Fig. 14A and Fig. 14B described, not with the visual tracker 1304, but combined with historical tracking data. In 1422, the tracking and fusion system 1306 is designed to convert new image plane orbits into ground plane orbits, as previously described in the text. Fig. As described in section 13, in sections 1424 and 1426, the new ground-level paths are compared with older active and obsolete paths, respectively. Obsolete paths are those that have not yet been deleted as being too old and / or unreliable, but which also do not meet a classification criterion for active paths. The criteria for active and obsolete paths are generally distinguished based on factors such as the age of the tracking data, the object's position relative to the autonomous vehicle, and the tracking data confidence. For example, a path might be considered active if the tracking data has been updated within a threshold time period and / or if the path is assigned a tracking data confidence threshold based on the position and / or direction of movement of the object associated with the path.The "outdated path" criterion can also be based, at least in part, on a number of image frames captured without an object being tracked. In some embodiments, the level of tracking data confidence can also be based on the temporal consistency of the data from which the tracking data is derived.

[0099] In particular, at 1424, newly converted ground-level orbits are compared with active orbits. The new ground-level orbits that meet a match criterion with an older active orbit are merged at 1428 with the matching active orbit to form a single updated active orbit. In some embodiments, the tracking data of the older active orbit are merged with the tracking data of the new orbit by assigning the new tracking data to an orbit ID of the older active orbit. In this way, the sensor data assigned to the new ground-level orbit can be combined with the older tracking data, allowing the autonomous navigation system 1308 to use both current and historical data to more easily predict the future movement of the object.The matching criteria may vary depending on the operating conditions of the autonomous vehicle, but are generally based on the new ground-level orbits being close to an expected or projected location of the object based on historical tracking data associated with the active orbits.

[0100] The tracking and fusion system 1306 can optionally include a process at 1426 in which remaining, non-matching ground-level tracks are compared with older tracks that have an obsolete status. This allows some objects that were temporarily obscured by another object, glare, or the like to be re-identified. At 1430, the non-matching ground-level tracks that meet a match criterion with one of the older obsolete tracks are fused with data from the obsolete track to create a single updated active track, which is then forwarded to the autonomous navigation system 1308. In this way, the new sensor data can be combined with the older tracking data, allowing the autonomous navigation system 1308 to use both current and historical data to better predict the object's future movement.The matching criteria can vary depending on the autonomous vehicle's operating conditions, but generally rely on a new ground-level path being close to the object's expected location based on historical ground-level paths. It should be noted that matching outdated paths can be more problematic, as these paths are typically not updated, at least for a short time, making reliable prediction difficult. Therefore, to avoid mismatches, the matching criteria for merging current paths with older, outdated paths can be configured more conservatively than the matching criteria for active paths.In this way, the new sensor data can be reliably combined with the older tracking data, so that the autonomous navigation system 1308 can use both current and historical data to more easily predict the future behavior of the object in situations where one or more of the objects remain undetected for short periods of time.

[0101] If new ground paths are available that do not match any of the older active or obsolete paths, the non-matching ground plane paths without historical tracking data can be forwarded to the autonomous navigation system 1308 for use in predicting behavior to assist the autonomous vehicle's navigation at 1432.

[0102] The Fig. Figure 15 shows a flowchart that provides additional details at block 1408 of the information previously presented in the Fig. The process described in section 14 illustrates this. In particular, the mapping of the matching image plane trajectories and the active trajectories at 1502 involves the use of a processing circuit, such as processor 146 or 304, to execute instructions stored in a memory, such as main memory 306 or other storage media (e.g., storage device 310), which, using tracking data from the matching active trajectory, predict an initial position of the object at time T0. Since the tracking data assigned to the active trajectory may contain position data collected over the course of several frames, this predicted position can be more accurate than the detected position due to potential inaccuracies in measuring a single position.In some embodiments, the prediction is based on a previous position of the object both on the ground plane and in the image plane. In 1504, a second position of the object can be obtained from the image acquired by the processing circuit at time T0 (e.g., from the tracking data forming the recently converted image plane trajectory). If, in 1506, a determination is made that the detected second position of the object lies within a threshold distance of the predicted first position of the object, the position of the object at time T0 is recorded based on both the detected second position and the predicted first position.The threshold distance is based, at least in part, on one or more differences between a feature of the object as it appears in the detected image or images associated with the last object detection and the feature as it appeared in the images used to generate the associated active trajectory. Generally, the threshold distance is not affected by expected changes to the feature, such as changes in the object's orientation relative to the autonomous vehicle.

[0103] An example of a feature change that could reduce the threshold is a significant color change. While a color change might be due to a change in lighting, less leeway is generally given for a position change if other features, such as color, shape, or size, have changed significantly. In some embodiments, the registered position of the object at time T0 is determined by calculating a weighted average of the detected second position and the predicted first position. The weighting of this average can vary based on a number of factors, including the consistency of historical data, the quality of the image captured at time T0, and other factors that influence whether the predicted or the detected position is considered more likely to be the actual position of the object at time T0.While block 1408 of the visual tracker 1304 as in the . Fig. As described in 15, in some embodiments new tracking data may only contain the detected position, without taking into account the predicted position, before the tracking data for the image plane orbit is updated and the image plane orbit is sent to the detection and tracking system 1306, where the image plane orbit is converted into a ground plane orbit and then provided to the autonomous navigation system, which enables a control circuit to use the orbit information to avoid the object that is associated with the updated active orbit.

[0104] The Fig. Figures 16A to 16B show images captured by an object detection system on board an autonomous vehicle as it approached a traffic intersection. In particular, the Fig. 16A, how vehicle 1602 is tracked by the object detection system. A first rectangular marker 1604 indicates a detected position of vehicle 1602. A second marker 1606 indicates a projected position of vehicle 1602, based on previous activities of vehicle 1602. The object detection system also tracks vehicle 1608, as indicated by rectangular markers 1610 and 1612, which similarly represent the respective detected and projected positions of vehicle 1608. Vehicle 1614 contains only the rectangular marker 1616, which represents a detected position of vehicle 1614, as it recently moved from a position behind vehicle 1618 and no data is available to determine a projected position of vehicle 1614.Although numerous other vehicles and markings are depicted in this scene, and the object detection system could be designed to track all vehicles, for the sake of clarity only the movements of vehicles 1602, 1608 and 1614 are explained.

[0105] The Fig. Figure 16B shows how vehicle 1608 reappears after passing behind vehicle 1602. When a vehicle is obscured in this way, a track associated with the tracking system may temporarily fail to provide a projected position for that vehicle. In this particular example, vehicle 1606 was obscured as it entered an intersection, making the extrapolation of its motion so uncertain that the associated track was marked as obsolete due to the uncertainty of its trajectory, and in some embodiments, no projection is attempted. For example, vehicle 1608 could proceed through the intersection, stop at the intersection, or begin turning right at the intersection while vehicle 1608 is obscured.While the obstruction of vehicles is one reason why the system loses track of an object, other factors can also prevent the object detection system from continuing to track a particular object. For example, glare from the sun could prevent continuous tracking of vehicle 1606, as could distortion of the object due to inherent distortion in the lens of the optical sensor.

[0106] The Fig. Figures 17A to 17C show a top view of the [unclear] in the Fig. The intersection shown in Figures 16A to 16B, together with the autonomous vehicle 1702 and the optical sensor 1704, corresponds in particular to the Fig. 17A and Fig. 17C the in the Fig. Images 16A and 16B are shown. As previously described, the position of vehicles 1602 and 1608 can be seen in the images shown in the Fig. 16A and Fig. The images shown in 16B are converted into ground plane orbits, as in the Fig. 17A and Fig. Figure 17B illustrates this. Converting the image data into location-based tracking data allows the autonomous vehicle to more accurately predict and avoid other objects it is tracking by correlating the objects' positions with map data. Fusing the tracking data with the map data provides increased certainty regarding direction, speed, and likely behavior, as the map contains information such as speed limits, the number and position of lanes, the location and operation of traffic signals, the location of parking spaces, and the like. For example, a vehicle detected at a location known to be a parking space can be expected to maintain its position and is unlikely to move from it.

[0107] The Fig. 17B shows an intermediate position of vehicles between those in the Fig. 16A and Fig. 16B shown in the images. In particular, it shows Fig. 17B, ​​how vehicle 1602 can obscure vehicle 1608 from view if vehicle 1602 moves further in front of autonomous vehicle 1702. This type of object obscuration can be referred to as object occlusion. As explained previously, without a clear line of sight to vehicle 1608 or vehicle 1614 at that time, the object detection system can only project a probable position of vehicles 1608 and 1614. In the Fig. 17C means that vehicle 1602 no longer obstructs the line of sight to vehicle 1608, allowing the object detection system to reacquire vehicle 1608. At this point, both a detected and a projected position for vehicle 1608 are again available. As shown, no tracking information is available for vehicle 1614 after a time interval exceeds a threshold, preventing even the determination of a projected position because the projection accuracy is too low (e.g., below a predefined threshold accuracy metric).

[0108] The Fig. Figure 18 shows a flowchart illustrating a procedure for assigning newly detected objects to older, obsolete tracks, which is used at block 1408 of the Fig. This is described in section 14B. At 1802, new image plane paths that have not already been assigned to an active path are compared with each of the obsolete paths by a processing circuit, such as processor 146 or 304, which executes instructions stored in a memory, such as main memory 306 or other storage media (e.g., storage device 310). It should be noted that the processes performed at block 1408 for matching active and obsolete paths can also be performed simultaneously, in which case each of the new image plane paths would also be compared with the obsolete paths before a final correlation is performed. At 1804, the comparison can be used to identify which of the new image plane paths correspond to one of the obsolete paths.Identification may involve implementing a matching criterion based on whether differences in the tracking data associated with the new image plane orbit and the outdated orbit are consistent with a time interval between the acquisition of the most recent data associated with the outdated orbit and the tracking data used to generate the new image plane orbit. In 1806, tracking data from matching orbits can be merged to create an updated active image plane orbit that incorporates information from both orbits. In some embodiments, a predicted object position generated from tracking data of the outdated orbit can be combined with the object's detected position from the new ground plane orbit to improve object accuracy.This type of combination is typically performed when the outdated path has just become obsolete and / or when, due to the fact that the object detection system tracks multiple objects, occlusion of one object by another tracked object is to be expected, and the deviations between the detected position and the predicted position are within the expected tolerances. It should be noted that even if in the . Fig. 18 describes the system of correlating obsolete objects, that the object tracking system in some embodiments works without an analysis of obsolete paths and simply ignores all tracking data that exceed a threshold for age or confidence level.

[0109] The foregoing description describes embodiments of the invention with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the description and the drawings should be considered illustrative rather than limiting. The sole and exclusive scope of protection of the invention is defined by the claims. Any definitions expressly set forth herein for terms contained in such claims are intended to determine the meaning of such terms as used in the claims.Furthermore, if the term “further comprising / further comprising” is used in the preceding description or subsequent claims, that which follows this expression may be an additional step or entity or an additional sub-step / sub-entity of a previously described step or entity.

Citation Information

Patent Citations

  • Method for detecting an object in an area surrounding a motor vehicle with prediction of the movement of the object, camera system and motor vehicle

    DE102016114168A1

  • Method and processing unit for detecting objects based on asynchronous sensor data

    DE102016203472A1