Vehicle, method for a vehicle, and storage medium
By receiving environmental data on the autonomous vehicle to generate and annotate the geometric model of the feasible area in the map, the problem that traditional navigation methods cannot adapt to the dynamic environment is solved, real-time update of the environment and precise positioning of navigation is achieved.
Patent Information
- Application Number
- CN202210556225.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-07
- Filing Date
- 2019-10-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-10-29
AI Technical Summary
Existing autonomous vehicle navigation methods rely on traditional static maps and cannot accurately locate and adapt to changing environmental characteristics and dynamic road conditions, especially in large environments, which are difficult to scale.
Receive environmental data through sensors on the carrier, generate and annotate geometric models of the travelable areas in the map, combine semantic data to automatically annotate, and update the real-time map to adapt to environmental changes.
Real-time updates and precise positioning of dynamic environments are achieved, and the navigation accuracy and security of autonomous vehicles in complex environments are improved.
Smart Images

Figure CN115082914B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of October 29, 2019, the application number of 201911035567.1, and the invention name of "Automatic Annotation of Environmental Features in a Map during Navigation of a Vehicle". Technical Field
[0002] This specification generally relates to navigation planning for vehicles, and particularly to the automatic annotation of environmental features in a map during the navigation of a vehicle. Background Art
[0003] Autonomous vehicles (AVs) have the potential to transform transportation systems by reducing road traffic accidents, traffic congestion, parking congestion, and fuel consumption. However, conventional methods for AV navigation typically rely on traditional static maps and are insufficient for precisely positioning AVs. Conventional methods for generating static maps may not be able to address the problems of changing environmental features and dynamic road conditions. Moreover, traditional mapping methods may not be able to scale to large environments such as cities or states. Summary of the Invention
[0004] Techniques are provided for the automatic annotation of environmental features in a map during the navigation of a vehicle. The techniques include using one or more processors of a vehicle located in the environment to receive a map of the environment. One or more sensors of the vehicle receive sensor data and semantic data. The sensor data includes multiple features of the environment. Based on the sensor data, a geometric model of a feature among the multiple features is generated. Generating includes associating the feature with a drivable area in the environment. Drivable segments are extracted from the drivable area. The drivable segments are divided into multiple geometric blocks. Each geometric block corresponds to a characteristic of the drivable area, and the geometric model of the feature includes the multiple geometric blocks. The geometric model is annotated using the semantic data. The annotated geometric model is embedded in the map.
[0005] In one embodiment, one or more sensors of a vehicle located in the environment are used to generate sensor data including features of the environment. Using the sensor data, the features are mapped to a drivable area within a map of the environment. Using the drivable area within the map, a polygon containing multiple geometric blocks is extracted. Each geometric block corresponds to a drivable segment of the drivable area. Using semantic data extracted from the drivable area within the map, the polygon is annotated. Using one or more processors of the vehicle, the annotated polygon is embedded into the map.
[0006] In one embodiment, the sensor data includes LiDAR point cloud data.
[0007] In one embodiment, one or more sensors include a camera, and the sensor data further includes an image of multiple features of the environment.
[0008] In one embodiment, the vehicle is located at a spatio-temporal position within the environment, and the multiple features are associated with the spatio-temporal position.
[0009] In one embodiment, the feature represents multiple lanes of a drivable area that are oriented in the same direction, and each geometric block of the multiple geometric blocks represents a single lane of the drivable area.
[0010] In one embodiment, the feature represents the height of a drivable area, a curb located near the drivable area, or a center line separating two lanes of the drivable area.
[0011] In one embodiment, the drivable area includes a road segment, a parking space located on the road segment, a parking lot connected to the road segment, or an open space located within the environment.
[0012] In one embodiment, one or more sensors include a Global Navigation Satellite System (GNSS) sensor or an Inertial Measurement Unit (IMU). The method further includes determining the spatial position of the vehicle relative to the boundary of the drivable area using the sensor data.
[0013] In one embodiment, the generation of the geometric model includes classifying LiDAR point cloud data.
[0014] In one embodiment, the generation of the geometric model includes superimposing multiple geometric blocks onto the LiDAR point cloud data to generate a polygon that includes the union of the geometric blocks among the multiple geometric blocks.
[0015] In one embodiment, the polygon represents dividing the lanes of the drivable area into multiple lanes.
[0016] In one embodiment, the polygon represents merging multiple lanes of the drivable area into a single lane.
[0017] In one embodiment, the polygon represents the intersection of multiple lanes of the drivable area.
[0018] In one embodiment, the polygon represents a roundabout, which includes spatial positions on the drivable area for a vehicle to enter or leave the roundabout.
[0019] In one embodiment, the polygon represents a bend in the lanes of the drivable area.
[0020] In one embodiment, annotating the geometric model includes generating a computer-readable semantic annotation that combines the geometric model and semantic data. The method further includes sending the map with the computer-readable semantic annotation to a remote server or another vehicle.
[0021] In one embodiment, annotating a geometric model using semantic data is performed in a first operating mode, and the method further includes navigating a vehicle on a drivable area using a map in a second operating mode using a control module of the vehicle.
[0022] In one embodiment, the semantic data represents markings on the drivable area, road signs located within the environment, or traffic signals located within the environment.
[0023] In one embodiment, annotating the geometric model using semantic data includes extracting logical driving constraints of the vehicle associated with navigating the vehicle along the drivable area from the semantic data.
[0024] In one embodiment, the logical driving constraints include a traffic signal sequence, conditional left or right turns, or traffic directions.
[0025] In one embodiment, in the second operating mode, an annotated geometric model is extracted from the map. Navigating the vehicle includes sending commands to the throttle or brakes of the vehicle in response to the extraction of the annotated geometric model.
[0026] In one embodiment, a plurality of operating metrics associated with navigating the vehicle along the drivable area are stored. Using the map, the plurality of operating metrics are updated.
[0027] Using the updated plurality of operating metrics, rerouting information for the vehicle is determined. The rerouting information is sent to a server located within the environment or another vehicle.
[0028] In one embodiment, using a control module of the vehicle, the vehicle is navigated on the drivable area using the determined rerouting information.
[0029] In one embodiment, a computer-readable semantic annotation of a map is received from a second vehicle. The map is merged with the received computer-readable semantic annotation for transmission to a remote server.
[0030] In one embodiment, embedding the annotated geometric model into the map is performed in the first operating mode. The method further includes determining the drivable area from the annotated geometric model within the map.
[0031] In one embodiment, the vehicle includes one or more computer processors. One or more non-transitory storage media store instructions that, when executed by the one or more computer processors, cause any of the embodiments disclosed herein to be performed.
[0032] In one embodiment, one or more non-transitory storage media store instructions that, when executed by one or more computing devices, cause any of the embodiments disclosed herein to be performed.
[0033] In one embodiment, a method includes performing machine-executed operations involving instructions that, when executed by one or more computing devices, cause any of the embodiments disclosed herein to be performed. The machine-executed operations are at least one of sending the instructions, receiving the instructions, storing the instructions, or executing the instructions.
[0034] In one embodiment, one or more non-transitory storage media store a map of an environment. The map is generated by receiving, using one or more processors of a vehicle located in the environment, the map of the environment. Using one or more sensors of the vehicle, sensor data and semantic data are received. The sensor data includes multiple features of the environment. Based on the sensor data, a geometric model of a feature among the multiple features is generated. Generating includes associating the feature with a drivable area in the environment. Drivable segments are extracted from the drivable area. The drivable segments are divided into multiple geometric blocks. Each geometric block corresponds to a characteristic of the drivable area, and the geometric model of the feature includes the multiple geometric blocks. The geometric model is annotated using the semantic data. The annotated geometric model is embedded in the map.
[0035] In one embodiment, a vehicle includes a communication device configured to receive, using one or more processors of a vehicle located in the environment, a map of the environment. One or more sensors are configured to receive sensor data and semantic data. The sensor data includes multiple features of the environment. One or more processors are configured to generate, from the sensor data, a geometric model of a feature among the multiple features. Generating includes associating the feature with a drivable area in the environment. Drivable segments are extracted from the drivable area. The drivable segments are divided into multiple geometric blocks. Each geometric block corresponds to a characteristic of the drivable area, and the geometric model of the feature includes the multiple geometric blocks. The geometric model is annotated using the semantic data. The annotated geometric model is embedded in the map.
[0036] These and other aspects, features, and implementations may be expressed as a method, apparatus, system, component, program product, means or step for performing a function, and in other ways.
[0037] These and other aspects, features, and implementations will become apparent from the following description including the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 FIG. illustrates an example of an autonomous vehicle (AV) having autonomous capabilities according to one or more embodiments.
[0039] Figure 2 The figure illustrates an example "cloud" computing environment in accordance with one or more embodiments.
[0040] Figure 3 The figure illustrates a computer system in accordance with one or more embodiments.
[0041] Figure 4 The figure illustrates an example architecture for AV in accordance with one or more embodiments.
[0042] Figure 5 The figure illustrates examples of inputs and outputs that may be used by a sensing module in accordance with one or more embodiments.
[0043] Figure 6 The figure illustrates an example of a LiDAR system in accordance with one or more embodiments.
[0044] Figure 7 The figure illustrates a LiDAR system in operation in accordance with one or more embodiments.
[0045] Figure 8 Illustrates in more detail the operation of a LiDAR system in accordance with one or more embodiments.
[0046] Figure 9 The figure illustrates a block diagram of the relationship between the inputs and outputs of a planning module in accordance with one or more embodiments.
[0047] Figure 10 The figure illustrates a directed graph used in path planning in accordance with one or more embodiments.
[0048] Figure 11 The figure illustrates a block diagram of the inputs and outputs of a control module in accordance with one or more embodiments.
[0049] Figure 12 The figure illustrates a block diagram of the inputs, outputs, and components of a controller in accordance with one or more embodiments.
[0050] Figure 13 The figure illustrates a block diagram of an architecture for automatic annotation of environmental features in a map during navigation of a vehicle in accordance with one or more embodiments.
[0051] Figure 14 The figure illustrates an example environment for automatic annotation of environmental features in a map during navigation of a vehicle in accordance with one or more embodiments.
[0052] Figure 15 The figure illustrates example automatically annotated environmental features in a map during navigation of a vehicle in accordance with one or more embodiments.
[0053] Figure 16The figure shows an example automatically annotated map during the navigation of a vehicle according to one or more embodiments.
[0054] Figure 17 The figure shows a process of automatically annotating environmental features in a map during the navigation of a vehicle according to one or more embodiments.
[0055] Figure 18 The figure shows a process of automatically annotating environmental features in a map during the navigation of a vehicle according to one or more embodiments. Detailed Description
[0056] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention.
[0057] In the drawings, for ease of description, a particular arrangement or order of schematic elements is shown, such as those representing devices, modules, instruction blocks, and data elements. However, those skilled in the art should understand that the particular ordering or arrangement of the schematic elements in the drawings is not intended to imply a requirement for a particular processing order or sequence, or a separation of processes. Additionally, the inclusion of schematic elements in the drawings does not imply that such elements are required in all embodiments, or that the features represented by such elements in some embodiments may not be included in or combined with other elements.
[0058] Furthermore, in the drawings, in cases where connecting elements such as solid lines or dashed lines or arrows are used to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element does not imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, in cases where the connecting element represents the communication of signals, data, or instructions, those skilled in the art should understand that such an element may represent one or more signal paths (e.g., a bus) as needed to effect the communication.
[0059] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the described embodiments. However, it will be apparent to one of ordinary skill in the art that the described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0060] The following describes several features, each of which may be used independently of one another or in any combination with other features. However, any single feature may not solve any of the problems discussed above, or may only solve one of the problems discussed above. Some of the problems discussed above may not be fully solved by any one of the features described herein. Although headings are provided, information related to a particular heading but not found in the section with that heading may also be found elsewhere in this specification. The embodiments are described herein according to the following outline:
[0061] 1. General Overview
[0062] 2. System Overview
[0063] 3. Autonomous Vehicle Architecture
[0064] 4. Autonomous Vehicle Input
[0065] 5. Autonomous Vehicle Planning
[0066] 6. Autonomous Vehicle Control
[0067] 7. Architecture for Automatic Annotation of Environmental Features in a Map
[0068] 8. Environment for Automatic Annotation of Environmental Features in a Map
[0069] 9. Example Automatically Annotated Environmental Features in a Map
[0070] 10. Example Automatically Annotated Map
[0071] 11. Process for Automatic Annotation of Environmental Features in a Map
[0072] General Overview
[0073] A vehicle annotates a real-time map of the environment in which it is navigating. The vehicle annotates the real-time map by obtaining information about the environment and automatically annotating the information onto the real-time map. The real-time map is a real-time map or depiction of the environment that represents dynamic environmental changes or conditions. The vehicle accesses multiple portions of the map and updates those portions in real time. To annotate the map, the vehicle receives a map of the environment from a server or another vehicle. The vehicle uses one or more sensors of the vehicle to receive sensor data and semantic data. For example, the sensors can include lidar or cameras. The sensor data includes multiple features of the environment. Each feature represents one or more characteristics of the environment, such as physical or semantic aspects. For example, the feature can model the surface or structure of a curb or a road centerline. Based on the sensor data, the vehicle generates a geometric model of the features of the environment. To generate the geometric model, the vehicle associates the features with a drivable area in the environment. The vehicle extracts drivable segments from the drivable area. For example, the drivable segments can include lanes of a road. The vehicle separates the extracted drivable segments into multiple geometric blocks. Each geometric block corresponds to a characteristic of the drivable area, and the geometric model of the features includes multiple geometric blocks. The vehicle annotates the geometric model using the semantic data. The vehicle embeds the annotated geometric model into the real-time map to update the real-time map with changing conditions in the environment. For example, the annotation can inform other vehicles of new traffic patterns or new construction areas in the environment.
[0074] System Overview
[0075] Figure 1 The figure illustrates an example of an autonomous vehicle 100 with autonomous capabilities.
[0076] As used herein, the term "autonomous capabilities" refers to functions, features, or facilities that enable a vehicle to be operated partially or fully without real-time human intervention, including but not limited to fully autonomous vehicles, highly autonomous vehicles, and conditionally autonomous vehicles.
[0077] As used herein, an autonomous vehicle (AV) is a vehicle with autonomous capabilities.
[0078] As used herein, "vehicle" includes a device for transporting goods or people. For example, cars, buses, trains, airplanes, drones, trucks, ships, vessels, submersibles, airships, etc. A driverless car is an example of a vehicle.
[0079] As used herein, "trajectory" refers to a path or route for navigating an AV from a first spatio-temporal location to a second spatio-temporal location. In an embodiment, the first spatio-temporal location is referred to as an initial or starting location, and the second spatio-temporal location is referred to as a destination, final location, target, target location, or target site. In some examples, a trajectory is composed of one or more segments (e.g., road segments), and each segment is composed of one or more blocks (e.g., lanes or portions of intersections). In an embodiment, a spatio-temporal location corresponds to a real-world location. For example, a spatio-temporal location is a pick-up or drop-off location for picking up or dropping off people or goods.
[0080] As used herein, "(a) sensor(s)" includes one or more hardware components that detect information about the environment around the sensor. Some of the hardware components may include sensing components (e.g., image sensors, biometric sensors), transmitting and / or receiving components (e.g., laser or radio wave transmitters and receivers), electronic components (such as, analog-to-digital converters), data storage devices (such as, RAM and / or non-volatile storage), software or firmware components, and data processing components (such as, ASICs (application specific integrated circuits)), microprocessors, and / or microcontrollers.
[0081] As used herein, "scene description" is a data structure (e.g., a list) or data stream that includes one or more classified or identified objects detected by one or more sensors on an AV vehicle or provided by a source external to the AV.
[0082] As used herein, "road" is a physical area that can be traversed by a vehicle, and may correspond to a named passageway (e.g., a city street, an interstate highway, etc.), or may correspond to an unnamed passageway (e.g., a lane in a residential or office building, a section of a parking lot, a section of an open space, and a dirt road in a rural area, etc.). Since some vehicles (e.g., four-wheel pickups, sport utility vehicles, etc.) are capable of traversing various physical areas that are not specifically adapted for vehicle travel, "road" can be a physical area that is not officially defined as a passageway by any municipality or other government or administrative body.
[0083] As used herein, a "lane" is a portion of a road that can be traversed by a vehicle and can correspond to most or all of the space between lane markings, or can correspond to only some (e.g., less than 50%) of the space between lane markings. For example, a road with widely spaced lane markings can accommodate two or more vehicles between the markings such that one vehicle can pass another without crossing the lane markings and can thus be interpreted as having a lane narrower than the space between the lane markings, or having two lanes between the lane markings. In the absence of lane markings, lanes can also be interpreted. For example, lanes can be defined based on the physical characteristics of the environment, such as rocks and trees along a path in a rural area.
[0084] "One or more" includes: a function performed by one element, a function performed by more than one element, such as in a distributed manner, several functions performed by one element, several functions performed by several elements, or any combination of the above.
[0085] It will also be understood that although in some instances the terms first, second, etc. are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first contact member can be referred to as a second contact member, and similarly, a second contact member can be referred to as a first contact member, without departing from the scope of the various described embodiments. Both the first contact member and the second contact member are contact members, but they are not the same contact member.
[0086] The terms used in the description of the various embodiments herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes", "including", "comprises", and "comprising", when used in this application, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0087] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when" or "after" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if [stated condition or event] is detected" is optionally interpreted to mean "after determining" or "in response to determining" or "after detecting [stated condition or event]" or "in response to detecting [stated condition or event]".
[0088] As used herein, an AV system refers to the AV and an array of hardware, software, stored data, and real-time generated data that support the operation of the AV. In an embodiment, the AV system is incorporated into the AV. In an embodiment, the AV system is distributed across several locations. For example, some implementations of the software of the AV system are on a cloud computing environment similar to the cloud computing environment 300 described below Figure 3 on a cloud computing environment.
[0089] Generally, the techniques described herein apply to any vehicle having one or more autonomous capabilities, including fully autonomous vehicles, highly autonomous vehicles, and conditionally autonomous vehicles, such as so-called level 5, level 4, and level 3 vehicles, respectively (see SAE International Standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, which is incorporated herein by reference in its entirety for more details regarding the classification of vehicle autonomy levels). The techniques described in this document may also apply to partially autonomous vehicles and driver-assisted vehicles, such as so-called level 2 and level 1 vehicles (see SAE International Standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems). In an embodiment, one or more of the level 1, level 2, level 3, level 4, and level 5 vehicle systems can automate certain vehicle operations (e.g., steering, braking, and using maps) under certain operating conditions based on the processing of sensor inputs. The techniques described herein can benefit any level of vehicle within the range from fully autonomous vehicles to manually operated vehicles.
[0090] Reference Figure 1, the AV system 120 operates the AV 100 along a trajectory 198 through an environment 190 to a destination 199 (sometimes referred to as a final position), while avoiding objects (e.g., natural obstacles 191, vehicles 193, pedestrians 192, cyclists, and other obstacles) and complying with road rules (e.g., operating rules or driving preferences).
[0091] In an embodiment, the AV system 120 includes a device 101 that is configured to receive and act on operation commands from a computer processor 146. In an embodiment, the computing processor 146 is similar to the processor 304 described below with reference to Figure 3 Examples of the device 101 include a steering control 102, a brake 103, gears, an accelerator pedal or other acceleration control mechanisms, windshield wipers, side door locks, window controls, or turn indicators.
[0092] In an embodiment, the AV system 120 includes sensors 121 for measuring or inferring attributes of the state or condition of the AV 100, such as the position of the AV, linear velocity and acceleration, angular velocity and acceleration, and heading (e.g., the orientation of the front end of the AV 100). Examples of the sensors 121 are GNSS, an inertial measurement unit (IMU) that measures the linear acceleration and angular velocity of the vehicle, wheel speed sensors for measuring or estimating wheel slip ratio, wheel brake pressure or brake torque sensors, engine torque or wheel torque sensors, and steering angle and angular velocity sensors.
[0093] In an embodiment, the sensors 121 also include sensors for sensing or measuring attributes of the environment of the AV. For example, monocular or stereo video cameras 122, LiDAR 123, radar, ultrasonic sensors, time-of-flight (TOF) depth sensors, speed sensors, temperature sensors, humidity sensors, and precipitation sensors that employ the visible light, infrared, or thermal (or both) spectra.
[0094] In an embodiment, the AV system 120 includes a data storage unit 142 and a memory 144 for storing machine instructions associated with the computer processor 146 or data collected by the sensors 121. In an embodiment, the data storage unit 142 is similar to the ROM 308 or storage device 310 described below in connection with Figure 3 In an embodiment, the memory 144 is similar to the main memory 306 described below. In an embodiment, the data storage unit 142 and the memory 144 store historical information, real-time information, and / or predictive information related to the environment 190. In an embodiment, the stored information includes maps, driving performance, traffic jam updates, or weather conditions. In an embodiment, data related to the environment 190 is transmitted from a remotely located database 134 to the AV 100 via a communication channel.
[0095] In an embodiment, the AV system 120 includes a communication device 140 for transmitting to the AV 100 the measured or inferred attributes of the status and condition of other vehicles, such as position, linear and angular velocity, linear and angular acceleration, and linear and angular heading directions. These devices include vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication devices and devices for wireless communication via point-to-point or ad hoc networks or both. In an embodiment, the communication device 140 communicates across the electromagnetic spectrum (including radio and optical communication) or other media (e.g., air and acoustic media). The combination of vehicle-to-vehicle (V2V) communication and vehicle-to-infrastructure (V2I) communication (and in some embodiments, one or more other types of communication) is sometimes referred to as vehicle-to-everything (V2X) communication. V2X communication generally complies with one or more communication standards for communicating with or between autonomous vehicles.
[0096] In an embodiment, the communication device 140 includes a communication interface. For example, a wired, wireless, WiMAX, Wi-Fi, Bluetooth, satellite, cellular, optical, near field, infrared, or radio interface. The communication interface transfers data from a remotely located database 134 to the AV system 120. In an embodiment, the remotely located database 134 is embedded in a cloud computing environment 200 as shown in Figure 2 The communication interface transfers data collected from the sensors 121 or other data related to the operation of the AV 100 to the remotely located database 134. In an embodiment, the communication interface transfers information related to teleoperation to the AV 100. In some embodiments, the AV 100 communicates with other remote (e.g., "cloud") servers 136.
[0097] In an embodiment, the remotely located database 134 also stores and transmits digital data (e.g., stores data such as road and street locations). Such data is stored in the memory 144 located on the AV 100 or is transferred to the AV 100 from the remotely located database 134 via a communication channel.
[0098] In an embodiment, the remotely located database 134 stores and transmits historical information about the driving attributes of a vehicle (e.g., speed and acceleration profiles) that previously traveled along a trajectory 198 at a similar time of day. In one implementation, such data can be stored in the memory 144 located on the AV 100 or can be transferred to the AV 100 from the remotely located database 134 via a communication channel.
[0099] The computing device 146 located on the AV 100 generates control actions algorithmically based on real-time sensor data and prior information, allowing the AV system 120 to perform its autonomous driving capabilities.
[0100] In an embodiment, the AV system 120 includes computer peripherals 132 coupled to the computing device 146 for providing information and alerts to a user (e.g., a passenger or a remote user) of the AV 100 and receiving input from the user. In an embodiment, the peripherals 132 are similar to the display 312, input device 314, and cursor controller 316 discussed below with reference to Figure 3 The coupling is wireless or wired. Any two or more of the interface devices can be integrated into a single device.
[0101] Example cloud computing environment
[0102] Figure 2 An example “cloud” computing environment is shown. Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services). In a typical cloud computing system, one or more large cloud data centers house the machines used to deliver the services provided by the cloud. Now refer to Figure 2 , the cloud computing environment 200 includes cloud data centers 204a, 204b, and 204c, which are interconnected by a cloud 202. The data centers 204a, 204b, and 204c provide cloud computing services to computer systems 206a, 206b, 206c, 206d, 206e, and 206f connected to the cloud 202.
[0103] The cloud computing environment 200 includes one or more cloud data centers. Generally, a cloud data center (e.g., Figure 2 the cloud data center 204a shown in Figure 2 refers to the physical arrangement of servers that make up a cloud (e.g., Figure 3 the cloud 202 shown in
[0104] The cloud 202 includes cloud data centers 204a, 204b, and 204c and network and networking resources (e.g., networking equipment, nodes, routers, switches, and networking cables) that interconnect the cloud data centers 204a, 204b, and 204c and help facilitate access to cloud computing services by computing systems 206a-f. In an embodiment, the network represents any combination of one or more local area networks, wide area networks, or internetworks coupled using wired or wireless links, and the wired and wireless links are deployed using terrestrial or satellite connections. Data exchanged over the network is transmitted using any number of network layer protocols such as Internet Protocol (IP), Multiprotocol Label Switching (MPLS), Asynchronous Transfer Mode (ATM), and Frame Relay. Further, in embodiments where the network represents a combination of multiple subnets, different network layer protocols are used at each of the underlying subnets. In some embodiments, the network represents one or more interconnected internetworks such as the public Internet.
[0105] The computing systems 206a-f or cloud computing device consumers are connected to the cloud 202 via network links and network adapters. In an embodiment, the computing systems 206a-f are implemented as various computing devices such as servers, desktop computers, laptop computers, tablets, smartphones, Internet of Things (IoT) devices, autonomous vehicles (including cars, drones, shuttles, trains, buses, etc.), and consumer electronic products. In an embodiment, the computing systems 206a-f are implemented in other systems or as part of other systems.
[0106] Computer system
[0107] Figure 3 FIG. shows a computer system 300. In an implementation, the computer system 300 is a special-purpose computing device. A special-purpose computing device is hardwired to perform techniques, or includes digital electronic devices (such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)) that are persistently programmed to perform techniques, or may include one or more general-purpose hardware processors programmed to perform techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such special-purpose computing devices may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. In various embodiments, the special-purpose computing device is a desktop computer system, a portable computer system, a handheld device, a network device, or any other device that includes hardwired and / or program logic for implementing these techniques.
[0108] In an embodiment, computer system 300 includes a bus 302 or other communication mechanism for transferring information, and a hardware processor 304 coupled to the bus 304 for processing information. The hardware processor 304 is, for example, a general-purpose microprocessor. The computer system 300 also includes a main memory 306 (such as, a random access memory (RAM) or other dynamic storage device) coupled to the bus 302 for storing information and instructions for execution by the processor 304. In one implementation, the main memory 306 is used to store temporary variables or other intermediate information during the execution of instructions for execution by the processor 304. Such instructions, when stored in a non-transitory storage medium accessible to the processor 304, cause the computer system 300 to be presented as a special-purpose machine customized to perform the operations specified in these instructions.
[0109] In an embodiment, the computer system 300 further includes a read-only memory (ROM) 308 or other static storage device coupled to the bus 302 for storing static information and instructions for the processor 304. A storage device 310 is provided and coupled to the bus 302 for storing information and instructions, such as a magnetic disk, an optical disk, a solid state drive, or a three-dimensional cross-point memory.
[0110] In an embodiment, the computer system 300 is coupled via the bus 302 to a display 312 for displaying information to a computer user, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, or an organic light emitting diode (OLED) display. An input device 314 including alphanumeric and other keys is coupled to the bus 302 for transmitting information and command selections to the processor 304. Another type of user input device is a cursor controller 316 for transmitting direction information and command selections to the processor 304 and for controlling cursor movement on the display 312, such as a mouse, a trackball, a touch-enabled display, or cursor direction keys. The input device typically has two degrees of freedom along two axes (a first axis (e.g., the x-axis) and a second axis (e.g., the y-axis)), which allows the device to specify a position in a plane.
[0111] According to one embodiment, the techniques described herein are performed by the computer system 300 in response to one or more sequences of one or more instructions contained in the main memory 306 being executed by the processor 304. Such instructions are read into the main memory 306 from another storage medium, such as the storage device 310. Execution of the instruction sequences contained in the main memory 306 causes the processor 304 to perform the process steps described herein. In alternative embodiments, hardwired circuitry is used in place of, or in combination with, software instructions.
[0112] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs, magnetic disks, solid state drives, or three-dimensional cross-point memories, such as storage device 310. Volatile media include dynamic memories, such as main memory 306. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM, and EPROM, flash-EPROM, NV-RAM, or any other memory chip or cartridge memory.
[0113] Storage media are different from transmission media but can be used in conjunction with transmission media. Transmission media participate in transferring information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the wires that form bus 302. Transmission media can also take the form of acoustic waves or light waves, such as those generated during radio wave and infrared data communications.
[0114] In an embodiment, various forms of media are involved in carrying one or more sequences of one or more instructions to processor 304 for execution. For example, the instructions are initially carried on a disk or solid state drive of a remote computer. The remote computer loads the instructions into its dynamic memory and sends the instructions over a telephone line using a modem. A modem local to computer system 300 receives the data on the telephone line and converts the data to an infrared signal using an infrared transmitter. An infrared detector receives the data carried in the infrared signal, and appropriate circuitry places the data on bus 302. Bus 302 carries the data to main memory 306, from which processor 304 retrieves and executes the instructions. The instructions received by main memory 306 may optionally be stored on storage device 310 before or after being executed by processor 304.
[0115] The computer system 300 also includes a communication interface 318 coupled to the bus 302. The communication interface 318 provides a two-way data communication coupling to a network link 320 that is connected to a local network 322. For example, the communication interface 318 is an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection to a corresponding type of telephone line. As another example, the communication interface 318 is a LAN card for providing a data communication connection to a compatible local area network (LAN). In some implementations, a wireless link is also implemented. In any such implementation, the communication interface 318 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0116] The network link 320 typically provides data communication through one or more networks to other data devices. For example, the network link 320 provides a connection through the local network 322 to a host computer 324 or to a cloud data center or facility operated by an Internet service provider (ISP) 326. The ISP 326 in turn provides data communication services through the worldwide packet data communication network, now commonly referred to as the "Internet" 328. Both the local network 322 and the Internet 328 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link 320 and through the communication interface 318 are example forms of transmission media that carry digital data to and from the computer system 300. In an embodiment, the network 320 includes the cloud 202 or a portion of the cloud 202 described above.
[0117] The computer system 300 sends messages and receives data including program code through the (one or more) networks, the network link 320, and the communication interface 318. In an embodiment, the computer system 300 receives code for processing. The received code is executed by the processor 304 when it is received, and / or is stored in the storage device 310 and / or other non-volatile storage for later execution.
[0118] Autonomous Vehicle Architecture
[0119] Figure 4 The figure shows an example architecture 400 for an autonomous vehicle (e.g., Figure 1 the AV 100 shown in ). The architecture 400 includes a perception module 402 (sometimes referred to as a perception circuit), a planning module 404 (sometimes referred to as a planning circuit), a control module 406 (sometimes referred to as a control circuit), a localization module 408 (sometimes referred to as a localization circuit), and a database module 410 (sometimes referred to as a database circuit). Each module plays a role in the operation of the AV 100. The modules 402, 404, 406, 408, and 410 together can beFigure 1 A portion of the AV system 120 shown. In some embodiments, any one of the modules 402, 404, 406, 408, and 410 is a combination of computer software (e.g., executable code stored on a computer-readable medium) and computer hardware (e.g., one or more microprocessors, microcontrollers, application-specific integrated circuits (ASICs), hardware memory devices, other types of integrated circuits, other types of computer hardware, or any combination of these things or all of these things).
[0120] In use, the planning module 404 receives data representing the destination 412 and determines data representing a trajectory 414 (sometimes referred to as a route) along which the AV 100 can travel to reach (e.g., arrive at) the destination 412. To enable the planning module 404 to determine the data representing the trajectory 414, the planning module 404 receives data from the perception module 402, the positioning module 408, and the database module 410.
[0121] The perception module 402 uses one or more sensors 121 (e.g., also as Figure 1 shown therein) to identify nearby physical objects. The objects are classified (e.g., grouped into types such as pedestrians, bicycles, motor vehicles, traffic signs, etc.), and a scene description including the classified objects 416 is provided to the planning module 404.
[0122] The planning module 404 also receives data representing the AV position 418 from the positioning module 408. The positioning module 408 determines the AV position by calculating the position using data from the sensors 121 and data from the database module 410 (e.g., geographical data). For example, the positioning module 408 uses data from GNSS (Global Navigation Satellite System) sensors and uses geographical data to calculate the longitude and latitude of the AV. In an embodiment, the data used by the positioning module 408 includes: a high-precision map of road geometric attributes, a map describing the road network connectivity attributes, a map describing the road physical attributes (such as traffic speed, traffic volume, the number of vehicle and bicycle lanes, lane width, lane traffic direction, or lane marking type and location, or a combination thereof), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various other driving signals).
[0123] The control module 406 receives data representing the trajectory 414 and data representing the AV position 418, and operates the control functions 420a-c (e.g., steering, throttle, braking, ignition) of the AV in a manner that will cause the AV 100 to travel along the trajectory 414 to the destination 412. For example, if the trajectory 414 includes a left turn, the control module 406 will operate the control functions 420a-c in a manner such that the steering angle of the steering function will cause the AV 100 to turn left and the throttle and brakes will cause the AV 100 to stop and wait for passing pedestrians or vehicles before making the turn.
[0124] Autonomous vehicle input
[0125] Figure 5 The figure shows examples of the inputs 502a-d (e.g., Figure 4 ) used by the sensing module 402 ( Figure 1 ) and the outputs 504a-d (e.g., sensor data). One input 502a is a LiDAR (light detection and ranging) system (e.g., Figure 1 ) shown in. LiDAR is a technology that uses light (e.g., bursts of light such as infrared light) to obtain data about physical objects within its line of sight. The LiDAR system produces LiDAR data as the output 504a. For example, LiDAR data is a collection of 3D or 2D points (also known as a point cloud) that is used to construct a representation of the environment 190.
[0126] Another input 502b is a radar system. Radar is a technology that uses radio waves to obtain data about nearby physical objects. Radar can obtain data about objects that are not within the line of sight of the LiDAR system. The radar system 502b generates radar data as the output 504b. For example, radar data is one or more radio frequency electromagnetic signals that are used to construct a representation of the environment 190.
[0127] Another input 502c is a camera system. The camera system uses one or more cameras (e.g., a digital camera using a light sensor such as a charge coupled device (CCD)) to obtain information about nearby physical objects. The camera system generates camera data as output 504c. The camera data is typically in the form of image data (e.g., data in image data formats such as RAW, JPEG, PNG, etc.). In some examples, the camera system has multiple independent cameras, such as for stereoscopic observation (stereo vision), which enables the camera system to perceive depth. Although the objects perceived by the camera system are described as "nearby" in this article, this is relative to the AV. In use, the camera system can be configured to "see" distant objects, for example, objects located up to one kilometer or more in front of the AV. Accordingly, the camera system may have features such as sensors and lenses that are optimized to perceive distant objects.
[0128] Another input 502d is a traffic light detection (TLD) system. The TLD system uses one or more cameras to obtain information about traffic lights, street signs, and other physical objects that provide visual navigation information. The TLD system generates TLD data as an output 504d. The TLD data is typically in the form of image data (e.g., data in image data formats such as RAW, JPEG, PNG, etc.). The TLD system is different from a system that incorporates a camera in that the TLD system uses a camera with a wide field of view (e.g., using a wide-angle lens or a fisheye lens) to obtain information about as many physical objects as possible that provide visual navigation information, thereby enabling the AV 100 to access all relevant navigation information provided by these objects. For example, the viewing angle of the TLD system can be about 120 degrees or greater.
[0129] In some embodiments, the outputs 504a-d are combined using sensor fusion techniques. Thus, the individual outputs 504a-d are provided to other systems of the AV 100 (e.g., to Figure 4 The combined outputs can be provided to other systems in the form of a single combined output or multiple combined outputs of the same type (e.g., using the same combining technique, or combining the same outputs, or both) or different types (e.g., using different corresponding combining techniques, or combining different corresponding outputs, or both). In some embodiments, early fusion techniques are used. Early fusion techniques are characterized by combining the outputs before one or more data processing steps are applied to the combined outputs. In some embodiments, late fusion techniques are used. Late fusion techniques are characterized by combining the outputs after one or more data processing steps are applied to the individual outputs.
[0130] Figure 6The figure shows an example of a LiDAR system 602 (e.g., Figure 5 the input 502a shown in). The LiDAR system 602 emits light 604a-c from a light emitter 606 (e.g., a laser emitter). The rays emitted by the LiDAR system are typically not within the visible spectrum; for example, infrared rays are commonly used. Some of the emitted light 604b encounters a physical object 608 (e.g., a vehicle) and is reflected back to the LiDAR system 602. (The light emitted by the LiDAR system generally does not penetrate physical objects, e.g., physical objects that exist in a solid state.) The LiDAR system 602 also has one or more light detectors 610 that detect the reflected light. In an embodiment, one or more data processing systems associated with the LiDAR system generate an image 612 that represents the field of view 614 of the LiDAR system. The image 612 includes information representing the boundary 616 of the physical object 608. In this way, the image 612 is used to determine the boundary 616 of one or more physical objects near the AV.
[0131] Figure 7 The figure shows the LiDAR system 602 in operation. In the scenario shown in this figure, the AV 100 receives the camera system output 504c in the form of an image 702 and the LiDAR system output 504a in the form of LiDAR data points 704. In use, the data processing system of the AV100 compares the image 702 with the data points 704. Specifically, the physical objects 706 identified in the image 702 are also identified in the data points 704. In this way, the AV 100 perceives the boundaries of physical objects based on the contours and densities of the data points 704.
[0132] Figure 8 The operation of the LiDAR system 602 is shown in more detail. As described above, the AV 100 detects the boundaries of physical objects based on the characteristics of the data points detected by the LiDAR system 602. As Figure 8 shown, flat objects (such as the ground 802) will reflect the light 804a-d emitted by the LiDAR system 602 in a consistent manner. In other words, since the LiDAR system 602 emits rays at consistent intervals, the ground 802 will reflect the rays back to the LiDAR system 602 at the same consistent intervals. As the AV 100 travels on the ground 802, if there are no objects obstructing the road, the LiDAR system 602 will continue to detect the light reflected by the next valid ground point 806. However, if an object 808 obstructs the road, the light 804e-f emitted by the LiDAR system 602 will be reflected from points 810a-b in a manner inconsistent with the expected consistent manner. From this information, the AV 100 can determine that there is an object 808.
[0133] Path Planning
[0134] Figure 9 The block diagram 900 shows the relationship between the inputs and outputs of the planning module 404 (e.g., as shown in Figure 4 ). Generally, the output of the planning module 404 is a route 902 from a starting point 904 (e.g., a source location or an initial location) to an ending point 906 (e.g., a destination or a final location). The route 902 is typically defined by one or more segments. For example, a segment refers to a distance traveled on at least a portion of a street, road, highway, lane, or other physical area suitable for motor vehicle travel. In some examples, e.g., if the AV 100 is a vehicle with off-road capabilities (such as a four-wheel drive (4WD) or all-wheel drive (AWD) car, SUV, pickup truck, etc.), then the route 902 includes "off-road" segments such as unpaved roads or open spaces.
[0135] In addition to the route 902, the planning module also outputs lane-level route planning data 908. The lane-level route planning data 908 is used to traverse the segments of the route 902 based on the conditions of the segments at a particular time. For example, if the route 902 includes a multi-lane highway, then the lane-level route planning data 908 includes trajectory planning data 910 that the AV 100 can use to select a lane among multiple lanes, e.g., based on whether it is approaching an exit, whether one or more lanes in the lane have other vehicles, or other factors that change within a few minutes or less. Similarly, in some implementations, the lane-level route planning data 908 includes speed constraints 912 specific to a section of the route 902. For example, if the section includes pedestrians or unexpected traffic, the speed constraints 912 can limit the AV 100 to a driving speed lower than the expected speed (e.g., the speed based on the speed limit data for that section).
[0136] In an embodiment, the inputs to the planning module 404 include database data 914 (e.g., from a database module 410 as shown in Figure 4 ), current location data 916 (e.g., the AV location 418 as shown in Figure 4 ), destination data 918 (e.g., for a destination 412 as shown in Figure 4 ), and object data 920 (e.g., by an object module 420 as shown in Figure 4The classified object 416 sensed by the illustrated sensing module 402). In some embodiments, the database data 914 includes rules used in planning. The rules are specified using a formal language, e.g., using boolean logic. In any given situation encountered by the AV 100, at least some of the rules will apply to that situation. If a rule has a condition that is satisfied based on information available to the AV 100 (e.g., information about the surrounding environment), then the rule applies to the given situation. Rules can have priorities. For example, the rule "if the road is a highway, move to the leftmost lane" may have a lower priority than "if an exit is approaching within one mile, move to the rightmost lane".
[0137] Figure 10 The figure shows a directed graph 1000 used in route planning (e.g., performed by the planning module 404( Figure 4 ). Generally, a directed graph 1000 such as the directed graph shown in Figure 10 is used to determine a path between any starting point 1002 and ending point 1004. In the real world, the distance separating the starting point 1002 and the ending point 1004 can be relatively large (e.g., in two different urban areas) or can be relatively small (e.g., two intersections of adjacent city blocks or two lanes of a multi-lane road).
[0138] In an embodiment, the directed graph 1000 has nodes 1006a-d that represent different positions between the starting point 1002 and the ending point 1004 that can be occupied by the AV 100. In some examples, e.g., when the starting point 1002 and the ending point 1004 represent different urban areas, the nodes 1006a-d represent segments of a road. In some examples, e.g., when the starting point 1002 and the ending point 1004 represent different positions on the same road, the nodes 1006a-d represent different positions on that road. In this way, the directed graph 1000 includes information at different levels of granularity. In an embodiment, a directed graph with a high granularity is also a subgraph of another directed graph with a larger proportion. For example, in a directed graph where the starting point 1002 is far from the ending point 1004 (e.g., many miles apart), most of the information in the graph has a low granularity and the graph is based on the stored data, but it also includes some high-granularity information for the part of the graph that represents the physical positions within the field of view of the AV 100.
[0139] Nodes 1006a-d are different from objects 1008a-b, and objects 1008a-b cannot overlap with the nodes. In an embodiment, when the granularity is low, objects 1008a-b represent areas that cannot be traversed by a motor vehicle, e.g., areas without streets or roads. When the granularity is high, objects 1008a-b represent physical objects within the field of view of the AV 100, e.g., other motor vehicles, pedestrians, or other entities with which the AV 100 cannot share physical space. In an embodiment, some or all of objects 1008a-b are static objects (e.g., objects that do not change position, such as street lights or utility poles) or dynamic objects (e.g., objects that can change position, such as pedestrians or other vehicles).
[0140] Nodes 1006a-d are connected by edges 1010a-c. If two nodes 1006a-b are connected by an edge 1010a, the AV 100 can travel between one node 1006a and another node 1006b, e.g., without having to travel to an intermediate node before reaching the other node 1006b. (When we refer to the AV 100 traveling between nodes, we mean the AV 100 traveling between two physical locations represented by the corresponding nodes.) Edges 1010a-c are generally two-way in the sense that the AV 100 can travel from a first node to a second node or from the second node to the first node. In an embodiment, edges 1010a-c are one-way in the sense that the AV 100 can travel from a first node to a second node but the AV 100 cannot travel from the second node to the first node. Edges 1010a-c are one-way when they represent, e.g., one-way streets, individual lanes of a street, road, or highway, or other features that can only be traversed in one direction due to legal or physical constraints.
[0141] In an embodiment, the planning module 404 uses the directed graph 1000 to identify a path 1012 composed of nodes and edges between the starting point 1002 and the ending point 1004.
[0142] Edges 1010a-c have associated costs 1014a-b. Costs 1014a-b are values representing the resources that will be expended if the AV 100 selects that edge. A typical resource is time. For example, if one edge 1010a represents twice the physical distance of another edge 1010b, the associated cost 1014a of the first edge 1010a can be twice the associated cost 1014b of the second edge 1010b. Other factors affecting time include expected traffic, the number of intersections, speed limits, etc. Another typical resource is fuel economy. Two edges 1010a-b may represent the same physical distance, but one edge 1010a may require more fuel than the other edge 1010b (e.g., due to road conditions, expected weather, etc.).
[0143] When the planning module 404 identifies a path 1012 between a starting point 1002 and an ending point 1004, the planning module 404 typically selects a path optimized for cost, e.g., a path having the lowest total cost when the costs of the respective edges are added together.
[0144] Autonomous vehicle control
[0145] Figure 11 The figure shows (e.g., as shown in Figure 4 a block diagram 1100 of the inputs and outputs of the control module 406. The control module operates according to a controller 1102 that includes, for example, one or more processors similar to the processor 304 (e.g., one or more computer processors such as a microprocessor, or a microcontroller, or both), short-term and / or long-term data storage similar to the main memory 306, the ROM 1308, and the storage device 210 (e.g., a memory random access memory, or a flash memory, or both), and instructions stored in the memory that, when executed (e.g., by one or more processors), perform the operations of the controller 1102.
[0146] In an embodiment, the controller 1102 receives data representing a desired output 1104. The desired output 1104 typically includes speed, e.g., rate and forward direction. The desired output 1104 can be based on, for example, data received from the planning module 404 (e.g., as shown in Figure 4 ). Based on the desired output 1104, the controller 1102 generates data that can be used as a throttle input 1106 and a steering input 1108. The throttle input 1106 represents, for example, the magnitude by which to engage the throttle (e.g., acceleration control) of the AV 100 by engaging the steering pedal, or engaging another throttle control, to achieve the desired output 1104. In some examples, the throttle input 1106 also includes data that can be used to engage the brakes (e.g., deceleration control) of the AV 100. The steering input 1108 represents the steering angle, e.g., the angle at which the steering control of the AV (e.g., the steering wheel, the steering angle actuator, or other functions for controlling the steering angle) should be positioned to achieve the desired output 1104.
[0147] In an embodiment, the controller 1102 receives feedback that is used to adjust the inputs provided to the throttle and steering. For example, if the AV 100 encounters a disturbance 1110, such as a hill, the measured speed 1112 of the AV 100 is reduced to the expected output speed. In an embodiment, any measured output 1114 is provided to the controller 1102 such that, for example, necessary adjustments are performed based on the difference 1113 between the measured speed and the desired output. The measured outputs 1114 include the measured position 1116, the measured speed 1118 (including rate and orientation), the measured acceleration 1120, and other outputs measurable by the sensors of the AV 100.
[0148] In an embodiment, information about the disturbance 1110 is detected in advance (e.g., by sensors such as cameras or LiDAR sensors) and provided to the predictive feedback module 1122. The predictive feedback module 1122 then provides information to the controller 1102, which the controller 1102 can use to make adjustments accordingly. For example, if the sensors of the AV 100 detect (“see”) a hill, this information can be used by the controller 1102 to prepare to engage the throttle at an appropriate time to avoid significant deceleration.
[0149] Figure 12 The figure shows a block diagram 1200 of the inputs, outputs, and components of the controller 1102. The controller 1102 has a speed analyzer 1202 that affects the operation of the throttle / brake controller 1204. For example, depending on the feedback received by the controller 1102 and processed by the speed analyzer 1202, the speed analyzer 1202 uses the throttle / brake 1206 to instruct the throttle / brake controller 1204 to participate in acceleration or deceleration.
[0150] The controller 1102 also has a lateral tracking controller 1208 that affects the operation of the steering controller 1210. For example, depending on the feedback received by the controller 1102 and processed by the lateral tracking controller 1208, the lateral tracking controller 1208 instructs the steering controller 1204 to adjust the position of the steering angle actuator 1212.
[0151] The controller 1102 receives several inputs for determining how to control the throttle / brake 1206 and the steering angle actuator 1212. The planning module 404 provides information used by, for example, the controller 1102 to select a heading when the AV 100 starts operating and to determine which segment to traverse when the AV 100 reaches an intersection. The localization module 408 provides information to the controller 1102 that describes, for example, the current position of the AV 100, enabling the controller 1102 to determine whether the AV 100 is at a position expected based on the manner in which the throttle / brake 1206 and the steering angle actuator 1212 are being controlled. In an embodiment, the controller 1102 receives information from other inputs 1214 (e.g., information received from a database, a computer network, etc.).
[0152] Architecture for Automatic Annotation of Maps
[0153] Figure 13 FIG. shows a block diagram of an architecture 1300 for automatic annotation of environmental features in a map 1364 during navigation of an AV 1308, according to one or more embodiments. The architecture 1300 includes an environment 1304 in which the AV 1308 is located. The architecture 1300 also includes a remote server 1312 and one or more other vehicles 1316. The one or more other vehicles 1316 are other AVs, semi-autonomous vehicles, or non-autonomous vehicles that are navigating or parked outside or within the environment 1304. For example, the one or more other vehicles 1316 enter and leave the environment 1304 during navigation and navigate in other environments. The one or more other vehicles 1316 can be part of the traffic that the AV 1308 encounters on the roads of the environment 1304. In some embodiments, the one or more other vehicles 1316 belong to one or more AV fleets. In other embodiments, the architecture 1300 includes additional or fewer components compared to the components described herein. Similarly, the functions can be distributed among the components and / or different entities in a manner different from the manner described herein.
[0154] The environment 1304 can be an example of the environment 190 referenced above Figure 1 shown and described. The environment 1304 represents a geographic area, such as a state, town, neighborhood, or road network or segment. The environment 1304 includes the AV 1308 and one or more objects 1320, 1324. These objects are physical entities external to the AV 1308. In other embodiments, the environment 1304 includes more or fewer components than those described herein. Similarly, the functions can be distributed among the components and / or different entities in a manner different from the manner described herein.
[0155] Server 1312 stores data accessed by AV 1308 and vehicle 1316, and performs computations used by AV 1308 and vehicle 1316. Server 1312 can be Figure 1 an example of server 136 as shown. Server 1312 is communicatively coupled to AV 1308 and vehicle 1316. In one embodiment, server 1308 can be a "cloud" server more specifically described as server 136 above with respect to Figure 1 and Figure 2 . Parts of server 1308 can be implemented in software or hardware. For example, server 1308 or parts of server 1308 can be part of any of the following: a PC, a tablet PC, an STB, a smart phone, an Internet of Things (IoT) appliance, or any machine capable of executing instructions that specify actions to be taken by that machine. In one embodiment, server 1312 stores in real time a map 1364 of environment 1304 generated by AV 1308 or other antenna or computer device, and sends map 1364 to AV 1308 for updating or annotation. Server 1312 can also send parts of map 1364 to AV 1308 or vehicle 1316 to assist AV 1308 or vehicle 1316 in navigating environment 1304.
[0156] Vehicle 1316 is a non-autonomous, semi-autonomous, or autonomous vehicle located inside or outside environment 1304. In the Figure 13 embodiment, vehicle 1316 is shown located outside environment 1304, however, vehicle 1316 can enter or leave environment 1304. There are two operating modes: mapping mode and driving mode. In mapping mode, AV 1308 drives in environment 1304 to generate map 1364. In driving mode, AV 1308 or vehicle 1316 can receive a portion of the real-time map 1364 from server 1312 for navigation assistance. In driving mode, AV 1308 uses its sensors 1336 and 1340 to determine the characteristics of environment 1304 and annotate or update that portion of map 1364. In this way, changes or dynamic characteristics of environment 1304 (such as a temporary construction area or a traffic jam) are continuously updated on the real-time map 1364. In an embodiment, vehicle 1316 can generate portions of map 1364 or send portions of map 1364 to server 1312, and optionally to AV 1308.
[0157] Objects 1320, 1324 are located within environment 1304 outside of AV 1308, as referenced above with respect to Figure 4 and Figure 51304, a parking space located on the road segment, a parking lot connected to the road segment, or an open space within the environment 1304. In an embodiment, the object 1320 is a static portion or aspect of the environment 1304, such as a road segment, a traffic light, a building, a parking space located on the road segment, a highway exit or entrance ramp, multiple lanes of a drivable area of the environment 1304 oriented in the same direction, the height of the drivable area, a curb adjacent to the drivable area, or a center line separating two lanes of the drivable area. The drivable area includes a road segment of the environment 1304, a parking space located on the road segment, a parking lot connected to the road segment, or an open space within the environment 1304. In an embodiment, the drivable area includes off-road trails and other unmarked or undifferentiated paths that the AV 1308 can traverse. Static objects have a more permanent characteristic of the environment 1304 that does not change from day to day. In the driving mode, once the features representing the static characteristics are mapped, the AV 1308 can focus on navigating and mapping features representing more dynamic characteristics, such as another vehicle 1316.
[0158] In one embodiment, the object 1320 represents a road sign or traffic sign used for vehicle navigation, such as a marking on a road segment, a road sign located within the environment, a traffic signal, or a center line separating two lanes (which will instruct the driver regarding driving direction and lane changes). In driving mode, the AV 1308 can use the semantic information of such mapped features from the map 1364 to gain context and make navigation decisions. For example, the semantic information embedded in the feature representing the center line separating two lanes will instruct the AV 1308 not to get too close to the center line close to oncoming traffic.
[0159] In another embodiment, the objects 1320, 1324 have characteristics of the environment associated with vehicle maneuvers, such as a lane split into multiple lanes, a merge of multiple lanes into a single lane, an intersection of multiple lanes, a roundabout including a spatial location for a vehicle to enter or exit the roundabout, or a bend in the lanes of the road network. In driving mode, mapping features representing such characteristics invoke navigation algorithms within the AV 1308 navigation system to guide the AV 1308 in maneuvers such as a three-point turn, a lane merge, a U-turn, a Y-turn, a K-turn, or entering a roundabout. In another embodiment, the objects 1324 have more dynamic characteristics, such as another vehicle 1316, a pedestrian, or a cyclist. Mapping features representing dynamic characteristics instruct the AV 1308 to perform collision prediction and reduce driving aggressiveness when necessary. In mapping mode, the objects 1320, 1324 are classified by the AV 1308 (e.g., grouped into types such as pedestrians, cars, etc.), and data representing the classified objects 1320, 1324 is used for map generation. Reference above Figure 6 , Figure 7 and Figure 8The physical object 608, the boundary 616 of the physical object 608, the physical object 706, the ground 802, and the object 808 describe the objects 1320 and 1324 in more detail.
[0160] The AV 1308 includes a map annotator 1328, one or more sensors 1344 and 1348, and a communication device 1332. The AV 1308 is communicatively coupled to the server 1312 and optionally coupled to one or more vehicles 1316. The AV 1308 can be Figure 1 an example of the AV 100 in. In other embodiments, the AV 1308 includes additional or fewer components compared to those described herein. Similarly, the various functions can be distributed among the components and / or different entities in a manner different from that described herein.
[0161] One or more visual sensors 1336 sense the state of the environment 1304, such as the presence and structure of the objects 1320 and 1324, and send sensor data 1344 and semantic data representing the state to the map annotator 1328. The visual sensor 1336 can be an example of the sensors 122 - 123 shown and described above. The visual sensor 1336 is communicatively coupled to the map annotator 1328 to send the sensor data 1344 and the semantic data. The visual sensor 1336 includes: one or more monocular or stereo cameras in the visible light, infrared, or thermal (or both) spectra; LiDAR; radar; ultrasonic sensors; time-of-flight (TOF) depth sensors, and can include temperature sensors, humidity sensors, or precipitation sensors. Figure 1
[0162] In one embodiment, the visual sensor 1336 generates sensor data 1344, where the sensor data 1344 includes multiple features of the environment. A feature is a part of the sensor data 1344 that represents one or more characteristics of the object 1320 or 1324. For example, a feature can model a part of the surface or a part of the structure of the object 1324. The map annotator 1328 receives the sensor data 1344 and the semantic data and generates derived values (features) from the sensor data 1344. These features are intended to be an informative and non-redundant representation of the physical or semantic characteristics of the object 1320. The annotation of the map 1364 is performed by using the reduced representation (features) instead of the complete initial sensor data 1344. For example, the LiDAR sensor of the AV 1308 is used to irradiate the target object 1320 with pulsed laser light and measure the reflected pulse 1344. Then, the differences in the laser return time and wavelength can be used to create a digital 3-D representation (feature) of the target object 1320.
[0163] In one embodiment, the vision sensor 1336 is a spatially distributed intelligent camera or LiDAR device that can process sensor data 1344 of the environment 1304 from various perspectives and fuse it into a more useful data form than individual images. For example, the sensor data 1344 includes LiDAR point cloud data reflected from the target object 1320. In another example, the sensor data includes images of multiple features of the environment 1340. The sensor data 1344 is sent to the map annotator 1328 for image processing, communication, and storage functions. As previously referenced Figure 6 , Figure 7 and Figure 8 the input 502a-d, LiDAR system 602, lights 604a-c, light emitter 606, light detector 610, field of view 614, and lights 804a-d in more detail describe the vision sensor 1336. As previously referenced Figure 6 , Figure 7 and Figure 8 the output 504a-d, image 612, and LiDAR data points 704 in more detail describe the sensor data 1344.
[0164] In one embodiment, the AV 1308 is located at a spatio-temporal position within the environment 1304, and multiple features are associated with the spatio-temporal position. The spatio-temporal position includes geographical coordinates, the time associated with the AV 1308 located at the geographical coordinates, or the forward direction (direction orientation or pose) of the AV 1308 located at the geographical coordinates. In one embodiment, the first spatio-temporal position includes GNSS coordinates, a business name, a street address, or the name of a city or town.
[0165] Sensors 1336 and 1340 receive the sensor data 1344 and 1348 and semantic data associated with the environment 1304. In one embodiment, the semantic data represents markings on the drivable area, road signs located within the environment 1304, or traffic signals located within the environment 1304. For example, images captured by the camera within the sensor 1336 can be used for image processing to extract semantic data. The AV 1308 uses the semantic data to assign meanings to the features so that these features can be used for intelligent navigation decisions. For example, markings on the drivable area that direct traffic in a specific direction prompt the AV 1308 to drive in that direction. Road signs with "yield" markings or instructions located within the environment 1304 direct the AV 1308 to yield to oncoming traffic or merge traffic. Traffic signals located within the environment 1304 direct the AV 1308 to stop at a red light and proceed at a green light.
[0166] One or more odometry sensors 1340 sense the state of the AV 1308 relative to the environment 1304 and send odometry data 1348 representative of the state of the AV 1308 to the map annotator 1328. The odometry sensors 1340 can be examples of the sensors 121 described and shown above with reference to Figure 1 the sensors 121 shown and described. The odometry sensors 1340 are communicatively coupled to the map annotator 1328 to send the odometry data 1348. The odometry sensors 1340 include one or more GNSS sensors, an IMU that measures the linear acceleration and angular velocity of the vehicle, a vehicle speed sensor for measuring or estimating the wheel slip ratio, a wheel brake pressure or brake torque sensor, an engine torque or wheel torque sensor, or a steering angle and angular velocity sensor. The IMU is an electronic device that measures and reports the specific force, angular rate, or magnetic field around the AV. The IMU uses a combination of accelerometers, gyroscopes, or magnetometers. The IMU is used to maneuver the AV. When GNSS signals are unavailable (such as in a tunnel) or there is electronic interference, the IMU allows the GNSS receiver on the AV to operate. Odometry measurements include speed, acceleration, or steering angle. The AV uses the odometry data to provide a uniquely identified signature to distinguish different spatio-temporal locations in the environment.
[0167] In one embodiment, the odometry sensors 1340 use a combination of accelerometers, gyroscopes, or magnetometers to measure and report the spatio-temporal position, specific force, angular rate, or magnetic field around the AV 1308. In another embodiment, the odometry sensors 1340 generate odometry data 1348 that includes speed, steering angle, longitudinal acceleration, or lateral acceleration. The odometry sensors 1340 utilize the raw IMU measurements to determine the attitude, angular rate, linear velocity, and position relative to a global reference frame. In one embodiment, the odometry data 1348 reported by the IMU is used to determine the attitude, speed, and position by integrating the angular rate from the gyroscope to calculate the angular position.
[0168] The map annotator 1328 receives a map 1364 of the environment 1304 from the server 1312, generates a geometric model 1360 of the features of the environment 1304, and embeds an annotated version of the geometric model into the map 1364. The map annotator 1328 includes a geometric model generator 1352 and a geometric model annotator 1356. The map annotator 1328 is communicatively coupled to the vision sensors 1336 and the odometry sensors 1340 to receive sensor data 1344, semantic data, and odometry data 1348. The map annotator 1328 is communicatively coupled to the communication device 1332 to send the annotated map 1364 to the communication device 1332.
[0169] The map annotator 1328 can be Figure 4An example of the illustrated planning module 404 may alternatively be included within the planning module 404. The map annotator 1328 may be implemented in software or hardware. For example, the map annotator 1328 or a portion of the map annotator 1328 may be part of a PC, a tablet PC, an STB, a smart phone, an IoT appliance, or any machine capable of executing instructions that specify actions to be taken by a machine. In other embodiments, the map annotator 1328 includes more or fewer components than those described herein. Similarly, the various functions may be distributed among the components and / or different entities in ways different from those described herein.
[0170] The geometry model generator 1352 generates a geometry model 1360 of a feature among a plurality of features from the sensor data 1344. The geometry model generator 1352 may be implemented in software or hardware. For example, the geometry model generator 1352 or a portion of the geometry model generator 1352 may be part of a PC, a tablet PC, an STB, a smart phone, an IoT appliance, or any machine capable of executing instructions that specify actions to be taken by a machine.
[0171] The geometry model generator 1352 extracts features from the sensor data 1344. To extract features, in one embodiment, the geometry model generator 1352 generates a plurality of pixels from the sensor data 1344. The number of generated pixels N corresponds to the number of LiDAR beams of the vision sensor 1336, e.g., the number of lights 604a-c (LiDAR beams) from the light emitter 606 of the LiDAR system 602 in Figure 6 . The geometry model generator 1352 analyzes the depth difference between a first pixel among the plurality of pixels and an adjacent pixel of the first pixel to extract features. In one embodiment, the geometry model generator 1352 renders each pixel among the plurality of pixels according to a series of heights collected for the pixels. The geometry model generator 1352 determines the visible height of the target object 1320. The geometry model generator 1352 generates a digital model of the target object 1320 from the sensor data 1344 and segments the digital model into features using a sequence filter. When generating a plurality of pixels from the sensor data 1344, the geometry model generator 1352 takes into account the position uncertainty and orientation of the AV 1308.
[0172] The geometric model generator 1352 associates features with the drivable area within the environment 1304. In one embodiment, existing information from the real-time map 1364 and sensor data 1344 are combined to associate features with the drivable area. The geometric model generator 1352 detects the drivable area based on information learned from other areas of the environment 1304 from the real-time map. According to the sensor data 1344, the geometric model generator 1352 detects the drivable area by learning the current appearance of the road from the training area. For example, a Bayesian framework can be used to combine real-time map information and sensor data 1344 to associate features with the drivable area.
[0173] The geometric model generator 1352 extracts drivable segments from the drivable area. The drivable segments are associated with the extracted features based on the spatio-temporal position of the AV 1308. In one embodiment, the geometric model generator 1352 uses cues such as color and lane markings to extract drivable segments. The geometric model generator 1352 further uses the spatial information of LiDAR points to analyze the drivable area and associate segments of the drivable area with features. In one embodiment, the geometric model generator 1352 fuses the image data of the features and the LiDAR points to perform the association. The benefit and advantage of the fusion method is that it eliminates the need for a training step or manually labeled data. In this embodiment, the geometric model generator 1352 uses superpixels as the basic processing unit to combine sparse LiDAR points with image data. A superpixel is a polygonal part of a digital image and is larger than a normal pixel. Each superpixel is rendered with a uniform color and brightness. Superpixels are dense and contain color information that cannot be captured by the LiDAR sensor, while LiDAR points contain depth information that cannot be obtained by the camera. The geometric model generator 1352 performs superpixel classification to segment the drivable area within the environment 1304 and extract the drivable parts.
[0174] In one embodiment, the geometric model generator 1352 uses the above-described odometry data 1348 to determine the spatio-temporal position of the AV 1308 relative to the boundaries of the drivable area. The position coordinates include latitude, longitude, altitude, or coordinates relative to an external object (e.g., object 1320 or vehicle 1316). In one embodiment, the geometric model generator 1352 uses the odometry data 1348 to analyze the associated sensor data 1344 to estimate the change in the position of the AV 1308 over time relative to an initial spatio-temporal position. The geometric model generator 1352 extracts image feature points from the sensor data 1344 and tracks the image feature points in the image sequence to determine the position coordinates of the AV 1308. The geometric model generator 1352 defines features of interest and performs probabilistic matching of the features across image frames to construct an optical flow field for the AV 1308. The geometric model generator 1352 checks for potential tracking errors in the flow field vectors and removes any identified outliers. The geometric model generator 1352 estimates the motion of the vision sensor 1336 based on the constructed optical flow. Then, the estimated motion is used to determine the spatio-temporal position and orientation of the AV 1308. In this way, a drivable segment containing the spatio-temporal position of the AV 1308 is extracted from the drivable area.
[0175] The geometric model generator 1352 separates the extracted drivable segment into a plurality of geometric blocks. Each geometric block corresponds to a characteristic of the drivable area. The characteristics of the drivable area refer to the positioning of the physical elements of the road, which is defined by the alignment, profile, and cross-section of the drivable area. The alignment is defined as a series of horizontal tangents and curves of the drivable area. The profile is the vertical aspect of the road, including crests and sag curves. The cross-section describes the position and the number of lanes, bike lanes, and sidewalks, as well as their slopes. The cross-section also describes the drainage characteristics or the pavement structure of the drivable area.
[0176] In one embodiment, the generation of the geometric model 1360 includes classifying the sensor data 1344 (e.g., LiDAR point cloud data). For example, height-based segmentation can be used to interpolate a surface model from the LiDAR point cloud data 1344. In one embodiment, shape fitting is used to triangulate the surface from the LiDAR point cloud data 1344, and then each geometric block is segmented using a morphological interpretation of the shape of the interpolated surface. In another embodiment, spectral segmentation is used to separate the extracted drivable segment into a plurality of geometric blocks based on the intensity values or RGB values of the LiDAR point cloud data 1344. The geometric model 1360 of the features includes a plurality of separated geometric blocks.
[0177] This disclosure relates to embodiments for generating geometric models for different environmental objects and features. The geometric model of a lane 1360 includes a baseline path. The geometric model is the smallest geometric entity on which other elements are based. In one embodiment, the geometric model represents a plurality of lanes oriented in the same direction in a drivable area. Each geometric block among the plurality of geometric blocks represents a single lane in the drivable area. The geometric model is a lane block within a non-intersecting section. The lane block has the same traffic direction within its area, such that each lane has the same source edge and destination edge on the graph of the road network (e.g., as shown in Figure 10 ). Within the lane block, the number of lanes does not change. The plurality of lanes share lane dividers within the drivable area, and the lane block is represented as the union of lane blocks, e.g., (Lane 1) U (Lane 2) U... U (Lane N). In one embodiment, the feature represents a parking lot or a parking garage associated with the lane block.
[0178] The lane connector feature represents a connection from one edge of the source lane to the other edge of the destination lane (see Figure 10 ). The area of the lane connector is based on the width of the source edge and the destination edge along the curvature of the baseline path included in the lane connector. In one embodiment, the geometric model represents a block of lane connectors within an intersection. Each block of the lane connector has the same traffic direction, i.e., in the Figure 10 directed graph, each lane connector has the same source edge and destination edge. The number of lane connectors does not change within the intersection.
[0179] In one embodiment, the feature represents a Dubins path. A Dubins path is the shortest curve connecting two points in a two-dimensional Euclidean plane (X-Y plane). The Dubins path has constraints on the curvature of the path and the specified initial and final tangents of the path.
[0180] In one embodiment, the feature represents the height of the drivable area, a curb located near the drivable area, or a centerline separating two lanes of the drivable area. For example, the feature can represent super-elevation, such as the amount by which the outer edge of a curve on a road is inclined above the inner edge. The presence of a curb can be represented by the cross-section of the road. The cross-section includes the number of lanes, their widths and cross-slopes, and the presence of shoulders, curbs, sidewalks, gutters, ditches, and other road features. The centerline or road divider feature separates one lane block from another, e.g., the centerline or road divider separates two lane blocks with different traffic directions. The lane divider feature can separate two lanes with the same traffic direction.
[0181] In one embodiment, the feature represents a walkway or sidewalk on a road segment. For example, a crosswalk feature is represented by a geometric model superimposed on the road segment. The geometric model of the pedestrian observation area represents an area where the AV 1308 may encounter pedestrians even if the pedestrian observation area is not a crosswalk or sidewalk. Based on historical observations, the pedestrian observation area is extracted from the sensor data 1344.
[0182] In one embodiment, the generation of the geometric model 1360 includes superimposing a plurality of geometric blocks on the LiDAR point cloud data 1344 to generate a polygon including the union of the geometric blocks. For example, the LiDAR point cloud data 1344 is used as a base map on which the geometric blocks are superimposed. The positions, sizes, and shapes of the geometric blocks vary. The respective attributes of the geometric blocks are edited onto the LiDAR point cloud data 1344 to obtain the polygon and visualize the spatial distribution of the union of the geometric blocks. For example, the drivable area includes road segments, parking spaces located on the road segments, parking lots connected to the road segments, or open spaces within the environment. Thus, the geometric model representing a part of the drivable area is a polygon containing one or more road segments. The drivable area is represented as the union of road segments, e.g., (road segment (non-intersection)) U (road segment (intersection)). Thus, the polygon representation is the union of all road segment polygons. In one embodiment, the drivable area includes a feature representing an avoidance area. The avoidance area is an area containing another stopped vehicle.
[0183] In one embodiment, the polygon represents an "inout edge". The inout edge or two-way edge is Figure 10 the edge of a directed graph that separates non-crossing road segments and crossing road segments. The inout edge includes one or more passing edges in one or more traffic directions. If the inout edge has more than one traffic direction, the passing edges share nodes of the graph that are part of the road separator line. The inout edge is represented as a union, such as (passing edge 1) U (passing edge 2) U... U (passing edge N). Thus, the inout edge is represented by the union of all the passing edges it includes. Each inout edge is associated with a part of the drivable area. The passing edge is the edge that separates the lane block and the road block connector. The passing edge includes one or more lane edges. If the passing edge includes more than one lane edge, the included lane edges share nodes that are part of the lane separator line. Thus, the passing edge is represented as a union, such as (lane edge 1) U (lane edge 2) U... U (lane edge N), that is, the passing edge is the union of the lane edge lines it includes. The lane edge is the edge that separates the lane and the lane connector.
[0184] In one embodiment, the polygon represents a priority area. The priority area is a portion of the drivable area that is determined to have a higher priority than the current spatio - temporal position occupied by the AV 1308. Thus, other vehicles, cyclists, or pedestrians moving or stopping in the priority area have the right of way. The priority area has multiple signs, such as cars, bicycles, pedestrians, etc. If the priority area is marked with a specific sign, the vehicle or pedestrian has a higher priority than the AV 1308.
[0185] In one embodiment, the polygon represents a stop line of the AV 1308. The stop line is coupled to a stop area where the AV 1308 needs to stop. When encountering a stop line (such as a stop sign, yield sign, crosswalk, traffic signal, or right - turn or left - turn), the AV 1308 will stop. The stop line is associated with a block of the lane (when the stop line is a stop sign, yield sign, crosswalk, or traffic signal) or a road - block connector (when the stop line is a right - turn or left - turn). Each stop line is also associated with a crosswalk or a traffic signal. Each crosswalk stop line is associated with a crosswalk. Each traffic - signal stop line is associated with a traffic signal. Each right - turn or left - turn stop line is associated with a crosswalk.
[0186] In one embodiment, the polygon represents dividing a lane of the drivable area into multiple lanes, an intersection of multiple lanes, or merging multiple lanes into a single lane. The division or merging of the lanes is associated with an unobstructed line - of - sight area to enable the safe operation of the AV 1308. When merging, basic right - of - way rules are applied (yielding to the vehicle on the right or the major - road rule, depending on the location). The AV 1308 should be able to sense approaching traffic on the intersecting road at a point where the AV 1308 can adjust its speed or stop to yield to other traffic.
[0187] In one embodiment, an intersection includes a change in the number of lanes along a road segment, a fork, a merging of lanes, or an intersection of three or more roads. The polygon representing the intersection contains one or more road - block connectors. The intersection is represented as a union, such as (road - block connector 1) U (road - block connector 2) U... U (road - block connector N). The intersection polygon can include a three - way intersection (e.g., an intersection point between three road segments), a T - intersection when two arms form a road, a Y - intersection or a fork, or a four - way intersection or a crossroads.
[0188] In one embodiment, the polygon represents a roundabout, which includes spatial positions on the drivable area for the AV 1308 to enter or leave the roundabout. The roundabout features include a circular intersection or junction where road traffic flows in one direction around a central island. Roundabout features require entering traffic to yield to traffic already in the circle and optimally comply with various design rules to enhance safety. The roundabout polygon may be associated with features indicating tram or train lines or two-way traffic flow.
[0189] In one embodiment, the polygon represents a bend in a lane of the drivable area. For example, the polygon includes one or more geometric blocks oriented in multiple directions. The geometric blocks share road dividers. Intersections are represented as unions, such as (road block connector 1) U (road block connector 2) U... U (road block connector N). The polygon representation is the union of all road blocks included in the road bend. Road bends affect the line of sight available to the AV 1308, i.e., the length of the road ahead that the AV 1308 can see before being blocked by a hilltop or an obstacle inside a horizontal bend or intersection.
[0190] The geometric model annotator 1356 generates an annotated map 1364 of the environment 1304. The geometric model annotator 1356 is communicatively coupled to the vision sensor 1336 to receive sensor data 1344 and semantic data. The geometric model annotator 1356 is communicatively coupled to the odometry sensor 1340 to receive odometry data 1348. The geometric model annotator 1356 is communicatively coupled to the communication device 1332 to send the annotated map 1364. The geometric model annotator 1356 can be implemented in software or hardware. In one embodiment, the geometric model annotator 1356 or a part of the geometric model annotator 1356 can be a part of a PC, a tablet PC, a STB, a smart phone, an IoT appliance, or any machine capable of executing instructions specifying actions to be taken by a machine.
[0191] The geometric model annotator 1356 performs semantic annotation to label pixels or parts of the geometric model 1360 using semantic data. Thus, semantic labeling is different from the geometric segmentation discussed above, which extracts features of interest and generates a geometric model of the feature. On the other hand, semantic annotation provides an understanding of the features or pixels and parts of the geometric model 1360. In one embodiment, the geometric model annotator 1356 annotates the geometric model 1360 by generating a computer-readable semantic annotation that combines the geometric model 1360 and the semantic data. To perform the annotation, the geometric model annotator 1356 associates each pixel of the geometric model 1360 with a category label (such as "yield sign", "no right turn", or "parked vehicle").
[0192] In one embodiment, the geometric model annotator 1356 extracts logical driving constraints from semantic data. The logical driving constraints are associated with navigating the AV 1308 along a drivable area. For example, the logical driving constraints include traffic signal sequences, conditional left or right turns, or traffic directions. Other examples of logical driving constraints include reducing speed when the AV 1308 approaches a speed bump, moving to an exit lane or decelerating to yield to a passing vehicle when approaching an exit of a highway. The logical driving constraints can be annotated onto the geometric model 1360 in a computer-readable script, which can be executed by the planning module 404 or the control module 406 shown and described above with reference to Figure 4 The planning module 404 or the control module 406 shown and described above with reference to
[0193] The geometric model annotator 1356 embeds the annotated geometric model 1360 into the map 1364. In one embodiment, the geometric model annotator 1356 embeds the annotated geometric model 1360 by embedding computer-readable semantic annotations corresponding to the annotated geometric model 1360 in the map 1364. The geometric model annotator 1356 embeds the semantic annotations such that they become an information source that is easily interpretable, combinable, and reusable by the AV 1308, the server 1312, or the vehicle 1316. For example, the semantic annotations can be structured digital side notes that are not visible in the human-readable portion of the map 1364. Because the semantic annotations are machine-interpretable, the semantic annotations enable the AV 1308 to perform operations such as classifying, linking, inferring, searching, or filtering the map 1364. The geometric model annotator 1356 thus enriches the map content with computer-readable information by linking the map 1364 to the extracted environmental features.
[0194] In one embodiment, the map annotator 1328 receives computer-readable semantic annotations from the second vehicle 1316. The map annotator 1328 merges the real-time map 1364 with the received computer-readable semantic annotations for transmission to the remote server 1312. For example, the map annotator 1328 processes a bitmap image containing the received computer-readable semantic annotations. Regions of interest in the bitmap image are automatically extracted based on color information, and their geometric properties are determined. The result of this process is a structured description of the features that serves as an input to the map 1364.
[0195] In one embodiment, the geometric model annotator 1356 embeds the annotated geometric model 1360 into the map 1364 in a mapping operation mode or a driving operation mode. In the driving mode, the AV 1308 receives the real-time map 1364 from the server 1312. The AV 1308 determines, based on the annotated geometric model in the map 1364, as described above with reference to Figure 11The described and shown drivable area for navigation. The geometry model annotator 1356 annotates the geometry model 1360 with semantic data for newly sensed features. The control module 406 uses the map 1364 to navigate the AV 1308 within the drivable area.
[0196] The map 1364 enables autonomous driving using semantic features such as the presence of parking spaces or bypasses. In the driving mode, the AV 1308 combines its own sensor data with the semantic annotations from the map 1364 to observe the road conditions in real time. Thus, the AV 1308 can learn about road conditions more than 100 m in advance. The annotated map 1364 provides higher accuracy and reliability than the on-vehicle sensors alone. The map 1364 provides an estimate of the spatio-temporal position of the AV 1308, and the navigation system of the AV 1308 knows the expected destination. The AV 1308 determines local navigation goals within the line of sight of the AV 1308. The sensors of the AV 1308 use the map 1364 to determine the position to the road edge and generate a path to the local navigation goal. In the driving mode, the AV 1308 extracts the annotated geometry model 1360 from the real-time map 1364. The control module 406 navigates the AV 1308 by sending commands to the throttle or brakes 1206 of the AV 1308 in response to the extraction of the annotated geometry model 1360 (as described in detail above with reference to Figure 12 and shown).
[0197] In one embodiment, the AV 1308 stores a plurality of operation metrics associated with navigating the AV 1308 along the drivable area. As described above with reference to Figure 3As described and illustrated, the operational metrics are stored in the main memory 306, ROM 308, or storage device 310. The operational metrics associated with a driving segment may represent: multiple predicted collisions of the AV 1308 with the object 1320 when driving along the driving segment; multiple predicted stops of the AV 1308 when driving along the driving segment; or the maximum lateral clearance between the AV 1308 and the object 1320 when driving along the driving segment. The map annotator 1328 updates the multiple operational metrics using the map 1364. The planning module 404 uses the updated multiple operational metrics to determine rerouting information for the AV 1308. For example, the determination of the rerouting information includes selecting a driving segment associated with a number of predicted collisions less than a threshold (e.g., 1), selecting a driving segment associated with a number of predicted stops less than a threshold, or selecting a driving segment associated with a predicted lateral clearance between the AV 1308 and the object 1320 greater than a threshold. The communication device 1332 sends the rerouting information to the server 1312 or another vehicle 1316 located within the environment 1304. The control module 406 of the AV 1308 uses the determined rerouting information to navigate the AV 1308 over the drivable area.
[0198] The map annotator 1328 sends the map 1364 with computer-readable semantic annotations via the communication device 1332. The communication device 1332 also exchanges data with the server 1312, the passengers within the AV 1308, or other vehicles, such as instructions or measured or inferred attributes of the status and conditions of the AV 1308 or other vehicles. The communication device 1332 may be Figure 1 an example of the communication device 140 shown. The communication device 1332 is communicatively coupled to the server 1312 across a network and optionally coupled to the vehicle 1316. In an embodiment, the communication device 1332 communicates across the Internet, the electromagnetic spectrum (including radio communication and optical communication), or other media (e.g., air and acoustic media). Portions of the communication device 1332 may be implemented in software or hardware. For example, the communication device 1332 or a portion of the communication device 1332 may be a PC, a tablet PC, a STB, a smart phone, an IoT appliance, or a part of any machine capable of executing instructions that specify actions to be taken by a machine. The communication device 1332 was described in more detail above with reference to Figure 1 the communication device 140 in
[0199] The benefits and advantages of the embodiments disclosed herein are that, compared to traditional methods of directly scanning the point cloud space, navigating an AV using an annotated map is more accurate and computationally cheaper. The AV can effectively determine the positioning for real-time navigation. Navigating the AV using a real-time annotated map can improve the safety of passengers and pedestrians, reduce the wear and tear of the AV, reduce travel time, shorten the travel distance, etc. Increased safety is also obtained for other vehicles on the road network.
[0200] Fusing image data and LiDAR points to associate features with drivable areas eliminates the need for a training step or manual labeling of data. Additionally, annotating the geometric model with semantic data reduces errors that may result from sensor noise and the AV's estimation of its orientation and position. Other benefits and advantages of using a map annotator are that the AV can correctly distinguish between two or more plausible locations in the real-time map. Thus, the annotated features will be embedded at the correct spatio-temporal location in the map.
[0201] Example environment of automatic annotation of environmental features in a map
[0202] Figure 14 FIG. illustrates an example environment 1400 of automatic annotation of environmental features in a map 1364 during the navigation of an AV 1308 according to one or more embodiments. Environment 1400 includes a spatio-temporal location 1404 where the AV 1308 is located. Environment 1400 includes a road segment 1408, a curb 1412, a construction area 1416, a road marking 1420 (direction arrow), lane markings 1424 representing the boundaries of the lane in which the AV 1308 is located, and parked vehicles 1428 outside the lane.
[0203] The vision sensor 1336 of the AV 1308 generates sensor data 1344 representing the environment 1400 at the spatio-temporal location 1404. For example, the sensor data 1344 generated in Figure 14 describes the shape, structure, and position of the curb 1412, the construction area 1416, the road marking 1420, the lane markings 1424, or the parked vehicles 1428. The vision sensor 1336 includes one or more monocular or stereo cameras, infrared or thermal (or both) sensors, LiDAR 1436, radar, ultrasonic sensors, or time-of-flight (TOF) depth sensors, and may include temperature sensors, humidity sensors, or precipitation sensors. In one embodiment, the sensor data 1344 includes camera images, three-dimensional LiDAR point cloud data, or LiDAR point cloud data indexed by time.
[0204] The odometry sensor 1340 generates odometry data 1348 representing the operating state of the AV 1308. The odometry sensor 1340 of the AV 1308 includes one or more GNSS sensors, an IMU that measures the linear acceleration and angular velocity of the vehicle, a vehicle speed sensor for measuring or estimating the wheel slip ratio, a wheel brake pressure or brake torque sensor, an engine torque or wheel torque sensor, or a steering angle and angular velocity sensor. The odometry data 1348 is associated with the spatio-temporal position 1404 and the road segment 1408. In one embodiment, the map annotator 1328 uses the odometry data 1348 to determine the position 1404 and orientation of the AV 1308. The AV 1308 sends the determined position coordinates and orientation to the server 1312 or another vehicle 1316 to provide navigation assistance to other vehicles 1316.
[0205] The map annotator 1328 annotates the map 1364 of the environment 1400 using the sensor data 1344. To annotate the map 1364, the map annotator 1328 generates a geometric model 1360 of the environmental features from the sensor data 1344. For example, the environmental features may represent a construction area 1416 or a portion of a road marking 1420. In one embodiment, the extraction of the environmental features includes generating a plurality of pixels, where the number of the plurality of pixels corresponds to the number of LiDAR beams of the LiDAR 1436. The AV 1308 analyzes the depth difference between a first pixel among the plurality of pixels and an adjacent pixel of the first pixel to extract the environmental features.
[0206] To generate the geometric model 1360, the map annotator 1328 associates the features with the drivable area within the environment 1400. The map annotator 1328 extracts drivable segments from the drivable area, such as 1408. The map annotator 1328 divides the drivable segments into a plurality of geometric blocks, where each geometric block corresponds to the characteristics of the drivable area. The geometric model 1360 of the features includes the plurality of geometric blocks. In one embodiment, the map annotator 1328 generates the geometric model 1360 by superimposing the plurality of geometric blocks on the sensor data 1344 such as LiDAR point cloud data to generate a polygon including the union of the geometric blocks. The polygon models the features of the environment 1400, such as a roundabout or an intersection. In one embodiment, the map annotator 1328 generates the geometric model 1360 by mapping three-dimensional LiDAR point cloud data to a two-dimensional image. The map annotator 1328 uses an edge operator for edge detection to extract edge pixels representing the shape of the target construction area 1416 (object) from the two-dimensional image. The extracted edge pixels are assembled into the contour of the construction area 1416 (object) using a transformation to extract the polygon corresponding to the environmental features. The map annotator 1328 annotates the geometric model 1360 using semantic data and embeds the annotated geometric model 1360 into the map 1364.
[0207] In one embodiment, the map annotator 1328 associates environmental features with the odometry data 1348. This reduces the error caused by sensor noise during the estimation of the orientation and position 1404 of the AV 1308. The AV 1308 resolves such errors by using the ranging data 1348 in a closed loop. For example, the AV 1308 identifies when two different spatio-temporal positions (e.g., 1404 and 1432) are incorrectly associated and removes the relationship between the two different spatio-temporal positions 1404 and 1432. The map annotator 1328 embeds computer-readable semantic annotations corresponding to the geometric model 1360 within the real-time map 1364. The environmental features are thus represented by geometric shapes. For example, the shape can represent a construction area 1416.
[0208] In one embodiment, when the AV 1308 is at a second spatio-temporal position 1432 within the environment 1400, the vision sensor 1336 generates second sensor data representing the environment 1400. The AV 1308 associates the second sensor data with the second spatio-temporal position 1432 such that the second sensor data is distinct from the sensor data 1344. The AV 1308 thus generates uniquely identified signatures for all nodes (spatio-temporal positions) in the environment 1400. In one embodiment, the map annotator 1328 receives LiDAR point cloud data and generates a base map based on the LiDAR point cloud data. The map may have inaccuracies and errors in the form of misidentified loop closures. These errors are resolved by using the odometry measurement data 1348 in a closed loop to associate the second sensor data with the second spatio-temporal position 1432, e.g., identifying when two nodes (e.g., 1404 and 1432) are incorrectly related and removing the relationship between the two nodes.
[0209] Example automatically annotated real-time map for AV navigation
[0210] Figure 15 FIG. illustrates an example automatically annotated geometric model of environmental features in the real-time map 1500 during the navigation of the AV 1308 according to one or more embodiments. The map 1500 corresponds to the environment 1400 described and illustrated above. In map mode, the geometric model annotator 1356 embeds Figure 14 the annotated geometric model 1360 of the lane marking feature 1424 within, as computer-readable semantic annotations 1504 within the map 1500. The semantic annotations 1504 are readily understandable by Figure 14 and Figure 13The information sources that can be interpreted, combined, and reused by the AV 1308, the server 1312, or the vehicle 1316. In one embodiment, the semantic annotation 1504 is a machine-interpretable color-structured digital sidenote. The color or texture of the annotation 1504 enables the AV 1308 to perform operations such as classifying, linking, inferring, searching, or filtering the map 1500. In the driving mode, the color of the feature 1504 can indicate to the AV 1308 that the feature 1504 is part of the lane marking 1424. This enables the rule of staying within the lane to navigate the AV 1308.
[0211] The geometric model annotator 1356 thus enriches the map content with computer-readable information by linking the map 1500 to the extracted environmental features. For example, the geometric model generator 1352 associates the feature with the drivable area within the environment 1400. The feature can be the height or the bend in the road segment 1408. The geometric model generator 1352 extracts the drivable segments 1408 from the drivable area. The geometric model generator 1352 separates the drivable segment 1408 into multiple geometric blocks, such as 1508, 1524, etc. Each geometric block corresponds to the characteristics of the drivable area. For example, the geometric block can correspond to a part of the surface of the object 1324, a part of the structure, or a semantic description. Each geometric block is an informative and non-redundant representation of the physical or semantic characteristics of the drivable area. The geometric model of the feature includes multiple geometric blocks. The lanes oriented in the same direction of the road segment 1408 can have the same annotation and color as the geometric block 1508. The direction of the lane is represented by the semantic annotation 1532, which is Figure 14 a representation of the road marking 1420 in
[0212] In one embodiment, the geometric blocks are superimposed on the sensor data 1344 to generate polygons. For example, polygons such as the geometric shape 1512 are generated. The geometric shape 1512 represents a parked vehicle 1428. In the driving mode, the color or other annotation associated with the geometric shape 1512 indicates to the AV 1308 that the position of the parked vehicle 1428 is a parking space. Each time the AV 1308 subsequently accesses the road segment 1408 in the mapping mode, the AV 1308 searches for changes in the landscape and updates the map 1500. For example, the AV 1308 passes through the environment 1400 multiple times in the mapping mode, and in some cases may not observe the parked vehicle 1428. This is interpreted by the navigation system of the AV 1308 or the server 1312 as meaning that the position of the parked vehicle 1428 may be occupied by other vehicles, or not occupied at other times.
[0213] The map annotator 1328 embeds semantic metadata, markers, and custom annotations, such as 1516, in the map 1500. In the driving mode, the AV 1308 can use the annotation 1516 to determine the relationships between different spatio-temporal locations. For example, the AV 1308 determines that the annotation 1516 represents Figure 2 the curb 1412 in. Thus, the AV 1308 should drive in the drivable area indicated by the geometric block 1508 and should not drive in the spatio-temporal location 1520 on the other side of the curb.
[0214] In one embodiment, each location or place in the map 1500 is represented by latitude and longitude coordinates. In the driving mode, the AV 1308 can use a search-based method to determine the distance or path between two locations (such as Figure 14 1404 and 1432 in). By using semantic annotations to annotate each spatio-temporal location, traffic or driving information is added to and retrieved from the map 1500. For example, the color of the semantic annotation 1524 indicates that the AV 1308 should slow down or otherwise exercise caution, such as increasing the lateral distance from the construction area 1416 represented by the geometric shape 1528.
[0215] In some embodiments, one or more sensors 1336 of the AV 1308 are used to generate sensor data 1344 including the features of the environment 1304. For example, the mapping mode can be used to collect data (such as LIDAR point data) by the sensors 1336 on the AV 1308 for extracting the geometric profiles of the features of interest. The sensor data 1344 is used to map the features into the drivable area within the map 1500 of the environment 1304. Rules can be used to identify each feature of interest based on its geometric profile. The drivable area within the map 1500 is used to extract polygons including a plurality of geometric blocks. Each geometric block (e.g., 1508) corresponds to a drivable segment of the drivable area. Thus, the geometric profiles of the features of interest (such as boundaries, parked vehicles, traffic lights, road signs, etc.) can be extracted. The semantic data extracted from the drivable area within the map 1500 is used to annotate the polygons. The extracted features (such as the identified parking spaces) can be automatically identified based on the geometric profiles of the blocks and segments (polygon representation). Therefore, annotating the map 1500 is represented based on the identified features. One or more processors 146 of the AV 1308 are used to embed the annotated polygons in the map 1500.
[0216] In some embodiments, the extraction of polygons includes overlaying a plurality of geometric blocks onto sensor data 1344. A polygon representation is generated for each feature. The polygon representation is a union of line segments and / or blocks. A union of the overlaid plurality of geometric blocks is generated. Clusters of LIDAR point data are thereby classified as features (e.g., parked vehicles, intersections, etc.). In some embodiments, the drivable area within the map is used to extract semantic data. The semantic data includes logical driving constraints associated with navigating the AV 1308 along the drivable area. For example, an autonomous (mission) mode can be used by the AV 1308 to extract semantic annotations from the map 1500 representation to identify the position and pose of the AV 1308 relative to the map 1500 representation and to determine the drivable area.
[0217] In some embodiments, embedding an annotated polygon within the map 1500 includes generating a computer-readable semantic annotation from the annotated polygon. The computer-readable semantic annotation is inserted into the map 1500. In some embodiments, the AV 1308 receives a second annotated polygon from a second vehicle 1316. From the second annotated polygon, a second computer-readable semantic annotation is generated. The second computer-readable semantic annotation is embedded into the map 1500. When docking the AV 1308 for charging, the annotations and the partially updated map can be uploaded from the AV 1500 to the server 1312 over the network. The server 1312 can merge the updated real-time annotations from several different AVs into a real-time map representation.
[0218] In some embodiments, the sensor data 1344 includes three-dimensional LiDAR point cloud data. In some embodiments, one or more sensors 1336 include cameras, and the sensor data 1344 also includes images of features. In some embodiments, the polygon represents a plurality of lanes within the drivable area that are oriented in the same direction. Each of the plurality of geometric blocks represents a single lane of the drivable area. In some embodiments, the polygon represents the height of the drivable area, a curb located near the drivable area, or a centerline separating two lanes of the drivable area. The drivable area can include a road, a parking space located on the road, a parking lot or an open space connected to the road.
[0219] In some embodiments, sensor data 1344 is used to determine the spatial position of the AV 1308 relative to the boundaries of the drivable area. One or more sensors 1336 include a Global Navigation Satellite System (GNSS) sensor or an Inertial Measurement Unit (IMU). The determined spatial position of the AV 1308 relative to the boundaries of the drivable area is used to operate the AV 1308 within the drivable area. In one embodiment, a polygon representation divides the lanes of the drivable area into multiple lanes. In some embodiments, a polygon representation merges multiple lanes of the drivable area into a single lane. In some embodiments, a polygon representation represents the intersection of multiple lanes of the drivable area. In some embodiments, a polygon representation represents a roundabout, which includes spatial positions on the drivable area for the AV 1308 to enter or leave the roundabout.
[0220] In some embodiments, a polygon representation represents the curvature of the lanes of the drivable area. The control circuit 406 of the AV 1308 is used to operate the AV 1308 on the drivable area in an operating mode. The embedding of the annotated polygon is performed in a mapping mode. In some embodiments, semantic data represents markings on the drivable area, road signs located within the environment, or traffic signals located within the environment. In some embodiments, logical driving constraints include traffic signal sequences, conditional left or right turns, or traffic directions. In some embodiments, an annotated polygon is extracted from the map 1500. In response to the extraction of the annotated polygon, commands are sent to the throttle 420b or the brake 420c of the AV 1308.
[0221] Example Annotated Map
[0222] Figure 16 The figure shows an example automatically annotated map 1600 during the navigation of the AV 1308 according to one or more embodiments. The map 1600 represents the drivable area 1604 and the non-drivable area 1608 mapped and annotated by the AV 1308. The generation of a geometric model such as the geometry 1616 includes superimposing multiple geometric blocks onto the LiDAR point cloud data to generate a polygon including the union of the geometric blocks. The respective attributes of the multiple geometric blocks are edited onto the LiDAR point cloud data 1344 to obtain the polygon and visualize the spatial distribution of the union of the geometric blocks. The drivable area 1604 includes road segments, left turns 1612, intersections 1620, and curbs 1624.
[0223] The polygon 1612 represents the stop line of the AV 1308. The stop line is coupled to a stop area where the AV 1308 needs to stop. The AV 1308 retrieves the polygon 1612 from the map 1600 and identifies that it is encountering a left turn. The stop line is associated with Figure 13 the road block connectors shown. The left turn stop line is associated with a crosswalk.
[0224] The geometric shape 1616 represents multiple lanes oriented in the same direction within the drivable area 1604. Each of the multiple geometric blocks represents a single lane of the drivable area. The geometric shape 1616 is thus a lane block within a non-intersecting section. The lane block has the same traffic direction within its area, such that each lane has the same source edge and destination edge on the graph of the road network (e.g., as shown in Figure 10 ). Within the lane block, the number of lanes does not change. The multiple lanes share lane dividers within the drivable area, and the lane block is represented as the union of lane blocks, e.g., (Lane 1) U (Lane 2) U... U (Lane N). The centerline or road divider feature separates one lane block from another. The centerline or road divider separates lane blocks with different traffic directions. The lane divider feature may separate two lanes with the same traffic direction.
[0225] The intersection 1620 (polygon) represents the intersection of multiple lanes of the drivable area. This intersection 1620 is associated with a sight triangle without obstacles for the safe operation of the AV 1308. At the intersection 1620, basic right-of-way rules (yield to the vehicle on the right or the boulevard rule, depending on the location) will be applied. The AV 1308 should be able to sense approaching traffic on the intersecting road at a point where the AV 1308 can adjust its speed or stop to yield to other traffic. The intersection 1620 is defined as the intersection of three or more roads. The intersection 1620 (polygon) consists of one or more road block connectors. The intersection 1620 is represented as a union, such as (Road block connector 1) U (Road block connector 2) U... U (Road block connector N).
[0226] In one embodiment, the feature represents the height of the drivable area 1604, the curb 1624 located near the drivable area 1604, or the centerline separating two lanes of the drivable area 1604. The curb 1624 is represented by the cross-section of the road. The cross-section includes the number of lanes, their widths and cross slopes, and whether there are shoulders, curbs 1624, sidewalks, gutters, ditches, and other road features.
[0227] Process for automatic annotation of maps
[0228] Figure 17 FIG. illustrates a process 1700 for automatic annotation of environmental features in a map 1364 during the navigation of the AV 1308 according to one or more embodiments. In one embodiment, Figure 17The process can be performed by one or more components of the AV 1308, such as the map annotator 1328. In other embodiments, other entities, such as the remote server 1312, perform some or all of the steps of process 1700. Similarly, embodiments can include different and / or additional steps, or perform steps in a different order.
[0229] The AV 1308 uses one or more processors of the AV 1308 located within the environment 1304 to receive 1704 the map 1364 of the environment 1364. The environment 1304 represents a geographic area, such as a state, town, neighborhood, or road network or segment. The environment 1304 includes the AV 1308 and the objects 1320, 1324. These objects are physical entities external to the AV 1308.
[0230] The AV 1308 uses one or more sensors of the AV 1308 to receive 1708 sensor data 1344 and semantic data. The sensor data 1344 includes multiple features of the environment 1304. The sensors include one or more vision sensors 1336, such as monocular or stereo cameras in the visible light, infrared, or thermal (or both) spectra; LiDAR 1436; radar; ultrasonic sensors, or time-of-flight (TOF) depth sensors. In one embodiment, the sensor data 1344 includes camera images, three-dimensional LiDAR point cloud data, or LiDAR point cloud data indexed by time. The sensors also include an odometry sensor 1340 to generate odometry data 1348 representing the operating state of the AV 1308. The odometry sensor 1340 of the AV 1308 includes one or more GNSS sensors, an IMU that measures the linear acceleration and angular velocity of the vehicle, a vehicle speed sensor for measuring or estimating the wheel slip rate, a wheel brake pressure or brake torque sensor, an engine torque or wheel torque sensor, or a steering angle and angular velocity sensor.
[0231] The AV 1308 generates 1712 a geometric model of the features from the sensor data 1344 by associating the features with the drivable area 1604 in the environment 1304. In one embodiment, the AV 1308 uses statistical methods to associate the features (structures) with the drivable area 1604 based on the 2D structure of the sensor data 1344. The RGB pixel values of the features are extracted from the sensor data 1344 and combined using a statistical model to obtain a higher-order architecture. For example, based on the histogram of the road network, the texture of the drivable area is determined by minimizing the variance of small regions within the sensor data 1344.
[0232] The AV 1308 extracts drivable segments 1716 from the drivable area 1604. The drivable segments are correlated with the extracted features based on the spatio-temporal position of the AV 1308. In one embodiment, the AV 1308 uses cues such as color and lane markings to extract the drivable segments. The AV 1308 further uses the spatial information of the LiDAR points to analyze the drivable area 1604 and correlates the segments of the drivable area with features. In one embodiment, the AV 1308 fuses the image data of the features and the LiDAR points to perform the correlation.
[0233] The AV 1308 segments 1720 the drivable segments into multiple geometric blocks. Each geometric block corresponds to a characteristic of the drivable area 1604. The characteristics of the drivable area 1604 refer to the positioning of the physical elements of the road, which is defined by the alignment, profile, and cross-section of the drivable area. The alignment is defined as a series of horizontal tangents and curves of the drivable area 160. The profile is the vertical aspect of the road, including crests and sag curves. The cross-section illustrates the location and the number of lanes, bike lanes, or sidewalks. The geometric model 1360 of the features includes multiple geometric blocks.
[0234] The AV 1308 annotates the geometric model 1360 using semantic data. The AV 1308 performs semantic annotation to label pixels or parts of the geometric model 1360 using semantic data. The semantic annotation provides an understanding of the features or pixels and parts of the geometric model 1360. In one embodiment, the AV 1308 annotates the geometric model 1360 by generating a computer-readable semantic annotation that combines the geometric model 1360 and the semantic data. To perform the annotation, the AV 1308 associates each pixel of the geometric model 1360 with a class label (such as "yield sign", "no right turn", or "parked vehicle").
[0235] The AV 1308 embeds 1728 the annotated geometric model 1360 in the real-time map 1364. The AV 1356 embeds the semantic annotation such that it becomes an information source that is easy to interpret, combine, and reuse by the AV 1308, the server 1312, or the vehicle 1316. For example, the semantic annotation can be a structured digital side note that is not visible in the human-readable part of the map 1364. Since the semantic annotation is machine-interpretable, the semantic annotation enables the AV 1308 to perform operations such as classifying, linking, inferring, searching, or filtering the map 1364. The AV 1308 thus enriches the map content with computer-readable information by linking the map 1364 to the extracted environmental features.
[0236] Process for automatic annotation of maps
[0237] Figure 18The figure shows a process 1800 for automatically annotating environmental features in a map 1364 during navigation of an AV 1308 according to one or more embodiments. In one embodiment, Figure 18 the process may be performed by one or more components of the AV 1308. In other embodiments, other entities such as the remote server 1312 perform some or all of the steps of the process 1800. Similarly, embodiments may include different and / or additional steps, or perform steps in a different order.
[0238] The AV 1308 uses one or more sensors 1336 to generate 1804 sensor data 1344 that includes features of the environment 1304. The sensors include one or more vision sensors 1336, such as monocular or stereo cameras in the visible light, infrared, or thermal (or both) spectra; LiDAR 1436; radar; ultrasonic sensors, or time-of-flight (TOF) depth sensors. In one embodiment, the sensor data 1344 includes camera images, three-dimensional LiDAR point cloud data, or LiDAR point cloud data indexed by time. The environment 1304 represents a geographic area, such as a state, town, neighborhood, or road network or segment. The environment 1304 includes the AV 1308 as well as objects 1320, 1324. These objects are physical entities external to the AV 1308.
[0239] The AV 1308 uses the sensor data 1344 to map 1808 features into a drivable area within a map 1500 of the environment 1304. The drivable area may include segments of the environment 1304, parking spaces located on the segments, parking lots connected to the segments, or open spaces within the environment 1304. In an embodiment, the drivable area includes off-road trails and other unmarked or undifferentiated paths that the AV 1308 can traverse.
[0240] The AV 1308 uses the drivable area within the map 1500 to extract 1812 polygons that include a plurality of geometric blocks. Each geometric block corresponds to a drivable segment of the drivable area. The respective attributes of the geometric blocks can be edited onto the LiDAR point cloud data 1344 to derive the polygons and visualize the spatial distribution of the union of the geometric blocks. For example, the drivable area may include segments of a road, parking spaces located on the segments, parking lots connected to the segments, or open spaces within the environment. Thus, a geometric model representing a portion of the drivable area is a polygon that includes one or more segments of a road.
[0241] AV 1308 annotates polygons with semantic data extracted from drivable areas within map 1500. AV 1308 performs semantic annotation to label pixels or portions of the polygons with semantic data. The semantic annotation provides an understanding of the pixels and portions of the polygons. In one embodiment, AV 1308 annotates a polygon by generating a computer-readable semantic annotation that combines the polygon and the semantic data. To perform the annotation, AV 1308 associates each pixel of the polygon with a category label such as "yield sign", "no right turn", or "parked vehicle".
[0242] AV 1308 embeds 1828 the annotated polygons into map 1500 using one or more processors 146 of AV 1308. AV 1308 embeds the annotated polygons such that they become information sources that are easily interpretable, combinable, and reusable by AV 1308, server 1312, or vehicle 1316. For example, the annotated polygons can be structured digital side notes that are not visible in the human-readable portion of map 1500. Because the annotated polygons are machine-interpretable, the annotated polygons enable AV 1308 to perform operations such as classifying, linking, inferring, searching, or filtering map 1500. AV 1308 thus enriches the map content with computer-readable information by linking map 1500 to the extracted environmental features.
[0243] In the foregoing description, embodiments of the present invention have been described with reference to numerous specific details, which may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The sole and exclusive indicator of the scope of the present invention, which the applicant desires the scope of the present invention to cover, is the literal and equivalent scope of the claims issued from this application in the specific form in which such claims are issued, including any subsequent corrections. Any definition of terms expressly set forth herein for inclusion in such claims shall be construed to have the meaning such terms have when used in the claims. Additionally, when we use the phrase "further comprising" in the foregoing specification or the appended claims, the remainder of that phrase can be additional steps or entities, or sub-steps / sub-entities of the previously recited steps or entities.
Claims
1. A method for a vehicle, comprising: Receiving, by one or more processors of the vehicle located in the environment, a map of the environment; Receiving, by one or more sensors of the vehicle, sensor data, wherein the sensor data includes three-dimensional data, i.e., 3D data; Extracting semantic data and a plurality of features of the environment from the sensor data; Generating, by the one or more processors, a geometric model of a feature among the plurality of features based on the sensor data, the generating including: Associating, by the one or more processors, the feature with a drivable area within the environment; Extracting, by the one or more processors, at least one drivable segment from the drivable area; Dividing, by the one or more processors, the at least one drivable segment into a plurality of geometric blocks corresponding to characteristics of the drivable area based on an intensity value or a color model of the 3D data; and Overlaying, by the one or more processors, the plurality of geometric blocks on the 3D data to generate a polygon including a union of geometric blocks among the plurality of geometric blocks; Annotating, by the one or more processors, the geometric model using the semantic data and the polygon including the union of geometric blocks among the plurality of geometric blocks; and Embedding, by the one or more processors, the annotated geometric model in the map.
2. The method according to claim 1, wherein the vehicle is located at a spatio-temporal position within the environment, and the plurality of features are associated with the spatio-temporal position.
3. The method according to claim 1, wherein the feature represents a plurality of lanes of the drivable area, each lane among the plurality of lanes is oriented in the same direction, and each geometric block among the plurality of geometric blocks represents a single lane among the plurality of lanes.
4. The method according to claim 1, wherein the feature represents at least one of a height of the drivable area, a curb located near the drivable area, and a center line separating two lanes of the drivable area.
5. The method according to claim 1, wherein the drivable area includes at least one of: a road segment, a parking space located on the road segment, a parking lot connected to the road segment, and an open space within the environment.
6. The method according to claim 1, further comprising: Determining, using the sensor data, a spatial position of the vehicle relative to a boundary of the drivable area, the one or more sensors including a global navigation satellite system sensor, i.e., a GNSS sensor, or an inertial measurement unit, i.e., an IMU.
7. The method according to claim 1, wherein the polygon represents dividing the lanes of the drivable area into a plurality of lanes.
8. The method according to claim 1, wherein the polygon represents merging the plurality of lanes of the drivable area into a single lane.
9. The method according to claim 1, wherein the polygon represents an intersection of the plurality of lanes of the drivable area.
10. The method according to claim 1, wherein the polygon represents a roundabout, and the roundabout includes spatial positions on the drivable area for the vehicle to enter or leave the roundabout.
11. The method according to claim 1, wherein the polygon represents a bend in a lane of the drivable area.
12. The method according to claim 1, wherein annotating the geometric model includes generating a computer-readable semantic annotation that combines the geometric model with the semantic data, and the method further includes: Send the map with the computer-readable semantic annotations to a remote server or another vehicle.
13. The method according to claim 1, wherein annotating the geometric model using the semantic data is performed in a first operating mode, and the method further comprises: Use the control module of the vehicle to navigate the vehicle on the drivable area using the map in a second operating mode.
14. The method according to claim 1, wherein the semantic data represents at least one of markers on the drivable area, road signs located within the environment, and traffic signals located within the environment.
15. The method according to claim 1, wherein annotating the geometric model using the semantic data comprises: Extract logical driving constraints associated with navigating the vehicle within the drivable area from the semantic data.
16. The method according to claim 1, wherein the 3D data is light detection and ranging point cloud data, i.e., LiDAR point cloud data.
17. The method according to claim 1, wherein the color model includes red, green, and blue values, i.e., RGB values.
18. A vehicle, comprising: One or more computer processors; And One or more non-transitory storage media storing instructions that, when executed by the one or more computer processors, cause the one or more computer processors to: Receive a map of the environment in which the vehicle is located; Receive sensor data through one or more sensors of the vehicle, wherein the sensor data includes three-dimensional data, i.e., 3D data; Extract semantic data and multiple features of the environment from the sensor data; Generate a geometric model of a feature among the multiple features according to the sensor data, and the generation includes: Associate the feature with a drivable area in the environment; Extract at least one drivable segment from the drivable area; Based on the intensity value or color model of the 3D data, divide the at least one drivable segment into multiple geometric blocks corresponding to the characteristics of the drivable area; and Overlay the multiple geometric blocks on the 3D data to generate a polygon including the union of the geometric blocks among the multiple geometric blocks; Annotate the geometric model using the semantic data; and Embed the annotated geometric model in the map.
19. The vehicle according to claim 18, wherein the 3D data is light detection and ranging point cloud data, i.e., LiDAR point cloud data.
20. The vehicle according to claim 18, wherein the color model includes red, green, and blue values, i.e., RGB values.
21. One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to: Receive a map of the environment in which the vehicle is located; Receive sensor data through one or more sensors of the vehicle, wherein the sensor data includes three-dimensional data, i.e., 3D data; Extract semantic data and multiple features of the environment from the sensor data; Generate a geometric model of a feature among the multiple features based on the sensor data, the generating including: Associate the feature with a drivable area in the environment; Extract at least one drivable segment from the drivable area; Based on an intensity value or a color model of the 3D data, divide the at least one drivable segment into multiple geometric blocks corresponding to characteristics of the drivable area; and Superimpose the multiple geometric blocks on the 3D data to generate a polygon including a union of the geometric blocks among the multiple geometric blocks; Annotate the geometric model using the semantic data; and Embed the annotated geometric model in the map.
22. The one or more non-transitory computer-readable storage media according to claim 21, wherein the 3D data is light detection and ranging point cloud data, i.e., LiDAR point cloud data.
23. The one or more non-transitory computer-readable storage media according to claim 21, wherein the color model includes red, green, and blue values, i.e., RGB values.
24. A computer program product comprising computer program instructions that, when run by one or more processors, implement the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Method for updating digital maps
CN102197419A
Method, device and equipment used for marking map
CN108694882A