Systems and methods for navigating at a safe distance
The system uses camera-based image analysis and sensor data to ensure safe autonomous navigation by determining braking and acceleration rates, addressing safety and scalability challenges in autonomous vehicles.
Patent Information
- Application Number
- JP2024220846
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-11
- Filing Date
- 2024-12-17
- Publication Date
- 2026-01-28
- Estimated Expiration
- 2039-08-14
AI Technical Summary
Autonomous vehicles face challenges in navigating safely and efficiently while adhering to safety assurance requirements and scalability constraints, necessitating interpretable mathematical models for navigation systems that can handle various environmental factors and interactions with other objects.
The system utilizes cameras and processing devices to analyze images, combined with GPS and sensor data, to determine navigation operations, including braking rates, acceleration capacities, and stopping distances, ensuring safe navigation by considering the vehicle's and target's positions and movements.
The system enables safe and efficient navigation by accurately assessing potential collisions and obstacles, adhering to safety constraints, and scaling to millions of vehicles, thus enhancing the viability of autonomous driving technology.
Smart Images

Figure 0007808176000296 
Figure 0007808176000297 
Figure 0007808176000298
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 718,554, filed August 14, 2018, U.S. Provisional Patent Application No. 62 / 724,355, filed August 29, 2018, U.S. Provisional Patent Application No. 62 / 772,366, filed November 28, 2018, and U.S. Provisional Patent Application No. 62 / 777,914, filed December 11, 2018. All of the above applications are incorporated herein by reference in their entirety.
[0002] The present disclosure relates generally to autonomous vehicle navigation. Additionally, the present disclosure relates to systems and methods for navigating according to potential accident liability constraints. [Background technology]
[0003] As technology continues to evolve, the goal of fully autonomous vehicles capable of navigating roads becomes more realistic. An autonomous vehicle may need to consider various factors and, based on those factors, make appropriate decisions to safely and accurately reach its intended destination. For example, an autonomous vehicle may need to process and interpret visual information (e.g., information captured from a camera), information from radar or lidar, and may also use information obtained from other sources (e.g., a GPS device, a speed sensor, an accelerometer, a suspension sensor, etc.). At the same time, to navigate to its destination, an autonomous vehicle may also need to identify its position within a particular road (e.g., a particular lane within a multi-lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, observe traffic signals and signs, navigate from one road to another at appropriate intersections or interchanges, and respond to any other conditions that occur or develop during the vehicle's operation. Furthermore, the navigation system may have to comply with certain imposed constraints. In some cases, those constraints may relate to interactions between the host vehicle and one or more other objects, such as other vehicles or pedestrians. In other cases, the constraints may relate to responsibility rules to be followed when performing one or more navigation operations for the host vehicle.
[0004] In the field of autonomous driving, there are two important considerations for a viable autonomous vehicle system. The first consideration is the standardization of safety assurances, including requirements that all self-driving vehicles must meet to ensure safety, and how those requirements can be verified. The second consideration is scalability, because engineering solutions that cause rising costs may not scale to millions of vehicles and may hinder widespread or even moderate adoption of autonomous vehicles. Therefore, there is a need for interpretable mathematical models for safety assurances and the design of systems that can scale to millions of vehicles while still adhering to the safety assurance requirements. Summary of the Invention
[0005] Embodiments according to the present disclosure provide systems and methods for autonomous vehicle navigation. The disclosed embodiments may use cameras to provide autonomous vehicle navigation features. For example, according to embodiments of the present disclosure, the disclosed system may include one, two, or more cameras that monitor the vehicle's environment. The disclosed system may provide a navigation response, for example, based on analysis of images captured by one or more of the cameras. The navigation response may also consider other data, including, for example, global positioning system (GPS) data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data.
[0006] In one embodiment, a system for navigating a host vehicle is disclosed. The system may include at least one processing device programmed to: receive at least one image representing the environment of the host vehicle from the image capture device; determine a planned navigation operation for achieving a navigation goal of the host vehicle based on at least one driving policy; analyze the at least one image to identify a target vehicle within the environment of the host vehicle, the target vehicle's heading toward the host vehicle; determine a next-state distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed; determine a host vehicle braking rate, a host vehicle maximum acceleration capacity, and a current speed of the host vehicle; determine a current stopping distance of the host vehicle based on the host vehicle braking rate, the host vehicle maximum acceleration capacity, and the current speed of the host vehicle; determine the target vehicle's current speed, the target vehicle maximum acceleration capacity, and the target vehicle braking rate; determine the target vehicle's stopping distance based on the target vehicle braking rate, the target vehicle maximum acceleration capacity, and the current speed of the target vehicle; and perform the planned navigation operation if the determined next-state distance is greater than the sum of the host vehicle's stopping distance and the target vehicle's stopping distance.
[0007] In one embodiment, a method for navigating a host vehicle is disclosed. The method may include receiving at least one image representing an environment of the host vehicle from an image capture device, determining a planned navigation operation for achieving a navigation goal of the host vehicle based on at least one driving policy, analyzing the at least one image to identify a target vehicle within the environment of the host vehicle, wherein the target vehicle's heading is toward the host vehicle, determining a next-state distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed, determining a host vehicle braking rate, a host vehicle maximum acceleration capacity, and a current speed of the host vehicle, determining a current stopping distance of the host vehicle based on the host vehicle braking rate, the host vehicle maximum acceleration capacity, and the current speed of the host vehicle, determining the target vehicle's current speed, the target vehicle maximum acceleration capacity, and the target vehicle braking rate, determining the target vehicle's stopping distance based on the target vehicle braking rate, the target vehicle maximum acceleration capacity, and the current speed of the target vehicle, and performing the planned navigation operation if the determined next-state distance is greater than a sum of the host vehicle stopping distance and the target vehicle stopping distance.
[0008] In one embodiment, a system for navigating a host vehicle is disclosed, the system receiving at least one image representing an environment of the host vehicle from an image capture device, determining a planned navigation operation for achieving a navigation goal of the host vehicle based on at least one driving policy, analyzing the at least one image to identify a target vehicle within the environment of the host vehicle, determining a next-state lateral distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed, determining a maximum yaw rate capability of the host vehicle, a maximum change in turning radius capability of the host vehicle, and a current lateral speed of the host vehicle, and and a current lateral speed of the host vehicle; determining a lateral braking distance of the host vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in turning radius capability of the target vehicle; determining a lateral braking distance of the target vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in turning radius capability of the target vehicle; and performing the planned navigation operation if the determined next state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle.
[0009] In one embodiment, a method for navigating a host vehicle is disclosed, the method including steps of receiving at least one image representing an environment of the host vehicle from an image capture device, determining a planned navigation operation for achieving a navigation goal of the host vehicle based on at least one driving policy, analyzing the at least one image to identify a target vehicle within the environment of the host vehicle, determining a next-state lateral distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed, determining a maximum yaw rate capability of the host vehicle, a maximum change in turning radius capability of the host vehicle, and a current lateral velocity of the host vehicle, and determining a maximum yaw rate capability of the host vehicle. The method may include determining a lateral braking distance of the host vehicle based on the yaw rate capability, the maximum change in turning radius capability of the host vehicle, and the current lateral speed of the host vehicle; determining a current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in turning radius capability of the target vehicle; determining a lateral braking distance of the target vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in turning radius capability of the target vehicle; and performing the planned navigation operation if the determined lateral distance of the next state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle.
[0010] In one embodiment, a system for navigating a host vehicle near a pedestrian crossing is disclosed. The system may include at least one processing device programmed to: receive at least one image representing an environment of the host vehicle from an image capture device, detect a representation of a pedestrian crossing in the at least one image based on an analysis of the at least one image, determine whether a representation of a pedestrian appears in the at least one image based on the analysis of the at least one image, detect the presence of a traffic light in the environment of the host vehicle, determine whether the detected traffic light is associated with the host vehicle and the pedestrian crossing, determine a state of the detected traffic light, determine a proximity of the pedestrian to the detected crosswalk when a representation of a pedestrian appears in the at least one image, determine a planned navigation operation for navigating the host vehicle to the detected crosswalk based on at least one driving policy, wherein the determination of the planned navigation operation is further based on the determined state of the detected traffic light and the determined proximity of the pedestrian to the detected crosswalk, and cause one or more actuator systems of the host vehicle to implement the planned navigation operation.
[0011] In one embodiment, a method for navigating a host vehicle near a pedestrian crossing is disclosed. The method may include receiving at least one image representing an environment of the host vehicle from an image capture device, detecting a representation of a pedestrian crossing in the at least one image based on an analysis of the at least one image, determining whether a representation of a pedestrian appears in the at least one image based on the analysis of the at least one image, detecting the presence of a traffic light in the environment of the host vehicle, determining whether the detected traffic light is associated with the host vehicle and the pedestrian crossing, determining a state of the detected traffic light, determining a proximity of the pedestrian to the detected crosswalk when a representation of a pedestrian appears in the at least one image, determining a planned navigation operation for navigating the host vehicle to the detected crosswalk based on at least one driving policy, wherein determining the planned navigation operation is further based on the determined state of the detected traffic light and the determined proximity of the pedestrian to the detected crosswalk, and causing one or more actuator systems of the host vehicle to perform the planned navigation operation.
[0012] In one embodiment, a method for navigating a host vehicle near a pedestrian crossing is disclosed. The method may include receiving at least one image representing an environment of the host vehicle from an image capture device, detecting a start position and an end position of the pedestrian crossing, determining whether a pedestrian is present near the pedestrian crossing based on an analysis of the at least one image, detecting the presence of a traffic light in the environment of the host vehicle, determining whether the traffic light is associated with the host vehicle and the pedestrian crossing, determining a status of the traffic light, determining a navigation operation of the host vehicle near the detected pedestrian crossing based on the traffic light association, the determined status of the traffic light, the presence of a pedestrian near the pedestrian crossing, a selected shortest distance between the pedestrian and either the start position or the end position of the pedestrian crossing, and a motion vector of the pedestrian, and causing one or more actuator systems of the host vehicle to perform the navigation operation.
[0013] According to other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions executable by at least one processing device and that perform any of the steps and / or methods described herein.
[0014] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the scope of the claims. [Brief explanation of the drawings]
[0015] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments.
[0016] [Figure 1] 1 is a diagrammatic representation of an exemplary system according to the disclosed embodiments.
[0017] [Figure 2A] 1 is a side view representation of an exemplary vehicle including a system according to disclosed embodiments.
[0018] [Figure 2B] 2B is a top view representation of the vehicle and system shown in FIG. 2A according to a disclosed embodiment.
[0019] [Figure 2C] 1 is a top view representation of another embodiment of a vehicle including a system according to the disclosed embodiments.
[0020] [Figure 2D] 1 is a top view representation of yet another embodiment of a vehicle including a system according to the disclosed embodiments.
[0021] [Figure 2E] 1 is a top view representation of yet another embodiment of a vehicle including a system according to the disclosed embodiments.
[0022] [Figure 2F]1 is a diagrammatic representation of an exemplary vehicle control system according to the disclosed embodiments.
[0023] [Figure 3A] 1 is a diagrammatic representation of the interior of a vehicle including a rearview mirror and a user interface of a vehicle imaging system according to disclosed embodiments.
[0024] [Figure 3B] 1 is a diagram of an example of a camera mount configured to be positioned behind a rearview mirror, facing a vehicle windshield, according to a disclosed embodiment.
[0025] [Figure 3C] 3C is a diagram of the camera mount shown in FIG. 3B from a different perspective, according to a disclosed embodiment.
[0026] [Figure 3D] 1 is a diagram of an example of a camera mount configured to be positioned behind a rearview mirror, facing a vehicle windshield, according to a disclosed embodiment.
[0027] [Figure 4] FIG. 1 is an exemplary block diagram of a memory configured to store instructions for performing one or more operations in accordance with the disclosed embodiments.
[0028] [Figure 5A] 1 is a flowchart illustrating an exemplary process for generating one or more navigational responses based on monocular image analysis, according to disclosed embodiments.
[0029] [Figure 5B] 1 is a flowchart illustrating an exemplary process for detecting one or more vehicles and / or pedestrians in a set of images, according to disclosed embodiments.
[0030] [Figure 5C]1 is a flowchart illustrating an exemplary process for detecting road markings and / or lane geometry information in a set of images, according to disclosed embodiments.
[0031] [Figure 5D] 1 is a flowchart illustrating an exemplary process for detecting traffic lights in a set of images, according to disclosed embodiments.
[0032] [Figure 5E] 1 is a flowchart of an exemplary process for generating one or more navigational responses based on a vehicle path, according to disclosed embodiments.
[0033] [Figure 5F] 1 is a flowchart illustrating an example process for determining whether a leading vehicle is changing lanes, according to disclosed embodiments.
[0034] [Figure 6] 1 is a flowchart illustrating an example process for generating one or more navigational responses based on stereo image analysis, according to disclosed embodiments.
[0035] [Figure 7] 1 is a flowchart illustrating an exemplary process for generating one or more navigational responses based on an analysis of three sets of images, according to disclosed embodiments.
[0036] [Figure 8] 1 is a block diagram representation of modules that may be implemented by one or more specifically programmed processing devices of a navigation system for an autonomous vehicle, according to disclosed embodiments.
[0037] [Figure 9] 1 is a graph of navigation options in accordance with a disclosed embodiment;
[0038] [Figure 10] 1 is a graph of navigation options in accordance with a disclosed embodiment;
[0039] [Figure 11A] 1 illustrates a schematic diagram of navigation options for a host vehicle within a merge area, according to a disclosed embodiment. [Figure 11B] 1 illustrates a schematic diagram of navigation options for a host vehicle within a merge area, according to a disclosed embodiment. [Figure 11C] 1 illustrates a schematic diagram of navigation options for a host vehicle within a merge area, according to a disclosed embodiment.
[0040] [Figure 11D] 1 illustrates a diagrammatic representation of a double merge scenario in accordance with a disclosed embodiment;
[0041] [Figure 11E] 10 illustrates a graph of potentially useful options in a double merge scenario in accordance with disclosed embodiments.
[0042] [Figure 12] 1 shows a diagram of a representative image captured of a host vehicle's environment along with potential navigation constraints, in accordance with the disclosed embodiments.
[0043] [Figure 13] 1 illustrates a flowchart of an algorithm for navigating a vehicle according to a disclosed embodiment.
[0044] [Figure 14] 1 illustrates a flowchart of an algorithm for navigating a vehicle according to a disclosed embodiment.
[0045] [Figure 15] 1 illustrates a flowchart of an algorithm for navigating a vehicle according to a disclosed embodiment.
[0046] [Figure 16] 1 illustrates a flowchart of an algorithm for navigating a vehicle according to a disclosed embodiment.
[0047] [Figure 17A] 1 illustrates a diagram of a host vehicle navigating into a roundabout, according to a disclosed embodiment; [Figure 17B] 1 illustrates a diagram of a host vehicle navigating into a roundabout, according to a disclosed embodiment;
[0048] [Figure 18] 1 illustrates a flowchart of an algorithm for navigating a vehicle according to a disclosed embodiment.
[0049] [Figure 19] 1 illustrates an example of a host vehicle traveling on a multi-lane highway, according to disclosed embodiments.
[0050] [Figure 20A] 1 illustrates an example of a vehicle cutting in front of another vehicle, according to disclosed embodiments. [Figure 20B] 1 illustrates an example of a vehicle cutting in front of another vehicle, according to disclosed embodiments.
[0051] [Figure 21] 1 illustrates an example of a vehicle following another vehicle, according to disclosed embodiments.
[0052] [Figure 22] 1 illustrates an example of a vehicle exiting a parking lot and merging onto a potentially busy road, according to disclosed embodiments.
[0053] [Figure 23] 1 illustrates a vehicle traveling on a road, according to a disclosed embodiment;
[0054] [Figure 24A] 1 illustrates an example scenario in accordance with disclosed embodiments. [Figure 24B] 1 illustrates an example scenario in accordance with disclosed embodiments. [Figure 24C] 1 illustrates an example scenario in accordance with disclosed embodiments. [Figure 24D] 1 illustrates an example scenario in accordance with disclosed embodiments.
[0055] [Figure 25] 1 illustrates an example scenario in accordance with disclosed embodiments.
[0056] [Figure 26] 1 illustrates an example scenario in accordance with disclosed embodiments.
[0057] [Figure 27] 1 illustrates an example scenario in accordance with disclosed embodiments.
[0058] [Figure 28A] 1 illustrates an example scenario in which a vehicle is tailgating another vehicle, according to disclosed embodiments. [Figure 28B] 1 illustrates an example scenario in which a vehicle is tailgating another vehicle, according to disclosed embodiments.
[0059] [Figure 29A] 1 illustrates an example of a fault in an interruption scenario, according to disclosed embodiments. [Figure 29B] 1 illustrates an example of a fault in an interruption scenario, according to disclosed embodiments.
[0060] [Figure 30A] 1 illustrates an example of a fault in an interruption scenario, according to disclosed embodiments. [Figure 30B] 1 illustrates an example of a fault in an interruption scenario, according to disclosed embodiments.
[0061] [Figure 31A] 1 illustrates an example of a fault in a drift scenario, according to disclosed embodiments. [Figure 31B]1 illustrates an example of a fault in a drift scenario, according to disclosed embodiments. [Figure 31C] 1 illustrates an example of a fault in a drift scenario, according to disclosed embodiments. [Figure 31D] 1 illustrates an example of a fault in a drift scenario, according to disclosed embodiments.
[0062] [Figure 32A] 1 illustrates an example of a fault in a two-way traffic scenario, according to disclosed embodiments. [Figure 32B] 1 illustrates an example of a fault in a two-way traffic scenario, according to disclosed embodiments.
[0063] [Figure 33A] 1 illustrates an example of a fault in a two-way traffic scenario, according to disclosed embodiments. [Figure 33B] 1 illustrates an example of a fault in a two-way traffic scenario, according to disclosed embodiments.
[0064] [Figure 34A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 34B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0065] [Figure 35A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 35B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0066] [Figure 36A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 36B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0067] [Figure 37A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 37B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0068] [Figure 38A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 38B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0069] [Figure 39A] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments. [Figure 39B] 1 illustrates an example of a fault in a route priority scenario, according to disclosed embodiments.
[0070] [Figure 40A] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments. [Figure 40B] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments.
[0071] [Figure 41A] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments. [Figure 41B] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments.
[0072] [Figure 42A] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments. [Figure 42B] 1 illustrates an example of a fault in a traffic light scenario, according to disclosed embodiments.
[0073] [Figure 43A]1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 43B] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 43C] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments.
[0074] [Figure 44A] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 44B] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 44C] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments.
[0075] [Figure 45A] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 45B] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 45C] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments.
[0076] [Figure 46A] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 46B] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 46C] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments. [Figure 46D] 1 illustrates an example of a vulnerable road user (VRU) scenario, according to disclosed embodiments.
[0077] [Figure 47A] FIG. 1 is a diagram of fault points and appropriate responses according to disclosed embodiments.
[0078] [Figure 47B] FIG. 10 is an illustration of route priorities for routes of different geometries, in accordance with the disclosed embodiments.
[0079] [Figure 47C] 10A-10C are diagrams illustrating vertical ordering on routes of different geometries according to disclosed embodiments.
[0080] [Figure 47D] FIG. 10 is an illustration of safe longitudinal distance between vehicles according to disclosed embodiments.
[0081] [Figure 47E] FIG. 1 is a diagram of a situation in which a vehicle is unable to predict the path of another vehicle, in accordance with a disclosed embodiment.
[0082] [Figure 47F] FIG. 1 is an illustration of route priority at a traffic light, according to a disclosed embodiment;
[0083] [Figure 47G] FIG. 1 is a diagram of an exemplary unstructured route, consistent with the disclosed embodiments.
[0084] [Figure 47H] 10A-10C are diagrams of exemplary lateral behavior in unstructured situations, according to disclosed embodiments;
[0085] [Figure 47I] 10A-10C are diagrams of exposure and negligence times in a concealed region, according to disclosed embodiments;
[0086] [Figure 48A] 1 illustrates an example scenario of two vehicles traveling in opposite directions, according to disclosed embodiments.
[0087] [Figure 48B] 1 illustrates an example of a target vehicle traveling toward a host vehicle, according to disclosed embodiments.
[0088] [Figure 49] 1 illustrates an example of a host vehicle maintaining a safe longitudinal distance according to disclosed embodiments.
[0089] [Figure 50A] 1 provides a flowchart illustrating an exemplary process for maintaining a safe vertical distance, according to disclosed embodiments. [Figure 50B] 1 provides a flowchart illustrating an exemplary process for maintaining a safe vertical distance, according to disclosed embodiments.
[0090] [Figure 51A] 1 illustrates an example scenario of two vehicles laterally spaced apart from one another, according to disclosed embodiments.
[0091] [Figure 51B] 1 illustrates an example of a host vehicle maintaining a safe lateral distance according to disclosed embodiments.
[0092] [Figure 52A] 1 illustrates an example of a host vehicle performing a planned navigation operation, according to disclosed embodiments.
[0093] [Figure 52B] 1 illustrates an example of a host vehicle determining whether to perform a navigation operation.
[0094] [Figure 53A] 1 provides a flowchart illustrating an exemplary process for maintaining a safe lateral distance, according to disclosed embodiments. [Figure 53B] 1 provides a flowchart illustrating an exemplary process for maintaining a safe lateral distance, according to disclosed embodiments.
[0095] [Figure 54] 1 is a schematic diagram of a road including a pedestrian crossing, according to a disclosed embodiment;
[0096] [Figure 55A] 1 is a schematic diagram of possible navigation actions performed by a vehicle traveling along a road, according to a disclosed embodiment; [Figure 55B] 1 is a schematic diagram of possible navigation actions performed by a vehicle traveling along a road, according to a disclosed embodiment;
[0097] [Figure 55C] 1 illustrates an example of estimating distance from a pedestrian to a crosswalk according to a disclosed embodiment. [Figure 55D] 1 illustrates an example of estimating distance from a pedestrian to a crosswalk according to a disclosed embodiment.
[0098] [Figure 56] 1 is a flowchart illustrating a process for navigating a vehicle near a pedestrian crossing, according to a disclosed embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0099] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the components shown in the drawings, and the exemplary methods described herein may be modified by substituting, reordering, deleting, or adding steps of the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.
[0100] Autonomous Vehicle Overview
[0101] As used throughout this disclosure, the term “autonomous vehicle” refers to a vehicle that can implement at least one navigation change without driver input. A “navigation change” refers to one or more changes in the vehicle's steering, braking, or acceleration / deceleration. To be autonomous, a vehicle need not be fully automatic (e.g., fully operable without a driver or driver input). Rather, an autonomous vehicle includes a vehicle that can operate under driver control during certain periods of time and without driver control during other periods of time. An autonomous vehicle can also include a vehicle that controls only some aspects of vehicle navigation, such as steering (e.g., to maintain the vehicle course between vehicle lane constraints), or that controls some steering actions under certain circumstances (but not under all circumstances), but leaves other aspects (e.g., braking or braking under certain circumstances) to the driver. In some cases, an autonomous vehicle may handle some or all aspects of the vehicle's braking, speed control, and / or steering.
[0102] Because human drivers typically rely on visual cues and observations to control their vehicles, transportation infrastructure is built accordingly, with lane markings, traffic signs, and traffic lights designed to provide visual information to drivers. In light of these design features of transportation infrastructure, autonomous vehicles may include cameras and processing units that analyze visual information captured from the vehicle's environment. Visual information may include, for example, images depicting transportation infrastructure components (e.g., lane markings, traffic signs, traffic lights, etc.) and other obstacles (e.g., other vehicles, pedestrians, debris, etc.) observable by the driver. Furthermore, autonomous vehicles may also use stored information, such as information that provides a model of the vehicle's environment, when navigating. For example, a vehicle may use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data to provide information related to the vehicle's environment while the vehicle is traveling, and the vehicle (and other vehicles) may use the information to locate itself within the model. Some vehicles may also be capable of vehicle-to-vehicle communication, sharing information, alerting peer vehicles to hazards or changes in the vehicle's surroundings, etc.
[0103] System Overview
[0104] FIG. 1 is a block diagram representation of system 100 according to a disclosed exemplary embodiment. System 100 may include various components depending on particular implementation requirements. In some embodiments, system 100 may include processing unit 110, image acquisition unit 120, position sensor 130, one or more memory units 140, 150, map database 160, user interface 170, and wireless transceiver 172. Processing unit 110 may include one or more processing devices. In some embodiments, processing unit 110 may include application processor 180, image processor 190, or any other suitable processing device. Similarly, image acquisition unit 120 may include any number of image acquisition devices and components depending on the requirements of a particular application. In some embodiments, image acquisition unit 120 may include one or more image capture devices (e.g., cameras, CCDs, any other type of image sensor), such as image capture device 122, image capture device 124, image capture device 126, etc. System 100 may also include a data interface 128 that communicatively connects processing unit 110 to image acquisition unit 120. For example, data interface 128 may include any one or more wired and / or wireless links for transmitting image data acquired by image acquisition unit 120 to processing unit 110.
[0105] Wireless transceiver 172 may include one or more devices configured to exchange transmissions with one or more networks (e.g., cellular, Internet, etc.) over a wireless interface using radio frequencies, infrared frequencies, magnetic fields, or electric fields. Wireless transceiver 172 may send and / or receive data using any known standard (e.g., Wi-Fi, Bluetooth, Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions may include communications from the host vehicle to one or more remotely located servers. Such transmissions may also include communications (one-way or two-way) between the host vehicle and one or more target vehicles in the host vehicle's environment (e.g., to facilitate coordinating the host vehicle's navigation in light of or with target vehicles in the host vehicle's environment), as well as broadcast transmissions to unspecified recipients in the transmitting vehicle's vicinity.
[0106] Both application processor 180 and image processor 190 may include various types of hardware-based processing devices. For example, either or both of application processor 180 and image processor 190 may include a microprocessor, a preprocessor (such as an image preprocessor), a graphics processor, a central processing unit (CPU), support circuitry, a digital signal processor, an integrated circuit, memory, or any other type of device suitable for running applications and processing and analyzing images. In some embodiments, application processor 180 and / or image processor 190 may include any type of single-core or multi-core processor, mobile device microcontroller, central processing unit, etc. A variety of processing devices may be used, including, for example, processors available from manufacturers such as Intel®, AMD®, etc., and may include various architectures (e.g., x86 processor, ARM®, etc.).
[0107] In some embodiments, application processor 180 and / or image processor 190 may include any of the EyeQ series of processors available from Mobileye®. These processor designs include multiple processing units, each with its own local memory and instruction set. Such processors may include video inputs that receive image data from multiple image sensors and may also include video output capabilities. In one example, the EyeQ2® uses 90 nm-micron technology operating at 332 MHz. The EyeQ2® architecture consists of two floating-point, hyper-threaded 32-bit RISC CPUs (MIPS32® 34K® cores), five Vision Computation Engines (VCEs), three Vector Microcode Processors (VMP®), a Denali 64-bit mobile DDR controller, a 128-bit internal audio interconnect, dual 16-bit video input and 18-bit video output controllers, a 16-channel DMA, and several peripherals. The MIPS34K CPU manages five VCEs, three VMP™ processors and DMA, a second MIPS34K CPU and multi-channel DMA, and other peripherals. The five VCEs, three VMP™ processors, and the MIPS34K CPU can perform intensive vision calculations required by feature-rich bundled applications. In another example, the disclosed embodiments may use the EyeQ3™, a third-generation processor that is six times more powerful than the EyeQ2™. In another example, the EyeQ4™ and / or EyeQ5™ processors may be used with the disclosed embodiments. Of course, newer or future EyeQ processing devices may be used with the disclosed embodiments.
[0108] Any of the processing devices disclosed herein can be configured to perform a particular function. Configuring a processing device, such as any of the described EyeQ processors or other controllers or microprocessors, to perform a particular function may include programming computer-executable instructions and providing those instructions to the processing device for execution during operation of the processing device. In some embodiments, configuring a processing device may include directly programming architectural instructions into the processing device. In other embodiments, configuring a processing device may include storing executable instructions on a memory accessible by the processing device during operation. For example, the processing device may access the memory to retrieve and execute the stored instructions during operation. In any event, a processing device configured to perform the sensing, image analysis, and / or navigation functions disclosed herein represents a dedicated hardware-based system that controls multiple hardware-based components of a host vehicle.
[0109] 1 shows two separate processing devices included in processing unit 110, more or fewer processing devices may be used. For example, in some embodiments, a single processing device may be used to accomplish the tasks of application processor 180 and image processor 190. In other embodiments, these tasks may be performed by three or more processing devices. Furthermore, in some embodiments, system 100 may include one or more of processing units 110 without including other components, such as image acquisition unit 120.
[0110] The processing unit 110 may include various types of devices. For example, the processing unit 110 may include various devices such as a controller, an image preprocessor, a central processing unit (CPU), support circuits, a digital signal processor, an integrated circuit, memory, or any other type of device that processes and analyzes images. The image preprocessor may include a video processor that captures, digitizes, and processes images from an image sensor. The CPU may include any number of microcontrollers or microprocessors. The support circuits may be any number of circuits commonly known in the art, including cache, power supplies, clocks, and input / output circuits. The memory may store software that, when executed by the processor, controls the operation of the system. The memory may include databases and image processing software. The memory may include any number of random access memories, read-only memories, flash memories, disk drives, optical storage devices, tape storage devices, removable storage devices, and other types of storage devices. In one example, the memory may be separate from the processing unit 110. In another example, the memory may be integrated into the processing unit 110.
[0111] Each memory 140, 150 may contain software instructions that, when executed by a processor (e.g., application processor 180 and / or image processor 190), may control the operation of various aspects of system 100. These memory units may include various databases and image processing software, as well as trained systems such as neural networks or deep neural networks. The memory units may include random access memory, read-only memory, flash memory, disk drives, optical storage devices, tape storage devices, removable storage devices, and / or any other type of storage device. In some embodiments, memory units 140, 150 may be separate from application processor 180 and / or image processor 190. In other embodiments, these memory units may be integrated into application processor 180 and / or image processor 190.
[0112] Position sensor 130 may include any type of device suitable for determining a location associated with at least one component of system 100. In some embodiments, position sensor 130 may include a GPS receiver. Such a receiver may determine the location and velocity of a user by processing signals broadcast by Global Positioning System satellites. Position information from position sensor 130 may be provided to application processor 180 and / or image processor 190.
[0113] In some embodiments, system 100 may include components such as a speed sensor (e.g., a speedometer) for measuring the speed of vehicle 200. System 100 may also include one or more accelerometers (single-axis or multi-axis) for measuring the acceleration of vehicle 200 along one or more axes.
[0114] The memory units 140, 150 may contain a database or any other organized data indicating the locations of known landmarks. Sensory information of the environment (images, radar signals, depth information from lidar or stereo processing of two or more images, etc.) can be processed along with location information such as GPS coordinates and vehicle ego-motion to determine the vehicle's current position relative to known landmarks and refine the vehicle's position. Certain aspects of this technology are included in a location technology known as REM™, sold by the assignee of the present application.
[0115] User interface 170 may include any device suitable for providing information or receiving input from one or more users of system 100. In some embodiments, user interface 170 may include user input devices including, for example, a touchscreen, a microphone, a keyboard, a pointer device, a track wheel, a camera, a knob, buttons, etc. Using such input devices, a user may be able to provide information input or commands to system 100 by typing instructions or information, providing voice commands, selecting on-screen menu options using buttons, pointer or eye tracking, or through any other suitable technique for communicating information to system 100.
[0116] User interface 170 may comprise one or more processing devices configured to provide information to or receive information from a user and process that information, for example, for use by application processor 180. In some embodiments, such processing devices may execute instructions to recognize and track eye movements, receive and interpret voice commands, recognize and interpret touches and / or gestures made on a touchscreen, respond to keyboard entries or menu selections, etc. In some embodiments, user interface 170 may include a display, a speaker, a tactile device, and / or any other device that provides output information to a user.
[0117] Map database 160 may include any type of database that stores map data useful to system 100. In some embodiments, map database 160 may include data related to the locations in a reference coordinate system of various items, including roads, water features, geographic features, businesses, points of interest, restaurants, gas stations, etc. Map database 160 may store not only the locations of such items, but also descriptors related to those items, including, for example, names associated with any of the stored features. In some embodiments, map database 160 may be physically located with other components of system 100. Alternatively or additionally, map database 160 or portions thereof may be located remotely with respect to other components of system 100 (e.g., processing unit 110). In such embodiments, information from map database 160 may be downloaded to a network via a wired or wireless data connection (e.g., via a cellular network and / or the Internet, etc.). In some cases, map database 160 may store sparse data models including polynomial representations of specific road features (e.g., lane markings) or a target trajectory of the host vehicle. The map database 160 may also include stored representations of various recognized landmarks that may be used to determine or update the known position of the host vehicle relative to the target trajectory. The landmark representations may include data fields such as landmark type and landmark location, among other potential identifiers.
[0118] Image capture devices 122, 124, and 126 may each include any type of device suitable for capturing at least one image from an environment. Furthermore, any number of image capture devices may be used to obtain images for input to the image processor. Some embodiments may include only a single image capture device, while other embodiments may include two, three, or even four or more image capture devices. Image capture devices 122, 124, and 126 are further described below with reference to Figures 2B-2E.
[0119] One or more cameras (e.g., image capture devices 122, 124, and 126) may be part of a detection block included on the vehicle. Various other sensors may be included in the detection block, and any or all of the sensors may be utilized to develop the vehicle's sensed navigation state. In addition to cameras (forward-facing, side-facing, rear-facing, etc.), other sensors, such as radar, lidar, acoustic sensors, etc., may be included in the detection block. Additionally, the detection block may include one or more components configured to communicate, transmit, and receive information related to the vehicle's environment. For example, such components may include a wireless transceiver (e.g., RF) that may receive sensor-based information or any other type of information related to the host vehicle's environment from a source located remotely relative to the host vehicle. Such information may include sensor output information or related information received from vehicle systems other than the host vehicle. In some embodiments, such information may include information received from a remote computing device, a centralized server, etc. Furthermore, the cameras may take many different configurations, such as a single camera unit, multiple cameras, a camera cluster, a long FOV, a short FOV, a wide-angle, a fisheye, etc.
[0120] System 100 or various components of system 100 may be incorporated into a variety of different platforms. In some embodiments, system 100 may be included in vehicle 200, as shown in FIG. 2A. For example, vehicle 200 may include processing unit 110 and any other components of system 100, as described above with respect to FIG. 1. In some embodiments, vehicle 200 may include only a single image capture device (e.g., a camera), while in other embodiments, such as those discussed in connection with FIGS. 2B-2E, multiple image capture devices may be used. For example, as shown in FIG. 2A, either of image capture devices 122 and 124 of vehicle 200 may be part of an ADAS (Advanced Driver Assistance Systems) imaging suite.
[0121] The image capture device included in vehicle 200 as part of image acquisition unit 120 may be located in any suitable location. In some embodiments, as shown in Figures 2A-2E and 3A-3C, image capture device 122 may be located near the rearview mirror. This location may provide a line of sight similar to that of the driver of vehicle 200 and may assist in determining what the driver can and cannot see. While image capture device 122 may be located anywhere near the rearview mirror, placing image capture device 122 on the driver's side of the mirror may further assist in capturing images representative of the driver's field of view and / or line of sight.
[0122] Other locations for the image capture devices of image acquisition unit 120 can also be used. For example, image capture device 124 can be located on or within the bumper of vehicle 200. Such a location can be particularly suitable for image capture devices with a wide field of view. The line of sight of an image capture device located on the bumper can be different from the driver's line of sight, and thus the bumper image capture device and the driver do not always see the same object. Image capture devices (e.g., image capture devices 122, 124, and 126) can also be located in other locations. For example, the image capture devices can be located on one or both side mirrors of vehicle 200, on the roof of vehicle 200, on the hood of vehicle 200, in the trunk of vehicle 200, on the sides of vehicle 200, mounted on, positioned behind, or positioned in front of any window of vehicle 200, mounted near or behind the front and / or rear lights of vehicle 200, etc.
[0123] In addition to the image capture device, vehicle 200 may include various other components of system 100. For example, processing unit 110 may be integrated into the vehicle's engine control unit (ECU) or may be included in vehicle 200 separate from the ECU. Vehicle 200 may also be equipped with a position sensor 130, such as a GPS receiver, and vehicle 200 may also include a map database 160 and memory units 140 and 150.
[0124] As described above, wireless transceiver 172 may transmit and / or receive data via one or more networks (e.g., a cellular network, the Internet, etc.). For example, wireless transceiver 172 may upload data collected by system 100 to one or more servers and download data from one or more servers. Via wireless transceiver 172, system 100 may receive updates to data stored in map database 160, memory 140, and / or memory 150, for example, periodically or on demand. Similarly, wireless transceiver 172 may upload any data from system 100 (e.g., images captured by image acquisition unit 120, data received by position sensor 130, other sensors or vehicle control systems, etc.) and / or any data processed by processing unit 110 to one or more servers.
[0125] System 100 may upload data to a server (e.g., the cloud) based on a privacy level setting. For example, system 100 may implement a privacy level setting that regulates or limits the types of data (including metadata) that may uniquely identify the vehicle and / or the driver / owner of the vehicle that are sent to the server. Such settings may be set by a user via wireless transceiver 172, initialized by factory default settings, or set by data received by wireless transceiver 172, for example.
[0126] In some embodiments, system 100 may upload data according to a "high" privacy level, under a configuration setting, system 100 may transmit data (e.g., location information associated with a route, captured images, etc.) without any details about the specific vehicle and / or driver / owner. For example, when uploading data according to a "high" privacy level, system 100 may not include the vehicle identification number (VIN) or the name of the vehicle's driver or owner, but instead transmit data such as captured images and / or limited location information associated with a route.
[0127] Other privacy levels are also contemplated. For example, system 100 may transmit data to a server according to a “medium” privacy level, which may include additional information not included under a “high” privacy level, such as the make and / or model of the vehicle and / or vehicle type (e.g., passenger car, sport utility vehicle, truck, etc.). In some embodiments, system 100 may upload data according to a “low” privacy level. Under the “low” privacy level setting, system 100 may upload and include sufficient data to uniquely identify a particular vehicle, owner / driver, and / or some or all of the route traveled by the vehicle. Such “low” privacy level data may include, for example, one or more of: VIN, driver / owner name, vehicle's starting point prior to departure, vehicle's intended destination, vehicle make and / or model, vehicle type, etc.
[0128] Figure 2A is a side view representation of an exemplary vehicle imaging system according to a disclosed embodiment. Figure 2B is a top view representation of the embodiment shown in Figure 2A. As shown in Figure 2B, the disclosed embodiment may show a vehicle 200 including within its body a system 100 having a first image capture device 122 positioned near a rearview mirror and / or near a driver of the vehicle 200, a second image capture device 124 positioned on or within a bumper area (e.g., one of bumper areas 210) of the vehicle 200, and a processing unit 110.
[0129] As shown in Figure 2C, both image capture devices 122 and 124 may be positioned near the rearview mirror and / or near the driver of vehicle 200. Furthermore, while two image capture devices 122 and 124 are shown in Figures 2B and 2C, it should be understood that other embodiments may include three or more image capture devices. For example, in the embodiment shown in Figures 2D and 2E, a first image capture device 122, a second image capture device 124, and a third image capture device 126 are included in system 100 of vehicle 200.
[0130] 2D, image capture device 122 may be positioned near the rearview mirror and / or near the driver of vehicle 200, and image capture devices 124 and 126 may be positioned on a bumper area (e.g., one of bumper areas 210) of vehicle 200. Also, as shown in FIG. 2E, image capture devices 122, 124, and 126 may be positioned near the rearview mirror and / or near the driver's seat of vehicle 200. The disclosed embodiments are not limited to any particular number or configuration of image capture devices, and image capture devices may be positioned in any suitable location within and / or on vehicle 200.
[0131] It should be understood that the disclosed embodiments are not limited to vehicles and may be applicable in other contexts. It should also be understood that the disclosed embodiments are not limited to a particular type of vehicle 200 and may be applicable to all types of vehicles, including cars, trucks, trailers, and other types of vehicles.
[0132] First image capture device 122 may include any suitable type of image capture device. Image capture device 122 may include an optical axis. In one example, image capture device 122 may include an Aptina M9V024 WVGA sensor with a global shutter. In other embodiments, image capture device 122 may provide a resolution of 1280 x 960 pixels and may include a rolling shutter. Image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included to provide, for example, a desired focal length and field of view for the image capture device. In some embodiments, image capture device 122 may be associated with a 6 mm lens or a 12 mm lens. In some embodiments, image capture device 122 may be configured to capture an image having a desired field of view (FOV) 202, as shown in FIG. 2D . For example, image capture device 122 may be configured to have a conventional FOV, such as within the range of 40 degrees to 56 degrees, including a 46 degree FOV, a 50 degree FOV, a 52 degree FOV, or degrees greater than 52 degrees. Alternatively, image capture device 122 may be configured to have a narrower FOV, such as a 23 to 40 degree FOV, such as a 28 degree FOV or a 36 degree FOV. Additionally, image capture device 122 may be configured to have a wider FOV, such as a 100 to 180 degree FOV. In some embodiments, image capture device 122 may include a wide-angle bumper camera or a bumper camera with an FOV of up to 180 degrees. In some embodiments, image capture device 122 may be a 7.2 Mpixel image capture device with an aspect ratio of approximately 2:1 (e.g., H×V=3800×1900 pixels) with a horizontal FOV of approximately 100 degrees. Such an image capture device may be used instead of a three-dimensional image capture device configuration. Due to large lens distortion, the vertical FOV of such image capture devices can be much less than 50 degrees in implementations in which the image capture device uses a radially symmetric lens, for example, such lenses are not radially symmetric, thereby allowing a vertical FOV greater than 50 degrees with a horizontal FOV of 100 degrees.
[0133] First image capture device 122 may acquire multiple first images of a scene associated with vehicle 200. The multiple first images may each be acquired as a series of image scan lines, which may be captured using a rolling shutter. Each scan line may include multiple pixels.
[0134] The first image capture device 122 may have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate may refer to the rate at which the image sensor can acquire image data associated with each pixel included in a particular scan line.
[0135] Image capture devices 122, 124, and 126 may include any suitable type and number of image sensors, including, for example, CCD sensors or CMOS sensors. In one embodiment, a CMOS image sensor may be utilized with a rolling shutter, whereby each pixel in a row is read one at a time, and the scanning of the rows proceeds row by row until the entire image frame is captured. In some embodiments, the rows may be captured sequentially from top to bottom relative to the frame.
[0136] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124, and 126) may constitute a high-resolution imager and may have a resolution of greater than 5M pixels, greater than 7M pixels, greater than 10M pixels, or more.
[0137] The use of a rolling shutter can result in pixels in different rows being exposed and captured at different times, which can cause skew and other image artifacts in the captured image frame. On the other hand, if image capture device 122 is configured to operate using a global or synchronous shutter, all pixels can be exposed for the same amount of time during a common exposure period. As a result, image data in a frame collected from a system utilizing a global shutter represents a snapshot of the entire FOV (such as FOV 202) at a particular time. Conversely, when applying a rolling shutter, each row in a frame is exposed and data is captured at a different time. Therefore, moving objects can appear distorted with an image capture device having a rolling shutter. This phenomenon is described in more detail below.
[0138] Second image capture device 124 and third image capture device 126 may be any type of image capture device. Like first image capture device 122, each of image capture devices 124 and 126 may include an optical axis. In one embodiment, each of image capture devices 124 and 126 may include an Aptina M9V024 WVGA sensor with a global shutter. Alternatively, each of image capture devices 124 and 126 may include a rolling shutter. Like image capture device 122, image capture devices 124 and 126 may be configured to include various lenses and optical elements. In some embodiments, the lenses associated with image capture devices 124 and 126 may provide the same FOV as that associated with image capture device 122 (such as FOV 202) or a narrower FOV (such as FOVs 204 and 206). For example, image capture devices 124 and 126 may have a FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less than 20 degrees.
[0139] Image capture devices 124 and 126 may acquire second and third multiple images of a scene associated with vehicle 200. Each of the second and third multiple images may be acquired as second and third series of image scan lines, which may be captured using a rolling shutter. Each scan line or row may have a plurality of pixels. Image capture devices 124 and 126 may have second and third scan rates associated with acquiring each image scan line included in the second and third series.
[0140] Each image capture device 122, 124, and 126 may be positioned in any suitable location and in any suitable orientation relative to vehicle 200. The relative positions of image capture devices 122, 124, and 126 may be selected to facilitate fusing together information obtained from the image capture devices. For example, in some embodiments, the FOV associated with image capture device 124 (FOV 204) may partially or completely overlap with the FOV associated with image capture device 122 (e.g., FOV 202) and the FOV associated with image capture device 126 (e.g., FOV 206).
[0141] Image capture devices 122, 124, and 126 may be positioned on vehicle 200 at any suitable relative height. In one example, there may be a height difference between image capture devices 122, 124, and 126, which may provide sufficient parallax information to enable stereo analysis. For example, as shown in FIG. 2A, two image capture devices 122 and 124 are at different heights. There may also be a lateral displacement difference between image capture devices 122, 124, and 126, which may provide additional parallax information for stereo analysis by processing unit 110, for example. The lateral displacement difference may be calculated as d x In some embodiments, a forward or aft displacement (e.g., range displacement) may exist between image capture devices 122, 124, and 126. For example, image capture device 122 may be positioned 0.5 to 2 meters or more behind image capture device 124 and / or image capture device 126. This type of displacement may allow one of the image capture devices to cover a potential blind spot for the other image capture device.
[0142] Image capture device 122 may have any suitable resolution capability (e.g., number of pixels associated with the image sensor), and the resolution of the image sensor associated with image capture device 122 may be higher, lower, or the same as the resolution of the image sensors associated with image capture devices 124 and 126. In some embodiments, the image sensors associated with image capture device 122 and / or image capture devices 124 and 126 may have a resolution of 640x480, 1024x768, 1280x960, or any other suitable resolution.
[0143] The frame rate (e.g., the rate at which the image capture device acquires a set of pixel data for one image frame before moving on to acquire pixel data associated with the next image frame) may be controllable. The frame rate associated with image capture device 122 may be higher, lower, or the same as the frame rate associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 may depend on various factors that may affect the timing of the frame rate. For example, one or more of image capture devices 122, 124, and 126 may include a selectable pixel delay period imposed before or after acquisition of image data associated with one or more pixels of an image sensor within image capture devices 122, 124, and / or 126. Generally, image data corresponding to each pixel may be acquired according to the device's clock rate (e.g., one pixel per clock cycle). Furthermore, in embodiments including a rolling shutter, one or more of image capture devices 122, 124, and 126 may include a selectable horizontal blanking period imposed before or after acquisition of image data associated with a row of pixels of an image sensor in image capture devices 122, 124, and / or 126. Furthermore, one or more of image capture devices 122, 124, and / or 126 may include a selectable vertical blanking period imposed before or after acquisition of image data associated with an image frame of image capture devices 122, 124, and 126.
[0144] These timing controls may enable the frame rates associated with image capture devices 122, 124, and 126 to be synchronized even if the line scan rate of each image capture device is different. Additionally, as discussed in more detail below, among other factors (e.g., image sensor resolution, maximum line scan rate, etc.), these selectable timing controls may enable the synchronization of image capture from areas where the FOV of image capture device 122 overlaps with the FOV of one or more of image capture devices 124 and 126, even if the field of view of image capture device 122 differs from the FOV of image capture devices 124 and 126.
[0145] The frame rate timing for image capture devices 122, 124, and 126 may depend on the resolution of the associated image sensors. For example, assuming both devices have similar line scan rates, if one device includes an image sensor with a resolution of 640x480 and the other device includes an image sensor with a resolution of 1280x960, it will take longer to capture a frame of image data from the sensor with the higher resolution.
[0146] Another factor that may affect the timing of image data acquisition at image capture devices 122, 124, and 126 is the maximum line scan rate. For example, acquisition of a row of image data from the image sensors included in image capture devices 122, 124, and 126 requires some minimum amount of time. Assuming no pixel delay period is added, this minimum amount of time to acquire a row of image data will be related to the maximum line scan rate of a particular device. Devices that offer higher maximum line scan rates have the potential to provide higher frame rates than devices with lower maximum line scan rates. In some embodiments, one or both of image capture devices 124 and 126 may have a maximum line scan rate that is higher than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture devices 124 and / or 126 may be 1.25, 1.5, 1.75, or 2 times or more the maximum line scan rate of image capture device 122.
[0147] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may operate at a scan rate that is equal to or less than that maximum scan rate. The system may be configured so that one or both of image capture devices 124 and 126 operate at a line scan rate that is equal to the line scan rate of image capture device 122. In other examples, the system may be configured so that the line scan rate of image capture devices 124 and / or 126 may be 1.25, 1.5, 1.75, or 2 times or more the line scan rate of image capture device 122.
[0148] In some embodiments, image capture devices 122, 124, and 126 may be asymmetric. That is, the image capture devices may include cameras with different fields of view (FOV) and focal lengths. The fields of view of image capture devices 122, 124, and 126 may include, for example, any desired area of the environment of vehicle 200. In some embodiments, one or more of image capture devices 122, 124, and 126 may be configured to acquire image data from the environment in front of vehicle 200, the environment behind vehicle 200, the environment on either side of vehicle 200, or a combination thereof.
[0149] Additionally, the focal length associated with each image capture device 122, 124, and / or 126 may be selectable (e.g., by inclusion of an appropriate lens, etc.) so that each device captures images of objects at a desired distance range from vehicle 200. For example, in some embodiments, image capture devices 122, 124, and 126 may capture images of close-up objects within a few meters of the vehicle. Image capture devices 122, 124, 126 may also be configured to capture images of objects at greater distances from the vehicle (e.g., 25 m, 50 m, 100 m, 150 m, or more). Furthermore, the focal lengths of image capture devices 122, 124, and 126 may be selected such that one image capture device (e.g., image capture device 122) can capture images of objects relatively close to the vehicle (e.g., within 10 m or within 20 m), while the other image capture devices (e.g., image capture devices 124 and 126) can capture images of objects farther away from vehicle 200 (e.g., more than 20 m, more than 50 m, more than 100 m, more than 150 m, etc.).
[0150] According to some embodiments, the FOV of one or more of image capture devices 122, 124, and 126 may have a wide angle. For example, it may be advantageous to have an FOV of 140 degrees, especially for image capture devices 122, 124, and 126 that may be used to capture images of areas near vehicle 200. For example, image capture device 122 may be used to capture images of areas to the right or left of vehicle 200, and in such embodiments, it may be desirable for image capture device 122 to have a wide FOV (e.g., at least 140 degrees).
[0151] The field of view associated with each of image capture devices 122, 124, and 126 may depend on the respective focal lengths. For example, as the focal lengths increase, the corresponding field of view decreases.
[0152] Image capture devices 122, 124, and 126 may be configured to have any suitable field of view. In one particular example, image capture device 122 may have a horizontal FOV of 46 degrees, image capture device 124 may have a horizontal FOV of 23 degrees, and image capture device 126 may have a horizontal FOV between 23 and 46 degrees. In another example, image capture device 122 may have a horizontal FOV of 52 degrees, image capture device 124 may have a horizontal FOV of 26 degrees, and image capture device 126 may have a horizontal FOV between 26 and 52 degrees. In some embodiments, the ratio between the FOV of image capture device 122 and the FOV of image capture device 124 and / or image capture device 126 may vary between 1.5 and 2.0. In other embodiments, this ratio may vary between 1.25 and 2.25.
[0153] System 100 may be configured so that the field of view of image capture device 122 at least partially or completely overlaps the field of view of image capture device 124 and / or image capture device 126. In some embodiments, system 100 may be configured so that the fields of view of image capture devices 124 and 126, for example, fall within the field of view of image capture device 122 (e.g., are smaller than the field of view of image capture device 122) and share a common center with the field of view of image capture device 122. In other embodiments, image capture devices 122, 124, and 126 may capture adjacent FOVs or may have partially overlapping FOVs. In some embodiments, the fields of view of image capture devices 122, 124, and 126 may be aligned such that the center of image capture device 124 and / or 126, which has a narrower FOV, may be located in the bottom half of the field of view of device 122, which has a wider FOV.
[0154] 2F is a diagrammatic representation of an exemplary vehicle control system according to disclosed embodiments. As shown in FIG. 2F, vehicle 200 may include a throttle system 220, a braking system 230, and a steering system 240. System 100 may provide inputs (e.g., control signals) to one or more of throttle system 220, braking system 230, and steering system 240 via one or more data links (e.g., any wired and / or wireless link or link that transmits data). For example, based on analysis of images acquired by image capture devices 122, 124, and / or 126, system 100 may provide control signals to one or more of throttle system 220, braking system 230, and steering system 240 to navigate vehicle 200 (e.g., by causing it to accelerate, turn, shift lanes, etc.). Additionally, system 100 may receive inputs indicative of the operating conditions of vehicle 200 (e.g., speed, whether vehicle 200 is braking and / or turning, etc.) from one or more of throttle system 220, braking system 230, and steering system 24. Further details are provided below in connection with Figures 4-7.
[0155] As shown in FIG. 3A , vehicle 200 may also include a user interface 170 for interacting with a driver or passenger of vehicle 200. For example, user interface 170 in a vehicle application may include a touchscreen 320, knobs 330, buttons 340, and a microphone 350. A driver or passenger of vehicle 200 may also interact with system 100 using a steering wheel (e.g., located on or near the steering column of vehicle 200, including, for example, a turn signal handle), buttons (e.g., located on the steering wheel of vehicle 200), and the like. In some embodiments, microphone 350 may be positioned adjacent to rearview mirror 310. Similarly, in some embodiments, image capture device 122 may be located near rearview mirror 310. In some embodiments, user interface 170 may also include one or more speakers 360 (e.g., speakers of a vehicle audio system). For example, system 100 may provide various notifications (e.g., alerts) via speaker 360.
[0156] 3B-3D are diagrams of an exemplary camera mount 370 configured to be positioned behind a rearview mirror (e.g., rearview mirror 310) and opposite a vehicle windshield, according to disclosed embodiments. As shown in FIG. 3B , camera mount 370 may include image capture devices 122, 124, and 126. Image capture devices 124 and 126 may be positioned behind glare shield 380, which may be in direct contact with the windshield and may include a film and / or anti-reflective material composition. For example, glare shield 380 may be positioned to be aligned opposite a windshield having a matching slope. In some embodiments, each of image capture devices 122, 124, and 126 may be positioned behind glare shield 380, for example, as shown in FIG. 3D . The disclosed embodiments are not limited to any particular configuration of image capture devices 122, 124, and 126, camera mount 370, and glare shield 380. FIG. 3C is a view of the camera mount 370 shown in FIG. 3B from the front.
[0157] As will be appreciated by those skilled in the art having the benefit of this disclosure, many variations and / or modifications may be made to the above-disclosed embodiments. For example, not all components are essential to the operation of system 100. Furthermore, any component may be located in any suitable portion of system 100, and the components may be rearranged in various configurations while still providing the functionality of the disclosed embodiments. Accordingly, the configurations discussed above are examples, and regardless of the configuration described above, system 100 may provide a wide range of functionality for analyzing the surroundings of vehicle 200 and navigating vehicle 200 in response to the analysis.
[0158] As discussed in more detail below, according to various disclosed embodiments, system 100 may provide various features related to autonomous driving and / or driver assistance technologies. For example, system 100 may analyze image data, location data (e.g., GPS location information), map data, speed data, and / or data from sensors included in vehicle 200. System 100 may collect data for analysis, for example, from image acquisition unit 120, location sensor 130, and other sensors. Furthermore, system 100 may analyze the collected data to identify whether vehicle 200 should take a particular action and then automatically take the determined action without human intervention. For example, when vehicle 200 navigates without human assistance, system 100 may automatically control the braking, acceleration, and / or steering of vehicle 200 (e.g., by sending control signals to one or more of throttle system 220, braking system 230, and steering system 240). Furthermore, system 100 may analyze the collected data and issue warnings and / or alerts to vehicle occupants based on the analysis of the collected data. Further details regarding various embodiments provided by system 100 are provided below.
[0159] Forward-facing multi-imaging system
[0160] As discussed above, system 100 may provide driver assistance features using a multi-camera system. The multi-camera system may use one or more cameras facing forward of the vehicle. In other embodiments, the multi-camera system may include one or more cameras facing the side of the vehicle or the rear of the vehicle. In one embodiment, for example, system 100 may use a two-camera imaging system, where a first camera and a second camera (e.g., image capture devices 122 and 124) may be positioned at the front and / or side of the vehicle (e.g., vehicle 200). Other camera configurations are consistent with the disclosed embodiments, and the configurations disclosed herein are examples. For example, system 100 may include any number of camera configurations (e.g., 1, 2, 3, 4, 5, 6, 7, 8, etc.). Additionally, system 100 may include a "cluster" of cameras. For example, a cluster of cameras (including any suitable number, e.g., 1, 4, 8, etc., of cameras) can face forward relative to the vehicle, or can face in any other direction (e.g., rearward, sideways, diagonally, etc.). Thus, system 100 can include multiple clusters of cameras, with each cluster oriented in a particular direction to capture images from a particular region of the vehicle's environment.
[0161] The first camera may have a field of view that is larger, smaller, or partially overlaps the field of view of the second camera. Furthermore, the first camera may be connected to a first image processor to perform monocular image analysis of images provided by the first camera, and the second camera may be connected to a second image processor to perform monocular image analysis of images provided by the second camera. The outputs (e.g., processed information) of the first and second image processors may be combined. In some embodiments, the second image processor may receive images from both the first and second cameras and perform stereo analysis. In another embodiment, system 100 may use a three-camera imaging system, where each camera has a different field of view. Thus, such a system may make decisions based on information derived from objects at various distances both in front of and to the sides of the vehicle. References to monocular image analysis may refer to cases where image analysis is performed based on images captured from a single viewpoint (e.g., a single camera). Stereo image analysis may refer to cases where image analysis is performed based on two or more images captured with one or more image capture parameters changed. For example, captured images suitable for performing stereoscopic analysis may include images captured from two or more different positions, images captured from different fields of view, images captured using different focal lengths, images captured with parallax information, etc.
[0162] For example, in one embodiment, system 100 may implement a three-camera configuration using image capture devices 122-126. In such a configuration, image capture device 122 may provide a narrow field of view (e.g., 34 degrees or other value selected from the range of approximately 20 to 45 degrees), image capture device 124 may provide a wide field of view (e.g., 150 degrees or other value selected from the range of approximately 100 to approximately 180 degrees), and image capture device 126 may provide a medium field of view (e.g., 46 degrees or other value selected from the range of approximately 35 to approximately 60 degrees). In some embodiments, image capture device 126 may operate as the main or primary camera. Image capture devices 122-126 may be positioned substantially side-by-side (e.g., 6 cm apart) behind rearview mirror 310. Furthermore, in some embodiments, as discussed above, one or more of image capture devices 122-126 may be mounted behind a glare shield 380 that is flush with the windshield of vehicle 200. Such a shield may operate to minimize the effect of any reflections from the interior of the vehicle on the image capture devices 122-126.
[0163] 3B and 3C, the wide field of view camera (e.g., image capture device 124 in the example above) may be mounted lower than the narrow main field of view camera (e.g., image capture devices 122 and 126 in the example above). This configuration may provide a free line of sight from the wide field of view camera. To reduce reflections, the camera may be mounted near the windshield of vehicle 200 and may include a polarizer to attenuate reflected light.
[0164] A three-camera system may offer certain performance characteristics. For example, some embodiments may include the ability to verify the detection of an object by one camera based on the detection results from another camera. In the three-camera configuration discussed above, processing unit 110 may include, for example, three processing devices (e.g., three EyeQ series processor chips as discussed above), each dedicated to processing images captured by one or more of image capture devices 122-126.
[0165] In a three-camera system, a first processing device may receive images from both the primary camera and the narrow FOV camera and perform vision processing for the narrow FOV camera to detect, for example, other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the first processing device may calculate pixel discrepancies between the images from the primary camera and the narrow camera and create a 3D reconstruction of the environment of vehicle 200. The first processing device may then combine the 3D reconstruction with 3D map data or 3D information calculated based on information from another camera.
[0166] The second processing device may receive images from the primary camera and perform vision processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the second processing device may calculate camera displacement, calculate pixel discrepancies between successive images based on the displacement, and create a 3D reconstruction (e.g., structure-from-motion) of the scene. The second processing device may send the structure-from-motion based on the 3D reconstruction to the first processing device and combine the structure-from-motion with the stereoscopic 3D image.
[0167] The third processing device may receive images from the wide FOV camera and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing device may further execute additional processing instructions to analyze the images and identify moving objects in the images, such as vehicles changing lanes, pedestrians, etc.
[0168] In some embodiments, having image-based information streams captured and processed independently may provide an opportunity for redundancy in the system, such as using a first image capture device and images processed from that device to verify and / or supplement information obtained by capturing and processing image information from at least a second image capture device.
[0169] In some embodiments, system 100 may use two image capture devices (e.g., image capture devices 122 and 124) in providing navigation assistance to vehicle 200, and may use a third image capture device (e.g., image capture device 126) to provide redundancy and verify the analysis of data received from the other two image capture devices. For example, in such a configuration, image capture devices 122 and 124 may provide images for stereo analysis by system 100 for navigating vehicle 200, while image capture device 126 may provide images for monocular analysis by system 100 to provide redundancy and verification of information derived based on images captured from image capture device 123 and / or image capture device 124. That is, image capture device 126 (and corresponding processing device) may be considered to provide a redundant subsystem that provides a check on the analysis derived from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Additionally, in some embodiments, redundancy and validation of received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors, information received from one or more transceivers outside the vehicle, etc.).
[0170] Those skilled in the art will recognize that the above camera configurations, camera placements, camera numbers, camera locations, etc. are merely exemplary. These components, etc., described for an overall system, may be assembled and used in a variety of different configurations without departing from the scope of the disclosed embodiments. Further details regarding the use of multi-camera systems to provide driver assistance and / or autonomous vehicle functionality follow below.
[0171] 4 is an example functional block diagram of memory 140 and / or 150 that may be stored / programmed with instructions to perform one or more operations in accordance with the disclosed embodiments. While reference is made below to memory 140, those skilled in the art will recognize that instructions may be stored in memory 140 and / or 150.
[0172] 4, memory 140 may store a monocular image analysis module 402, a stereo image analysis module 404, a velocity and acceleration module 406, and a navigation response module 408. The disclosed embodiments are not limited to any particular configuration of memory 140. Furthermore, application processor 180 and / or image processor 190 may execute instructions stored in any of modules 402-408 included in memory 140. Those skilled in the art will understand that references to processing unit 110 in the following discussion may refer to application processor 180 and image processor 190 individually or collectively. Accordingly, any steps of the following processes may be performed by one or more processing devices.
[0173] In one embodiment, monocular image analysis module 402 may store instructions (e.g., computer vision software) that, when executed by processing unit 110, perform monocular image analysis of a set of images acquired by one of image capture devices 122, 124, and 126. In some embodiments, processing unit 110 may combine information from the set of images with additional sensory information (e.g., information from radar) to perform the monocular image analysis. As described below in connection with FIGS. 5A-5D , monocular image analysis module 402 may include instructions for detecting a set of features in the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazards, and any other features associated with the vehicle's environment. Based on the analysis, system 100 (e.g., via processing unit 110) may cause one or more navigational responses in vehicle 200, such as turns, lane shifts, and acceleration changes, as discussed below in connection with navigation response module 408.
[0174] In one embodiment, monocular image analysis module 402 may store instructions (e.g., computer vision software) that, when executed by processing unit 110, perform monocular image analysis of a set of images acquired by one of image capture devices 122, 124, and 126. In some embodiments, processing unit 110 may combine information from the set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform the monocular image analysis. As described below in connection with FIGS. 5A-5D , monocular image analysis module 402 may include instructions for detecting a set of features in the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazards, and any other features associated with the vehicle's environment. Based on the analysis, system 100 may cause one or more navigational responses in vehicle 200, such as a turn, lane shift, acceleration change, etc., as discussed below in connection with determining a navigational response (e.g., by processing unit 110).
[0175] In one embodiment, stereo image analysis module 404 may store instructions (e.g., computer vision software) that, when executed by processing unit 110, perform stereo image analysis of first and second sets of images acquired by a combination of image capture devices selected from image capture devices 122, 124, and 126. In some embodiments, processing unit 110 may combine information from the first and second sets of images with additional sensory information (e.g., information from radar) to perform stereo image analysis. For example, stereo image analysis module 404 may include instructions to perform stereo image analysis based on the first set of images acquired by image capture device 124 and the second set of images acquired by image capture device 126. As described below in connection with FIG. 6 , stereo image analysis module 404 may include instructions to detect a set of features in the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and hazards. Based on the analysis, processing unit 110 may cause one or more navigation responses in vehicle 200, such as turns, lane shifts, and acceleration changes, as described below in connection with navigation response module 408. Additionally, in some embodiments, stereo image analysis module 404 may implement techniques related to trained systems (such as neural networks or deep neural networks) or untrained systems.
[0176] In one embodiment, speed and acceleration module 406 may store software configured to analyze data received from one or more computational and electromechanical devices within vehicle 200 configured to alter the speed and / or acceleration of vehicle 200. For example, processing unit 110 may execute instructions associated with speed and acceleration module 406 to calculate a target speed of vehicle 200 based on data derived from the execution of monocular image analysis module 402 and / or stereo image analysis module 404. Such data may include, for example, target position, speed, and / or acceleration, the position and / or speed of vehicle 200 relative to nearby vehicles, pedestrians, or road objects, and the position information of vehicle 200 relative to lane markings on the road. Additionally, processing unit 110 may calculate a target speed of vehicle 200 based on sensory input (e.g., information from radar) and input from other systems of vehicle 200, such as throttle system 220, braking system 230, and / or steering system 240 of vehicle 200. Based on the calculated target speed, the processing unit 110 may send electronic signals to the throttle system 220, the braking system 230 and / or the steering system 240 of the vehicle 200 to trigger a change in speed and / or acceleration, for example, by physically releasing the brakes or easing the accelerator of the vehicle 200.
[0177] In one embodiment, the navigation response module 408 may be executable by the processing unit 110 and may store software that determines a desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include position and speed information associated with nearby vehicles, pedestrians, and road objects, as well as target position information for the vehicle 200. Furthermore, in some embodiments, the navigation response may be based (partially or fully) on map data, a predetermined position of the vehicle 200, and / or a relative speed or relative acceleration between the vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 may also determine a desired navigation response based on sensory input (e.g., information from radar) and inputs from other systems of the vehicle 200, such as the throttle system 220, the braking system 230, and the steering system 240 of the vehicle 200. Based on the desired navigation response, processing unit 110 may send electronic signals to throttle system 220, braking system 230, and steering system 240 of vehicle 200 to trigger the desired navigation response, for example, by turning the steering wheel of vehicle 200 to achieve a predetermined angle of rotation. In some embodiments, processing unit 110 may use the output of navigation response module 408 (e.g., the desired navigation response) as input to the execution of velocity and acceleration module 406 to calculate a change in velocity of vehicle 200.
[0178] Additionally, any of the modules disclosed herein (e.g., modules 402, 404, and 406) may implement techniques related to trained systems (such as neural networks or deep neural networks) or untrained systems.
[0179] 5A is a flowchart illustrating an example process 500A for generating one or more navigational responses based on monocular image analysis, according to a disclosed embodiment. At step 510, processing unit 110 may receive a plurality of images via data interface 128 between processing unit 110 and image acquisition unit 120. For example, a camera (such as image capture device 122 having field of view 202) included in image acquisition unit 120 may capture a plurality of images of an area in front of vehicle 200 (or, for example, to the side or rear of the vehicle) and transmit them to processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). Processing unit 110 may execute monocular image analysis module 402 to analyze the plurality of images at step 520, as described in further detail below in connection with FIGS. 5B-5D . By performing the analysis, processing unit 110 may detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, and traffic lights.
[0180] Processing unit 110 may also execute monocular image analysis module 402 in step 520 to detect various road hazards, such as truck tire parts, fallen road signs, loose cargo, and small animals. Road hazards may vary in structure, shape, size, and color, making such hazards more difficult to detect. In some embodiments, processing unit 110 may execute monocular image analysis module 402 to perform multi-frame analysis on multiple images to detect road hazards. For example, processing unit 110 may estimate camera movement between consecutive image frames and calculate pixel discrepancies between frames to build a 3D map of the road. Processing unit 110 may then use the 3D map to detect the road surface and hazards present on the road surface.
[0181] At step 530, processing unit 110 may execute navigation response module 408 to cause vehicle 200 to make one or more navigational responses based on the analysis performed at step 520 and the techniques described above in connection with FIG. 4 . The navigational responses may include, for example, a turn, a lane shift, and an acceleration change. In some embodiments, processing unit 110 may use data derived from execution of speed and acceleration module 406 to cause one or more navigational responses. Furthermore, multiple navigational responses may occur simultaneously, sequentially, or any combination thereof. For example, processing unit 110 may cause vehicle 200 to cross a lane and then, for example, accelerate, for example, by sequentially sending control signals to steering system 240 and throttle system 220 of vehicle 200. Alternatively, processing unit 110 may cause vehicle 200 to brake and simultaneously shift lanes, for example, by simultaneously sending control signals to braking system 230 and steering system 240 of vehicle 200.
[0182] FIG. 5B is a flowchart illustrating an example process 500B for detecting one or more vehicles and / or pedestrians in a set of images according to a disclosed embodiment. Processing unit 110 may execute monocular image analysis module 402 to perform process 500B. In step 540, processing unit 110 may identify a set of candidate objects representing possible vehicles and / or pedestrians. For example, processing unit 110 may scan one or more images, compare the images with one or more predetermined patterns, and identify locations within each image that may contain a target object (e.g., a vehicle, a pedestrian, or portions thereof). The predetermined patterns may be specified to achieve a low rate of “false hits” and a low rate of “misses.” For example, processing unit 110 may use a low similarity threshold to the predetermined patterns to identify candidate objects as possible vehicles or pedestrians. By doing so, processing unit 110 may reduce the probability of missing (e.g., not identifying) a candidate object representing a vehicle or pedestrian.
[0183] At step 542, processing unit 110 may filter the set of candidate objects to exclude certain candidates (e.g., irrelevant or less relevant objects) based on classification criteria. Such criteria may be derived from various characteristics associated with object types stored in a database (e.g., a database stored in memory 140). The characteristics may include the object's shape, dimensions, texture, and location (e.g., relative to vehicle 200), etc. Thus, processing unit 110 may use one or more sets of criteria to reject false candidates from the set of candidate objects.
[0184] At step 544, processing unit 110 may analyze multiple image frames to identify whether objects in the set of candidate images represent vehicles and / or pedestrians. For example, processing unit 110 may track detected candidate objects across successive frames and accumulate frame-by-frame data associated with the detected objects (e.g., size, position relative to vehicle 200, etc.). Additionally, processing unit 110 may estimate parameters of the detected objects and compare the object's frame-by-frame position data to predicted positions.
[0185] In step 546, processing unit 110 may construct a set of measurements of the detected objects. Such measurements may include, for example, position, velocity, and acceleration values (relative to vehicle 200) associated with the detected objects. In some embodiments, processing unit 110 may construct the measurements based on estimation techniques using a series of time-based observations, such as a Kalman filter or linear quadratic estimation (LQE), and / or modeling data available for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). A Kalman filter may be based on measurements of the object's scale, where the scale measurement is proportional to the time to impact (e.g., the amount of time it takes vehicle 200 to reach the object). Thus, by performing steps 540-546, processing unit 110 may identify vehicles and pedestrians appearing in the set of captured images and derive information (e.g., position, velocity, size) associated with the vehicles and pedestrians. Based on the identification and derived information, processing unit 110 may cause one or more navigation responses in vehicle 200, as described above in connection with FIG. 5A .
[0186] In step 548, processing unit 110 may perform optical flow analysis of one or more images to reduce the probability of detecting “false hits” and missing candidate objects representing vehicles or pedestrians. Optical flow analysis may refer to, for example, analyzing movement patterns, separate from road surface movement, for vehicle 200 in one or more images associated with other vehicles and pedestrians. Processing unit 110 may calculate the movement of candidate objects by observing different positions of the object across multiple image frames captured at different times. Processing unit 110 may use the position and time values as inputs to a mathematical model to calculate the movement of candidate objects. Thus, optical flow analysis may provide another method for detecting vehicles and pedestrians in the vicinity of vehicle 200. Processing unit 110 may perform optical flow analysis in combination with steps 540-546 to provide redundancy for detecting vehicles and pedestrians and increase the reliability of system 100.
[0187] 5C is a flowchart illustrating an example process 500C for detecting road marks and / or lane geometry information in a set of images according to the disclosed embodiments. Processing unit 110 may execute monocular image analysis module 402 to perform process 500C. In step 550, processing unit 110 may detect a set of objects by scanning one or more images. To detect lane mark segments, lane geometry information, and other related road marks, processing unit 110 may filter the set of objects to exclude those determined to be irrelevant (e.g., small holes, small rocks, etc.). In step 552, processing unit 110 may group together segments detected in step 550 that belong to the same road or lane mark. Based on the grouping, processing unit 110 may develop a model, such as a mathematical model, to represent the detected segments.
[0188] At step 554, processing unit 110 may construct a set of measurements associated with the detected segment. In some embodiments, processing unit 110 may create a projection of the detected segment from the image plane onto the real-world plane. The projection may be characterized using a third-order polynomial with coefficients corresponding to physical properties such as the position, slope, curvature, and curvature derivative of the detected road. In generating the projection, processing unit 110 may take into account road surface variations and pitch and roll rates associated with vehicle 200. Additionally, processing unit 110 may model road height by analyzing the position and motion cues present on the road surface. Furthermore, processing unit 110 may estimate pitch and roll rates associated with vehicle 200 by tracking a set of feature points in one or more images.
[0189] In step 556, processing unit 110 may perform a multi-frame analysis, for example, by tracking the detected segment across successive image frames and accumulating frame-by-frame data associated with the detected segment. As processing unit 110 performs the multi-frame analysis, the set of measurements constructed in step 554 may become more reliable and may be associated with an increasingly higher degree of confidence. Thus, by performing steps 550-556, processing unit 110 may identify road marks appearing in the set of captured images and derive lane geometry information. Based on the identification and derived information, processing unit 110 may cause one or more navigational responses in vehicle 200, as described above in connection with FIG. 5A .
[0190] In step 558, processing unit 110 may consider additional information sources to further develop a safety model of vehicle 200 in the vehicle's surroundings. Processing unit 110 may use the safety model to define situations in which system 100 may safely perform autonomous control of vehicle 200. To develop the safety model, in some embodiments, processing unit 110 may consider the positions and movements of other vehicles, detected road edges and barriers, and / or general road shape descriptions extracted from map data (such as data from map database 160). By considering additional information sources, processing unit 110 may provide redundancy in detecting road marks and lane geometry and increase the reliability of system 100.
[0191] FIG. 5D is a flowchart illustrating an example process 500D for detecting traffic lights in a set of images according to disclosed embodiments. Processing unit 110 may execute monocular image analysis module 402 to perform process 500D. In step 560, processing unit 110 may scan the set of images and identify objects that appear at locations in the images that are likely to contain traffic lights. For example, processing unit 110 may filter the identified objects to construct a set of candidate objects that excludes objects that are unlikely to correspond to traffic lights. Filtering may be based on various characteristics associated with traffic lights, such as shape, size, texture, and location (e.g., relative to vehicle 200). Such characteristics may be based on many examples of traffic lights and traffic control signals and may be stored in a database. In some embodiments, processing unit 110 may perform multi-frame analysis on the set of candidate objects that reflect possible traffic lights. For example, processing unit 110 may track the candidate objects across consecutive image frames, estimate the real-world locations of the candidate objects, and filter out moving objects (which are unlikely to be traffic lights). In some embodiments, processing unit 110 may perform color analysis on the candidate object to identify the relative location of the detected color represented within the potential traffic light.
[0192] In step 562, processing unit 110 may analyze the geometry of the intersection. The analysis may be based on any combination of (i) the number of lanes detected on either side of vehicle 200, (ii) marks (such as arrow marks) detected on the road, and (iii) a description of the intersection extracted from map data (such as data from map database 160). Processing unit 110 may perform the analysis using information derived from execution of monocular analysis module 402. In addition, processing unit 110 may identify correspondences between traffic lights detected in step 560 and lanes appearing near vehicle 200.
[0193] As vehicle 200 approaches the intersection, in step 564, processing unit 110 may update a confidence level associated with the analyzed intersection geometry and detected traffic lights. For example, the number of traffic lights estimated to appear at the intersection compared to the number that actually appear at the intersection may affect the confidence level. Thus, based on the confidence level, processing unit 110 may delegate control to the driver of vehicle 200 to improve the safety situation. By performing steps 560-564, processing unit 110 may identify traffic lights appearing in the set of captured images and analyze the intersection geometry information. Based on the identification and analysis, processing unit 110 may cause one or more navigation responses in vehicle 200, as described above in connection with FIG. 5A .
[0194] 5E is a flowchart of an example process 500E for generating one or more navigation responses in vehicle 200 based on a vehicle path, according to disclosed embodiments. In step 570, processing unit 110 may construct an initial vehicle path associated with vehicle 200. The vehicle path may be represented using a set of points represented by coordinates (x, y), and the distance d between any two points in the set of points may be expressed as i may be in the range of 1 to 5 meters. In one embodiment, processing unit 110 may construct an initial vehicle path using two polynomials, such as left and right road polynomials. Processing unit 110 may calculate the geometric midpoint between the two polynomials and, if there is a predetermined offset (a zero offset may correspond to driving in the center of the lane), offset each point in the resulting vehicle path by the predetermined offset (e.g., a smart lane offset). The offset may be in a direction perpendicular to the segment between any two points in the vehicle path. In another embodiment, processing unit 110 may use one polynomial and an estimated lane width to offset each point in the vehicle path by half the estimated lane width plus the predetermined offset (e.g., a smart lane offset).
[0195] In step 572, processing unit 110 may update the vehicle path constructed in step 570. Processing unit 110 may update the distance d k is the distance d i A higher resolution may be used to reconstruct the vehicle path constructed in 570 so that the distance d is shorter than k may be in the range of 0.1 to 0.3 meters. Processing unit 110 may reconstruct the vehicle path using a parabolic spline algorithm, which may result in a cumulative distance vector S corresponding to the total length of the vehicle path (i.e., based on the set of points representing the vehicle path).
[0196] In step 574, processing unit 110 calculates the look-ahead point ((x l ,z l ) in coordinates. The processing unit 110 may extract look-ahead points from the cumulative distance vector S, and the look-ahead points may be associated with a look-ahead distance and a look-ahead time. The look-ahead distance may have a lower bound range of 10 to 20 meters and may be calculated as the product of the speed of the vehicle 200 and the look-ahead time. For example, as the speed of the vehicle 200 decreases, the look-ahead distance may also decrease (e.g., until a lower bound is reached). The look-ahead time, which may range from 0.5 to 1.5 seconds, may be inversely proportional to the gain of one or more control loops associated with producing a navigation response in the vehicle 200, such as a heading error tracking control loop. For example, the gain of the heading error tracking control loop may depend on the bandwidth of the yaw rate loop, the steering actuator loop, and the vehicle lateral dynamics. Thus, the higher the gain of the heading error tracking control loop, the shorter the look-ahead time.
[0197] In step 576, processing unit 110 may determine a heading error and yaw rate command based on the look-ahead point identified in step 574. Processing unit 110 may calculate the arctangent of the look-ahead point, e.g., arctan(x l / z l) to identify the heading error. Processing unit 110 may determine the yaw rate command as the product of the heading error and a high-level control gain. The high-level control gain may be equal to (2 / look ahead time) if the look ahead distance is not at a lower limit. If the look ahead distance is at a lower limit, the high-level control gain may be equal to (2*velocity of vehicle 200 / look ahead distance).
[0198] 5F is a flowchart illustrating an example process 500F for determining whether a leading vehicle is changing lanes, according to the disclosed embodiments. In step 580, processing unit 110 may determine navigation information associated with the leading vehicle (e.g., a vehicle traveling in front of vehicle 200). For example, processing unit 110 may determine the position, velocity (e.g., direction and speed), and / or acceleration of the leading vehicle using the techniques described above in connection with FIGS. 5A and 5B. Processing unit 110 may also determine one or more road polynomials, lookahead points (associated with vehicle 200), and / or snail trails (e.g., a set of points describing the path taken by the leading vehicle) using the techniques described above in connection with FIG. 5E.
[0199] In step 582, processing unit 110 may analyze the navigation information identified in step 580. In one embodiment, processing unit 110 may calculate the distance (e.g., along the trail) between the snail trail and the road polynomial. If the difference in this distance along the trail exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters for straight roads, 0.3 to 0.4 meters for gently curving roads, and 0.5 to 0.6 meters for sharply curving roads), processing unit 110 may determine that the leading vehicle is likely changing lanes. If multiple vehicles are detected traveling in front of vehicle 200, processing unit 110 may compare the snail trails associated with each vehicle. Based on the comparison, processing unit 110 may determine that a vehicle whose snail trail does not match the snail trails of the other vehicles is likely changing lanes. Processing unit 110 may further compare the curvature of the snail trail (associated with the leading vehicle) with the expected curvature of the road segment along which the leading vehicle is traveling. The expected curvature may be extracted from map data (e.g., data from map database 160), road polynomials, snail trails of other vehicles and prior knowledge about the road, etc. If the difference between the snail trail curvature and the expected curvature of the road segment exceeds a predetermined threshold, processing unit 110 may determine that the leading vehicle is likely changing lanes.
[0200] In another embodiment, processing unit 110 may compare the instantaneous position of the leading vehicle to a look-ahead point (associated with vehicle 200) over a specific time period (e.g., 0.5-1.5 seconds). If the difference in distance and the cumulative sum of discrepancies between the instantaneous position of the leading vehicle and the look-ahead point during the specific time period exceeds a predetermined threshold (e.g., 0.3-0.4 meters for straight roads, 0.7-0.8 meters for gently curving roads, and 1.3-1.7 meters for sharply curving roads), processing unit 110 may determine that the leading vehicle is likely changing lanes. In another embodiment, processing unit 110 may analyze the geometry of the snail trail by comparing the lateral distance traveled along the trail to the expected curvature of the snail trail. The expected radius of curvature is calculated as: (δz 2 +δ x 2 ) / 2 / (δ x ) in which δ x represents the lateral movement distance, and δ z represents the longitudinal movement distance. If the difference between the lateral movement distance and the expected curvature exceeds a predetermined threshold (e.g., 500-700 meters), processing unit 110 may determine that the leading vehicle is likely changing lanes. In another embodiment, processing unit 110 may analyze the position of the leading vehicle. If the position of the leading vehicle obscures the road polynomial (e.g., the leading vehicle overlaps the road polynomial), processing unit 110 may determine that the leading vehicle is likely changing lanes. If the position of the leading vehicle is such that another vehicle is detected ahead of the leading vehicle and the snail trails of the two vehicles are not parallel, processing unit 110 may determine that the (closer) leading vehicle is likely changing lanes.
[0201] In step 584, processing unit 110 may determine whether leading vehicle 200 is changing lanes based on the analysis performed in step 582. For example, processing unit 110 may make that determination based on a weighted average of the individual analyses performed in step 582. Under such a scheme, for example, a determination by processing unit 110 that the leading vehicle is likely changing lanes based on a particular type of analysis may be assigned a value of “1” (with a “0” representing a determination that the leading vehicle is unlikely to be changing lanes). Different weights may be assigned to different analyses performed in step 582, and the disclosed embodiments are not limited to any particular combination of analyses and weights. Furthermore, in some embodiments, the analysis may utilize a trained system (e.g., a machine learning or deep learning system), which may, for example, estimate a future path beyond the vehicle's current location based on images captured at the current location.
[0202] 6 is a flowchart illustrating an example process 600 for generating one or more navigational responses based on stereo image analysis, according to disclosed embodiments. In step 610, processing unit 110 may receive first and second pluralities of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122 and 124 having fields of view 202 and 204) may capture first and second pluralities of images of an area ahead of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first and second pluralities of images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0203] In step 620, processing unit 110 may execute stereo image analysis module 404 to perform stereo image analysis of the first and second plurality of images to create a 3D map of the road in front of the vehicle and detect features in the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and road hazards. The stereo image analysis may be performed similarly to the steps described above in connection with FIGS. 5A-5D . For example, processing unit 110 may execute stereo image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road marks, traffic lights, road hazards, etc.) in the first and second plurality of images, filter out a subset of the candidate objects based on various criteria, perform multi-frame analysis, construct measurements, and identify confidence levels for the remaining candidate objects. In performing the above steps, processing unit 110 may consider information from both the first and second plurality of images, rather than information from only one set of images. For example, processing unit 110 may analyze differences in pixel-level data (or other subsets of data from the two streams of captured images) of a candidate object that appears in both the first and second plurality of images. As another example, processing unit 110 may estimate the position and / or velocity (e.g., relative to vehicle 200) of a candidate object by observing that the object appears in one of the plurality of images but not in the other, or other differences that may exist for objects that appear in the two image streams. For example, the position, velocity, and / or acceleration relative to vehicle 200 may be determined based on the trajectory, location, movement characteristics, etc. of features associated with the object that appears in one or both of the image streams.
[0204] In step 630, processing unit 110 may execute navigation response module 408 to generate one or more navigational responses in vehicle 200 based on the analysis performed in step 620 and the techniques described above in connection with FIG. 4. The navigational responses may include, for example, turns, lane shifts, acceleration changes, speed changes, braking, etc. In some embodiments, processing unit 110 may generate one or more navigational responses using data derived from execution of speed and acceleration module 406. Furthermore, multiple navigational responses may be performed simultaneously, sequentially, or any combination thereof.
[0205] 7 is a flowchart illustrating an example process 700 for generating one or more navigational responses based on the analysis of three sets of images, according to disclosed embodiments. In step 710, processing unit 110 may receive first, second, and third pluralities of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) may capture first, second, and third pluralities of images of an area in front of and / or to the sides of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first, second, and third pluralities of images via three or more data interfaces. For example, each of image capture devices 122, 124, and 126 may have an associated data interface that communicates data to processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0206] At step 720, processing unit 110 may analyze the first, second, and third plurality of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and road hazards. The analysis may be performed similar to the steps described above in connection with FIGS. 5A-5D and 6. For example, processing unit 110 may perform monocular image analysis on each of the first, second, and third plurality of images (e.g., via execution of monocular image analysis module 402 and based on the steps described above in connection with FIGS. 5A-5D). Alternatively, processing unit 110 may perform stereo image analysis on the first and second plurality of images, the second and third plurality of images, and / or the first and third plurality of images (e.g., via execution of stereo image analysis module 404 and based on the steps described above in connection with FIG. 6). Processed information corresponding to the analysis of the first, second, and / or third plurality of images may be combined. In some embodiments, processing unit 110 may perform a combination of monocular image analysis and stereo image analysis. For example, processing unit 110 may perform monocular image analysis on the first plurality of images (e.g., via execution of monocular image analysis module 402) and stereo image analysis on the second and third plurality of images (e.g., via execution of stereo image analysis module 404). The configuration of image capture devices 122, 124, and 126—including their respective positions and fields of view 202, 204, and 206—may affect the type of analysis performed on the first, second, and third plurality of images. The disclosed embodiments are not limited to a particular configuration of image capture devices 122, 124, and 126 or the type of analysis performed on the first, second, and third plurality of images.
[0207] In some embodiments, processing unit 110 may perform tests on system 100 based on the images acquired and analyzed in steps 710 and 720. Such tests may provide an indicator of the overall performance of system 100 with a particular configuration of image capture devices 122, 124, and 126. For example, processing unit 110 may identify the rate of "false hits" (e.g., when system 100 incorrectly determines the presence of a vehicle or pedestrian) and "misses."
[0208] In step 730, processing unit 110 may cause one or more navigation responses in vehicle 200 based on information derived from two of the first, second, and third pluralities of images. The selection of two of the first, second, and third pluralities of images may depend on various factors, such as, for example, the number, type, and size of objects detected in each of the multiple images. Processing unit 110 may make the selection based on the quality and resolution of the images, the effective field of view reflected in the images, the number of captured frames, and the extent to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which the object appears, the proportion in which the object appears in each such frame, etc.).
[0209] In some embodiments, processing unit 110 may select information derived from two of the first, second, and third pluralities of images by identifying the extent to which information derived from one image source is consistent with information derived from other image sources. For example, processing unit 110 may combine processed information (whether monocular analysis, stereo analysis, or any combination of the two) derived from each of image capture devices 122, 124, and 126 to identify visual indicators (e.g., lane markings, detected vehicles and / or their positions and / or paths, detected traffic lights, etc.) that are consistent across images captured from each of image capture devices 122, 124, and 126. Processing unit 110 may also filter out information that is inconsistent across captured images (e.g., a vehicle changing lanes, a lane model showing a vehicle too close to vehicle 200, etc.). Thus, processing unit 110 may select information derived from two of the first, second, and third pluralities of images based on the identification of consistent and inconsistent information.
[0210] The navigational responses may include, for example, turns, lane shifts, and acceleration changes. Processing unit 110 may generate one or more navigational responses based on the analysis performed in step 720 and the techniques described above in connection with FIG. 4 . Processing unit 110 may also generate one or more navigational responses using data derived from execution of velocity and acceleration module 406. In some embodiments, processing unit 110 may generate one or more navigational responses based on the relative position, relative velocity, and / or relative acceleration between vehicle 200 and an object detected in any of the first, second, and third plurality of images. The multiple navigational responses may be performed simultaneously, sequentially, or any combination thereof.
[0211] Reinforcement learning and trained navigation systems
[0212] The following sections discuss autonomous driving, along with systems and methods for achieving autonomous control of a vehicle, whether the vehicle is fully autonomous (self-driving vehicle) or partially autonomous (e.g., one or more drivers assisting the system or function). As shown in FIG. 8 , the autonomous driving task can be divided into three main modules, including a sensing module 801, a driving policy module 803, and a control module 805. In some embodiments, modules 801, 803, and 805 can be stored within memory unit 140 and / or memory unit 150 of system 100, and / or modules 801, 803, and 805 (or portions thereof) can be stored remotely from system 100 (e.g., stored in a server accessible to system 100 via wireless transceiver 172). Additionally, any of the modules disclosed herein (e.g., modules 801, 803, and 805) can implement techniques related to trained systems (e.g., neural networks or deep neural networks) or untrained systems.
[0213] The detection module 801, which may be implemented using the processing unit 110, can handle various tasks related to detecting navigation conditions within the host vehicle's environment. Such tasks may depend on inputs from various sensors and detection systems associated with the host vehicle. These inputs may include images or image streams from one or more onboard cameras, GPS location information, accelerometer output, user feedback, user input to one or more user interface devices, radar, lidar, etc. Detections, which may include data from cameras and / or any other available sensors along with map information, can be collected, analyzed, and organized into "detection states" that describe information extracted from a scene within the host vehicle's environment. The detection states may include detection information related to target vehicles, lane markings, pedestrians, traffic lights, road geometry, lane shape, obstacles, distances to other objects / vehicles, relative speeds, and relative accelerations, among other potential detection information. Supervised machine learning can be performed to generate detection state outputs based on the detection data provided to the detection module 801. The output of the sensing module may represent the sensed navigation “state” of the host vehicle, which may be sent to the driving policy module 803 .
[0214] While the sensed states may be developed based on image data received from one or more cameras or image sensors associated with the host vehicle, the sensed states used for navigation may be developed using any suitable sensor or combination of sensors. In some embodiments, the sensed states may be developed without using captured image data. Indeed, any of the navigation principles described herein may be applicable to sensed states developed based on captured image data as well as sensed states developed using other non-image-based sensors. The sensed states may also be determined by sources external to the host vehicle. For example, the sensed states may be developed completely or partially based on information received from a source remote from the host vehicle (e.g., based on sensor information, processed state information, etc. shared from other vehicles, shared from a central server, or any other source of information relevant to the host vehicle's navigation state).
[0215] The driving policy module 803, which will be described in more detail below and may be implemented using the processing unit 110, may implement a desired driving policy for determining one or more navigation actions for the host vehicle to take in response to sensed navigation conditions. When there are no other agents (e.g., target vehicles or pedestrians) in the host vehicle's environment, the sensed conditions input to the driving policy module 803 may be processed in a relatively straightforward manner. When the sensed conditions require negotiation with one or more other agents, the task becomes more complex. Techniques used to generate the output of the driving policy module 803 may include reinforcement learning (which will be described in more detail below). The output of the driving policy module 803 may include at least one navigation action for the host vehicle, and may include, among other potential desired navigation actions, a desired acceleration (which may lead to an updated speed of the host vehicle), a desired yaw rate of the host vehicle, and a desired trajectory.
[0216] Based on the output from the driving policy module 803, the control module 805, which may also be implemented using the processing unit 110, may develop control instructions for one or more actuators or controlled devices associated with the host vehicle. Such actuators and devices may include an accelerator, one or more steering controls, brakes, signal transmitters, displays, or any other actuators or devices that may be controlled as part of a navigation operation associated with the host vehicle. Aspects of control theory may be used to generate the output of the control module 805. To implement the desired navigation goals or requirements of the driving policy module 803, the control module 805 may be responsible for developing and outputting instructions to controllable components of the host vehicle.
[0217] Returning to the driving policy module 803, in some embodiments, the driving policy module 803 can be implemented using a trained system trained by reinforcement learning. In other embodiments, the driving policy module 803 can be implemented without machine learning methods by using specified algorithms to "manually" address various scenarios that may arise during autonomous navigation. However, such an approach, while feasible, may result in an overly simplistic driving policy and may lack the flexibility of a trained system based on machine learning. A trained system may be better equipped to handle complex navigation situations, be better able to determine whether a taxi is parked or has stopped to pick up or drop off a passenger, be better able to determine whether a pedestrian is about to cross the road in front of the host vehicle, be better able to balance defensiveness against the unexpected behavior of other drivers, be better able to navigate busy roads containing target vehicles and / or pedestrians, be better able to determine when to suspend certain navigation rules or augment others, be better able to anticipate undetected but expected conditions (e.g., whether a pedestrian appears from behind a car or obstacle), etc. Trained systems based on reinforcement learning may be better equipped to deal with continuous and high-dimensional state spaces as well as continuous action spaces.
[0218] Training the system using reinforcement learning may include learning a driving policy to map from sensed states to navigation actions, where the driving policy is a function π:S→A, where S is a set of states,
number
[0219] The system can be trained by exposing it to various navigation conditions, having the system apply a policy, and providing rewards (based on a reward function designed to reward desired navigation behavior). Based on reward feedback, the system can "learn" the policy and become trained in producing desired navigation behavior. For example, a learning system can learn the current state s t ∈S and set a policy
number
[0220] The goal of reinforcement learning (RL) is to find a policy π. At time t, given state s t It is located in the t A reward function r that measures the immediate quality of the t However, at time t, the action a tTaking an action affects the environment and therefore the value of future states. As a result, when deciding which action to take, one should not only consider the current reward, but also future rewards. In some cases, if the system determines that a higher reward can be realized in the future by taking a choice that has a lower reward now, then the system should take that action even if that particular action is associated with a lower reward than another choice available. To formalize this, suppose that a policy π and an initial state s are
number
[0221]
number
[0222] Rather than restricting the horizon to T, we can discount future rewards and define the following for some fixed γ∈(0,1):
[0223]
number
[0224] In any case, the optimal policy is
[0225]
number
[0226] The solution is the expected value over the initial state s.
[0227] There are several possible methodologies for training a driving policy system. For example, the system can use imitation techniques (e.g., behavior cloning) that learn from state / action pairs, where the actions are those selected by a good agent (e.g., a human) in response to specific observed states. Suppose a human driver is observed. The observations provide a basis for training the driving policy system (s t ,a t ) format (s t is the state, a t Many examples of π(s) can be obtained, observed, and used. For example, t )≒a t We can use supervised learning to learn the policy π so that ||π(s t )-a t It is often impossible to learn functions where || is very small. Furthermore, even small errors can accumulate over time to produce large errors.
[0228] Another technique that can be used is policy-based learning, where the policy is expressed in parametric form and can be directly optimized using a suitable optimization technique (e.g., stochastic gradient descent).
number
[0229] The system can also be trained by value-based learning (learning the Q-function or V-function). The optimal value function V * Suppose we can learn a good approximation to . An optimal policy can be constructed (e.g., by using the Bellman equation). Some versions of value-based learning can be performed offline (referred to as "off-policy" training). Some disadvantages of value-based approaches can arise from their strong reliance on Markov assumptions and the required approximation of complex functions (approximating a value function can be more difficult than approximating a policy directly).
[0230] Another technique may involve model-based learning and planning (learning the state transition probabilities and solving an optimization problem to find the optimal V). A combination of these techniques can also be used to train a learning system. In this approach, the dynamics of the process, i.e. (s t ,a t ) and the next state s t+1 Once this function is learned, we can solve an optimization problem to find the policy π whose value is optimal. This is called "planning." One advantage of this approach is that the learning part is supervised, and the triples (s t ,a t ,s t+1) can be applied offline by observing the learning process. As with the "imitation" approach, one disadvantage of this approach can be that small errors in the learning process can accumulate, resulting in a poorly performing policy.
[0231] Another approach for training the driving policy module 803 may include decomposing the driving policy function into semantically significant components. Doing so allows for manual implementation of parts of the policy, which may guarantee the safety of the policy, and for other parts of the policy to be implemented using reinforcement learning methods, which may enable adaptability to many scenarios, a human-like balance between defensive / aggressive behavior, and human-like negotiation with other drivers. From a technical perspective, reinforcement learning methods can combine several methodologies to provide a tractable training procedure, where most of the training can be performed using recorded data or a self-built simulator.
[0232] In some embodiments, training of the driving policy module 803 can utilize an "options" mechanism. To illustrate this, consider a simple scenario of driving policy for a two-lane highway. In a direct RL approach, we define the state as
number
[0233] Autonomous Cruise Control (ACC) policy, ACC :S→A: This policy always outputs a yaw rate of 0 and only changes speed to ensure smooth and accident-free driving.
[0234] ACC+Left policy, o L:S→A: The longitudinal commands for this policy are the same as the ACC commands. The yaw rate is a simple implementation of centering the vehicle toward the center of the left lane while ensuring safe lateral movement (e.g., not moving left if there is a car on the left).
[0235] ACC+Right Policy, o R :S→A:o L Same as above, but the vehicle may be centered towards the center of the right lane.
[0236] These policies can be called "alternatives." Policies that depend on these "alternatives" and select alternatives o : S → O, where O is the set of available choices. In some cases, O = {o ACC ,o L ,o R For all s for which} holds,
number
[0237] In practice, a policy function can be decomposed into a choice graph 901 as shown in FIG. 9. Another example of a choice graph 1000 is shown in FIG. 10. A choice graph may represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There is a special node called the root node 903 of the graph. This node has no input nodes. The decision process begins at the root node and traverses the graph until it reaches a "leaf" node, which refers to a node with no output decision line. As shown in FIG. 9, leaf nodes may include, for example, nodes 905, 907, and 909. When a leaf node is encountered, the driving policy module 803 may output acceleration and steering commands related to the desired navigation behavior associated with the leaf node.
[0238] For example, an interior node, such as nodes 911, 913, or 915, may implement a policy for selecting a child from among its available choices. The set of available children of an interior node includes all of the nodes related to the particular interior node by decision lines. For example, interior node 913, shown in Figure 9 as "Merge," includes three child nodes 909, 915, and 917 ("Stay," "Overtake Right," and "Overtake Left," respectively) that are each connected to node 913 by decision lines.
[0239] Allowing nodes to adjust their position within the hierarchy of the graph of choices allows for flexibility in the decision-making system. For example, any node may be allowed to declare itself as "critical." Each node may implement a function "is critical" that outputs "true" if the node is within the critical section of its policy implementation. For example, a node responsible for takeover may declare itself as critical during operation. This may impose constraints on the set of available children of node u, which may include all nodes v that are children of node u and for which there is a path from v to a leaf node that passes through all nodes designated as critical. Such an approach may, on the one hand, allow for the declaration of a desired path through the graph at each time step, while, on the other hand, may preserve the stability of the policy, especially while the critical parts of the policy are being implemented.
[0240] By defining a graph of choices, the problem of learning a driving policy π:S→A can be decomposed into the problem of defining a policy for each node in the graph, where the policy at an internal node should be selected from among its available child nodes. For some of the nodes, individual policies can be implemented manually (e.g., by an if-then algorithm that specifies a set of actions depending on observed states), while for others, the policy can be implemented using a trained system constructed by reinforcement learning. The choice between a manual approach or a trained / learned approach can depend on the safety aspects relevant to the task and its relative simplicity. The graph of choices can be constructed in a way that some of the nodes are easily implemented, while other nodes can rely on trained models. Such an approach can ensure the safe operation of the system.
[0241] The following discussion provides further details regarding the role of the choice graph of Figure 9 within driving policy module 803. As discussed above, the input to driving policy module is a "sensed state" that outlines, for example, a map of the environment obtained from available sensors. The output of driving policy module 803 is a set of aspirations (optionally together with a set of hard constraints) that define a trajectory as a solution to an optimization problem.
[0242] As described above, the graph of choices represents a hierarchical set of decisions organized as a DAG. There is a special node called the "root" of the graph. The root node is the only node with no incoming edges (e.g., decision lines). The decision process starts at the root node and traverses the graph until it reaches a "leaf" node, i.e., a node with no outgoing edges. Each interior node should implement a policy for selecting one of its available children. Every leaf node should implement a policy that defines a set of aspirations (e.g., a set of navigation goals for the host vehicle) based on the entire path from the root to the leaf. The set of aspirations, together with a set of hard constraints that are directly determined based on the sensed conditions, establish an optimization problem whose solution is the vehicle's trajectory. The hard constraints can be used to further enhance the safety of the system, and the aspirations can be used to provide driving comfort and human-like driving behavior of the system. The trajectory provided as the solution to the optimization problem therefore defines the commands to be applied to steering, braking, and / or engine actuators to achieve the trajectory.
[0243] Returning to FIG. 9 , option graph 901 represents an option graph for a two-lane highway with a merging lane (meaning that at some point a third lane merges into the highway's right or left lane). Root node 903 first determines whether the host vehicle is in a simple road scenario or approaching a merging scenario. This is an example of a decision that can be made based on the detected conditions. Simple road node 911 includes three child nodes: stay node 909, overtake left node 917, and overtake right node 915. Stay refers to a situation in which the host vehicle wants to continue traveling in the same lane. The stay node is a leaf node (no outgoing edges / lines). Thus, the stay node defines a set of wishes. A first wish defined by this node may include, for example, a desired lateral position as close as possible to the center of the current lane of travel. There may also be a wish to navigate smoothly (e.g., within a predetermined or allowable maximum acceleration). The stay node may also define how the host vehicle reacts to other vehicles. For example, a stationary node can survey detected target vehicles and assign each a semantic meaning that can be translated into components of a trajectory.
[0244] Various semantic meanings can be assigned to target vehicles in the host vehicle's environment. For example, in some embodiments, the semantic meanings can include any of the following indications: 1) Unrelated: indicates that the detected vehicle in the scene is not currently related; 2) Next Lane: indicates that the detected vehicle is in an adjacent lane and should maintain an appropriate offset relative to that vehicle (the exact offset can be calculated in an optimization problem that constructs a trajectory given the aspirations and hard constraints and can possibly be vehicle-dependent; the leaf of the graph of choices that the host vehicle lands on sets the semantic type of the target vehicle that defines the aspirations for the target vehicle); 3) Yield: the host vehicle attempts to yield to the detected target vehicle (e.g., by slowing down) (especially if the host vehicle determines that the target vehicle is likely to cut into the host vehicle's lane); 4) Lead: the host vehicle attempts to accept and comply with the right-of-way, e.g., by accelerating; 5) Follow: the host vehicle wishes to follow this target vehicle and maintain a smooth ride; 6) Overtake Left / Right: this means that the host vehicle wishes to initiate a lane change to the left or right lane. Left overtaking node 917 and right overtaking node 915 are interior nodes that have not yet determined a wish.
[0245] The next node in the options graph 901 is a gap selection node 919. This node may be responsible for selecting a gap between two target vehicles in a particular target lane that the host vehicle wishes to enter. By selecting a node of the form IDj, for some value of j, the host vehicle arrives at a leaf that specifies a desire for the trajectory optimization problem, e.g., a maneuver the host vehicle wishes to perform to reach the selected gap. Such maneuvers may include first accelerating / braking in the current lane and proceeding to the target lane at a suitable time to enter the selected gap. If the gap selection node 919 cannot find a suitable gap, it proceeds to an abort node 921 that specifies a desire to return to the center of the current lane and cancel the overtaking.
[0246] Returning to the merge node 913, as the host vehicle approaches the merge, the host vehicle has several options that may depend on the particular situation. For example, as shown in Figure 11A, the host vehicle 1105 is traveling along a two-lane road without detecting any other target vehicles in the main or merge lane 1111 of the two-lane road. In this situation, the driving policy module 803 may select a stay node 909 upon reaching the merge node 913. That is, if it does not detect any target vehicles merging onto the road, it may be desirable to remain in its current lane.
[0247] 11B, the situation is slightly different, where the host vehicle 1105 detects one or more target vehicles 1107 entering the main road 1112 from the merging lane 1111. In this situation, when the driving policy module 803 encounters the merging node 913, the driving policy module 803 may decide to initiate a left-hand overtaking maneuver to avoid the merging situation.
[0248] 11C , the host vehicle 1105 encounters one or more target vehicles 1107 entering the main road 1112 from a merging lane 1111. The host vehicle 1105 also detects a target vehicle 1109 traveling in a lane adjacent to the host vehicle's lane. The host vehicle also detects one or more target vehicles 1110 traveling in the same lane as the host vehicle 1105. In this situation, the driving policy module 803 may decide to adjust the speed of the host vehicle 1105 to yield to the target vehicle 1107 and advance ahead of the target vehicle 1115. This may be achieved, for example, by proceeding to the gap selection node 919, which selects the gap between ID0 (vehicle 1107) and ID1 (vehicle 1115) as the appropriate merge gap. The appropriate gap for the merge situation then defines the objective of the trajectory planner optimization problem.
[0249] As discussed above, nodes in the graph of options can declare themselves "critical," and such declaration can ensure that a selected option passes through the critical nodes. Formally, each node can implement a function IsCritical. After making a forward pass from root to leaf on the graph of options and solving the trajectory planner optimization problem, a backward pass can be made from the leaf to the root. Along this backward pass, the IsCritical function of all nodes in the path can be called, and a list of all critical nodes can be saved. In the forward path corresponding to the next time slot, the driving policy module 803 can be requested to select a path from the root node to the leaf nodes that passes through all critical nodes.
[0250] 11A-11C can be used to illustrate the potential benefits of this approach. For example, in a situation where an overtaking maneuver is initiated and the driving policy module 803 reaches the leaf corresponding to IDk, it is undesirable to select, for example, the stay node 909 if the host vehicle is in the middle of the overtaking maneuver. To avoid this jumpiness, the IDj node can designate itself as critical. During the maneuver, the success of the trajectory planner can be monitored, and the function IsCritical returns a "true" value if the overtaking maneuver proceeds as intended. This approach can ensure that the overtaking maneuver continues within the next time frame (rather than jumping to another, potentially inconsistent maneuver before completing the initially selected maneuver). On the other hand, if maneuver monitoring indicates that the selected maneuver is not proceeding as intended, or if the maneuver becomes unnecessary or impossible, the function IsCritical can return a "false" value. This can allow the gap selection node to select a different gap within the next time frame or to abort the overtaking maneuver entirely. On the one hand, this approach may allow the declaration of a desired path on the graph of alternatives at each time step, while on the other hand, it may help promote stability of the policy during critical parts of the execution.
[0251] Hard constraints, which are described in more detail below, can be distinguished from navigation aspirations. For example, hard constraints can ensure safe driving by applying an additional layer of filtering of planned navigation behavior. The involved hard constraints, which can be manually programmed and defined, can be determined from sensed conditions rather than by using a trained system built on reinforcement learning. However, in some embodiments, the trained system can learn applicable hard constraints to apply and follow. Such an approach can encourage the driving policy module 803 to arrive at a selected action that already complies with the applicable hard constraints, thereby reducing or eliminating selected actions that may later require modification to comply with the applicable hard constraints. Nevertheless, as a redundant safety measure, hard constraints can be applied to the output of the driving policy module 803 even when the driving policy module 803 has been trained to take into account certain hard constraints.
[0252] There are many examples of potential hard constraints. For example, hard constraints can be defined in relation to guardrails at the edge of the road. Under no circumstances is the host vehicle allowed to cross the guardrail. Such a rule imposes a hard lateral constraint on the trajectory of the host vehicle. Another example of a hard constraint can include road bumps (e.g., speed control bumps), which can impose hard constraints on driving speed before or while traversing the bump. Hard constraints can be considered safety-critical and therefore can be defined manually rather than relying solely on a trained system to learn the constraints during training.
[0253] In contrast to strict constraints, the goal of a desire may be to enable or achieve a comfortable driving experience. As discussed above, one example of a desire may include a goal to position the host vehicle at a lateral position within the lane that corresponds to the center of the host vehicle's lane. Another desire may include the ID of a gap to move into. Note that the host vehicle need not be strictly centered in the lane; instead, a desire to be as close to the center of the lane as possible can ensure that the host vehicle is likely to move to the center of the lane if it deviates from the center of the lane. A desire may not be safety-critical. In some embodiments, a desire may require negotiation with other drivers and pedestrians. One approach for constructing a desire may utilize a graph of options, and policies implemented within at least some nodes of the graph may be based on reinforcement learning.
[0254] For nodes in the graph of choices 901 or 1000 that are implemented as nodes that are trained based on learning, the training process may involve decomposing the problem into a supervised learning phase and a reinforcement learning phase. In the supervised learning phase:
number
number
number
number
[0255] A key element that can be provided in some scenarios is a differentiable path from future losses / rewards back to a decision about action. In the structure of a graph of choices, the implementation of choices, including safety constraints, is usually not differentiable. To overcome this problem, the selection of children at nodes in the learned policy can be probabilistic. That is, a node can output a probability vector p, which assigns a probability to be used in selecting each of the children of a particular node. Suppose a node has k children, and a (1) ,...,a (k) Let be the behavior of the path from each child to the leaf. Thus, the resulting predicted behavior is
number
number
[0256] s t , a t Given
number
[0257] Additionally, in some embodiments, the system may implement a multi-agent approach. For example, the system may consider data from various sources and / or imagery captured from multiple angles. Additionally, some disclosed embodiments may result in energy savings because prediction of events that do not directly involve the host vehicle but may have an impact on the host vehicle may be considered, and prediction of events that may result in unpredictable situations involving other vehicles may also be a consideration (e.g., radar may "see through" to unavoidable events that affect the preceding vehicle and the host vehicle, as well as predicting the high likelihood of such events).
[0258] Trained system with imposed navigation constraints
[0259] In the context of autonomous driving, a significant concern is how to ensure that the learned policy of a trained navigation network is safe. In some embodiments, constraints can be used to train a driving policy system, so that actions selected by the trained system may already take into account applicable safety constraints. Additionally, in some embodiments, an additional layer of safety can be provided by subjecting the selected actions of the trained system to one or more strict constraints implicated by the particular sensed scene in the host vehicle's environment. Such an approach can ensure that actions taken by the host vehicle are limited to those confirmed to satisfy applicable safety constraints.
[0260] At its core, a navigation system may include a learning algorithm based on a policy function that maps observed states to one or more desired actions. In some implementations, the learning algorithm is a deep learning algorithm. The desired actions may include at least one action expected to maximize an expected reward for the vehicle. In some cases, the actual action taken by the vehicle may correspond to one of the desired actions, while in other cases, the actual action taken may be determined based on the observed states, the one or more desired actions, and non-learning hard constraints (e.g., safety constraints) imposed on the learning navigation engine. These constraints may include no-driving zones surrounding various types of detected objects (e.g., target vehicles, pedestrians, static objects on or in the road, moving objects on or in the road, guardrails, etc.). In some cases, the size of the zones may vary based on the detected movement (e.g., speed and / or direction) of the detected objects. Other constraints may include maximum travel speed when passing through a pedestrian influence zone, maximum deceleration (to accommodate the spacing of target vehicles behind the host vehicle), mandatory stops at detected crosswalks or railroad crossings, etc.
[0261] Strict constraints used with a system trained by machine learning can provide a degree of safety in autonomous driving that may exceed that obtainable based on the output of the trained system alone. For example, a machine learning system can be trained using a set of desired constraints as training guidelines, and the trained system can then be configured with applicable navigation constraint limits and select actions that adhere to those limits in response to sensed navigation conditions. However, the trained system still has some flexibility when selecting a navigation action, and therefore there may be at least some situations in which the action selected by the trained system may not strictly adhere to the relevant navigation constraints. Therefore, to require that the selected action strictly adhere to the relevant navigation constraints, non-machine learning components that ensure strict application of the relevant navigation constraints outside the learning / training framework can be used to combine, compare, filter, adjust, modify, etc. the output of the trained system.
[0262] The following discussion provides further details on the trained system and the potential benefits (especially from a safety perspective) that can be gained from combining the trained system with algorithmic components outside of the training / learning framework. As discussed above, reinforcement learning objectives with policies can be optimized by stochastic gradient ascent. The objective (e.g., expected reward) is
number
[0263] In machine learning scenarios, objectives that include expectations can be used. However, such objectives, without being bound by navigation constraints, may not return behavior that is strictly bound by those constraints. For example, for a trajectory that represents a rare "turn" event (e.g., an accident) that should be avoided,
number
number
number
number
[0264]
number
number
number
number
[0265] Lemma: π o Let p be a policy, p and r be scalars, so that with probability p,
number
number
[0266]
number
[0267] holds, and the last approximation applies when r≧1 / p.
[0268] This explanation is in the form
number
number
number
number
number
number
number
number
[0269] The double merge navigation situation shown in FIG. 11D provides an example that further illustrates these concepts. In a double merge, vehicles approach merge area 1130 from both the left and right sides. From each side, a decision can be made whether a vehicle, such as vehicle 1133 or vehicle 1135, will merge into the opposite lane of merge area 1130. In congested traffic, successfully executing a double merge can require significant negotiation skill and experience, and can be difficult to perform heuristically or brute force by enumerating all possible trajectories that all agents in the scene can take. In this double merge example, a set of desires D suitable for the double merge operation can be defined. D is the Cartesian product of the following sets: D=[0, v max ]xLx{g,t,o} n where [0,v max ] is the desired target speed of the host vehicle, L={1,1.5,2,2.5,3,3.5,4} is the desired lateral position in lanes, where integers indicate the center of the lane and fractions indicate the lane boundaries, and {g,t,o} are classification labels assigned to each of the other n vehicles. If the host vehicle should yield to the other vehicle, it can be assigned "g," if the host vehicle should gain right-of-way relative to the other vehicle, it can be assigned "t," or if the host vehicle should maintain an offset distance relative to the other vehicle, it can be assigned "o."
[0270] The following is a set of wishes (v, l, c1,..., c n )∈D can be transformed into a cost function over driving trajectories. A driving trajectory is a set of (x1,y1),...,(x k ,yk ), and (x i ,y i ) is the position (lateral, longitudinal) of the host vehicle (in egocentric units) at time τ·i. In some experiments, τ=0.1 seconds and k=10. Of course, other values can be chosen. The cost assigned to a trajectory can include a weighted sum of the individual costs assigned to the desired speed, lateral position, and labels assigned to each of the other n vehicles.
[0271] Desired velocity v∈[0,v max ], the cost of the trajectory relative to the velocity
number
[0272] Given a desired lateral position, the cost associated with the desired lateral position
number
[0273] where dist(x,y,l) is the distance from point (x,y) to lane position l. Regarding the cost due to other vehicles, for any other vehicle, (x ' 1,y ' 1),...,(x ' k ,y ' k ) can represent other vehicles in the egocentric unit of the host vehicle, and i can be expressed as (x i ,y i ) and (x ' j ,y ' j) can be the earliest point j exists such that the distance between them is small. If there is no such point, i can be set as i = ∞. If another vehicle is classified as "giving way", it may be desirable for τi > τj + 0.5, meaning that the host vehicle reaches that same point at least 0.5 seconds after the other vehicle reaches the trajectory intersection. A possible formula for converting the above constraint to cost is [τ(ji) + 0.5] + is.
[0274] Similarly, if another car is classified as "taking the road", it may be desirable for τj > τi + 0.5, which means that the cost [τ(ij) + 0.5] + If another vehicle is classified as "offset", it may be desirable for i = ∞, meaning that the host vehicle's trajectory and the offset vehicle's trajectory do not intersect. This condition can be converted into a cost by penalizing the distance between the trajectories.
[0275] Assigning weights to each of these costs yields a single objective function π for the trajectory planner. (T) Costs that promote smooth driving can be added to the objective. Hard constraints can be added to the objective to ensure functional safety of the track. For example, (x i、 y i ) may be prohibited from leaving the road, and (x i、 y i ) is the trajectory point (x ' j、 y ' j ) about (x ' j、 y ' j ) may be prohibited from approaching.
[0276] In summary, policy θcan be decomposed into an agnostic mapping from states to sets of desires and a mapping from desires to actual trajectories. The latter mapping is not based on learning and can be implemented by solving an optimization problem whose costs depend on desires and whose hard constraints can guarantee the functional safety of the policy.
[0277] The following discussion describes the agnostic mapping from states to sets of desires. As explained above, to comply with functional safety, systems relying solely on reinforcement learning must rely on reward-based learning.
number
[0278] For various reasons, the decision can be further decomposed into semantically meaningful components. For example, the size of D may be large and / or continuous. In the double confluence scenario described above with respect to Figure 11D, D = [0, v max ]xLx{g,t,o} n ) holds. In addition, the gradient estimator is
number
[0279] Returning to the concept of a choice graph, a choice graph that may represent the double-confluence scenario shown in FIG. 11D is shown in FIG. 11E. As discussed above, a choice graph may represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There may be a special node in the graph called a "root" node 1140, which is the only node with no incoming edges (e.g., decision lines). The decision process may start from the root node and traverse the graph until it reaches a "leaf" node, i.e., a node with no outgoing edges. Each internal node may implement a policy function that selects one child from among its available children. There may be a predefined mapping from a set of traversals on the choice graph to a set of wishes D. In other words, a traversal on the choice graph may be automatically converted to a wish in D. Given a node v in the graph, a parameter vector θ v may specify a policy for selecting the children of v. If θ is a v If the connection of θ is v By traversing the graph from root to leaf, selecting child nodes using the policy defined by
number
[0280] In the double merge option graph 1139 of FIG. 11E, root node 1140 can first determine whether the host vehicle is in a merge area (e.g., area 1130 of FIG. 11D) or whether the host vehicle is instead approaching a merge area and needs to prepare for a possible merge. In either case, the host vehicle may need to decide whether to change lanes (e.g., to the left or right) or whether to remain in its current lane. If the host vehicle decides to change lanes, it may need to continue and determine whether conditions are suitable for a lane-changing maneuver (e.g., at “go” node 1142). If changing lanes is not possible, the host vehicle can attempt to “push” toward the desired lane (e.g., at node 1144) by aiming to stay on a lane marking. Alternatively, the host vehicle may choose to “stay” in the same lane (e.g., at node 1146). This process can determine the lateral position of the host vehicle in a natural way. For example,
[0281] This may allow the desired lateral position to be determined in a natural way. For example, if the host vehicle is changing lanes from lane 2 to lane 3, the "go" node may set the desired lateral position to 3, the "stay" node may set the desired lateral position to 2, and the "push" node may set the desired lateral position to 2.5. The host vehicle may then decide whether to maintain the "same" speed (node 1148), "accelerate" (node 1150), or "slow down" (node 1152). The host vehicle may then enter a "chain" structure 1154 that examines other vehicles and sets their semantic meanings to values in the tuple {g,t,o}. This process may set aspirations for other vehicles. The parameters of all nodes in this chain may be shared (similar to recurrent neural networks).
[0282] A potential benefit of the choice is the interpretability of the results. Another potential benefit is that we can exploit the decomposable structure of the set D, thus allowing us to choose the policy at each node from a small number of possibilities. In addition, the structure may allow us to reduce the variance of the policy gradient estimator.
[0283] As discussed above, the length of an episode in a double-merge scenario may be approximately T = 250 steps. This value (or any other appropriate value depending on the particular navigation scenario) may allow sufficient time to acknowledge the consequences of the host vehicle's actions (e.g., if the host vehicle decides to change lanes in preparation for a merge, the host vehicle acknowledges the benefit only after successfully completing the merge). On the other hand, driving dynamics require the host vehicle to make decisions at a sufficiently fast frequency (e.g., 10 Hz in the above case).
[0284] The graph of choices may allow the effective value of T to be reduced in at least two ways. First, given high-level decisions, rewards can be established for lower-level decisions while taking shorter episodes into account. For example, if the host vehicle has already selected the “change lane” and “go” nodes, a policy for assigning semantic meaning to vehicles can be learned by looking at 2-3 second episodes (meaning T would be 20-30 instead of 250). Second, for high-level decisions (such as whether to change lanes or stay in the same lane), the host vehicle may not need to make decisions every 0.1 seconds. Instead, the host vehicle may be able to make decisions less frequently (e.g., every second) or may implement a “choice termination” function, where gradients are calculated only after each termination of a choice. In either case, the effective value of T may be an order of magnitude smaller than its original value. Overall, the estimators at all nodes can rely on values of T that are an order of magnitude smaller than the original 250 steps, which may immediately translate into smaller variance.
[0285] As discussed above, strict constraints can promote safer driving and can be of several different types. For example, static strict constraints can be determined directly from sensed conditions. These can include speed bumps, speed limits, road curvature, intersections, etc. in the host vehicle's environment, which can involve one or more constraints on the vehicle's speed, heading, acceleration, braking (deceleration), etc. Strict static constraints can also include semantic free space, where the host vehicle is prohibited, for example, from going outside the free space and from navigating too close to physical barriers. Strict static constraints can also limit (e.g., prohibit) maneuvers that do not comply with various aspects of the vehicle's kinematic motion; for example, static strict constraints can be used to prohibit maneuvers that could lead the host vehicle to roll over, skid, or otherwise lose control.
[0286] Strict constraints can also be vehicle-related. For example, a constraint can be used that requires a vehicle to maintain a longitudinal distance of at least 1 meter to other vehicles and a lateral distance of at least 0.5 meters from other vehicles. Constraints can also be applied to prevent the host vehicle from maintaining a collision course with one or more other vehicles. For example, time τ can be a measure of time based on a particular scene. The predicted trajectories of the host vehicle and one or more other vehicles from the current time to time τ can be considered. If the two trajectories intersect,
number
number
number
number
[0287] The time τ for tracking the trajectories of the host vehicle and one or more other vehicles can vary, except that in intersection scenarios where speeds may be low, τ can be longer and can be determined so that the host vehicle enters and exits the intersection in less than τ seconds.
[0288] Of course, applying strict constraints to the trajectories of vehicles requires that their trajectories be predicted. For a host vehicle, predicting a trajectory may be relatively straightforward, since the host vehicle generally already understands and, in fact, plans its intended trajectory at any given time. For other vehicles, predicting their trajectories may be less straightforward. For other vehicles, baseline calculations for determining a predicted trajectory may depend on the current speed and heading of the other vehicles, which may be determined, for example, based on analysis of image streams captured by one or more cameras and / or other sensors (radar, lidar, acoustics, etc.) onboard the host vehicle.
[0289] However, there may be some exceptions that simplify the problem or at least provide more confidence in the predicted trajectory of another vehicle. For example, on structured roads where there are lane designations and yielding rules may exist, the trajectory of another vehicle may be based at least in part on the other vehicle's position relative to the lane and on applicable yielding rules. Thus, in some situations, when there is an observed lane structure, vehicles in adjacent lanes may be assumed to adhere to lane boundaries. That is, a host vehicle may assume that vehicles in adjacent lanes will stay in their own lane unless there is observed evidence (e.g., signal lights, strong lateral movement, movement across lane boundaries) that indicates the vehicle in the adjacent lane will cut into the host vehicle's lane.
[0290] Other circumstances may also provide clues as to the expected trajectory of other vehicles. For example, at stop signs, traffic lights, roundabouts, etc., where the host vehicle may have the right-of-way, it can be assumed that other vehicles will respect that right-of-way. Thus, unless there is evidence that the rule has been observed to be broken, it can be assumed that other vehicles will proceed along a trajectory that respects the right-of-way held by the host vehicle.
[0291] Hard constraints may also be applied with respect to pedestrians in the host vehicle's environment. For example, a buffer distance for pedestrians may be established such that the host vehicle is prohibited from navigating any closer than a specified buffer distance to any observed pedestrian. The pedestrian buffer distance may be any suitable distance. In some embodiments, the buffer distance may be at least one meter to an observed pedestrian.
[0292] Similar to the vehicle situation, hard constraints can also be applied to the relative motion between a pedestrian and a host vehicle. For example, the pedestrian's trajectory (based on heading and speed) can be monitored against the host vehicle's predicted trajectory. Given a particular pedestrian's trajectory, for every point p on the trajectory, t(p) may represent the time it takes the pedestrian to reach point p. To maintain a required buffer distance of at least one meter from the pedestrian, t(p) must be greater than the time the host vehicle reaches point p (with enough time difference so that the host vehicle passes in front of the pedestrian by at least one meter), or (e.g., if the host vehicle brakes to yield to the pedestrian), t(p) must be less than the time the host vehicle reaches point p. Furthermore, in the latter example, a hard constraint may require the host vehicle to arrive at point p sufficiently later than the pedestrian so that it can pass behind the pedestrian and maintain the required buffer distance of at least one meter. Of course, there may be exceptions to the hard pedestrian constraint. For example, if the host vehicle has the right-of-way or is moving very slowly and there is no observed evidence of pedestrians refusing to yield to the host vehicle or navigating towards the host vehicle, the strict constraints on pedestrians can be relaxed (e.g., to a narrower buffer of at least 0.75 meters or 0.50 meters).
[0293] In some cases, constraints may be relaxed if it is determined that they cannot all be satisfied. For example, in situations where a road is too narrow to allow a desired distance (e.g., 0.5 meters) from both curves or from curves and parked vehicles, one or more of the constraints may be relaxed if there are mitigating factors. For example, if there are no pedestrians (or other objects) on the sidewalk, one may proceed slowly at 0.1 meters from a curve. In some embodiments, constraints may be relaxed if doing so improves the user experience. For example, to avoid potholes, constraints may be relaxed to allow a vehicle to navigate closer to lane edges, curves, or pedestrians than would normally be permitted. Furthermore, when determining which constraints to relax, in some embodiments, the one or more constraints decided to be relaxed are those deemed to have the least adverse effect on safety. For example, constraints regarding how close a vehicle can travel to a curve or concrete barrier may be relaxed before relaxing constraints dealing with proximity to other vehicles. In some embodiments, pedestrian constraints may be the last to be relaxed, or may never be relaxed depending on the circumstances.
[0294] FIG. 12 illustrates an example of a scene that may be captured and analyzed during navigation of a host vehicle. For example, the host vehicle may include the navigation system (e.g., system 100) described above that may receive multiple images representing the host vehicle's environment from a camera (e.g., at least one of image capture device 122, image capture device 124, and image capture device 126) associated with the host vehicle. The scene illustrated in FIG. 12 is an example of one of the images that may be captured at time t from the environment of the host vehicle traveling in lane 1210 along predicted trajectory 1212. The navigation system may include at least one processing device (e.g., including any of the EyeQ processors or other devices described above) that is specifically programmed to receive the multiple images and analyze the images to determine an action in response to the scene. Among other things, the at least one processing device may implement detection module 801, driving policy module 803, and control module 805 shown in FIG. 8. The detection module 801 may be responsible for collecting and outputting image information collected from the cameras and providing that information in the form of identified navigation states to the driving policy module 803, which may constitute a trained navigation system that has been trained by machine learning methods such as supervised learning or reinforcement learning. Based on the navigation state information provided by the detection module 801 to the driving policy module 803, the driving policy module 803 may generate a desired navigation action to be performed by the host vehicle in response to the identified navigation states (e.g., by implementing the option graph approach described above).
[0295] In some embodiments, at least one processing device, such as using control module 805, can directly translate the desired navigation behavior into navigation commands. However, in other embodiments, hard constraints can be applied to test the desired navigation behavior provided by driving policy module 803 against various predetermined navigation constraints that may be implicated by the scene and the desired navigation behavior. For example, if driving policy module 803 outputs a desired navigation behavior that causes the host vehicle to follow trajectory 1212, that navigation behavior can be tested against one or more hard constraints related to various aspects of the host vehicle's environment. For example, captured image 1201 can determine curve 1213, pedestrian 1215, target vehicle 1217, and static objects (e.g., an overturned box) present in the scene. Each of these can be associated with one or more hard constraints. For example, curve 1213 can be associated with a static constraint that prevents the host vehicle from navigating into or over the curve onto sidewalk 1214. The curve 1213 may also be associated with a road barrier envelope that defines a predetermined distance (e.g., a buffer zone) extending from and along the curve (e.g., 0.1 meter, 0.25 meter, 0.5 meter, 1 meter, etc.) that defines a no-navigation area for the host vehicle. Of course, static constraints may be associated with other types of roadside boundaries (e.g., guardrails, concrete pillars, cones, pylons, any other type of roadside barrier).
[0296] It should be noted that distance and ranging can be determined by any suitable method. For example, in some embodiments, distance information can be provided by an onboard radar and / or lidar system. Alternatively, or in addition, distance information can be derived by analyzing one or more images captured of the host vehicle's environment. For example, the number of pixels of a recognized object represented in an image can be determined and compared to the known field of view and focal length geometry of the image capture device to determine scale and distance. For example, velocity and acceleration can be determined by observing the change in scale between objects from image to image over a known time interval. This analysis can indicate how quickly an object is moving away from or approaching the host vehicle, as well as the direction of movement toward or away from the host vehicle. Traversal velocity can be determined by analyzing the positional change in the object's X coordinate from one image to another over a known time period.
[0297] The pedestrian 1215 may be associated with a pedestrian envelope that defines a buffer zone 1216. In some cases, the imposed strict constraint may prohibit the host vehicle from navigating within a distance of 1 meter from the pedestrian 1215 (in any direction relative to the pedestrian). The pedestrian 1215 may also define the location of a pedestrian influence zone 1220. This influence zone may be associated with a constraint that limits the speed of the host vehicle within the influence zone. The influence zone may extend 5 meters, 10 meters, 20 meters, etc. from the pedestrian 1215. A different speed limit may be associated with each class of influence zone. For example, within the area 1 meter to 5 meters from the pedestrian 1215, the host vehicle may be limited to a first speed (e.g., 10 mph, 20 mph, etc.) that may be less than the speed limit within the pedestrian influence zone extending from 5 meters to 10 meters. Any class for the various stages of the influence zone may be used. In some embodiments, the first stage can be narrower than 1 meter to 5 meters and may span only 1 meter to 2 meters. In other embodiments, the first stage of the influence zone may extend from 1 meter (the boundary of the non-navigable zone around the pedestrian) to a distance of at least 10 meters. The second stage may then extend from 10 meters to at least about 20 meters. The second stage may be associated with a maximum travel speed of the host vehicle that exceeds the maximum travel speed associated with the first stage of the pedestrian influence zone.
[0298] One or more static object constraints may also be implicated by the detected scene in the host vehicle's environment. For example, in image 1201, at least one processing device may detect a static object, such as a box 1219, in the road. The detected static object may include various objects, such as a tree, a pole, a road sign, and / or an object in the road. One or more predefined navigation constraints may be associated with the detected static object. For example, such constraints may include a static object envelope that defines a buffer zone around the object within which host vehicle navigation may be prohibited. At least a portion of the buffer zone may extend a predetermined distance from the edge of the detected static object. For example, in the scene represented by image 1201, a buffer zone of at least 0.1 meters, 0.25 meters, 0.5 meters, or more may be associated with box 1219, such that the host vehicle passes to the right or left of the box by at least some distance (e.g., the buffer zone distance) to avoid a collision with the detected static object.
[0299] The predefined hard constraints may also include one or more target vehicle constraints. For example, a target vehicle 1217 may be detected in image 1201. One or more hard constraints may be used to ensure that the host vehicle does not collide with the target vehicle 1217. In some cases, the target vehicle envelope may be associated with a single buffer zone distance. For example, the buffer zone may be defined by a distance of one meter surrounding the target vehicle in all directions. The buffer zone may define an area extending at least one meter from the target vehicle within which the host vehicle is prohibited from navigating.
[0300] However, the envelope surrounding the target vehicle 1217 need not be defined by a fixed buffer distance. In some cases, the predetermined strict constraint associated with the target vehicle (or any other movable object detected in the host vehicle's environment) may depend on the orientation of the host vehicle relative to the detected target vehicle. For example, in some cases, the distance of a longitudinal buffer zone (e.g., extending from the target vehicle to the front or rear of the host vehicle, such as when the host vehicle is traveling toward the target vehicle) may be at least 1 meter. The distance of a lateral buffer zone (e.g., extending from the target vehicle to either side of the host vehicle, such as when the host vehicle is traveling in the same or opposite direction as the target vehicle, such that the side of the host vehicle passes directly beside the side of the target vehicle) may be at least 0.5 meter.
[0301] As explained above, other constraints may also be involved by detecting target vehicles or pedestrians in the host vehicle's environment. For example, the predicted trajectories of the host vehicle and target vehicle 1217 may be considered, and when the two trajectories intersect (e.g., at intersection 1230), the hard constraint is:
number
[0302] Other hard constraints may also be used. For example, in at least some cases, a maximum deceleration rate of the host vehicle may be used. This maximum deceleration rate may be determined based on the detected distance to a target vehicle that is following the host vehicle (e.g., using images collected from a rear-facing camera). Hard constraints may include mandatory stops at detected crosswalks or railroad crossings or other applicable constraints.
[0303] If analysis of the scene in the host vehicle's environment indicates that one or more predefined navigation constraints may be implicated, those constraints can be imposed on one or more planned navigation operations of the host vehicle. For example, if analysis of the scene results in the driving policy module 803 returning a desired navigation operation, the desired navigation operation can be tested against one or more implicated constraints. If it is determined that the desired navigation operation violates any aspect of the implicated constraints (e.g., if the desired navigation operation brings the host vehicle within 0.7 meters of the pedestrian 1215 when a predefined strict constraint requires the host vehicle to stay at least 1.0 meter from the pedestrian 1215), at least one modification to the desired navigation operation can be made based on the one or more predefined navigation constraints. Adjusting the desired navigation operation in this manner can result in the host vehicle's actual navigation operation complying with the constraints implicated by the particular scene detected in the host vehicle's environment.
[0304] After determining the actual navigation behavior of the host vehicle, the navigation behavior can be implemented by causing adjustment of at least one navigation actuator of the host vehicle in response to the determined actual navigation behavior of the host vehicle, which may include at least one of a steering mechanism, a brake, or an accelerator of the host vehicle.
[0305] Prioritized Constraints
[0306] As explained above, various strict constraints can be used with the navigation system to ensure safe operation of the host vehicle. Constraints can include, among other things, a minimum safe driving distance to a pedestrian, target vehicle, road barrier, or detected object, a maximum travel speed when passing through a detected pedestrian's influence zone, or a maximum deceleration rate for the host vehicle. These constraints can be imposed by a trained system that is trained based on machine learning (supervised, reinforcement, or a combination thereof), but can also be useful in untrained systems (e.g., using algorithms that directly address expected situations that arise within the host vehicle's environmental scene).
[0307] In any case, there may be a hierarchy of constraints. In other words, some navigation constraints may take precedence over other navigation constraints. Thus, if a situation arises in which no navigation maneuver is available that would satisfy all involved constraints, the navigation system can determine an available navigation maneuver that first fulfills the highest priority constraint. For example, the system may have the vehicle avoid pedestrians first, even if navigating to avoid pedestrians results in a collision with another vehicle or object detected in the road. In another example, the system may have the vehicle navigate a curve to avoid pedestrians.
[0308] FIG. 13 shows a flowchart illustrating an algorithm for implementing a hierarchy of involved constraints determined based on an analysis of a scene in the host vehicle's environment. For example, in step 1301, at least one processor associated with the navigation system (e.g., an EyeQ processor, etc.) can receive multiple images representing the host vehicle's environment from an onboard camera of the host vehicle. In step 1303, a navigation state associated with the host vehicle can be determined by analyzing the one or more images representing the scene in the host vehicle's environment. For example, the navigation state may indicate, among other characteristics of the scene, that the host vehicle is traveling along a two-lane road 1210, as shown in FIG. 12, that a target vehicle 1217 is proceeding through an intersection ahead of the host vehicle, that a pedestrian 1215 is waiting to cross the road along which the host vehicle is traveling, and that there is an object 1219 ahead in the host vehicle's lane.
[0309] In step 1305, one or more navigation constraints implicated by the host vehicle's navigation state may be determined. For example, the at least one processing device may analyze a scene in the host vehicle's environment represented by one or more captured images, and then determine one or more navigation constraints implicated by objects, vehicles, pedestrians, etc. recognized by image analysis of the captured images. In some embodiments, the at least one processing device may determine at least a first predefined navigation constraint and a second predefined navigation constraint implicated by the navigation state, where the first predefined navigation constraint may be different from the second predefined navigation constraint. For example, the first navigation constraint may relate to one or more target vehicles detected in the host vehicle's environment, and the second navigation constraint may relate to pedestrians detected in the host vehicle's environment.
[0310] In step 1307, the at least one processing device may determine a priority associated with the constraints identified in step 1305. In the illustrated example, a second predefined navigation constraint related to pedestrians may have a higher priority than a first predefined navigation constraint related to the target vehicle. While the priority associated with a navigation constraint may be determined or assigned based on various factors, in some embodiments, the priority of a navigation constraint may relate to its relative importance from a safety perspective. For example, while it may be important to obey or satisfy all implemented navigation constraints in as many situations as possible, some constraints may be associated with a stronger safety risk than other constraints and therefore may be assigned a higher priority. For example, a navigation constraint requiring the host vehicle to maintain at least one meter of distance from pedestrians may have a higher priority than a constraint requiring the host vehicle to maintain at least one meter of distance from the target vehicle. This may be because a collision with a pedestrian may have more serious consequences than a collision with another vehicle. Similarly, maintaining a gap between the host vehicle and the target vehicle may have a higher priority than constraints requiring the host vehicle to avoid boxes in the road, drive below a certain speed over speed bumps, or expose the occupants of the host vehicle to less than a maximum acceleration level.
[0311] While the driving policy module 803 is designed to maximize safety by satisfying navigation constraints implied by a particular scene or navigation state, it may be physically impossible to satisfy all of the constraints implied by a given situation. In such a situation, the priority of each implied constraint may be used to determine which of the implied constraints should be satisfied first, as shown in step 1309. Continuing with the above example, in a situation where both the pedestrian clearance constraint and the target vehicle clearance constraint cannot be satisfied and only one of the constraints can be satisfied, the higher priority of the pedestrian clearance constraint may result in the constraint being satisfied before attempting to maintain the gap to the target vehicle. Thus, in a normal situation, when both the first predefined navigation constraint and the second predefined navigation constraint can be satisfied, as shown in step 1311, the at least one processing device may determine, based on the identified navigation state of the host vehicle, a first navigation operation for the host vehicle that satisfies both the first predefined navigation constraint and the second predefined navigation constraint. However, in other situations where all involved constraints cannot be satisfied, as shown in step 1313, if both the first predefined navigation constraint and the second predefined navigation constraint cannot be satisfied, the at least one processing device may determine a second navigation operation for the host vehicle based on the identified navigation state that satisfies the second predefined navigation constraint (i.e., the constraint with a higher priority) but does not satisfy the first predefined navigation constraint (which has a lower priority than the second navigation constraint).
[0312] Next, in step 1315, the at least one processing device may cause adjustment of at least one navigation actuator of the host vehicle in response to the determined first navigation operation or the determined second navigation operation for the host vehicle to implement the determined navigation operation for the host vehicle. As with the previous example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.
[0313] Relaxation of constraints
[0314] As discussed above, navigation constraints can be imposed for safety. Constraints can include, among other things, a minimum safe driving distance to a pedestrian, target vehicle, road barrier, or detected object, a maximum travel speed when passing within a detected pedestrian's influence zone, or a maximum deceleration rate for the host vehicle. These constraints can be imposed in a learning navigation system or a non-learning navigation system. In certain situations, these constraints can be relaxed. For example, if the host vehicle slows down or stops near a pedestrian and proceeds slowly to communicate its intent to pass the pedestrian, the pedestrian's response can be detected from the acquired imagery. If the pedestrian's response is to remain still or cease moving (and / or eye contact with the pedestrian is detected), it can be understood that the pedestrian recognizes the navigation system's intent to pass them. In such situations, the system can relax one or more default constraints and enforce less strict constraints (e.g., allowing the vehicle to navigate within 0.5 meters of the pedestrian rather than within a more strict 1-meter boundary).
[0315] FIG. 14 shows a flowchart for implementing control of a host vehicle based on relaxing one or more navigation constraints. In step 1401, at least one processing device may receive a plurality of images representing an environment of the host vehicle from a camera associated with the host vehicle. Analyzing the images in step 1403 may enable identifying a navigation state associated with the host vehicle. In step 1405, the at least one processor may determine navigation constraints associated with the navigation state of the host vehicle. The navigation constraints may include a first predefined navigation constraint implicated by at least one aspect of the navigation state. In step 1407, analyzing the plurality of images may determine the presence of at least one navigation constraint relaxing factor.
[0316] Navigation constraint mitigating factors may include any suitable indicator that one or more navigation constraints may be suspended, modified, or otherwise relaxed in at least one aspect. In some embodiments, at least one navigation constraint mitigating factor may include a determination (based on image analysis) that a pedestrian's eyes are looking in the direction of the host vehicle. In that case, the pedestrian may be more safely assumed to be aware of the host vehicle. As a result, there may be greater confidence that the pedestrian will not engage in unexpected behavior that would cause the pedestrian to move into the path of the host vehicle. Other constraint mitigating factors may also be used. For example, at least one navigation constraint mitigating factor may include a pedestrian determined to be not moving (e.g., one estimated to be unlikely to enter the path of the host vehicle) or a pedestrian determined to be slowing in their movement. Navigation constraint mitigating factors may also include more complex behaviors, such as a pedestrian determined to be not moving when the host vehicle stops and then resuming movement. In such a situation, the pedestrian may be assumed to understand that the host vehicle has the right-of-way, and the stopping pedestrian may indicate their intention to yield to the host vehicle. Other situations that may cause one or more constraints to be relaxed include the type of curb (e.g., a low curb or a gentle slope may allow for a relaxed distance constraint), the absence of pedestrians or other objects on the sidewalk, vehicles that do not have their engines running may have a relaxed distance, or situations in which pedestrians are facing away from and / or away from the area the host vehicle is heading into.
[0317] Upon identifying the presence of a navigation constraint relaxer (e.g., in step 1407), a second navigation constraint may be determined or developed in response to detecting the constraint relaxer. This second navigation constraint may differ from the first navigation constraint and may include at least one relaxed characteristic compared to the first navigation constraint. The second navigation constraint may include a newly generated constraint based on the first navigation constraint, the newly generated constraint including at least one modification that relaxes the first constraint in at least one respect. Alternatively, the second constraint may constitute a pre-defined constraint that is less strict than the first navigation constraint in at least one respect. In some embodiments, this second constraint may be reserved for use only in situations where a constraint relaxer is identified within the host vehicle's environment. Whether the second constraint is newly generated or selected entirely or partially from a set of available pre-defined constraints, application of the second navigation constraint in place of a stricter first navigation constraint (which would be applied absent detection of the associated navigation constraint relaxer) may be referred to as constraint relaxation and may be implemented in step 1409.
[0318] If at least one constraint relaxation factor is detected in step 1407 and at least one constraint is relaxed in step 1409, a navigation action for the host vehicle may be determined in step 1411. The navigation action for the host vehicle may be based on the identified navigation state and may satisfy the second navigation constraint. In step 1413, the navigation action may be implemented by causing adjustment of at least one navigation actuator of the host vehicle in response to the determined navigation action.
[0319] As discussed above, the use of navigation constraints and relaxed navigation constraints can be used with a trained (e.g., by machine learning) navigation system or an untrained navigation system (e.g., a system programmed to respond with predetermined actions in response to specific navigation conditions). When using a trained navigation system, the availability of relaxed navigation constraints for a particular navigation situation may represent a mode switch from the response of the trained system to the response of the untrained system. For example, a trained navigation network may determine an original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle may be different from the navigation action that satisfies the first navigation constraint. Rather, the action taken may satisfy a second, more relaxed navigation constraint and may be an action developed by an untrained system (e.g., in response to detecting specific conditions in the host vehicle's environment, such as the presence of a navigation constraint relaxer).
[0320] There are many examples of navigation constraints that may be relaxed in response to detecting constraint relaxers within the host vehicle's environment. For example, if a default navigation constraint includes a buffer zone associated with a detected pedestrian, and at least a portion of the buffer zone extends a predetermined distance from the detected pedestrian, a relaxed navigation constraint (either newly developed, recalled from memory from a predetermined set, or generated as a relaxed version of an existing constraint) may include a different or modified buffer zone. For example, the different or modified buffer zone may have a shorter distance to the detected pedestrian than the original or unmodified buffer zone to the detected pedestrian. As a result, if an appropriate constraint relaxer is detected within the host vehicle's environment, the host vehicle may be permitted to navigate closer to the detected pedestrian, taking into account the relaxed constraint.
[0321] Relaxed characteristics of navigation constraints may include a reduced width of a buffer zone associated with at least one pedestrian, as described above, but may also include a reduced width of a buffer zone associated with a target vehicle, a detected object, a roadside barrier, or any other object detected in the host vehicle's environment.
[0322] The at least one relaxed characteristic may also include other types of modifications in the navigation constraint characteristics. For example, the relaxed characteristic may include a speed increase associated with the at least one predefined navigation constraint. The relaxed characteristic may also include an increase in the maximum allowable deceleration / acceleration associated with the at least one predefined navigation constraint.
[0323] As noted above, constraints may be relaxed in certain situations, while navigation constraints may be augmented in other situations. For example, in some situations, the navigation system may determine that conditions justify augmenting the normal set of navigation constraints. Such augmentation may include adding new constraints to the set of default constraints or adjusting one or more aspects of the default constraints. This addition or adjustment may result in more conservative navigation relative to the set of default constraints applicable under normal driving conditions. Conditions that may justify augmenting constraints may include sensor failure, adverse environmental conditions (rain, snow, fog, or other conditions associated with reduced visibility or reduced vehicle traction), etc.
[0324] FIG. 15 shows a flowchart for implementing control of a host vehicle based on augmenting one or more navigation constraints. In step 1501, at least one processing device may receive a plurality of images representing the host vehicle's environment from a camera associated with the host vehicle. Analyzing the images in step 1503 may enable identifying a navigation state associated with the host vehicle. In step 1505, the at least one processor may determine navigation constraints associated with the host vehicle's navigation state. The navigation constraints may include a first predefined navigation constraint implicated by at least one aspect of the navigation state. In step 1507, analyzing the plurality of images may determine the presence of at least one navigation constraint augmentation factor.
[0325] The involved navigation constraints may include any of the navigation constraints discussed above (e.g., with respect to FIG. 12 ) or any other suitable navigation constraints. Navigation constraint augmentation factors may include any indicator that one or more navigation constraints may be supplemented / augmented in at least one aspect. Supplementing or augmenting navigation constraints may be done on a set-by-set basis (e.g., by adding a new navigation constraint to a set of predefined constraints) or on a constraint-by-constraint basis (e.g., modifying a particular constraint so that the modified constraint is more restrictive than the original, or adding a new constraint corresponding to a predefined constraint, where the new constraint is more restrictive in at least one aspect). Additionally or alternatively, supplementing or augmenting navigation constraints may refer to selecting from among a set of predefined constraints based on a hierarchy. For example, a set of augmented constraints may be provided for selection based on whether navigation augmentation factors are detected in the environment of or for the host vehicle. Under normal conditions, when no augmentation factors are detected, the involved navigation constraints may be derived from constraints applicable to normal conditions. On the other hand, if one or more constraint augmentation factors are detected, the constraints involved can be derived from augmented constraints that are generated or predetermined for the one or more augmentation factors, which can be more restrictive in at least one aspect than the corresponding constraints applicable under normal conditions.
[0326] In some embodiments, the at least one navigation constraint augmentation factor may include detecting (e.g., based on image analysis) the presence of ice, snow, or water on a road surface in the host vehicle's environment. This determination may be based, for example, on detecting areas that are more reflective than expected on a dry road (e.g., indicating ice or water on the road), white areas on the road surface indicating the presence of snow, shadows on the road consistent with the presence of longitudinal grooves in the road (e.g., tire tracks in snow), water droplets or small flakes of ice / snow on the host vehicle's windshield, or any other suitable indicator of the presence of water or ice / snow on the road surface.
[0327] At least one navigation constraint augmentation factor may also include detecting small particles on the exterior surface of the host vehicle's windshield. Such small particles may impair image quality of one or more image capture devices associated with the host vehicle. While described with respect to the host vehicle's windshield in association with a camera mounted behind the host vehicle's windshield, detecting small particles on other surfaces associated with the host vehicle (e.g., a camera lens or lens cover, a headlight lens, a rear window, a taillight lens, or any other surface of the host vehicle visible to (or detected by a sensor of) an image capture device) may also indicate the presence of a navigation constraint augmentation factor.
[0328] Navigation constraint augmentation factors may also be detected as characteristics of one or more image capture devices. For example, a detected degradation in image quality of one or more images captured by an image capture device (e.g., a camera) associated with the host vehicle may also constitute a navigation constraint augmentation factor. The degradation in image quality may be associated with a hardware failure or partial hardware failure associated with the image capture device or an assembly associated with the image capture device. Such degradation in image quality may also be caused by environmental conditions. For example, the presence of smoke, fog, rain, snow, etc. in the air surrounding the host vehicle may contribute to degradation in image quality with respect to roads, pedestrians, target vehicles, etc. that may be in the host vehicle's environment.
[0329] Navigation constraint augmentation factors may also relate to other aspects of the host vehicle. For example, in some circumstances, navigation constraint augmentation factors may include detected failure or partial failure of a system or sensor associated with the host vehicle. Such augmentation factors may include, for example, detected failure or partial failure of a speed sensor, GPS receiver, accelerometer, camera, radar, lidar, brakes, tires, or any other system associated with the host vehicle that may affect the host vehicle's ability to navigate in relation to navigation constraints associated with the host vehicle's navigation state.
[0330] If the presence of a navigation constraint augmentation factor is identified (e.g., in step 1507), a second navigation constraint may be determined or developed in response to detecting the constraint augmentation factor. This second navigation constraint may differ from the first navigation constraint and may include at least one characteristic that augments the first navigation constraint. The second navigation constraint may be more restrictive than the first navigation constraint because detecting a constraint augmentation factor in the host vehicle's environment or associated with the host vehicle may indicate that at least one navigation capability of the host vehicle may be degraded compared to normal operating conditions. Such degraded capability may include reduced road traction (e.g., ice, snow, or water on the road, low tire pressure, etc.), impaired visibility (e.g., rain, snow, dust, smoke, fog, etc. that reduces captured image quality), impaired detection capability (e.g., sensor failure or partial failure, degraded sensor performance, etc.), or any other degradation in the host vehicle's ability to navigate in response to the detected navigation conditions.
[0331] If at least one constraint augmentation factor is detected in step 1507 and at least one constraint is augmented in step 1509, a navigation action for the host vehicle may be determined in step 1511. The navigation action for the host vehicle may be based on the identified navigation state and may satisfy the second navigation (i.e., augmented) constraint. In step 1513, the navigation action may be implemented by causing adjustment of at least one navigation actuator of the host vehicle in response to the determined navigation action.
[0332] As discussed above, the use of navigation constraints and augmented navigation constraints can be used with a trained (e.g., by machine learning) navigation system or an untrained navigation system (e.g., a system programmed to respond with a predetermined action in response to a particular navigation situation). When using a trained navigation system, the availability of augmented navigation constraints for a particular navigation situation can represent a mode switch from the response of the trained system to the response of the untrained system. For example, a trained navigation network can determine an original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle can be different from the navigation action that satisfies the first navigation constraint. Rather, the action taken can satisfy an augmented second navigation constraint and can be an action developed by an untrained system (e.g., in response to detecting a particular condition in the host vehicle's environment, such as the presence of a navigation constraint augmentation factor).
[0333] There are many examples of navigation constraints that may be generated, supplemented, or augmented in response to detecting constraint augmentation factors within the host vehicle's environment. For example, if a default navigation constraint includes a buffer zone associated with a detected pedestrian, object, vehicle, etc., and at least a portion of the buffer zone extends a predetermined distance from the detected pedestrian / object / vehicle, an augmented navigation constraint (either newly developed, recalled from memory from a predetermined set, or generated as an augmented version of an existing constraint) may include a different or modified buffer zone. For example, the different or modified buffer zone may have a longer distance to the pedestrian / object / vehicle than the original or unmodified buffer zone for the detected pedestrian / object / vehicle. As a result, if an appropriate constraint augmentation factor is detected within or relative to the host vehicle's environment, the host vehicle may be forced to navigate further away from the detected pedestrian / object / vehicle, taking the augmented constraint into account.
[0334] The at least one augmented characteristic may also include other types of modifications in the navigation constraint characteristics. For example, the augmented characteristic may include a speed reduction associated with the at least one predefined navigation constraint. The augmented characteristic may also include a reduction in the maximum allowable deceleration / acceleration associated with the at least one predefined navigation constraint.
[0335] Navigation based on long-term planning
[0336] In some embodiments, the disclosed navigation system can not only respond to detected navigational conditions in the host vehicle's environment, but also determine one or more navigational actions based on long-term planning. For example, the system can consider the potential impact on future navigational conditions of one or more navigational actions available as options for navigating with respect to the detected navigational conditions. Considering the effect of available actions on future conditions may enable the navigation system to determine navigational actions based not only on the currently detected navigational conditions but also on long-term planning. Navigation using long-term planning techniques may be particularly applicable when one or more reward functions are used by the navigation system as a technique for selecting navigational actions from among available options. Potential rewards may be analyzed with respect to available navigational actions that can be performed in response to the detected current navigational conditions of the host vehicle. However, potential rewards may also be analyzed in relation to actions that can be performed in response to future navigational conditions that are predicted to result from the available actions for the current navigational conditions. As a result, in some cases, the disclosed navigation system may select a navigational action in response to the detected navigational conditions, even if the selected navigational action may not result in the highest reward among the available actions that can be performed in response to the current navigational conditions. This may be particularly true when the system determines that the selected action may result in a future navigation state that triggers one or more potential navigation actions that offer higher rewards than the selected action or, in some cases, any of the actions available for the current navigation state. This principle may be more simply expressed as taking a less advantageous action now in order to result in a higher-reward option in the future. Thus, the disclosed navigation system, capable of long-term planning, may select a suboptimal short-term action when long-term predictions indicate that a short-term loss in reward may result in an increase in long-term reward.
[0337] In general, autonomous driving applications may involve a series of planning problems in which a navigation system can determine immediate actions to optimize a long-term objective. For example, when a vehicle faces a merging situation at a roundabout, the navigation system can determine an immediate acceleration or braking command to begin navigation into the roundabout. The immediate action in response to the navigation state detected at the roundabout may include an acceleration or braking command depending on the detected state, while the long-term objective is to successfully merge, and the long-term effect of the selected command is the success or failure of the merge. The planning problem can be addressed by decomposing the problem into two phases. First, supervised learning can be applied to predict the near future based on the present (assuming that the predictors are differentiable with respect to the current representation). Second, a recurrent neural network can be used to model the complete trajectory of the agent, with unaccounted factors modeled as (additional) input nodes. This may enable the solution to the long-term planning problem using supervised learning methods and direct optimization on the recurrent neural network. Such an approach may also enable the learning of robust policies by incorporating adversarial elements into the environment.
[0338] Two of the most fundamental elements of an autonomous driving system are sensing and planning. Sensing deals with finding a compact representation of the current state of the environment, while planning deals with deciding which actions to take to optimize a future objective. Supervised machine learning methods are useful for solving the sensing problem. Machine learning algorithm frameworks can also be used for the planning part, especially reinforcement learning (RL) frameworks as described above.
[0339] RL can be executed through a series of successive rounds. In round t, the planner (also known as the agent or driving policy module 803) creates a state s representing the agent and the environment. t∈S. Then the planner takes action a t ∈A. After performing the action, the agent receives an immediate reward
number
number
number
[0340] Supervised learning (SL) can be considered a special case of RL, where s is chosen from some distribution over S. t is sampled, and the reward function is r t =-l(a t ,y t ), where l is the loss function, and the learning side can t y, the (possibly noisy) value of the optimal action to take when tObserve the value of . There may be some differences between the model of general RL and the specific case of SL, and these differences may make the general RL problem more difficult.
[0341] In some SL situations, the actions (or predictions) made by the learner may not have any effect on the environment. t+1 and a t and are independent. This can have two important implications. First, in SL, the samples (s1, y1),...,(s m ,y m ) can be collected in advance, and only then can we begin to search for a policy (or predictor) that has good accuracy for the sample. In contrast, in RL, t+1 Typically, depends on the action taken (and on previous states), which in turn depends on the policy used to generate the action. This ties the data generation process to the policy learning process. Second, in SL, actions do not affect the environment, so the effect of a on the performance of π is small. t The contribution of the selection of a is local. t only affects the value of the immediate reward. In contrast, in RL, actions taken in round t can have long-term effects on reward values in future rounds.
[0342] In SL, the reward shape r t =-l(a t ,y t ) and knowledge of the "correct" solution yt, a t can provide complete knowledge of the rewards for all possible choices, which is tThis may allow for the calculation of the derivative of the reward with respect to . In contrast, in RL, the "one-shot" value of the reward may be all that can be observed for a particular choice of action taken. This can be called "bandit" feedback. In an RL-based system, if only "bandit" feedback is available, the system may not always know whether the action taken was the best action to take, which is one of the most important reasons why "exploration" is necessary as part of long-term navigation planning.
[0343] Many RL algorithms rely, at least in part, on mathematically explicit models of Markov decision processes (MDPs). The Markov assumption states that s t and a t Given s t+1 The distribution of π is completely determined. This yields a closed form for the cumulative reward of a given policy in terms of its stationary distribution over the states of the MDP. The stationary distribution of a policy can be expressed as a solution to a linear programming problem. This gives rise to two families of algorithms: 1) optimization with respect to the primal problem, which can be called policy search, and 2) optimization with respect to the dual problem, whose variable is called the value function Vπ. The value function determines the expected cumulative reward if the MDP starts from an initial state s, from which actions are chosen according to π. The quantity involved is the state-action value function Q π (s, a), which determines the cumulative reward given starting from state s, immediately chosen action a, and actions chosen from there according to π. This Q-function can lead to a characterization of the optimal policy (using the Bellman equation). Specifically, this Q-function can be shown to be a deterministic function from S to A (indeed, the optimal policy can be characterized as a "greedy" policy with respect to the optimal Q-function).
[0344] One potential advantage of MDP models is that they allow us to combine the future with the present using a Q-function. For example, given that the host vehicle is currently in state s, Q πThe value of (s, a) may indicate the effect of performing action a on the future. Thus, the Q-function provides a local measure of the quality of action a, which can make RL problems more similar to SL scenarios.
[0345] Many RL algorithms approximate the V-function or Q-function in some way. Value iteration algorithms, such as Q-learning algorithms, can exploit the fact that the V-function and Q-function of the optimal policy can be fixed points of some operators obtained from the Bellman equation. Actor-critic policy iteration algorithms aim to learn a policy in an iterative way, where at iteration t, the "critic" is
number
[0346] Despite the mathematical simplicity of MDPs and the convenience of switching to a Q-function representation, this approach can have some limitations. For example, in some cases, an approximation of a Markovian behavioral state may be all that can be found. Furthermore, state transitions may depend not only on the agent's behavior but also on the behavior of other players in the environment. For example, in the ACC example above, the dynamics of the autonomous vehicle may be Markovian, but the next state may depend on the behavior of drivers of other cars, which is not necessarily Markovian. One possible solution to this problem is to use partially observable MDPs, in which there are Markovian states, but what is visible are observations that are distributed according to hidden states.
[0347] A more straightforward approach could consider game-theoretic generalizations of MDPs (e.g., in the framework of stochastic games). Indeed, algorithms for MDPs can be generalized to multi-agent games (e.g., minimax Q-learning or Nash Q-learning). Other approaches could include explicit modeling of other players and vanishing regret learning algorithms. Learning in a multi-agent setting can be more complex than learning in a single-agent setting.
[0348] A second limitation of the Q-function representation can arise from departing from a tabular setting. A tabular setting is when the number of states and actions is small, so that Q can be represented as a table with |S| rows and |A| columns. However, if the natural representation of S and A involves Euclidean space and the state and action spaces are discretized, the number of states / actions can be exponential in scale. In such cases, adopting a tabular setting may be impractical. Instead, the Q-function can be approximated by some function from a parametric hypothesis class (e.g., neural networks of a particular architecture). For example, a deep Q-network (DQN) learning algorithm can be used. In DQN, the state space can be continuous, but the action space can remain a small, discrete set. Techniques for accommodating continuous action spaces are possible, but they may rely on approximating the Q-function. In either case, the Q-function can be complex and sensitive to noise and therefore difficult to learn.
[0349] A different approach can be to address RL problems using recurrent neural networks (RNNs). In some cases, RNNs can be combined with concepts from multi-agent games and robustness to adversarial environments from game theory. Furthermore, this approach can be explicitly independent of the Markov assumption.
[0350] In the following, we describe in more detail an approach for navigation by prediction-based planning, in which the state space S is defined as:
number
number
[0351] As a first step in this approach, it can be observed that there are interesting problems where the bandit nature of the reward does not matter. For example, the reward value for ACC applications (discussed in more detail below) can be differentiable with respect to the current state and action. In fact, even if the reward is given in a "bandit" fashion,
number
number
number
number
number
[0352] Similar concepts can be used to address the connection between the past and the future. For example:
number
number
number
number
number
number
number
[0353] As discussed above, a learning system can benefit from robustness to hostile environments, such as the host vehicle's environment, which may include multiple other drivers who may behave in unexpected ways. t In a model that does not impose probabilistic assumptions on v t We can consider an environment where μ is chosen in an adversarial way. t We can place constraints on ||μ, which would otherwise make the planning problem difficult or even impossible for the adversary. One natural constraint is that ||μ t It may be necessary to require that || be bounded by a constraint.
[0354] Robustness against hostile environments can be useful in autonomous driving applications. tChoosing , can even accelerate the learning process, as it can focus the learning system towards a robust optimal policy. A simple game can be used to illustrate this concept. The states are
number
number
[0355] Such an approach can be applied to virtually all possible navigation situations. Below we describe the approach applied to one example: adaptive cruise control (ACC). In an ACC problem, the host vehicle may attempt to maintain a sufficient distance to a target vehicle ahead (e.g., 1.5 seconds to the target vehicle). Another goal may be to drive as smoothly as possible while maintaining a desired gap. A model representing this situation can be defined as follows: The state space is
number
number
number
number
[0356] The complete dynamics of the system can be described by the following equation:
number
[0357] This can be written as the sum of two vectors.
number
[0358] The first vector is the predictable part and the second vector is the unpredictable part. The reward for round t is defined as:
number
[0359] The first term may introduce a penalty for non-zero acceleration, thus encouraging smooth driving. The second term is the acceleration of the target vehicle x t Distance to and desired distance
number
[0360] Implementing the techniques outlined above, the host vehicle's navigation system may select an action in response to observed conditions (e.g., by operation of the driving policy module 803 in the navigation system's processing unit 110). The selected action may be based not only on an analysis of rewards associated with available responsive actions for the sensed navigation condition, but also on consideration and analysis of future conditions, potential actions in response to the future conditions, and rewards associated with the potential actions.
[0361] 16 illustrates an algorithmic approach to navigation based on detection and long-term planning. For example, in step 1601, at least one processing device 110 of a navigation system for a host vehicle may receive a plurality of images. These images may capture scenes representative of the host vehicle's environment and may be provided by any of the image capture devices (e.g., cameras, sensors, etc.) described above. Analyzing one or more of these images in step 1603 may enable the at least one processing device 110 to identify a current navigation state associated with the host vehicle (as described above).
[0362] In steps 1605, 1607, and 1609, various potential navigation actions may be determined depending on the sensed navigation conditions (e.g., to complete a merge, smoothly follow a leading vehicle, overtake a target vehicle, avoid an object in the road, slow down for a detected stop sign, avoid an incoming target vehicle, or complete any other navigation action that may further the navigation goals of the system). These potential navigation actions (e.g., from the first navigation action to the Nth available navigation action) may be determined based on the sensed conditions and the long-term goals of the navigation system.
[0363] For each determined potential navigation action, the system may determine an expected reward. The expected reward may be determined according to any of the techniques described above and may include analyzing the particular potential action against one or more reward functions. Expected rewards 1606, 1608, and 1610 may be determined for each (e.g., first, second, and Nth) potential navigation action determined in steps 1605, 1607, and 1609, respectively.
[0364] In some cases, the host vehicle's navigation system may make a selection among available potential actions based on values associated with expected rewards 1606, 1608, and 1610 (or any other type of indicator of expected rewards). For example, in some situations, the action that results in the highest expected reward may be selected.
[0365] In other cases, particularly where the navigation system engages in long-term planning to determine navigation actions for the host vehicle, the system may not select the potential action that will result in the highest expected reward. Rather, the system may look to the future and analyze whether there may be an opportunity to realize a higher reward later if a low-reward action is selected based on the current navigation state. For example, future states may be determined for any or all of the potential actions determined in steps 1605, 1607, and 1609. Each future state determined in steps 1613, 1615, and 1617 may represent a future navigation state that is expected to occur based on the current navigation state as modified by the respective potential action (e.g., the potential actions determined in steps 1605, 1607, and 1609).
[0366] For each future state predicted in steps 1613, 1615, and 1617, one or more future actions (as navigation options available depending on the determined future state) may be determined and evaluated. In steps 1619, 1621, and 1623, for example, a value or any other type of indicator of expected reward associated with one or more of the future actions may be developed (e.g., based on one or more reward functions). The expected reward associated with one or more future actions may be evaluated by comparing the values of the reward functions associated with the respective future actions or by comparing any other indicator associated with the expected reward.
[0367] In step 1625, the navigation system for the host vehicle may select a navigation action for the host vehicle based not only on potential actions identified for the current navigation state (e.g., in steps 1605, 1607, and 1609), but also on expected rewards determined as a result of available future potential actions depending on predicted future states (e.g., determined in steps 1613, 1615, and 1617) and based on comparing expected rewards. The selection in step 1625 may be based on the option and reward analysis performed in steps 1619, 1621, and 1623.
[0368] The selection of a navigation action in step 1625 may be based solely on comparing expected rewards associated with alternative future actions. In this case, the navigation system may select an action for the current state based solely on comparing expected rewards resulting from actions for potential future navigation states. For example, the system may select the potential action identified in steps 1605, 1607, or 1609 associated with the highest future reward value as determined by the analysis in steps 1619, 1621, and 1623.
[0369] The selection of a navigation action in step 1625 may be based solely on comparing current action alternatives (as described above). In this situation, the navigation system may select the potential action identified in steps 1605, 1607, or 1609 that is associated with the highest expected reward 1606, 1608, or 1610. This selection may be made with little or no consideration of future expected rewards for available navigation actions depending on future or anticipated future navigation states.
[0370] On the other hand, in some cases, the selection of a navigation action in step 1625 may be based on comparing expected rewards associated with both future action options and current action options. This may, in fact, be one of the principles of navigation based on long-term planning. To realize potentially higher rewards for subsequent navigation actions expected to become available depending on future navigation states, the expected rewards for future actions may be analyzed to determine whether any may justify selecting an action with a lower reward depending on the current navigation state. As an example, the value or other indicator of expected reward 1606 may indicate the highest expected reward among rewards 1606, 1608, and 1610. On the other hand, expected reward 1608 may indicate the lowest expected reward among rewards 1606, 1608, and 1610. Rather than simply selecting the potential action determined in step 1605 (i.e., the action that yields the highest expected reward 1606), an analysis of the future state, potential future actions, and future rewards may be used in making the selection of a navigation action in step 1625. In one example, it may be determined in step 1621 that the reward identified (in response to at least one future action for the future state determined in step 1615 based on the second potential action determined in step 1607) is likely to be higher than expected reward 1606. Based on this comparison, the second potential action determined in step 1607 may be selected over the first potential action determined in step 1605, even though expected reward 1606 is higher than expected reward 1608. In one example, the potential navigation action determined in step 1605 may include merging in front of the detected target vehicle, while the potential navigation action determined in step 1607 may include merging behind the target vehicle.Although the expected reward 1606 for merging in front of the target vehicle may be higher than the expected reward 1608 associated with merging behind the target vehicle, it may be determined that merging behind the target vehicle may result in a future state where there may be an option for action that offers a higher potential reward than the expected reward 1606, 1608 or other rewards based on actions available depending on the current sensed navigation state.
[0371] The selection among the potential actions in step 1625 may be based on any suitable comparison of expected rewards (or any other metric or indicator of benefit associated with one potential action over another). In some cases, as described above, a second potential action may be selected in preference to a first potential action if the second potential action is predicted to result in at least one future action associated with a higher expected reward than the reward associated with the first potential action. In other cases, more complex comparisons may be used. For example, rewards associated with action options depending on predicted future states may be compared to multiple expected rewards associated with the determined potential actions.
[0372] In some scenarios, actions and expected rewards based on predicted future states may influence the selection of potential actions for the current state if at least one of the future actions is expected to result in a reward higher than any of the expected rewards (e.g., expected rewards 1606, 1608, 1610, etc.) expected as a result of the potential actions for the current state. In some cases, the future action option that results in the highest expected reward (e.g., among the expected rewards associated with potential actions for the sensed current state and the expected rewards associated with potential future action options for potential future navigation states) may be used as a guide for selecting a potential action for the current navigation state. That is, after identifying the future action options that result in the highest expected reward (or a reward above a predetermined threshold, etc.), the potential action that leads to the future state associated with the identified future action that results in the highest expected reward may be selected in step 1625.
[0373] In other cases, the selection of available actions may be based on a determined difference between the expected rewards. For example, if the difference between the expected reward associated with the future action determined in step 1621 and expected reward 1606 exceeds the difference between expected reward 1608 and expected reward 1606 (assuming a difference with a + sign), the second potential action determined in step 1607 may be selected. In another example, if the difference between the expected reward associated with the future action determined in step 1621 and the expected reward associated with the future action determined in step 1619 exceeds the difference between expected reward 1608 and expected reward 1606, the second potential action determined in step 1607 may be selected.
[0374] Several examples have been described for selecting among potential actions for the current navigation state. However, any other suitable comparison technique or criteria can be used to select available actions by long-term planning based on an analysis of actions and rewards spanning predicted future states. In addition, while FIG. 16 shows two layers in the long-term planning analysis (e.g., a first layer that considers rewards resulting from potential actions for the current state and a second layer that considers rewards resulting from alternative future actions depending on predicted future states), analysis based on more layers may be possible. For example, rather than basing the long-term planning analysis on one or two layers, three, four, or even more layers of analysis can be used when selecting among available potential actions depending on the current navigation state.
[0375] After selecting from among the potential actions responsive to the sensed navigation state, in step 1627, the at least one processor may cause adjustment of at least one navigation actuator of the host vehicle responsive to the selected potential navigation action. The navigation actuator may include any suitable device for controlling at least one aspect of the host vehicle. For example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.
[0376] Navigation based on others' inferred aggression
[0377] The target vehicle can be monitored by analyzing the acquired image stream to determine indicators of driving aggressiveness. Aggression is described herein as a qualitative or quantitative parameter, but other characteristics, i.e., detected attention level (potential driver deficiencies, distractions—such as a cell phone or drowsiness), can also be used. In some cases, the target vehicle can be deemed to have a defensive attitude, while in other cases, the target vehicle can be deemed to have a more aggressive attitude. Navigation behavior can be selected or developed based on the indicators of aggressiveness. For example, in some cases, the target vehicle's relative speed, relative acceleration, increase in relative acceleration, pursuit distance, etc., relative to the host vehicle can be tracked to determine whether the target vehicle is aggressive or defensive. If the target vehicle is determined to have a level of aggressiveness above a threshold, for example, the host vehicle may be inclined to yield to the target vehicle. The target vehicle's level of aggressiveness can also be determined based on the target vehicle's determined behavior toward one or more obstacles in the path or near the target vehicle (e.g., a preceding vehicle, an obstacle in the road, a traffic light, etc.).
[0378] As an introduction to this concept, an example experiment involving a host vehicle merging into a roundabout is described, where the navigation goal is to exit through the roundabout. This situation can begin with the host vehicle reaching the roundabout entrance and can end with the host vehicle reaching the roundabout exit (e.g., the second exit). Success can be measured based on whether the host vehicle maintains a safe distance from all other vehicles at all times, whether the host vehicle completes the route as quickly as possible, and whether the host vehicle follows a policy of smooth acceleration. In this discussion, N TTarget vehicles can be randomly placed on a roundabout. To model the mixture of hostile and typical behavior, the target vehicles can be modeled with an "aggressive" driving policy with probability p, so that when the host vehicle attempts to merge in front of the target vehicle, the aggressive target vehicle accelerates. The target vehicles can be modeled with a "defensive" driving policy with probability 1-p, so that the target vehicle slows down and lets the host vehicle merge. In this experiment, p=0.5, and information about the type of other driver may not be provided to the host vehicle's navigation system. The type of other driver may be randomly selected at the start of the episode.
[0379] The navigation state can be represented as the host vehicle (agent)'s velocity and position and the target vehicle's position, velocity, and acceleration. Keeping track of the target's acceleration can be important to distinguish between aggressive and defensive drivers based on the current state. All target vehicles can travel on a one-dimensional curve that outlines the roundabout's path. The host vehicle can travel on its own one-dimensional curve that intersects with the target vehicle's curve at the junction, which is the origin of both curves. To model reasonable driving, the absolute value of all vehicle accelerations can be capped by a constant. Since reverse driving is not allowed, the velocity can also be passed through a ReLU. Note that by not allowing reverse driving, the agent cannot regret its past actions, which may require long-term planning.
[0380] As explained above, the next state s t+1 is the predictable part
number
number
number
[0381] FIG. 18 shows a flowchart illustrating an example of an algorithm for navigating a host vehicle based on the predicted aggressiveness of other vehicles. In the example of FIG. 18, a level of aggressiveness associated with at least one target vehicle can be inferred based on the target vehicle's observed behavior toward objects in the target vehicle's environment. For example, in step 1801, at least one processing device (e.g., processing device 110) of the host vehicle's navigation system can receive multiple images representing the host vehicle's environment from a camera associated with the host vehicle. In step 1803, analyzing one or more of the received images can enable at least one processor to identify a target vehicle (e.g., vehicle 1703) within the host vehicle's environment. In step 1805, analyzing one or more of the received images can enable at least one processing device to identify at least one obstacle for the target vehicle within the host vehicle's environment. The objects may include debris in the road, a stoplight / traffic light, a pedestrian, another vehicle (e.g., a vehicle moving in front of the target vehicle, a parked vehicle, etc.), a box in the road, a road barrier, a curve, or any other type of object that may be encountered in the environment of the host vehicle. In step 1807, analyzing one or more of the received images may enable the at least one processing device to determine at least one navigation characteristic of the target vehicle relative to at least one identified obstacle to the target vehicle.
[0382] To develop an appropriate navigation response to the target vehicle, various navigation characteristics can be used to infer the level of aggressiveness of the detected target vehicle. For example, such navigation characteristics may include the relative acceleration between the target vehicle and at least one identified obstacle, the distance of the target vehicle from the obstacle (e.g., the trailing distance of a target vehicle behind another vehicle), and / or the relative speed between the target vehicle and the obstacle, etc.
[0383] In some embodiments, the navigation characteristics of the target vehicle may be determined based on output from sensors associated with the host vehicle (e.g., radar, speed sensors, GPS, etc.). However, in some cases, the navigation characteristics of the target vehicle may be determined based partially or completely on analyzing images of the host vehicle's environment. For example, the image analysis techniques described, by way of example, above and in U.S. Pat. No. 9,168,868, incorporated herein by reference, may be used to recognize the target vehicle within the host vehicle's environment. Monitoring the position of the target vehicle within captured images over a period of time and / or monitoring the position within captured images of one or more features associated with the target vehicle (e.g., taillights, headlights, bumpers, wheels, etc.) may enable the relative distance, speed, and / or acceleration between the target vehicle and the host vehicle or between the target vehicle and one or more other objects in the host vehicle's environment to be determined.
[0384] The level of aggressiveness of an identified target vehicle can be inferred from any suitable observed navigation characteristic or any combination of observed navigation characteristics of the target vehicle. For example, the determination of aggressiveness can be based on any observed characteristic and one or more predetermined threshold levels or any other suitable qualitative or quantitative analysis. In some embodiments, a target vehicle can be deemed aggressive if it is observed to be pursuing a host vehicle or another vehicle at a distance less than a predetermined aggressive distance threshold. On the other hand, a target vehicle can be deemed defensive if it is observed to be pursuing a host vehicle or another vehicle at a distance greater than the predetermined aggressive distance threshold. The predetermined aggressive distance threshold need not be the same as the predetermined defensive distance threshold. Additionally, either or both of the predetermined aggressive distance threshold and the predetermined defensive distance threshold can include ranges rather than clear boundaries. Furthermore, neither the predetermined aggressive distance threshold nor the predetermined defensive distance threshold need be fixed. Rather, these values or ranges can shift over time, and various thresholds / threshold ranges can be applied based on the observed characteristics of the target vehicle. For example, the threshold applied may depend on one or more other characteristics of the target vehicle. A higher observed relative speed and / or acceleration may justify the application of a larger threshold / threshold range. Conversely, a lower relative speed and / or acceleration, including zero relative speed and / or acceleration, may justify the application of a smaller distance threshold / threshold range when making an offensive / defensive inference.
[0385] The aggressive / defensive inference may also be based on relative speed and / or relative acceleration thresholds. A target vehicle may be considered aggressive if its observed relative speed and / or its relative acceleration relative to another vehicle is above a predetermined level or range. A target vehicle may be considered defensive if its observed relative speed and / or its relative acceleration relative to another vehicle is below a predetermined level or range.
[0386] While the aggressive / defensive determination may be based solely on any observed navigation characteristic, the determination may also depend on any combination of observed characteristics. For example, as noted above, in some cases, a target vehicle may be deemed aggressive based solely on being observed to be following another vehicle at a distance less than a certain threshold or range. However, in other cases, a target vehicle may be deemed aggressive if it is following another vehicle by less than a predetermined amount (which may be the same as or different from the threshold applied when the determination is based solely on distance) and has a relative speed and / or relative acceleration above a predetermined amount or range. Similarly, a target vehicle may be deemed defensive based solely on being observed to be following another vehicle at a distance greater than a certain threshold or range. However, in other cases, a target vehicle may be deemed defensive if it is following another vehicle by more than a predetermined amount (which may be the same as or different from the threshold applied when the determination is based solely on distance) and has a relative speed and / or relative acceleration below a predetermined amount or range. The system 100 may take offensive / defensive action, for example, if the vehicle exceeds 0.5G acceleration or deceleration (e.g., a jerk of 5 m / s), if the vehicle changes lanes or has 0.5G lateral acceleration on a curve, if the vehicle forces another vehicle to do any of the above, if the vehicle changes lanes and forces another vehicle to give way exceeding 0.3G deceleration or a jerk of 3 m / s, and / or if the vehicle changes two lanes without stopping.
[0387] It should be understood that a reference to a quantity above a range may indicate that the quantity exceeds or falls within all values associated with the range. Similarly, a reference to a quantity below a range may indicate that the quantity falls below or falls within all values associated with the range. In addition, while the described examples of making aggressive / defensive inferences have been described with respect to distance, relative acceleration, and relative velocity, any other suitable quantities may be used. For example, calculations may be used to determine time to collision or any indirect indicator of the target vehicle's distance, acceleration, and / or velocity. It should also be noted that while the above examples focus on target vehicles relative to other vehicles, aggressive / defensive inferences may be made by observing the target vehicle's navigation characteristics relative to any other type of obstacle (e.g., pedestrians, road barriers, traffic lights, debris, etc.).
[0388] Returning to the example shown in FIGS. 17A and 17B , as host vehicle 1701 approaches a roundabout, a navigation system including at least one processing device thereof may receive an image stream from a camera associated with the host vehicle. Based on one or more analyses of the received images, any of target vehicles 1703, 1705, 1706, 1708, and 1710 may be identified. Further, the navigation system may analyze one or more navigation characteristics of the identified target vehicles. The navigation system may recognize that a gap between target vehicles 1703 and 1705 represents a first opportunity for a potential merge into the roundabout. The navigation system may analyze target vehicle 1703 to determine an indicator of aggressiveness associated with target vehicle 1703. If target vehicle 1703 is deemed aggressive, the host vehicle's navigation system may decide to yield to vehicle 1703 rather than merge in front of vehicle 1703. On the other hand, if the target vehicle 1703 is deemed defensive, the host vehicle's navigation system may attempt to complete the merging maneuver in front of the vehicle 1703 .
[0389] When the host vehicle 1701 reaches a roundabout, at least one processing device of the navigation system can analyze the captured image to determine navigation characteristics associated with the target vehicle 1703. For example, based on the image, it may be determined that the vehicle 1703 is following the vehicle 1705 at a distance that provides sufficient clearance for the host vehicle 1701 to safely enter. Indeed, it may be determined that the vehicle 1703 is following the vehicle 1705 at a distance that exceeds an aggressive distance threshold; therefore, based on this information, the host vehicle's navigation system may be inclined to identify the target vehicle 1703 as defensive. However, in some situations, multiple navigation characteristics of the target vehicle may be analyzed in making the aggressive / defensive determination, as discussed above. Expanding the analysis, the host vehicle's navigation system may determine that the target vehicle 1703 is following the target vehicle 1705 at a non-aggressive distance, but that the vehicle 1703 has a relative speed and / or relative acceleration with respect to the vehicle 1705 that exceeds one or more thresholds associated with aggressive behavior. Indeed, host vehicle 1701 may determine that target vehicle 1703 is accelerating relative to vehicle 1705, closing the gap between vehicles 1703 and 1705. Based on further analysis of relative speed, acceleration, and distance (as well as the rate at which the gap between vehicles 1703 and 1705 is closing), host vehicle 1701 may determine that target vehicle 1703 is behaving aggressively. Thus, although there may be sufficient gap for the host vehicle to navigate safely, host vehicle 1701 may anticipate that merging in front of target vehicle 1703 will result in an aggressively navigating vehicle directly behind the host vehicle. Furthermore, if host vehicle 1701 merges in front of vehicle 1703, it may be expected that target vehicle 1703 will continue to accelerate toward host vehicle 1701 or proceed toward host vehicle 1701 at a non-zero relative velocity, based on behavior observed by image analysis or other sensor output. Such a situation may be undesirable from a safety standpoint and may also cause discomfort to the occupants of the host vehicle.For this reason, as shown in FIG. 17B, host vehicle 1701 may decide to yield to vehicle 1703 and merge into the roundabout behind vehicle 1703 and ahead of vehicle 1710, which is deemed defensive based on one or more analyses of its navigation characteristics.
[0390] Returning to Figure 18, in step 1809, at least one processing device of the host vehicle's navigation system may determine a navigation action for the host vehicle (e.g., merge in front of vehicle 1710 and behind vehicle 1703) based on the at least one identified navigation characteristic of the target vehicle relative to the identified obstacle. To implement the navigation action (step 1811), the at least one processing device may cause adjustment of at least one navigation actuator of the host vehicle in response to the determined navigation action. For example, the brakes may be applied to yield to vehicle 1703 in Figure 17A, and the accelerator may be applied along with steering the wheels of the host vehicle to turn the host vehicle into the roundabout behind vehicle 1703 as shown in Figure 17B.
[0391] As described in the examples above, navigation of the host vehicle can be based on the navigation characteristics of the target vehicle relative to another vehicle or object. In addition, navigation of the host vehicle can be based solely on the navigation characteristics of the target vehicle without specific reference to another vehicle or object. For example, in step 1807 of FIG. 18 , analyzing multiple images captured from the host vehicle's environment can enable determining at least one navigation characteristic of the identified target vehicle that is indicative of a level of aggressiveness associated with the target vehicle. Navigation characteristics can include speed, acceleration, etc., which do not require reference to another object or target vehicle to make an aggressive / defensive determination. For example, observed acceleration and / or speed associated with the target vehicle that exceeds a predetermined threshold or falls within or exceeds a range of values can indicate aggressive behavior. Conversely, observed acceleration and / or speed associated with the target vehicle that is below a predetermined threshold or falls within or exceeds a range of values can indicate defensive behavior.
[0392] Of course, in some examples, observed navigation characteristics (e.g., position, distance, acceleration, etc.) may be referenced to the host vehicle to make an aggressive / defensive determination. For example, observed navigation characteristics of a target vehicle indicative of a level of aggressiveness associated with the target vehicle may include an increase in relative acceleration between the target vehicle and the host vehicle, a trailing distance of the target vehicle behind the host vehicle, a relative speed between the target vehicle and the host vehicle, etc.
[0393] Navigation based on accident liability constraints
[0394] As described in the previous section, planned navigation actions can be tested against predetermined constraints to ensure compliance with certain rules. In some embodiments, this concept can be extended to the consideration of potential accident liability. As discussed below, a primary goal of autonomous navigation is safety. Because absolute safety may be impossible (e.g., because at least a particular host vehicle under autonomous control cannot control other vehicles around it, but can only control its own actions), using potential accident liability as a consideration in autonomous navigation and as a constraint on planned actions can help ensure that a particular autonomous vehicle does not take any action that is deemed unsafe (e.g., an action that could result in potential accident liability being attributed to the host vehicle). If the host vehicle only takes actions that are determined to be safe and not to result in an accident for which the host vehicle is attributable or liable, a desired level of accident avoidance (e.g., 10 per driving hour) can be achieved. -9 (less than) can be achieved.
[0395] Challenges posed by current approaches to autonomous driving include a lack of safety guarantees (or at least an inability to provide a desired level of safety) and a lack of scalability. Consider the problem of ensuring multi-agent safe driving. Because society is unlikely to tolerate machine-induced traffic fatalities, an acceptable level of safety is paramount to the acceptance of autonomous vehicles. While the goal may be to avoid accidents entirely, this may be impossible because multiple agents are typically involved in accidents and accidents occur only due to the negligence of other agents. For example, as shown in FIG. 19 , host vehicle 1901 is traveling on a multi-lane highway. Host vehicle 1901 can control its own behavior relative to target vehicles 1903, 1905, 1907, and 1909, but cannot control the behavior of the target vehicles surrounding it. As a result, for example, if vehicle 1905 suddenly cuts into the host vehicle's lane on a collision course with the host vehicle, host vehicle 1901 may be unable to avoid an accident with at least one of the target vehicles. To address this challenge, the typical response from autonomous vehicle experts is to use a data-driven approach where safety validation becomes more rigorous as data is collected over more miles driven.
[0396] But to understand the problematic nature of a data-driven approach to safety, let's say the fatality rate caused by an accident per hour of (human) driving is 10 -6 For society to accept machines replacing humans in driving tasks, the fatality rate must be triple digits, i.e., 10 per hour. -9 It is reasonable to assume that the probability of fatality should be reduced to 10. This estimate is similar to the fatality rate assumed for airbags and aviation standards. For example, -9 is the probability that a wing will spontaneously detach from an aircraft in mid-air. However, attempting to guarantee safety using a data-driven approach that provides more confidence proportional to the accumulated number of miles driven is not practical. -9 The amount of data required to guarantee a fatality rate of 109 The total number of miles traveled is proportional to the number of hours of data, which is roughly 30 billion miles. Furthermore, multi-agent systems interact with their environment and are unlikely to be verifiable offline (unless a realistic simulator is available that emulates real-world human driving, with all its richness and complexities, such as reckless driving; however, the problem of validating a simulator is even more difficult than creating a safe autonomous vehicle agent). Any changes to the planning and control software would require the collection of new data on the same scale, which is clearly unmanageable and impractical. Furthermore, developing systems with data will always face a lack of interpretability and explainability of the actions taken; if an autonomous vehicle (AV) has a fatal accident, we must understand why. As a result, a model-based approach to safety is needed, but existing "functional safety" and ASIL requirements in the automotive industry are not designed to address multi-agent environments.
[0397] The second major challenge in developing a safe driving model for autonomous vehicles is the need for scalability. The underlying premise of AVs goes beyond "making the world better" and is based rather on the premise that driverless mobility can be sustained at lower costs than driverless mobility. This premise is always tied to the concept of scalability in terms of supporting mass production of AVs (in the millions) and, more importantly, supporting the trivial incremental cost of enabling them to operate in new cities. Thus, computational and sensing costs are significant, and when AVs are mass-produced, validation costs and the ability to operate "everywhere," rather than just in a select few cities, are also necessary requirements to sustain the business.
[0398] The problem with current approaches is their "brute force" thinking along three axes: (i) the required "computational density," (ii) the way in which high-definition maps are defined and created, and (iii) the required sensor specifications. Brute force approaches defy scalability and shift the burden to a future where infinite on-board computation is ubiquitous, the cost of building and maintaining HD maps is trivial and scalable, and novel, ultra-advanced sensors are developed and commercialized for vehicle grades at trivial cost. While a future where any of the above is realized is indeed plausible, achieving all of the above impacts is likely a low-probability event. Therefore, there is a need to provide a formal model that integrates safety and scalability within AV programs that is both socially acceptable and scalable to support millions of vehicles operating anywhere in the developed world.
[0399] The disclosed embodiments represent a solution that can provide target safety levels (and even exceed safety goals) and scale to systems involving millions (or more) of autonomous vehicles. On the safety front, a model called "Responsibility-Sensitive Safety" (RSS) is introduced, which formalizes the concept of "accidental fault" as interpretable and explainable, incorporating a sense of "responsibility" into the behavior of robotic agents. The RSS definition is agnostic depending on how it is implemented, which is an important feature for furthering the goal of creating a convincing global safety model. RSS is motivated by the idea that agents play asymmetric roles in accidents (as seen in Figure 19), and in such accidents, typically only one of the agents is responsible for the accident and therefore responsible for it. The RSS model also includes a formal treatment of "careful driving" under limited detection conditions where not all agents are always visible (e.g., due to occlusion). One main goal of the RSS model is to ensure that agents never cause accidents for which they are "at fault" or responsible. A model can only be useful if it is accompanied by an efficient policy (e.g., a function that maps "detected states" to actions) that conforms to the RSS. For example, an action that seems harmless now may lead to a catastrophe in the distant future (the "butterfly effect"). The RSS can help to construct a set of local constraints for the near-term future that can guarantee (or at least virtually guarantee) that no future accidents will occur as a result of the host vehicle's actions.
[0400] Another contribution revolves around the introduction of a "semantic" language consisting of units, measurements and action spaces, and a specification of how these are incorporated into AV planning, sensing and operation. To get an idea of what semantics is like, consider how a person taking a driving course in this context might be instructed to think about a "driving policy". These instructions might be non-geometric ("drive 13.7 metres at current speed, then drive 0.8 m / s"). 2(Instructions do not take the form "accelerate at a rate of 1000.") Rather, these instructions are semantic in nature ("follow the car in front" or "overtake the car on the left"). The typical language of human driving policy is in terms of longitudinal and lateral targets rather than geometric units of acceleration vectors. A formal semantic language can be useful in several ways: it relates to the computational complexity of planning, which does not grow exponentially with time and number of agents; it relates to how safety and comfort interact; it relates to how the sensing computation is defined; and it relates to the specification of sensor modalities and how they interact within a fusion methodology. Fusion methodologies (based on semantic languages) can be implemented in a variety of ways, including: 5 While the RSS model only performs offline validation over a dataset of driving data of the order of 10 hours, the -9 This can ensure that a fatality rate of 100% is achieved.
[0401] For example, in a reinforcement learning setting, the number of trajectories to be examined at any given time is limited to 10 4 A Q-function (e.g., a function that evaluates the long-term quality of performing an action a∈A when an agent is in state s∈S) on a semantic space constrained by π(s)=argmax aWe can define a Q(s, a) (which may be a choice of Q(s, a)). The signal-to-noise ratio in this space is likely to be high, allowing effective machine learning methods to successfully model the Q function. For detection computations, semantics can make it possible to distinguish between errors that affect safety and errors that affect driving comfort. We define a PAC model (probabilistically approximately correct (PAC)), borrowing Valiant's PAC learning terminology) for detection that is tied to the Q function and show how to incorporate measurement errors into the plan in a way that is RSS-compliant but allows for optimization of driving comfort. Because other standard error measures, such as error relative to a global coordinate system, may not conform to the PAC detection model, the semantic language may be important to the success of this model's performance aspects. Additionally, the semantic language may be a key enabler for defining HD maps that can be constructed using low-bandwidth detection data and therefore crowdsourced and support scalability.
[0402] In summary, the disclosed embodiments may include a formal model that spans key elements of AVs: sensing, planning, and action. This model may help ensure that there are no incidents where the AV is at fault from a planning perspective. Furthermore, the PAC sensing model allows the described fusion methodology to require only a very reasonable amount of offline data set to comply with the described safety model, even with sensing errors. Furthermore, this model may combine safety and scalability through a semantic language, thereby providing a complete methodology for safe and scalable AVs. Finally, it should be noted that developing an approved safety model that is adopted by industry and regulatory agencies may be a necessary condition for the success of AVs.
[0403] The RSS model may generally follow the classic sense-plan-act robotic control method. The sense system may be responsible for understanding the current state of the host vehicle's environment. The plan portion, which may be referred to as a "driving policy" and may be implemented as a set of hard-coded instructions by a trained system (e.g., a neural network) or a combination, may be responsible for determining what the next best move is in terms of available options to achieve a driving goal (e.g., how to move from the left lane to the right lane to exit a highway). The act portion is responsible for implementing the plan (e.g., a system of actuators and one or more controllers to steer, accelerate, and / or brake the vehicle to perform a selected navigation maneuver, etc.). The embodiments described below focus primarily on the sense and plan portions.
[0404] Accidents can result from sensing or planning errors. Planning is a multi-agent endeavor, as other road users (human and mechanical) react to the AV's actions. The described RSS model is designed to address safety, particularly for the planning portion. This can be called multi-agent safety. With statistical methods, estimation of the probability of planning errors can be performed "online." That is, billions of miles would have to be driven with each software update to yield an acceptable estimated level of planning error frequency. This is clearly infeasible. Alternatively, the RSS model can provide 100% assurance (or virtually 100% assurance) that the planning module will not make mistakes for which the AV is at fault (the concept of "fault" is formally defined). The RSS model can also provide an efficient means for its validation that does not rely on online testing.
[0405] Errors in detection systems may be easier to verify because detection can be independent of vehicle motion, and therefore the probability of a significant detection error can be verified using "offline" data. 9Even collecting offline data over time of driving is difficult. As part of the description of the disclosed detection system, we have described a fusion technique that can be validated using a significantly smaller amount of data.
[0406] The described RSS system can also be scalable to millions of vehicles, for example, the described semantic driving policies and applied safety constraints can meet detection and mapping requirements that even today's technology can scale to millions of vehicles.
[0407] A fundamental building block of such a system is a safety definition, which is a minimum standard that an AV system may need to adhere to. The following technical lemma shows that statistical methods for verifying AV systems are infeasible even for verifying simple claims such as "the system will have N accidents per hour." This implies that only model-based safety definitions are viable tools for verifying AV systems. Lemma 1 Let X be a probability space and A be the event such that Pr(A)=p1<0.1.
number
number
number
[0408] To get a perspective on typical values of such probabilities, -9 If you want to reduce the accident probability by 10% for a specific AV system, -8 Assume that the system only provides probabilities of 10 8 Even if time is taken to operate, there is a certain probability that the validation process will not show that the system is unsafe.
[0409] Finally, note that this problem concerns disabling a single, specific, risky AV system. The complete solution cannot be viewed as a single system, as new versions, bug fixes, and updates will be required. Each change, even a single line of code, creates a new system from the perspective of the verifier. Therefore, a statistically verified solution must be done online across new samples after every small modification or change to account for shifts in the distribution of states observed and reached by the new system. Repeatedly and systematically obtaining such a large number of samples (with a certain probability of still failing to verify the system) is infeasible.
[0410] Furthermore, any statistical claim must be formalized for measurement. Asserting statistical properties over the number of accidents a system has is significantly weaker than claiming that "the system operates in a safe manner." To say that, we need to formally define what safety is.
[0411] Absolute security is impossible
[0412] An action a performed by car c can be considered absolutely safe if it is possible for the action to be followed at some future point without causing an accident. For example, observing a simple driving scenario such as the one shown in Figure 19 reveals that absolute safety is impossible. From the perspective of vehicle 1901, no action can guarantee that surrounding vehicles will not collide with it. It is also impossible to solve the problem by forbidding autonomous vehicles from being in such a situation. Since all highways with more than two lanes will at some point invite this problem, forbidding this scenario would require the vehicle to stay in the garage. These implications seem disappointing at first glance. Nothing is absolutely safe. However, as is evident from the fact that human drivers do not adhere to the requirement of absolute safety, the requirement for absolute safety as defined above may be too strict. Rather, humans act according to a concept of safety that relies on responsibility.
[0413] responsibility-sensitive safety
[0414] An important aspect missing from the concept of absolute safety is the asymmetric nature of most accidents: it is usually one of the drivers who is responsible for the collision and therefore at fault. In the example of Figure 19, for example, the middle vehicle 1901 is not at fault when the left-hand vehicle 1909 suddenly rear-ends it. To formalize the fact that the middle vehicle 1901 is not at fault, the behavior of AV 1901, which stays in its lane, can be considered safe. To do so, we describe a formal concept of "accident fault" or accident liability that can serve as a premise for safe driver law.
[0415] As an example, consider two cars c driving side by side at the same speed along a straight road. f ,c r Consider the simple case of the car in front, c f However, suppose that an obstacle appears on the road and the car suddenly brakes to avoid it. f not keeping a sufficient distance from ris unable to react in time and f crash into the rear of the r It is clear that the situation is different and it is the responsibility of the vehicle behind to maintain a safe distance from the vehicle in front and to be prepared for unexpected but reasonable braking.
[0416] Next, we consider a broader set of scenarios: driving on multi-lane roads where vehicles can freely change lanes, cutting into other vehicles' paths, driving at various speeds, etc. To simplify the following discussion, we consider a straight road in a plane whose horizontal and vertical axes are the x- and y-axes, respectively. This can be achieved under moderate conditions by defining a homomorphism between actual curved roads and straight roads. In addition, we consider a discrete space-time. The definition can help distinguish two intuitively distinct sets of cases: simple ones without significant lateral maneuvers and more complex ones involving lateral movement.
[0417] Definition 1 (Car Zone) The zone of car c is the range [c x,left ,c x,right ]×[±∞], where c x,left , c x,right are the positions of the leftmost and rightmost corners of c.
[0418] Definition 2 (Cut-in) If car c1 (car 2003 in Figures 20A and 20B) does not intersect with the zone of car c0 (car 2001 in Figures 20A and 20B) at time t-1, but does intersect with it at time t, then car c1 cuts into the zone of car c0 at time t.
[0419] A further distinction can be made between the front / rear of a zone. The term "direction of cutting in" can refer to movement in the direction of the boundary of the relevant zone. These definitions can define cases involving lateral movement. In simple cases where such an occurrence does not occur, such as the simple case of one vehicle following another, a safe longitudinal distance is defined.
[0420] Definition 3 (Safe Longitudinal Distance) Car c r (Car 2103) and c rAnother vehicle in the area ahead of f The longitudinal distance 2101 (Fig. 21) between the vehicle (2105) is safe with respect to the response time p, and c f Any braking command a made by |a| max,brake About c r If applies its maximum brake from time p to a complete stop, then c r is c f does not collide with
[0421] Lemma 2 below shows that c r , c f speed, ...
Claims
1. 1. A system for navigating a host vehicle, the system comprising: receiving at least one image representing an environment of the host vehicle from an image capture device; determining a planned navigation operation for achieving a navigation goal for the host vehicle based on at least one driving policy; analyzing the at least one image to identify a target vehicle within the environment of the host vehicle; determining a next-state lateral distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed; determining a maximum yaw rate capability of the host vehicle, a maximum change in turning radius capability of the host vehicle, and a current lateral velocity of the host vehicle; determining a lateral braking distance of the host vehicle based on the maximum yaw rate capability of the host vehicle, the maximum change in turning radius capability of the host vehicle, and the current lateral velocity of the host vehicle; determining a current lateral velocity of the target vehicle, a maximum yaw rate capability of the target vehicle, and a maximum change in turning radius capability of the target vehicle; determining a lateral braking distance for the target vehicle based on the current lateral velocity of the target vehicle, a maximum yaw rate capability of the target vehicle, and a maximum change in turning radius capability of the target vehicle; performing the planned navigation operation if the determined next state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle; 1. A system comprising: at least one processing device programmed to:
2. 2. The system of claim 1, wherein the lateral braking distance of the host vehicle corresponds to a distance required for the host vehicle lateral velocity to reach zero at the maximum yaw rate capability of the host vehicle and the maximum change in the turning radius capability of the host vehicle after a lateral acceleration of the host vehicle as a result of a turning maneuver made toward the target vehicle starting from the current lateral velocity of the host vehicle and at the maximum yaw rate capability of the host vehicle and the maximum change in the turning radius capability of the host vehicle during a response time associated with the host vehicle.
3. 3. The system of claim 1, wherein the lateral braking distance of the target vehicle corresponds to a distance required for the target vehicle lateral velocity to reach zero at a maximum change in the target vehicle's maximum yaw rate capability and turning radius capability following a lateral acceleration of the target vehicle as a result of a turning maneuver by the target vehicle towards the host vehicle performed starting from the current lateral velocity of the target vehicle and at a maximum change in the target vehicle's maximum yaw rate capability and turning radius capability during a response time associated with the target vehicle.
4. 4. The system of claim 1, wherein the at least one processing device is programmed to implement the planned navigation operation if the planned navigation operation results in a safe longitudinal distance between the host vehicle and the target vehicle, regardless of the lateral distance of the next state.
5. the safe longitudinal distance is greater than or equal to the sum of the longitudinal stopping distance of the host vehicle and the longitudinal stopping distance of the target vehicle; the stopping distance of the host vehicle is based on a maximum longitudinal braking capability of the host vehicle and a current longitudinal speed of the host vehicle, the stopping distance of the host vehicle including a host vehicle acceleration distance corresponding to a distance that the host vehicle can travel at a maximum acceleration capability of the host vehicle during a response time associated with the host vehicle starting from the current longitudinal speed of the host vehicle; 5. The system of claim 4, wherein the stopping distance of the target vehicle is based on a maximum longitudinal braking capability of the target vehicle and a current longitudinal speed of the target vehicle, and the stopping distance of the target vehicle includes a target vehicle acceleration distance corresponding to a distance that the target vehicle can travel at a maximum acceleration capability of the target vehicle during a response time associated with the target vehicle starting from the current longitudinal speed of the target vehicle.
6. 6. The system of claim 1, wherein at least one of a maximum yaw rate capability of the target vehicle and a maximum change in turning radius capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through analysis of the at least one image.
7. The system of claim 6 , wherein the at least one characteristic of the target vehicle includes one or more of a vehicle type, a vehicle size, or a vehicle model.
8. The system of any one of claims 1 to 7, wherein the maximum yaw rate capability of the host vehicle and the maximum change in turning radius capability of the host vehicle are based on predetermined constraints of the host vehicle.
9. A system according to any preceding claim, wherein the current lateral velocity of the target vehicle is determined based on the output of a radar or lidar unit associated with the host vehicle.
10. 10. The system of claim 1, wherein at least one of the maximum yaw rate capability of the host vehicle or the maximum change in turning radius capability of the host vehicle is determined based on a detected condition of a road surface.
11. 11. The system of claim 1, wherein at least one of a maximum yaw rate capability of the target vehicle or a maximum change in turning radius capability of the target vehicle is determined based on a sensed condition of a road surface.
12. 12. The system of claim 1, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in turning radius capability of the target vehicle is determined based on detected weather conditions.
13. The system of any one of claims 1 to 12, wherein the planned navigation operation includes at least one of a lane change maneuver, a merging maneuver, a through maneuver, a parking lot navigation maneuver, or a turning maneuver.
14. 14. The system of claim 1, wherein the at least one processing device is configured to perform the planned navigation operation if the determined next state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle by at least a predetermined minimum distance.
15. 15. The system of claim 14, wherein the predetermined minimum distance corresponds to a predetermined lateral separation distance maintained between the host vehicle and other vehicles.
16. 16. The system of claim 15, wherein the predetermined lateral separation distance is at least 1 meter.
17. 1. A method for navigating a host vehicle, the method comprising: receiving at least one image representing an environment of the host vehicle from an image capture device; determining a planned navigation operation for achieving a navigation goal for the host vehicle based on at least one driving policy; analyzing the at least one image to identify target vehicles within the environment of the host vehicle; determining a next-state lateral distance between the host vehicle and the target vehicle that will occur if the planned navigation operation is performed; determining a maximum yaw rate capability of the host vehicle, a maximum change in turning radius capability of the host vehicle, and a current lateral velocity of the host vehicle; determining a lateral braking distance of the host vehicle based on the maximum yaw rate capability of the host vehicle, the maximum change in turning radius capability of the host vehicle, and the current lateral velocity of the host vehicle; determining a current lateral velocity of the target vehicle, a maximum yaw rate capability of the target vehicle, and a maximum change in turning radius capability of the target vehicle; determining a lateral braking distance for the target vehicle based on the current lateral velocity of the target vehicle, a maximum yaw rate capability of the target vehicle, and a maximum change in turning radius capability of the target vehicle; performing the planned navigation operation if the determined next state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle; A method comprising:
18. 18. The method of claim 17, wherein the lateral braking distance of the host vehicle corresponds to a distance required for a host vehicle lateral velocity to reach zero at the maximum yaw rate capability of the host vehicle and a maximum change in the turning radius capability of the host vehicle after a lateral acceleration of the host vehicle as a result of a turning maneuver made toward the target vehicle starting from the current lateral velocity of the host vehicle and at the maximum yaw rate capability of the host vehicle and a maximum change in the turning radius capability of the host vehicle during a response time associated with the host vehicle.
19. 19. The method of claim 17 or 18, wherein the lateral braking distance of the target vehicle corresponds to a distance required for a target vehicle lateral velocity to reach zero at a maximum change in the target vehicle's maximum yaw rate capability and turning radius capability following a lateral acceleration of the target vehicle as a result of a turning maneuver by the target vehicle towards the host vehicle performed starting from the current lateral velocity of the target vehicle and at a maximum change in the target vehicle's maximum yaw rate capability and turning radius capability during a response time associated with the target vehicle.
20. A method described in any one of claims 17 to 19, wherein if the planned navigation operation results in a safe longitudinal distance between the host vehicle and the target vehicle, a step of implementing the planned navigation operation is executed regardless of the lateral distance of the next state.
21. the safe longitudinal distance is greater than or equal to the sum of the longitudinal stopping distance of the host vehicle and the longitudinal stopping distance of the target vehicle; the stopping distance of the host vehicle is based on a maximum longitudinal braking capability of the host vehicle and a current longitudinal speed of the host vehicle, the stopping distance of the host vehicle including a host vehicle acceleration distance corresponding to a distance that the host vehicle can travel at a maximum acceleration capability of the host vehicle during a response time associated with the host vehicle starting from the current longitudinal speed of the host vehicle; 21. The method of claim 20, wherein the stopping distance of the target vehicle is based on a maximum longitudinal braking capability of the target vehicle and a current longitudinal speed of the target vehicle, and the stopping distance of the target vehicle comprises a target vehicle acceleration distance corresponding to a distance the target vehicle can travel at a maximum acceleration capability of the target vehicle during a response time associated with the target vehicle starting from the current longitudinal speed of the target vehicle.
22. 22. The method of any one of claims 17 to 21, wherein at least one of a maximum yaw rate capability of the target vehicle and a maximum change in turning radius capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through analysis of the at least one image.
23. The method of claim 22 , wherein the at least one characteristic of the target vehicle includes one or more of a vehicle type, a vehicle size, or a vehicle model.
24. The method of any one of claims 17 to 23, wherein the maximum yaw rate capability of the host vehicle and the maximum change in turning radius capability of the host vehicle are based on predetermined constraints of the host vehicle.
25. A method according to any one of claims 17 to 24, wherein the current lateral velocity of the target vehicle is determined based on the output of a radar or lidar unit associated with the host vehicle.
26. 26. The method of any one of claims 17 to 25, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in turning radius capability of the target vehicle is determined based on detected road surface conditions.
27. 27. The method of any one of claims 17 to 26, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in turning radius capability of the target vehicle is determined based on sensed weather conditions.
28. The method of any one of claims 17 to 27, wherein the planned navigation maneuver includes at least one of a lane change maneuver, a merging maneuver, a through maneuver, a parking lot navigation maneuver, or a turning maneuver.
29. A method described in any one of claims 17 to 28, wherein the step of performing the planned navigation operation is executed when the lateral distance of the determined next state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle by at least a predetermined minimum distance.
30. 30. The method of claim 29, wherein the predetermined minimum distance corresponds to a predetermined lateral separation distance maintained between the host vehicle and other vehicles.
Citation Information
Patent Citations
Object detection system
JP2012160103A