Navigation system and method for navigating a host vehicle

By combining camera image analysis, GPS data and sensor data, the autonomous vehicle system can effectively solve the problem of safe and accurate arrival of the destination during navigation, and achieve constraints and management of potential accident liability, improving the interpretability and scalability of the system.

CN115384486BActive Publication Date: 2025-05-20MOBILEYE VISION TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211011663.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-11
Filing Date
2019-03-20
Publication Date
2025-05-20
Estimated Expiration
2039-03-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively address the challenge of autonomous vehicles considering safe and accurate access to their destinations during navigation, especially when dealing with multiple sources of information and complying with potential accident liability constraints.

Method used

By using a camera to provide autonomous vehicle navigation features, combining Global Positioning System (GPS) data, sensor data and map data, the image is analyzed to identify target vehicle and environmental features, determine the vehicle's maximum braking capacity and current status, and then implement planned navigation actions.

Benefits of technology

It realizes safe and accurate navigation of autonomous vehicles in complex environments, improves constraints and management of potential accident liability, and enhances the interpretability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115384486B_ABST
    Figure CN115384486B_ABST
Patent Text Reader

Abstract

The present disclosure provides a navigation system and method for navigating a main vehicle. The system receives a sensor output from a sensor indicating the motion of the main vehicle, generated at a first time, the first time being later than a data acquisition time when the measurement based on the sensor output is acquired and earlier than a second time when the sensor output is received by a processor; generates a prediction of the motion of the main vehicle based on the sensor output and an estimate of how the motion of the main vehicle changes between the data acquisition time and the motion prediction time; determines a planned navigation action of the main vehicle based on the navigation target of the main vehicle and based on the prediction of the motion of the main vehicle; generates a navigation command for implementing the planned navigation action and provides it to an actuation system of the main vehicle so that it receives the navigation command at a third time, the third time being later than the second time and earlier than or substantially equal to the actuation time of the actuation system's response to the command; the motion prediction time is after the data acquisition time and earlier than or equal to the actuation time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention titled "Systems and Methods for Navigating a Vehicle", with the application date of March 20, 2019, application number 201980002445.4 (International Application Number PCT / IB2019 / 000281).

[0002] Cross - reference to related applications

[0003] This application claims the priority benefits of U.S. Provisional Patent Application No. 62 / 645,479, filed on March 20, 2018; U.S. Provisional Patent Application No. 62 / 646,579, filed on March 22, 2018; U.S. Provisional Patent Application No. 62 / 718,554, filed on August 14, 2018; U.S. Provisional Patent Application No. 62 / 724,355, filed on August 29, 2018; U.S. Provisional Patent Application No. 62 / 772,366, filed on November 28, 2018; and U.S. Provisional Patent Application No. 62 / 777,914, filed on December 11, 2018. All of the above applications are hereby incorporated by reference in their entirety. Technical Field

[0004] This disclosure generally relates to autonomous vehicle navigation. In addition, this disclosure relates to systems and methods for navigating subject to potential accident liability constraints. Background Art

[0005] With the continuous progress of technology, the goal of fully autonomous vehicles that can navigate on roadways is approaching. Autonomous vehicles may need to consider a wide variety of factors and make appropriate decisions based on those factors to safely and accurately reach a desired destination. For example, autonomous vehicles may need to process and interpret visual information (e.g., information captured from cameras), information from radar or lidar, and may also use information obtained from other sources (e.g., from GPS devices, speed sensors, accelerometers, suspension sensors, etc.). At the same time, to navigate to a destination, an autonomous vehicle may also need to identify its location within a particular roadway (e.g., a particular lane within a multi - lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, observe traffic signals and signs, travel from one road to another at appropriate intersections or junctions, and respond to any other situations that occur or progress during the operation of the vehicle. In addition, a navigation system may need to comply with certain imposed constraints. In some cases, these constraints can involve interactions between the host vehicle and one or more other objects (such as other vehicles, pedestrians, etc.). In other cases, the constraints may involve liability rules to be followed when performing one or more navigation actions for the host vehicle.

[0006] In the field of autonomous driving, there are two important considerations for a viable autonomous vehicle system. The first is the standardization of safety assurance, including the requirements that each self-driving car must meet to ensure safety, and how to verify these requirements. The second is scalability, because engineering solutions that result in release costs will not scale to millions of cars and may prevent the widespread adoption of autonomous cars or even less widespread adoption. Therefore, an interpretable mathematical model for safety assurance and a system design that complies with safety assurance requirements while being scalable to millions of cars are needed. Summary of the Invention

[0007] Embodiments consistent with the present disclosure provide systems and methods for autonomous vehicle navigation. The disclosed embodiments may use cameras to provide autonomous vehicle navigation features. For example, consistent with the disclosed embodiments, the disclosed system may include one, two, or more cameras that monitor the vehicle environment. The disclosed system may provide a navigation response based on, for example, an analysis of images captured by one or more cameras. The navigation response may also consider other data, including, for example, global positioning system (GPS) data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data.

[0008] In one embodiment, a system for navigating a host vehicle may include at least one processing device programmed to receive at least one image representative of the host vehicle environment. The at least one image may be received from an image capture device. The at least one processing device may be programmed to determine a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy. The at least one processing device may be further programmed to analyze the at least one image to identify a target vehicle in the host vehicle environment and to determine a next state distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken. The at least one processing device may be programmed to determine a maximum braking capacity of the host vehicle, a maximum acceleration capacity of the host vehicle, and a current speed of the host vehicle. The at least one processing device may also be programmed to determine a current stopping distance of the host vehicle based on the current maximum braking capacity of the host vehicle, the current maximum acceleration capacity of the host vehicle, and the current speed of the host vehicle. The at least one processing device may be further programmed to determine a current speed of the target vehicle and to assume a maximum braking capacity of the target vehicle based on at least one identified characteristic of the target vehicle. The at least one processing device may also be programmed to implement the planned navigation action if the determined current stopping distance of the host vehicle is less than the sum of the determined next state distance and a distance traveled by the target vehicle determined based on the current speed of the target vehicle and the assumed maximum braking capacity of the target vehicle.

[0009] In one embodiment, a system for navigating a host vehicle may include at least one processing device. The at least one processing device may be programmed to receive at least one image representative of the host vehicle's environment from an image capture device. The at least one processing device may also be programmed to determine planned navigation actions for achieving a navigation goal of the host vehicle. The planned navigation actions may be based on at least one driving strategy. The at least one processing device may be programmed to analyze the at least one image to identify a target vehicle in the host vehicle's environment. The at least one processing device may further be programmed to determine a next-state distance between the host vehicle and the target vehicle that would result if the planned navigation actions were taken. The at least one processing device may be programmed to determine a current speed of the host vehicle and a current speed of the target vehicle. The at least one processing device may be programmed to assume a maximum braking rate capability of the target vehicle based on at least one identified characteristic of the target vehicle. The at least one processing device may further be programmed to implement the planned navigation actions if, for the determined current speed of the host vehicle and at a predetermined sub-maximum braking rate less than the maximum braking rate capability of the host vehicle, the host vehicle can stop within a host vehicle stopping distance that is less than the sum of the determined next-state distance and a target vehicle travel distance determined based on the current speed of the target vehicle and the assumed maximum braking rate capability of the target vehicle.

[0010] In one embodiment, a system for navigating a host vehicle may include at least one processing device. The at least one processing device may be programmed to receive at least one image representative of the host vehicle's environment from an image capture device. The at least one processing device may also be programmed to determine planned navigation actions for achieving a navigation goal of the host vehicle. The planned navigation actions may be based on at least one driving strategy. The at least one processing device may be programmed to analyze the at least one image to identify a target vehicle in the host vehicle's environment. The at least one processing device may further be programmed to determine a next-state distance between the host vehicle and the target vehicle that would result if the planned navigation actions were taken. The at least one processing device may be programmed to determine a current speed of the host vehicle. The at least one processing device may further be programmed to determine a current speed of the target vehicle and assume a maximum braking rate capability of the target vehicle based on at least one identified characteristic of the target vehicle. The at least one processing device may be programmed to implement the planned navigation actions if, for the determined current speed of the host vehicle and for a predetermined braking rate curve, the host vehicle can stop within a host vehicle stopping distance that is less than the sum of the determined next-state distance and a target vehicle travel distance determined based on the current speed of the target vehicle and the assumed maximum braking rate capability of the target vehicle, wherein the predetermined braking rate curve gradually increases from a sub-maximum braking rate of the host vehicle to the maximum braking rate.

[0011] In one embodiment, a system for braking a host vehicle may include at least one processing device programmed to perform one or more operations. The at least one processing device may be programmed to receive an output representative of the host vehicle environment from at least one sensor. The at least one processing device may be further programmed to detect a target vehicle in the host vehicle environment based on the output. The at least one processing device may be programmed to determine a current speed of the host vehicle and a current distance between the host vehicle and the target vehicle. At least based on the current speed of the host vehicle and the current distance between the host vehicle and the target vehicle, the at least one processor may be programmed to determine whether a braking condition exists. If it is determined that a braking condition exists, the at least one processor may be programmed to cause a braking device associated with the host vehicle to be applied according to a pre - defined braking curve, the pre - defined braking curve including segments that start at a sub - maximum braking rate of the host vehicle and gradually increase to the maximum braking rate of the host vehicle.

[0012] In one embodiment, an autonomous system for selectively superseding a human driver's control of a host vehicle may include at least one processing device. The at least one processing device may be programmed to receive at least one image representative of the host vehicle environment from an image capture device and, based on an analysis of the at least one image, detect at least one obstacle in the host vehicle environment. The at least one processing device may be programmed to monitor driver input to at least one of a throttle control, a brake control, or a steering control associated with the host vehicle. The at least one processing device may also be programmed to determine whether the driver input will cause the host vehicle to navigate within a proximity buffer zone relative to the at least one obstacle. If the at least one processing device determines that the driver input will not cause the host vehicle to navigate within a proximity buffer zone relative to the at least one obstacle, the at least one processing device may be programmed to allow the driver input to cause corresponding changes in one or more host vehicle motion control systems. If the at least one processing device determines that the driver input will cause the host vehicle to navigate within a proximity buffer zone relative to the at least one obstacle, the at least one processing device may be programmed to prevent the driver input from causing corresponding changes in one or more host vehicle motion control systems.

[0013] In one embodiment, a navigation system for navigating a host autonomous vehicle according to at least one navigation goal of the host vehicle may include at least one processor. The at least one processor may be programmed to receive sensor output from one or more sensors indicating at least one aspect of the motion of the host vehicle relative to the host vehicle environment. The sensor output may be generated at a first time that is later than the data acquisition time at which the measurement or data acquisition on which the sensor output is based was acquired and earlier than a second time at which the sensor output is received by the at least one processor. The at least one processor may be programmed to, for a motion prediction time, generate a prediction of at least one aspect of the motion of the host vehicle based at least in part on the received sensor output and an estimate of how at least one aspect of the motion of the host vehicle changes over the time interval between the data acquisition time and the motion prediction time. The at least one processor may be programmed to determine a planned navigation action of the host vehicle based at least in part on at least one navigation goal of the host vehicle and based on the generated prediction of at least one aspect of the motion of the host vehicle. The at least one processor may be further configured to generate a navigation command for implementing at least a portion of the planned navigation action. The at least one processor may be programmed to provide the navigation command to at least one actuation system of the host vehicle. The navigation command may be provided such that the at least one actuation system receives the navigation command at a third time that is later than the second time and earlier than or substantially equal to an actuation time at which components of the at least one actuation system respond to the received command. In some embodiments, the motion prediction time is after the data acquisition time and earlier than or equal to the actuation time.

[0014] In one embodiment, a system for navigating a host vehicle includes: at least one processing device programmed to: receive at least one image representative of the environment of the host vehicle from an image capture device; determine a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy; analyze the at least one image to identify a target vehicle in the environment of the host vehicle; determine a next state distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determine a maximum braking ability of the host vehicle, a maximum acceleration ability of the host vehicle, and a current speed of the host vehicle, wherein the maximum braking ability of the host vehicle is determined based on at least one factor associated with the host vehicle or the environment of the host vehicle; determine a current stopping distance of the host vehicle based on the current maximum braking ability of the host vehicle, the current maximum acceleration ability of the host vehicle, and the current speed of the host vehicle; determine a current speed of the target vehicle and assume a maximum braking ability of the target vehicle based on at least one identified characteristic of the target vehicle; and implement the planned navigation action if the determined current stopping distance of the host vehicle is less than the sum of the determined next state distance and a distance traveled by the target vehicle based on the current speed of the target vehicle and the assumed maximum braking ability of the target vehicle.

[0015] In one embodiment, a method for navigating a host vehicle includes: receiving, from an image capture device, at least one image representative of an environment of the host vehicle; determining, based on at least one driving strategy, a planned navigation action for achieving a navigation goal of the host vehicle; analyzing the at least one image to identify a target vehicle in the environment of the host vehicle; determining a next state distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determining a maximum braking ability of the host vehicle, a maximum acceleration ability of the host vehicle, and a current speed of the host vehicle, wherein the maximum braking ability of the host vehicle is determined based on at least one factor associated with the host vehicle or the environment of the host vehicle; determining a current stopping distance of the host vehicle based on the current maximum braking ability of the host vehicle, the current maximum acceleration ability of the host vehicle, and the current speed of the host vehicle; determining a current speed of the target vehicle and assuming a maximum braking ability of the target vehicle based on at least one identified characteristic of the target vehicle; and implementing the planned navigation action if the determined current stopping distance of the host vehicle is less than the sum of the determined next state distance and a distance traveled by the target vehicle determined based on the current speed of the target vehicle and the assumed maximum braking ability of the target vehicle.

[0016] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions executable by at least one processing device and perform any of the steps and / or methods described herein.

[0017] The foregoing general description and the following detailed description are merely exemplary and explanatory and are not restrictive of the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings incorporated in and constituting a part of this disclosure illustrate various disclosed embodiments.

[0019] In the drawings:

[0020] Figure 1 is an illustrative representation of an example system consistent with the disclosed embodiments.

[0021] Figure 2A is an illustrative side view representation of an example vehicle including a system consistent with the disclosed embodiments.

[0022] Figure 2B is consistent with the disclosed embodiments Figure 2A illustrative top view representation of the vehicle and system shown in.

[0023] Figure 2C is an illustrative top view representation of another embodiment of a vehicle including a system consistent with the disclosed embodiments.

[0024] Figure 2D It is an illustrative top - view representation of another embodiment of a vehicle including a system consistent with the disclosed embodiments.

[0025] Figure 2E It is an illustrative top - view representation of another embodiment of a vehicle including a system consistent with the disclosed embodiments.

[0026] Figure 2F It is an illustrative representation of an example vehicle control system consistent with the disclosed embodiments.

[0027] Figure 3A It is an illustrative representation of the interior of a vehicle consistent with the disclosed embodiments, including a rear - view mirror and a user interface for a vehicle imaging system.

[0028] Figure 3B It is an illustration of an example of a camera mount configured to be positioned behind a rear - view mirror and against a vehicle windshield consistent with the disclosed embodiments.

[0029] Figure 3C It is consistent with the disclosed embodiments Figure 3B An illustration of the camera mount shown in different perspectives.

[0030] Figure 3D It is an illustration of an example of a camera mount configured to be located behind a rear - view mirror and against a vehicle windshield consistent with the disclosed embodiments.

[0031] Figure 4 It is an example block diagram of a memory configured to store instructions for performing one or more operations consistent with the disclosed embodiments.

[0032] Figure 5A It is a flowchart showing an example process for causing one or more navigation responses based on monocular image analysis consistent with the disclosed embodiments.

[0033] Figure 5B It is a flowchart showing an example process for detecting one or more vehicles and / or pedestrians in a set of images consistent with the disclosed embodiments.

[0034] Figure 5C It is a flowchart showing an example process for detecting road markings and / or lane geometry information in a set of images consistent with the disclosed embodiments.

[0035] Figure 5D It is a flowchart showing an example process for detecting traffic lights in a set of images consistent with the disclosed embodiments.

[0036] Figure 5Eis a flow chart illustrating an example process for eliciting one or more navigation responses based on a vehicle path consistent with the disclosed embodiments.

[0037] Figure 5F is a flow chart illustrating an example process for determining whether a leading vehicle is changing lanes, consistent with the disclosed embodiments.

[0038] Figure 6 is a flow chart illustrating an example process for inducing one or more navigation responses based on stereoscopic image analysis consistent with the disclosed embodiments.

[0039] Figure 7 is a flow chart illustrating an example process for eliciting one or more navigation responses based on analysis of three sets of images consistent with the disclosed embodiments.

[0040] Figure 8 is a block diagram representation of modules that may be implemented by one or more specially programmed processing devices of a navigation system of an autonomous vehicle consistent with the disclosed embodiments.

[0041] Figure 9 is a diagram of navigation options consistent with the disclosed embodiments.

[0042] Figure 10 is a diagram of navigation options consistent with the disclosed embodiments.

[0043] Figure 11A 、 Figure 11B and Figure 11C A schematic representation of navigation options for a host vehicle in a merge zone consistent with the disclosed embodiments is provided.

[0044] Figure 11D An illustrative description of a dual-lane merge scenario consistent with the disclosed embodiments is provided.

[0045] Figure 11E A diagram of options that are potentially useful in a double-merge scenario consistent with the disclosed embodiments is provided.

[0046] Figure 12 A diagram capturing a representative image of a host vehicle's environment and potential navigation constraints is provided consistent with the disclosed embodiments.

[0047] Figure 13 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.

[0048] Figure 14 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.

[0049] Figure 15An algorithm flowchart for navigating a vehicle is provided that is consistent with the disclosed embodiments.

[0050] Figure 16 An algorithm flowchart for navigating a vehicle is provided that is consistent with the disclosed embodiments.

[0051] Figure 17A and 17B An illustrative illustration of a host vehicle navigating into a roundabout is provided that is consistent with the disclosed embodiments.

[0052] Figure 18 An algorithm flowchart for navigating a vehicle is provided that is consistent with the disclosed embodiments.

[0053] Figure 19 An example of a host vehicle traveling on a multi-lane highway is shown that is consistent with the disclosed embodiments.

[0054] Figure 20A and 20B An example of a vehicle cutting in front of another vehicle is shown that is consistent with the disclosed embodiments.

[0055] Figure 21 An example of a vehicle following another vehicle is shown that is consistent with the disclosed embodiments.

[0056] Figure 22 An example of a vehicle leaving a parking lot and merging into a potentially busy road is shown that is consistent with the disclosed embodiments.

[0057] Figure 23 A vehicle traveling on a road is shown that is consistent with the disclosed embodiments.

[0058] Figure 24A - 24D Four example scenarios are shown that are consistent with the disclosed embodiments.

[0059] Figure 25 An example scenario is shown that is consistent with the disclosed embodiments.

[0060] Figure 26 An example scenario is shown that is consistent with the disclosed embodiments.

[0061] Figure 27 An example scenario is shown that is consistent with the disclosed embodiments.

[0062] Figure 28A and 28B An example of a scenario where a vehicle is following another vehicle is shown that is consistent with the disclosed embodiments.

[0063] Figure 29A and 29BShows an example blame in a cut-in scenario consistent with the disclosed embodiments.

[0064] Figure 30A and 30B Shows an example blame in a cut-in scenario consistent with the disclosed embodiments.

[0065] Figure 31A - 31D Shows an example blame in a drift scenario consistent with the disclosed embodiments.

[0066] Figure 32A and 32B Shows an example blame in a two-way traffic scenario consistent with the disclosed embodiments.

[0067] Figure 33A and 33B Shows an example blame in a two-way traffic scenario consistent with the disclosed embodiments.

[0068] Figure 34A and 34B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0069] Figure 35A and 35B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0070] Figure 36A and 36B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0071] Figure 37A and 37B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0072] Figure 38A and 38B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0073] Figure 39A and 39B Shows an example blame in a route priority scenario consistent with the disclosed embodiments.

[0074] Figure 40A and 40B Shows an example blame in a traffic light scenario consistent with the disclosed embodiments.

[0075] Figure 41A and 41B Shows an example blame in a traffic light scenario consistent with the disclosed embodiments.

[0076] Figure 42A and 42B illustrates an example attribution in a traffic light scenario consistent with the disclosed embodiments.

[0077] Figure 43A - 43C illustrates an example vulnerable road user (VRU) scenario consistent with the disclosed embodiments.

[0078] Figure 44A - 44C illustrates an example vulnerable road user (VRU) scenario consistent with the disclosed embodiments.

[0079] Figure 45A - 45C illustrates an example vulnerable road user (VRU) scenario consistent with the disclosed embodiments.

[0080] Figure 46A - 46D illustrates an example vulnerable road user (VRU) scenario consistent with the disclosed embodiments.

[0081] Figure 47A and 47B illustrates an example scenario where a vehicle follows another vehicle consistent with the disclosed embodiments.

[0082] Figure 48 is a flowchart illustrating an exemplary process for navigating a host vehicle consistent with the disclosed embodiments.

[0083] Figure 49A - 49D illustrates an example scenario where a vehicle follows another vehicle consistent with the disclosed embodiments.

[0084] Figure 50 is a flowchart illustrating an exemplary process for braking a host vehicle consistent with the disclosed embodiments.

[0085] Figure 51 is a flowchart illustrating an exemplary process for navigating a host vehicle consistent with the disclosed embodiments.

[0086] Figure 52A - 52D illustrates an example proximity buffer of a host vehicle consistent with the disclosed embodiments.

[0087] Figure 53A and 53B illustrates an example scenario including a proximity buffer consistent with the disclosed embodiments.

[0088] Figure 54A and 54B illustrates an example scenario including a proximity buffer consistent with the disclosed embodiments.

[0089] Figure 55A flowchart for selectively replacing a human driver's control of a host vehicle is provided that is consistent with the disclosed embodiments.

[0090] Figure 56 A flowchart illustrates an exemplary process for navigating a host vehicle that is consistent with the disclosed embodiments.

[0091] Figure 57A - 57C An example scenario is shown that is consistent with the disclosed embodiments.

[0092] Figure 58 A flowchart illustrates an exemplary process for navigating a host vehicle that is consistent with the disclosed embodiments. DETAILED DESCRIPTION

[0093] The following detailed description refers to the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the following description to refer to the same or like parts. Although several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, components shown in the drawings may be replaced, added, or modified, and the exemplary methods described herein may be modified by replacing, reordering, removing, or adding steps of the disclosed methods. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.

[0094] Overview of Autonomous Vehicles

[0095] As used throughout this disclosure, the term "autonomous vehicle" refers to a vehicle that is capable of effecting at least one navigation change without driver input. A "navigation change" refers to one or more of a vehicle's steering, braking, or accelerating / decelerating. By autonomous, it is meant that the vehicle does not need to be fully automatic (e.g., fully operable without a driver or without driver input). Instead, autonomous vehicles include those vehicles that are capable of operating under a driver's control during some time periods and capable of operating without driver control during other time periods. Autonomous vehicles may also include vehicles that control only certain aspects of vehicle navigation (such as steering (e.g., maintaining a vehicle's route between vehicle lane constraints) or certain steering operations in some cases (but not all cases)), but may leave other aspects to the driver (such as braking or braking in some cases). In some instances, an autonomous vehicle may handle some or all aspects of a vehicle's braking, speed control, and / or steering.

[0096] Since human drivers typically rely on visual cues and observations to control a vehicle, traffic infrastructure has been built with lane markings, traffic signs, and traffic lights designed to provide visual information to drivers. Given these design features of the traffic infrastructure, an autonomous vehicle can include a camera and a processing unit that analyzes visual information captured from the vehicle's environment. The visual information can include, for example, images representing components of the traffic infrastructure (e.g., lane markings, traffic signs, traffic lights, etc.) and other obstacles (e.g., other vehicles, pedestrians, debris, etc.) that can be observed by a driver. Additionally, an autonomous vehicle can also use stored information, such as information that provides a model of the vehicle's environment during navigation. For example, a vehicle can use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data to provide information related to its environment while the vehicle is in motion, and the vehicle (and other vehicles) can use this information to locate itself on the model. Some vehicles are also capable of communicating between vehicles, sharing information, changing the hazards of companion vehicles or changes around the vehicle, etc.

[0097] System Overview

[0098] Figure 1 is a block diagram representation of system 100 consistent with the disclosed example embodiments. Depending on the requirements of a particular implementation, system 100 can include various components. In some embodiments, system 100 can include a processing unit 110, an image acquisition unit 120, a position sensor 130, one or more memory units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. The processing unit 110 can include one or more processing devices. In some embodiments, the processing unit 110 can include an application processor 180, an image processor 190, or any other suitable processing device. Similarly, depending on the requirements of a particular application, the image acquisition unit 120 can include any number of image acquisition devices and components. In some embodiments, the image acquisition unit 120 can include one or more image capture devices (e.g., cameras, charge-coupled devices (CCDs), or any other type of image sensor), such as image capture device 122, image capture device 124, and image capture device 126. System 100 can also include a data interface 128 that communicatively connects the processing unit 110 to the image acquisition unit 120. For example, the data interface 128 can include any wired and / or wireless one or more links for transmitting image data acquired by the image acquisition unit 120 to the processing unit 110.

[0099] The wireless transceiver 172 may include one or more devices configured to exchange transmissions over an air interface to one or more networks (e.g., cellular, Internet, etc.) by using radio frequency, infrared frequency, magnetic field, or electric field. The wireless transceiver 172 may use any well-known standard to send and / or receive data (e.g., Wi-Fi, Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions may include communications from the host vehicle to one or more remotely located servers. Such transmissions may also include (one-way or two-way) communications between the host vehicle and one or more target vehicles in the environment of the host vehicle (e.g., to facilitate coordinating the navigation of the host vehicle in view of or in conjunction with the target vehicles in the environment of the host vehicle), or even broadcast transmissions to unspecified receivers in the vicinity of the transmitting vehicle.

[0100] Both the application processor 180 and the image processor 190 may include various types of hardware-based processing devices. For example, either or both of the application processor 180 and the image processor 190 may include a microprocessor, a pre-processor (such as an image pre-processor), a graphics processor, a central processing unit (CPU), auxiliary circuitry, a digital signal processor, an integrated circuit, a memory, or any other type of device suitable for running applications and suitable for image processing and analysis. In some embodiments, the application processor 180 and / or the image processor 190 may include any type of single-core or multi-core processor, mobile device microcontroller, central processing unit, etc. Various processing devices may be used, including, for example, processors available from manufacturers such as etc., and may include various architectures (e.g., x86 processors, etc.).

[0101] In some embodiments, the application processor 180 and / or the image processor 190 may include any EyeQ series processor chip available from These processor designs each include multiple processing units with local memory and instruction sets. Such processors may include video inputs for receiving image data from multiple image sensors and may also include video output capabilities. In one example, uses 90 nanometer-micron technology operating at 332 megahertz. The architecture consists of two floating-point hyper-threaded 32-bit RISC CPUs ( cores), five vision computing engines (VCEs), three vector microcode processors It consists of a Denali 64-bit mobile DDR controller, a 128-bit internal Sonics Interconnect, a dual 16-bit video input and 18-bit video output controller, 16-channel DMA, and several peripheral devices. The MIPS34K CPU manages these five VCEs and three VMPs TM and DMA, a second MIPS34K CPU, and multi-channel DMA, as well as other peripheral devices. These five VCEs and three and the MIPS34K CPU can perform the intensive visual computations required for multifunctional bundled applications. In another example, as a third-generation processor and six times stronger than can be used in the disclosed embodiments. In other examples, can be used in the disclosed embodiments and / or and / or Of course, any newer or future EyeQ processing device can also be used with the disclosed embodiments.

[0102] Any processing device disclosed herein can be configured to perform certain functions. Configuring a processing device (such as any of the described EyeQ processors or other controllers or microprocessors) to perform certain functions can include programming computer-executable instructions and making the processing device accessible to these instructions for execution during its operation. In some embodiments, configuring the processing device can include directly programming the processing device using architectural instructions. In other embodiments, configuring the processing device can include storing the executable instructions on a memory accessible to the processing device during operation. For example, the processing device can access this memory during operation to obtain and execute the stored instructions. In either case, a processing device configured to perform the sensing, image analysis, and / or navigation functions disclosed herein represents a dedicated hardware-based system that controls multiple hardware-based components of the host vehicle.

[0103] Although Figure 1 depicts two separate processing devices included in processing unit 110, more or fewer processing devices can be used. For example, in some embodiments, a single processing device can be used to complete the tasks of application processor 180 and image processor 190. In other embodiments, these tasks can be performed by more than two processing devices. Additionally, in some embodiments, system 100 can include one or more processing units 110 and not include other components such as image acquisition unit 120.

[0104] The processing unit 110 may include various types of devices. For example, the processing unit 110 may include various devices such as a controller, an image pre-processor, a central processing unit (CPU), auxiliary circuits, a digital signal processor, an integrated circuit, a memory, or any other type of device for image processing and analysis. The image pre-processor may include a video processor for capturing, digitizing, and processing images from an image sensor. The CPU may include any number of microcontrollers or microprocessors. The auxiliary circuits may be any number of circuits well known in the art, including caches, power supplies, clocks, and input / output circuits. The memory may store software that controls the operation of the system when executed by the processor. The memory may include a database and image processing software. The memory may include any number of random access memories, read-only memories, flash memories, disk drives, optical storage, tape storage, removable storage, and other types of storage. In one example, the memory may be separate from the processing unit 110. In another example, the memory may be integrated into the processing unit 110.

[0105] Each of the memories 140, 150 may include software instructions that, when executed by a processor (e.g., the application processor 180 and / or the image processor 190), may control the operation of various aspects of the system 100. For example, these memory units may include various databases and image processing software, as well as trained systems such as neural networks and deep neural networks. The memory units may include random access memory, read-only memory, flash memory, disk drives, optical storage, tape storage, removable storage, and / or any other type of storage. In some embodiments, the memory units 140, 150 may be separate from the application processor 180 and / or the image processor 190. In other embodiments, these memory units may be integrated into the application processor 180 and / or the image processor 190.

[0106] The position sensor 130 may include any type of device suitable for determining the position associated with at least one component of the system 100. In some embodiments, the position sensor 130 may include a GPS receiver. Such a receiver may determine the user's position and speed by processing signals broadcast by global positioning system satellites. The position information from the position sensor 130 may be made available to the application processor 180 and / or the image processor 190.

[0107] In some embodiments, the system 100 may include components such as a speed sensor (e.g., a speedometer) for measuring the rate of the vehicle 200. The system 100 may also include one or more accelerometers (single-axis or multi-axis) for measuring the acceleration of the vehicle 200 along one or more axes.

[0108] Memory units 140, 150 may include a database, or data organized in any other form, indicating the locations of known landmarks. Sensing information of the environment (such as images from lidar or stereo processing of two or more images, radar signals, depth information) may be processed together with location information (such as GPS coordinates, the ego-motion of the vehicle, etc.) to determine the current position of the vehicle relative to known landmarks and improve the vehicle position. Some aspects of this technology are incorporated into a positioning technology called REM TM and is sold by the assignee of this application.

[0109] The user interface 170 may include any device suitable for providing information to or receiving input from one or more users of the system 100. In some embodiments, the user interface 170 may include user input devices, including for example a touch screen, a microphone, a keyboard, a pointer device, a trackball, a camera, a knob, buttons, etc. Using such input devices, a user can provide information input or commands to the system 100 by typing instructions or information, providing voice commands, using buttons, pointers or eye-tracking capabilities to select menu options on the screen, or by any other technique suitable for transmitting information to the system 100.

[0110] The user interface 170 may be equipped with one or more processing devices configured to provide and receive information from the user and process the information for use by, for example, the application processor 180. In some embodiments, such processing devices may execute instructions to identify and track eye movements, receive and interpret voice commands, identify and interpret touches and / or gestures made on the touch screen, respond to keyboard input or menu selections, etc. In some embodiments, the user interface 170 may include a display, a speaker, a haptic device, and / or any other device for providing output information to the user.

[0111] The map database 160 can include any type of database for storing map data useful to the system 100. In some embodiments, the map database 160 can include data related to the locations of various items in a reference coordinate system, the various items including roads, water features, geographical features, commercial areas, points of interest, restaurants, gas stations, etc. The map database 160 can store not only the locations of these items, but also descriptors related to these items, including, for example, names associated with any stored features. In some embodiments, the map database 160 can be physically co-located with other components of the system 100. Alternatively or additionally, the map database 160 or a portion thereof can be located remotely relative to other components of the system 100 (e.g., the processing unit 110). In such an embodiment, information from the map database 160 can be downloaded via a wired or wireless data connection to a network (e.g., via a cellular network and / or the Internet, etc.). In some cases, the map database 160 can store a sparse data model that includes polynomial representations of certain road features (e.g., lane markings) or the target trajectory of the host vehicle. The map database 160 can also include representations of various identified landmarks that can be used to determine or update the known position of the host vehicle relative to the target trajectory. The landmark representations can include data fields such as landmark type, landmark location, and other potential identifiers.

[0112] The image capture devices 122, 124, and 126 can each include any type of device suitable for capturing at least one image from the environment. Additionally, any number of image capture devices can be used to obtain images for input to the image processor. Some embodiments can include only a single image capture device, while other embodiments can include two, three, or even four, or more image capture devices. Further description of the image capture devices 122, 124, and 126 will be provided below with reference to Figures 2B to 2E be provided below.

[0113] One or more cameras (e.g., image capture devices 122, 124, and 126) can be part of a sensing block included on a vehicle. The sensing block can include various other sensors, and any or all of these sensors can be relied upon to form a sensed navigation state of the vehicle. In addition to cameras (front, side, rear, etc.), other sensors such as radar, lidar, and acoustic sensors can be included in the sensing block. Additionally, the sensing block can include one or more components configured to transmit and send / receive information related to the vehicle's environment. For example, such a component can include a wireless transceiver (RF, etc.) that can receive sensor-based information or any other type of information related to the vehicle's environment from a source located remotely relative to the host vehicle. This information can include sensor output information or related information received from vehicle systems other than the host vehicle. In some embodiments, this information can include information received from remote computing devices, central servers, etc. Furthermore, cameras can be in many different configurations: a single camera unit, multiple cameras, camera clusters, long FOV, short FOV, wide angle, fisheye, etc.

[0114] System 100 or its various components can be incorporated into a variety of different platforms. In some embodiments, system 100 can be included on vehicle 200, as Figure 2A shown. For example, vehicle 200 can be equipped with processing unit 110 of system 100 and any other components as described above with respect to Figure 1 . In some embodiments, vehicle 200 can be equipped with only a single image capture device (e.g., a camera), while in other embodiments, such as those discussed in conjunction with Figures 2B to 2E , multiple image capture devices can be used. For example, Figure 2A either of image capture devices 122 and 124 of vehicle 200 shown in

[0115] can be part of an ADAS (Advanced Driver Assistance System) imaging set. The image capture device included on vehicle 200 and that is part of image acquisition unit 120 can be placed in any suitable location. In some embodiments, as Figures 2A to 2E , and Figures 3A to 3C shown, image capture device 122 can be located near the rearview mirror. This location can provide a line of sight similar to that of the driver of vehicle 200, which can assist in determining what is visible and non-visible to the driver. Image capture device 122 can be placed at any location near the rearview mirror, and placing image capture device 122 on the driver's side of the mirror can also assist in obtaining an image representative of the driver's field of view and / or line of sight.

[0116] Other locations of the image capture device of the image acquisition unit 120 may also be used. For example, the image capture device 124 may be located on or in the bumper of the vehicle 200. Such a location may be particularly suitable for an image capture device with a wide field of view. The line of sight of the image capture device located on the bumper may be different from the driver's line of sight, and thus, the bumper image capture device and the driver may not always see the same object. The image capture devices (e.g., image capture devices 122, 124, and 126) may also be located in other locations. For example, the image capture device may be located on or in one or both of the side mirrors of the vehicle 200, on the roof of the vehicle 200, on the hood of the vehicle 200, on the trunk of the vehicle 200, on the side of the vehicle 200, mounted on any window of the vehicle 200, placed behind any window of the vehicle 200, or placed in front of any window, and in or near lighting devices mounted on the front and / or rear of the vehicle 200.

[0117] In addition to the image capture device, the vehicle 200 may also include various other components of the system 100. For example, the processing unit 110 may be included on the vehicle 200, integrated or separated from the vehicle's engine control unit (ECU). The vehicle 200 may also be equipped with a position sensor 130 such as a GPS receiver, and may also include a map database 160 and memory units 140 and 150.

[0118] As discussed earlier, the wireless transceiver 172 may transmit and / or receive data via one or more networks (e.g., cellular networks, the Internet, etc.). For example, the wireless transceiver 172 may upload the data collected by the system 100 to one or more servers and download data from one or more servers. Via the wireless transceiver 172, the system 100 may receive, for example, periodic updates or on-demand updates to the data stored in the map database 160, memory 140, and / or memory 150. Similarly, the wireless transceiver 172 may upload any data from the system 100 (e.g., images captured by the image acquisition unit 120, data received by the position sensor 130 or other sensors, vehicle control systems, etc.) and / or any data processed by the processing unit 110 to one or more servers.

[0119] The system 100 may upload data to a server (e.g., upload to the cloud) based on privacy level settings. For example, the system 100 may implement privacy level settings to specify or limit the types of data (including metadata) that can uniquely identify the vehicle and / or the driver / owner of the vehicle and that are sent to the server. These settings may be set by the user via, for example, the wireless transceiver 172, may be set by factory default settings, or may be initialized by data received by the wireless transceiver 172.

[0120] In some embodiments, the system 100 may upload data according to a "high" privacy level, and with settings set, the system 100 may transmit data (e.g., location information related to a journey, captured images, etc.) without any details about a specific vehicle and / or driver / owner. For example, when uploading data according to a "high" privacy setting, the system 100 may not include a vehicle identification number (VIN) or the name of the driver or owner of the vehicle, and may instead transmit data (such as captured images and / or restricted location information related to a journey).

[0121] Other privacy levels may also be considered. For example, the system 100 may transmit data to a server according to a "medium" privacy level, and may include additional information not included at the "high" privacy level, such as the model and / or make and / or type of vehicle (e.g., passenger vehicle, sport utility vehicle, truck, etc.). In some embodiments, the system 100 may upload data according to a "low" privacy level. Under the "low" privacy level setting, the system 100 may upload data and include information sufficient to uniquely identify a specific vehicle, owner / driver, and / or part or whole of the journey the vehicle has traveled. For example, such "low" privacy level data may include one or more of the following: VIN, driver / owner name, origin of the vehicle before departure, desired destination of the vehicle, model and / or make of the vehicle, vehicle type, etc.

[0122] Figure 2A is an illustrative side view representation of an example vehicle imaging system consistent with the disclosed embodiments. Figure 2B is Figure 2A an illustrative top view illustration of the embodiment shown in Figure 2B As shown, the disclosed embodiments may include a vehicle 200 that includes the system 100 in its body, the system 100 having a first image capture device 122 located near the vehicle 200's rearview mirror and / or near the driver, a second image capture device 124 located above or within the bumper region of the vehicle 200 (e.g., one of the bumper regions 210), and a processing unit 110.

[0123] As Figure 2C shown, both image capture devices 122 and 124 may be located near the vehicle 200's rearview mirror and / or near the driver. Additionally, although Figure 2B and Figure 2C show two image capture devices 122 and 124, it should be understood that other embodiments may include more than two image capture devices. For example, in Figure 2Dand Figure 2E In the embodiment shown in Figure 2E , the first image capture device 122, the second image capture device 124, and the third image capture device 126 are included in the system 100 of the vehicle 200.

[0124] As Figure 2D shown, the image capture device 122 may be located near the rearview mirror of the vehicle 200 and / or near the driver, and the image capture devices 124 and 126 may be located above or within the bumper area of the vehicle 200 (e.g., one of the bumper areas 210). And as Figure 2E shown, the image capture devices 122, 124, and 126 may be located near the rearview mirror of the vehicle 200 and / or near the driver's seat. The disclosed embodiments are not limited to any specific number and configuration of image capture devices, and the image capture devices may be located in any suitable location within or on the vehicle 200.

[0125] It should be understood that the disclosed embodiments are not limited to vehicles and may be applied in other scenarios. It should also be understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and may be applicable to all types of vehicles, including cars, trucks, trailers, and other types of vehicles.

[0126] The first image capture device 122 may include any suitable type of image capture device. The image capture device 122 may include an optical axis. In one example, the image capture device 122 may include an Aptina M9V024WVGA sensor with a global shutter. In other embodiments, the image capture device 122 may provide a resolution of 1280×960 pixels and may include a rolling shutter. The image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included, for example, to provide a desired focal length and field of view for the image capture device. In some embodiments, the image capture device 122 may be associated with a 6 - millimeter lens or a 12 - millimeter lens. In some embodiments, as Figure 2DAs shown, the image capture device 122 can be configured to capture an image with a desired field of view (FOV) 202. For example, the image capture device 122 can be configured to have a conventional FOV, such as in the range of 40 degrees to 56 degrees, including a 46-degree FOV, a 50-degree FOV, a 52-degree FOV, or a larger FOV. Alternatively, the image capture device 122 can be configured to have a narrow FOV in the range of 23 to 40 degrees, such as a 28-degree FOV or a 36-degree FOV. In addition, the image capture device 122 can be configured to have a wide FOV in the range of 100 to 180 degrees. In some embodiments, the image capture device 122 can include a wide-angle bumper camera or a camera with an FOV of up to 180 degrees. In some embodiments, the image capture device 122 can be a 7.2M (megapixel) image capture device with an aspect ratio of approximately 2:1 (e.g., H×V = 3800×1900 pixels) and a horizontal FOV of approximately 100 degrees. Such an image capture device can be used to replace three image capture device configurations. Due to significant lens distortion, in embodiments where the image capture device uses a radially symmetric lens, the vertical FOV of such an image capture device can be significantly less than 50 degrees. For example, such a lens can be non-radially symmetric, which would allow the vertical FOV to be greater than 50 degrees for a 100-degree horizontal FOV.

[0127] The first image capture device 122 can acquire a plurality of first images of a scene associated with the vehicle 200. Each of the plurality of first images can be acquired as a series of image scan lines, which can be captured using a rolling shutter. Each scan line can include a plurality of pixels.

[0128] The first image capture device 122 can have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate can refer to the rate at which the image sensor can acquire image data associated with each pixel included in a particular scan line.

[0129] The image capture devices 122, 124, and 126 can include any suitable type and number of image sensors, e.g., including CCD sensors or CMOS sensors, etc. In one embodiment, a CMOS image sensor and a rolling shutter can be employed such that each pixel in a row is read one at a time and the scanning of the rows continues on a row-by-row basis until an entire image frame has been captured. In some embodiments, the rows can be captured sequentially from the top to the bottom of the frame.

[0130] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124, and 126) may constitute a high-resolution imager and may have a resolution greater than 5M pixels, 7M pixels, 10M pixels, or more pixels.

[0131] The use of a rolling shutter may cause pixels in different rows to be exposed and captured at different times, which may cause distortions and other image artifacts in the captured image frame. On the other hand, when the image capture device 122 is configured to operate with a global or synchronous shutter, all pixels may be exposed for the same amount of time and during a common exposure period. As a result, the image data in a frame collected from a system employing a global shutter represents a snapshot of the entire FOV (such as FOV 202) at a particular time. In contrast, in a rolling shutter application, each row in the frame is exposed and data is captured at different times. Thus, in an image capture device with a rolling shutter, moving objects may appear distorted. This phenomenon will be described in more detail below.

[0132] The second image capture device 124 and the third image capture device 126 may be any type of image capture device. Similar to the first image capture device 122, each of the image capture devices 124 and 126 may include an optical axis. In one embodiment, each of the image capture devices 124 and 126 may include an Aptina M9V024WVGA sensor with a global shutter. Alternatively, each of the image capture devices 124 and 126 may include a rolling shutter. Similar to the image capture device 122, the image capture devices 124 and 126 may be configured to include various lenses and optical elements. In some embodiments, the lenses associated with the image capture devices 124 and 126 may provide FOVs (such as FOVs 204 and 206) that are equal to or narrower than the FOV (such as FOV 202) associated with the image capture device 122. For example, the image capture devices 124 and 126 may have an FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less.

[0133] The image capture devices 124 and 126 may acquire a plurality of second and third images of the scene associated with the vehicle 200. Each of the plurality of second and third images may be acquired as a second series of image scan lines and a third series of image scan lines, which may be captured using a rolling shutter. Each scan line or row may have a plurality of pixels. The image capture devices 124 and 126 may have a second scan rate and a third scan rate associated with the acquisition of each image scan line included in the second and third series.

[0134] Each of the image capture devices 122, 124, and 126 can be placed at any suitable position and orientation relative to the vehicle 200. The relative positions of the image capture devices 122, 124, and 126 can be selected to assist in fusing the information obtained from the image capture devices. For example, in some embodiments, the FOV associated with the image capture device 124 (such as FOV 204) may partially or completely overlap with the FOV associated with the image capture device 122 (such as FOV 202) and the FOV associated with the image capture device 126 (such as FOV 206).

[0135] The image capture devices 122, 124, and 126 can be located at any suitable relative height on the vehicle 200. In one example, there can be a height difference between the image capture devices 122, 124, and 126, which can provide sufficient parallax information to enable stereoscopic analysis. For example, as Figure 2A shown, two of the image capture devices 122 and 124 are at different heights. For example, there can also be a lateral displacement difference between the image capture devices 122, 124, and 126, giving additional parallax information for the stereoscopic analysis of the processing unit 110. As Figure 2C and Figure 2D shown, the difference in lateral displacement can be represented by d x In some embodiments, there may be a forward or backward displacement (e.g., a range displacement) between the image capture devices 122, 124, and 126. For example, the image capture device 122 can be located 0.5 to 2 meters or more behind the image capture device 124 and / or the image capture device 126. This type of displacement can enable one of the image capture devices to cover potential blind spots of the other image capture device(s).

[0136] The image capture device 122 can have any suitable resolution capability (e.g., the number of pixels associated with the image sensor), and the resolution of the image sensor(s) associated with the image capture device 122 can be higher, lower, or the same as the resolution of the image sensor(s) associated with the image capture devices 124 and 126. In some embodiments, the image sensor(s) associated with the image capture device 122 and / or the image capture devices 124 and 126 can have a resolution of 640×480, 1024×768, 1280×960, or any other suitable resolution.

[0137] The frame rate (e.g., at which the image capture device acquires a set of pixel data for an image frame and then proceeds to capture pixel data associated with the next image frame) can be controllable. The frame rate associated with image capture device 122 can be higher, lower, or the same as the frame rates associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 can depend on various factors that may affect the timing of the frame rate. For example, one or more of image capture devices 122, 124, and 126 can include selectable pixel delay periods that are applied before or after acquiring image data associated with one or more pixels of an image sensor in image capture devices 122, 124, and / or 126. Generally, image data corresponding to each pixel can be acquired according to the clock rate for the device (e.g., one pixel per clock cycle). Additionally, in embodiments including a rolling shutter, one or more of image capture devices 122, 124, and 126 can include selectable horizontal blanking periods that are applied before or after acquiring image data associated with a row of pixels of an image sensor in image capture devices 122, 124, and / or 126. Further, one or more of image capture devices 122, 124, and 126 can include selectable vertical blanking periods that are applied before or after acquiring image data associated with an image frame of image capture devices 122, 124, and 126.

[0138] These timing controls can enable synchronization of the frame rates associated with image capture devices 122, 124, and 126, even if the line scan rates of each are different. Additionally, as will be discussed in more detail below, these selectable timing controls and other factors (e.g., image sensor resolution, maximum line scan rate, etc.) can enable synchronization of image capture from regions where the FOV of image capture device 122 overlaps one or more FOVs of image capture devices 124 and 126, even if the field of view of image capture device 122 is different from the FOVs of image capture devices 124 and 126.

[0139] The frame rate timing in image capture devices 122, 124, and 126 can depend on the resolution of the associated image sensor. For example, assuming similar line scan rates for two devices, if one device includes an image sensor with a resolution of 640×480 and the other device includes an image sensor with a resolution of 1280×960, it takes more time to acquire one frame of image data from the higher resolution sensor.

[0140] Another factor that may affect the timing of image data acquisition in image capture devices 122, 124, and 126 is the maximum line scan rate. For example, it takes a certain minimum amount of time to acquire one line of image data from the image sensors included in image capture devices 122, 124, and 126. Assuming no pixel delay periods are added, this minimum amount of time for acquiring one line of image data will be related to the maximum line scan rate for a particular device. Devices that provide a higher maximum line scan rate have the potential to provide a higher frame rate than devices with a lower maximum line scan rate. In some embodiments, one or more of image capture devices 124 and 126 may have a maximum line scan rate that is higher than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture device 124 and / or 126 may be 1.25, 1.5, 1.75, or 2 times or more than the maximum line scan rate of image capture device 122.

[0141] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may operate at a scan rate that is less than or equal to its maximum scan rate. The system may be configured such that one or more of image capture devices 124 and 126 operate at a line scan rate that is equal to the line scan rate of image capture device 122. In other instances, the system may be configured such that the line scan rate of image capture device 124 and / or image capture device 126 may be 1.25, 1.5, 1.75, or 2 times or more than the line scan rate of image capture device 122.

[0142] In some embodiments, image capture devices 122, 124, and 126 may be asymmetric. That is, they may include cameras with different fields of view (FOV) and focal lengths. For example, the fields of view of image capture devices 122, 124, and 126 may include any desired regions of the environment of vehicle 200. In some embodiments, one or more of image capture devices 122, 124, and 126 may be configured to acquire image data from the environment in front of vehicle 200, behind vehicle 200, to the side of vehicle 200, or a combination thereof.

[0143] In addition, the focal length associated with each of the image capture devices 122, 124, and / or 126 may be selectable (e.g., by including an appropriate lens, etc.) such that each device captures an image of an object at a desired distance range relative to the vehicle 200. For example, in some embodiments, the image capture devices 122, 124, and 126 may capture images of nearby objects within a few meters of the vehicle. The image capture devices 122, 124, and 126 may also be configured to capture images of objects at a greater distance from the vehicle (e.g., 25 meters, 50 meters, 100 meters, 150 meters, or more). Additionally, the focal lengths of the image capture devices 122, 124, and 126 may be selected such that one image capture device (e.g., image capture device 122) can capture images of objects relatively close to the vehicle (e.g., within 10 meters or within 20 meters), while other image capture devices (e.g., image capture devices 124 and 126) can capture images of objects farther from the vehicle 200 (e.g., greater than 20 meters, 50 meters, 100 meters, 150 meters, etc.).

[0144] According to some embodiments, the FOV of one or more of the image capture devices 122, 124, and 126 may be wide-angle. For example, an FOV of 140 degrees may be advantageous, particularly for the image capture devices 122, 124, and 126 that may be used to capture images of areas near the vehicle 200. For example, the image capture device 122 may be used to capture images of areas to the right or left of the vehicle 200, and in these embodiments, it may be desirable for the image capture device 122 to have a wide FOV (e.g., at least 140 degrees).

[0145] The field of view associated with each of the image capture devices 122, 124, and 126 may depend on the respective focal length. For example, as the focal length increases, the corresponding field of view decreases.

[0146] The image capture devices 122, 124, and 126 may be configured to have any suitable field of view. In one particular example, the image capture device 122 may have a horizontal FOV of 46 degrees, the image capture device 124 may have a horizontal FOV of 23 degrees, and the image capture device 126 may have a horizontal FOV between 23 degrees and 46 degrees. In another example, the image capture device 122 may have a horizontal FOV of 52 degrees, the image capture device 124 may have a horizontal FOV of 26 degrees, and the image capture device 126 may have a horizontal FOV between 26 degrees and 52 degrees. In some embodiments, the ratio of the FOV of the image capture device 122 to the FOV of the image capture device 124 and / or the image capture device 126 may vary from 1.5 to 2.0. In other embodiments, the ratio may vary between 1.25 and 2.25.

[0147] System 100 may be configured such that the field of view of image capture device 122 at least partially or fully overlaps with the field of view of image capture device 124 and / or image capture device 126. In some embodiments, System 100 may be configured such that the fields of view of image capture devices 124 and 126, for example, fall within (e.g., are narrower than) the field of view of image capture device 122 and share a common center with the field of view of image capture device 122. In other embodiments, image capture devices 122, 124, and 126 may capture adjacent FOVs, or may have partial overlap in their FOVs. In some embodiments, the fields of view of image capture devices 122, 124, and 126 may be aligned such that the center of the narrower FOV image capture device 124 and / or 126 may be located in the lower half of the field of view of the wider FOV device 122.

[0148] Figure 2F is an illustrative representation of an example vehicle control system consistent with the disclosed embodiments. As Figure 2F indicated, vehicle 200 may include a throttle regulation system 220, a braking system 230, and a steering system 240. System 100 may provide inputs (e.g., control signals) to one or more of throttle regulation system 220, braking system 230, and steering system 240 via one or more data links (e.g., any wired and / or wireless link for transmitting data). For example, based on the analysis of images acquired by image capture devices 122, 124, and / or 126, System 100 may provide control signals to one or more of throttle regulation system 220, braking system 230, and steering system 240 to navigate vehicle 200 (e.g., by causing acceleration, steering, lane changes, etc.). Additionally, System 100 may receive inputs from one or more of throttle regulation system 220, braking system 230, and steering system 240 indicating the operating conditions of vehicle 200 (e.g., speed, whether vehicle 200 is braking and / or steering, etc.). Further details are provided below in connection with Figures 4 to 7 provide further details.

[0149] As Figure 3AAs shown, vehicle 200 may also include a user interface 170 for interacting with the driver or passengers of vehicle 200. For example, the user interface 170 in a vehicle application may include a touch screen 320, a knob 330, buttons 340, and a microphone 350. The driver or passengers of vehicle 200 may also use a handle (e.g., located on or near the steering column of vehicle 200, including, for example, a turn signal handle), buttons (e.g., located on the steering wheel of vehicle 200), etc. to interact with system 100. In some embodiments, the microphone 350 may be located adjacent to the rearview mirror 310. Similarly, in some embodiments, the image capture device 122 may be located near the rearview mirror 310. In some embodiments, the user interface 170 may also include one or more speakers 360 (e.g., speakers of a vehicle audio system). For example, system 100 may provide various notifications (e.g., alerts) via the speakers 360.

[0150] Figures 3B to 3D is an illustration of an example camera mount 370 configured to be located behind a rearview mirror (e.g., rearview mirror 310) and opposite a vehicle windshield, consistent with the disclosed embodiments. As Figure 3B shown, the camera mount 370 may include image capture devices 122, 124, and 126. The image capture devices 124 and 126 may be located behind the sun visor 380, where the sun visor 380 may be flush with the vehicle windshield and include a composite of thin film and / or anti-reflective material. For example, the sun visor 380 may be placed such that it is aligned with the vehicle windshield having a matching bevel. In some embodiments, each of the image capture devices 122, 124, and 126 may be located behind the sun visor 380, e.g., as depicted in Figure 3D . The disclosed embodiments are not limited to any particular configuration of the image capture devices 122, 124, and 126, the camera mount 370, and the sun visor 380. Figure 3C is Figure 3B an illustration of the camera mount 370 from a front perspective as shown.

[0151] As will be understood by those skilled in the art who benefit from this disclosure, many variations and / or modifications may be made to the foregoing disclosed embodiments. For example, not all components are necessary for the operation of system 100. Additionally, any component may be located in any suitable part of system 100 and the components may be rearranged into various configurations while providing the functions of the disclosed embodiments. Accordingly, the foregoing configurations are exemplary, and regardless of the configurations discussed above, system 100 may provide a wide range of functions to analyze the surroundings of vehicle 200 and navigate vehicle 200 in response to that analysis.

[0152] As discussed in more detail below and in accordance with various disclosed embodiments, system 100 may provide various features regarding autonomous driving and / or driver assistance technologies. For example, system 100 may analyze image data, location data (e.g., GPS location information), map data, rate data, and / or data from sensors included in vehicle 200. System 100 may collect data from, for example, image acquisition unit 120, location sensor 130, and other sensors for analysis. Additionally, system 100 may analyze the collected data to determine whether vehicle 200 should take a certain action and then automatically take the determined action without human intervention. For example, when vehicle 200 is navigating without human intervention, system 100 may automatically control the braking, acceleration, and / or steering of vehicle 200 (e.g., by sending control signals to one or more of throttle adjustment system 220, braking system 230, and steering system 240). Further, system 100 may analyze the collected data and issue warnings and / or alerts to vehicle occupants based on the analysis of the collected data. Additional details regarding various embodiments provided by system 100 are provided below.

[0153] Forward Multi - Imaging System

[0154] As discussed above, system 100 may provide driving assistance functions using a multi-camera system. The multi-camera system may use one or more cameras facing the front of the vehicle. In other embodiments, the multi-camera system may include one or more cameras facing the side or the rear of the vehicle. In one embodiment, for example, system 100 may use a dual-camera imaging system, where a first camera and a second camera (e.g., image capture devices 122 and 124) may be located in front of and / or at the side of the vehicle (e.g., vehicle 200). Other camera configurations are consistent with the disclosed embodiments, and the configurations disclosed herein are exemplary. For example, system 100 may include configurations with any number of cameras (e.g., one, two, three, four, five, six, seven, eight, etc.). Additionally, system 100 may include camera "clusters". For example, a camera cluster (including any suitable number of cameras, such as one, four, eight, etc.) may be forward-facing relative to the vehicle or may face any other direction (e.g., rearward, lateral, angled, etc.). Thus, system 100 may include multiple camera clusters, where each cluster is oriented in a specific direction to capture images from a specific area of the vehicle environment.

[0155] The first camera may have a field of view that is greater than, less than, or partially overlaps the field of view of the second camera. Additionally, the first camera may be connected to a first image processor to perform monocular image analysis on the images provided by the first camera, and the second camera may be connected to a second image processor to perform monocular image analysis on the images provided by the second camera. The outputs (e.g., processed information) of the first and second image processors may be combined. In some embodiments, the second image processor may receive images from both the first camera and the second camera to perform stereo analysis. In another embodiment, the system 100 may use a three-camera imaging system, where each camera has a different field of view. Thus, such a system may make decisions based on information obtained from objects at varying distances in front of and to the sides of the vehicle. References to monocular image analysis may refer to instances of performing image analysis based on images captured from a single viewpoint (e.g., from a single camera). Stereo image analysis may refer to instances of performing image analysis based on two or more images captured with one or more variations in the image capture parameters. For example, images captured suitable for performing stereo image analysis may include images captured from two or more different positions, from different fields of view, using different focal lengths, with disparity information, etc.

[0156] For example, in one embodiment, the system 100 may implement a three-camera configuration using image capture devices 122 through 126. In such a configuration, the image capture device 122 may provide a narrow field of view (e.g., 34 degrees or other values selected from a range of approximately 20 degrees to 45 degrees, etc.), the image capture device 124 may provide a wide field of view (e.g., 150 degrees or other values selected from a range of approximately 100 degrees to approximately 180 degrees), and the image capture device 126 may provide an intermediate field of view (e.g., 46 degrees or other values selected from a range of approximately 35 degrees to approximately 60 degrees). In some embodiments, the image capture device 126 may serve as the primary or base camera. The image capture devices 122 through 126 may be located behind the rearview mirror 310 and be substantially side-by-side (e.g., 6 centimeters apart). Additionally, in some embodiments, as discussed above, one or more of the image capture devices 122 through 126 may be mounted behind the sun visor 380 that is flush with the windshield of the vehicle 200. Such shielding may serve to reduce any reflections from inside the vehicle that affect the image capture devices 122 through 126.

[0157] In another embodiment, as described above in connection with Figure 3B and 3CAs discussed, the wide - field - of - view camera (e.g., the image capture device 124 in the above example) can be mounted lower than the narrow - field - of - view camera and the main - field - of - view camera (e.g., the image capture devices 122 and 126 in the above example). This configuration can provide a clear line of sight from the wide - field - of - view camera. To reduce reflections, the camera can be mounted close to the windshield of the vehicle 200, and a polarizer can be included on the camera to attenuate the reflected light.

[0158] The three - camera system can provide certain performance characteristics. For example, some embodiments can include the ability to verify the detection of an object by one camera based on the detection results from another camera. In the three - camera configuration discussed above, the processing unit 110 can include, for example, three processing devices (e.g., three EyeQ series processor chips as discussed above), where each processing device is dedicated to processing the images captured by one or more of the image capture devices 122 - 126.

[0159] In the three - camera system, the first processing device can receive images from both the main camera and the narrow - field - of - view camera, and perform visual processing on the narrow FOV camera to, for example, detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the first processing device can calculate the disparity of pixels between the images from the main camera and the narrow camera, and create a 3D reconstruction of the environment of the vehicle 200. Then the first processing device can combine the 3D reconstruction with 3D map data, or combine the 3D reconstruction with 3D information calculated based on information from another camera.

[0160] The second processing device can receive images from the main camera, and perform visual processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the second processing device can calculate the camera displacement, and calculate the disparity of pixels between consecutive images based on this displacement, and create a 3D reconstruction of the scene (e.g., structure from motion). The second processing device can send the structure - from - motion based on the 3D reconstruction to the first processing device for combination with the stereo 3D image.

[0161] The third processing device can receive images from the wide FOV camera, and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing device can also execute additional processing instructions to analyze the images to identify moving objects in the images, such as vehicles changing lanes, pedestrians, etc.

[0162] In some embodiments, enabling the flow of image-based information to be captured and processed independently can provide opportunities for providing redundancy in the system. Such redundancy can include, for example, using a first image capture device and the images processed from that device to verify and / or supplement the information obtained by capturing and processing image information from at least a second image capture device.

[0163] In some embodiments, system 100 uses two image capture devices (e.g., image capture devices 122 and 124) in providing navigation assistance for vehicle 200, and uses a third image capture device (e.g., image capture device 126) to provide redundancy and verify the analysis of data received from the other two image capture devices. For example, in such a configuration, image capture devices 122 and 124 can provide images for stereoscopic analysis by system 100 to navigate vehicle 200, while image capture device 126 can provide images for monocular analysis by system 100 to provide redundancy and verification of the information obtained based on the images captured from image capture device 122 and / or image capture device 124. That is, image capture device 126 (and the corresponding processing device) can be regarded as providing a redundant subsystem for providing an inspection of the analysis obtained from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Additionally, in some embodiments, the redundancy and verification of the received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors, information received from one or more transceivers external to the vehicle, etc.).

[0164] Those skilled in the art will recognize that the above camera configurations, camera placements, number of cameras, camera positions, etc. are merely examples. Without departing from the scope of the disclosed embodiments, these components and other components described with respect to the overall system can be assembled and used in a variety of different configurations. Further details regarding the use of multi-camera systems to provide driver assistance and / or autonomous vehicle functions are provided below.

[0165] Figure 4 is an exemplary functional block diagram of memories 140 and / or 150 that can store / program instructions for performing one or more operations consistent with the embodiments of the present disclosure. Although memory 140 is referred to below, those skilled in the art will recognize that the instructions can be stored in memory 140 and / or memory 150.

[0166] As Figure 4As shown, the memory 140 may store a monocular image analysis module 402, a stereo image analysis module 404, a speed and acceleration module 406, and a navigation response module 408. The disclosed embodiments are not limited to any particular configuration of the memory 140. Additionally, the application processor 180 and / or the image processor 190 may execute instructions stored in any of the modules 402 to 408 included in the memory 140. Those skilled in the art will understand that in the following discussion, references to the processing unit 110 may refer to the application processor 180 and the image processor 190 individually or collectively. Thus, any of the following processing steps may be performed by one or more processing devices.

[0167] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the set of images with additional sensing information (e.g., information from radar) to perform monocular image analysis. As described below in connection with Figures 5A to 5D what is described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such features as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, the system 100 (e.g., via the processing unit 110) may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in connection with the navigation response module 408.

[0168] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the set of images with additional sensing information (e.g., information from radar, lidar, etc.) to perform monocular image analysis. As described below in connection with Figures 5A to 5D what is described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such features as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, the system 100 (e.g., via the processing unit 110) may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in connection with determining the navigation response.

[0169] In one embodiment, the stereoscopic image analysis module 404 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform stereoscopic image analysis on a first set and a second set of images obtained by a combination of image capture devices selected from among the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the first set and the second set of images with additional sensing information (e.g., information from radar) to perform stereoscopic image analysis. For example, the stereoscopic image analysis module 404 may include instructions for performing stereoscopic image analysis based on a first set of images obtained by the image capture device 124 and a second set of images obtained by the image capture device 126. As described below in connection with Figure 6 what is described, the stereoscopic image analysis module 404 may include instructions for detecting a set of features within the first set and the second set of images, such features as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and the like. Based on this analysis, the processing unit 110 may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in connection with the navigation response module 408. Additionally, in some embodiments, the stereoscopic image analysis module 404 may implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.

[0170] In one embodiment, the speed and acceleration module 406 may store software configured to analyze data received from one or more computing and electromechanical devices in the vehicle 200 configured to cause a change in the speed and / or acceleration of the vehicle 200. For example, the processing unit 110 may execute instructions associated with the speed and acceleration module 406 to calculate a target rate of the vehicle 200 based on data obtained from the execution of the monocular image analysis module 402 and / or the stereoscopic image analysis module 404. Such data may include, for example, target position, speed and / or acceleration, the position and / or rate of the vehicle 200 relative to nearby vehicles, pedestrians, or road objects, position information of the vehicle 200 relative to the lane markings of the road, and the like. Additionally, the processing unit 110 may calculate the target rate of the vehicle 200 based on sensing input (e.g., information from radar) and inputs from other systems in the vehicle 200, such as the throttle regulation system 220, the braking system 230, and / or the steering system 240. Based on the calculated target rate, the processing unit 110 may transmit electrical signals to the throttle regulation system 220, the braking system 230, and / or the steering system 240 of the vehicle 200, for example, by physically depressing the brakes or releasing the accelerator of the vehicle 200, to trigger a change in speed and / or acceleration.

[0171] In one embodiment, the navigation response module 408 may store software that can be executed by the processing unit 110 to determine a desired navigation response based on data obtained from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include position and velocity information associated with nearby vehicles, pedestrians, and road objects, target position information of the vehicle 200, and the like. Additionally, in some embodiments, the navigation response may be (partially or fully) based on map data, a predetermined position of the vehicle 200, and / or the relative velocity or relative acceleration between the vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 may also determine a desired navigation response based on sensing inputs (e.g., information from radar) and inputs from other systems of the vehicle 200, such as the throttle regulation system 220, the braking system 230, and the steering system 240 of the vehicle 200. Based on the desired navigation response, the processing unit 110 may transmit electrical signals to the throttle regulation system 220, the braking system 230, and the steering system 240 of the vehicle 200 to trigger the desired navigation response, e.g., by rotating the steering wheel of the vehicle 200 to achieve a predetermined angular rotation. In some embodiments, the processing unit 110 may use the output of the navigation response module 408 (e.g., the desired navigation response) as an input to the execution of the speed and acceleration module 406 for calculating a change in the speed of the vehicle 200.

[0172] Furthermore, any module disclosed herein (e.g., modules 402, 404, and 406) may implement techniques associated with a trained system, such as a neural network or a deep neural network, or an untrained system.

[0173] Figure 5A is a flowchart showing an example process 500A for causing one or more navigation responses based on monocular image analysis consistent with the disclosed embodiments. At step 510, the processing unit 110 may receive a plurality of images via a data interface 128 between the processing unit 110 and the image acquisition unit 120. For example, a camera included in the image acquisition unit 120 (such as the image capture device 122 having a field of view 202) may capture a plurality of images of the area in front of the vehicle 200 (e.g., or the side or rear of the vehicle) and transmit them to the processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). At step 520, the processing unit 110 may execute the monocular image analysis module 402 to analyze the plurality of images, as further described in detail below in conjunction with Figures 5B to 5D By performing this analysis, the processing unit 110 may detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, etc.

[0174] At step 520, the processing unit 110 may also execute the monocular image analysis module 402 to detect various road hazards, such as for example parts of a truck tire, a fallen road sign, loose cargo, small animals, etc. Road hazards may vary in structure, shape, size, and color, which may make the detection of these hazards more difficult. In some embodiments, the processing unit 110 may execute the monocular image analysis module 402 to perform multi-frame analysis on the plurality of images to detect road hazards. For example, the processing unit 110 may estimate the camera motion between consecutive image frames and calculate the disparity in pixels between the frames to construct a 3D map of the road. Then, the processing unit 110 may use the 3D map to detect the road surface and the hazards present on the road surface.

[0175] At step 530, the processing unit 110 may execute the navigation response module 408 to cause one or more navigation responses based on the analysis performed in step 520 and as described above in connection with Figure 4 the techniques described. Navigation responses may include, for example, steering, lane changes, acceleration changes, etc. In some embodiments, the processing unit 110 may use data obtained from the execution of the speed and acceleration module 406 to cause one or more navigation responses. Additionally, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof. For example, the processing unit 110 may cause the vehicle 200 to change a lane and then accelerate by, for example, sequentially transmitting control signals to the steering system 240 and the throttle adjustment system 220 of the vehicle 200. Alternatively, the processing unit 110 may cause the vehicle 200 to brake and change lanes simultaneously by, for example, simultaneously transmitting control signals to the braking system 230 and the steering system 240 of the vehicle 200.

[0176] Figure 5B is a flowchart of an example process 500B for detecting one or more vehicles and / or pedestrians in a set of images consistent with the disclosed embodiments. The processing unit 110 may execute the monocular image analysis module 402 to implement process 500B. At step 540, the processing unit 110 may determine a set of candidate objects representing possible vehicles and / or pedestrians. For example, the processing unit 110 may scan one or more images, compare the image with one or more predetermined patterns, and identify possible locations within each image that may contain an object of interest (e.g., a vehicle, a pedestrian, or a part thereof). The predetermined patterns may be designed in such a way as to achieve a high "false hit" rate and a low "miss" rate. For example, the processing unit 110 may use a low similarity threshold with the predetermined patterns to identify candidate objects as possible vehicles or pedestrians. Doing so may allow the processing unit 110 to reduce the likelihood of missing (e.g., not identifying) candidate objects representing vehicles or pedestrians.

[0177] In step 542, the processing unit 110 may filter the set of candidate objects based on classification criteria to exclude certain candidates (e.g., irrelevant or less relevant objects). Such criteria may be derived from various attributes associated with the types of objects stored in a database (e.g., a database stored in the memory 140). The attributes may include object shape, size, texture, location (e.g., relative to the vehicle 200), etc. Thus, the processing unit 110 may use one or more sets of criteria to reject spurious candidates from the set of candidate objects.

[0178] In step 544, the processing unit 110 may analyze multiple frames of images to determine whether the objects in the set of candidate objects represent vehicles and / or pedestrians. For example, the processing unit 110 may track the detected candidate objects across consecutive frames and accumulate frame-by-frame data associated with the detected objects (e.g., size, location relative to the vehicle 200, etc.). Additionally, the processing unit 110 may estimate the parameters of the detected objects and compare the frame-by-frame position data of the object with the predicted position.

[0179] In step 546, the processing unit 110 may build a set of measurements for the detected objects. Such measurements may include, for example, position, velocity, and acceleration values (relative to the vehicle 200) associated with the detected objects. In some embodiments, the processing unit 110 may build the measurements based on estimation techniques such as Kalman filters or linear quadratic estimation (LQE) that use a series of time-based observations and / or based on modeling data available for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). The Kalman filter may be based on measurements of the scale of the object, where the scale measurement is proportional to the time to collision (e.g., the amount of time for the vehicle 200 to reach the object). Thus, by performing steps 540 to 546, the processing unit 110 may identify vehicles and pedestrians that appear within the set of captured images and obtain information associated with the vehicles and pedestrians (e.g., location, speed, size). Based on the identification and the obtained information, the processing unit 110 may cause one or more navigation responses in the vehicle 200, as described above in conjunction with Figure 5A as described.

[0180] At step 548, the processing unit 110 may perform optical flow analysis on one or more images to reduce the likelihood of detecting "false hits" and missing candidate objects representing vehicles or pedestrians. Optical flow analysis may refer to, for example, analyzing motion patterns in one or more images that are associated with other vehicles and pedestrians relative to vehicle 200 and are distinct from road surface motion. The processing unit 110 may calculate the motion of a candidate object by observing the different positions of the object across multiple image frames captured at different times. The processing unit 110 may use the position and time values as inputs to a mathematical model for calculating the motion of the candidate object. Thus, optical flow analysis may provide an alternative method for detecting vehicles and pedestrians near vehicle 200. The processing unit 110 may perform optical flow analysis in combination with steps 540 to 546 to provide redundancy in detecting vehicles and pedestrians and to improve the reliability of system 100.

[0181] Figure 5C is a flowchart showing an example process 500C for detecting road markings and / or lane geometry information in a set of images consistent with the disclosed embodiments. The processing unit 110 may execute the monocular image analysis module 402 to implement process 500C. At step 550, the processing unit 110 may detect a set of objects by scanning one or more images. To detect segments of lane markings, lane geometry information, and other related road markings, the processing unit 110 may filter the set of objects to exclude those determined to be irrelevant (e.g., small potholes, small stones, etc.). At step 552, the processing unit 110 may group together the segments detected in step 550 that belong to the same road marking or lane marking. Based on this grouping, the processing unit 110 may generate a model representing the detected segments, such as a mathematical model.

[0182] At step 554, the processing unit 110 may construct a set of measurements associated with the detected segments. In some embodiments, the processing unit 110 may create a projection of the detected segments from the image plane to the real-world plane. The projection may be characterized using a cubic polynomial having coefficients corresponding to physical attributes such as the position, slope, curvature, and curvature derivative of the detected road, for example. In generating the projection, the processing unit 110 may consider variations in the road surface, as well as the pitch rate and roll rate associated with vehicle 200. Additionally, the processing unit 110 may model the road elevation by analyzing position and motion cues present on the road surface. Additionally, the processing unit 110 may estimate the pitch rate and roll rate associated with vehicle 200 by tracking a set of feature points in one or more images.

[0183] At step 556, the processing unit 110 may perform multi-frame analysis by, for example, tracking the detected segments across consecutive image frames and accumulating frame-by-frame data associated with the detected segments. Since the processing unit 110 performs multi-frame analysis, the set of measurements constructed in step 554 may become more reliable and associated with an increasingly high confidence level. Thus, by performing steps 550 to 556, the processing unit 110 may identify road markings present in the set of captured images and obtain lane geometry information. Based on this identification and the obtained information, the processing unit 110 may cause one or more navigation responses in the vehicle 200, as described above in connection with Figure 5A as described.

[0184] At step 558, the processing unit 110 may consider additional information sources to further generate a safety model of the environment in which the vehicle 200 is located. The processing unit 110 may use this safety model to define the environment in which the system 100 may perform autonomous control of the vehicle 200 in a safe manner. To generate this safety model, in some embodiments, the processing unit 110 may consider the positions and movements of other vehicles, detected curbs and guardrails, and / or general road shape descriptions extracted from map data (such as data from the map database 160). By considering additional information sources, the processing unit 110 may provide redundancy for detecting road markings and lane geometry and increase the reliability of the system 100.

[0185] Figure 5D is a flowchart showing an example process 500D for detecting traffic lights in a set of images. The processing unit 110 may execute the monocular image analysis module 402 to implement the process 500D. At step 560, the processing unit 110 may scan the set of images and identify objects at positions in the images that may contain traffic lights. For example, the processing unit 110 may filter the identified objects to construct a set of candidate objects, excluding those objects that are unlikely to correspond to traffic lights. The filtering may be performed based on various attributes associated with traffic lights, such as shape, size, texture, position (e.g., relative to the vehicle 200), etc. Such attributes may be based on multiple examples of traffic lights and traffic control signals and stored in a database. In some embodiments, the processing unit 110 may perform multi-frame analysis on the set of candidate objects that reflect possible traffic lights. For example, the processing unit 110 may track candidate objects across consecutive image frames, estimate the real-world positions of the candidate objects, and filter out those moving objects (which are unlikely to be traffic lights). In some embodiments, the processing unit 110 may perform color analysis on the candidate objects and identify the relative positions of the detected colors that appear within the possible traffic lights.

[0186] In step 562, the processing unit 110 may analyze the geometry of the intersection. This analysis may be based on any combination of the following: (i) the number of lanes detected on either side of the vehicle 200, (ii) the markings detected on the road (such as arrow markings), and (iii) the description of the intersection extracted from map data (e.g., data from the map database 160). The processing unit 110 may perform the analysis using the information obtained from the execution of the monocular analysis module 402. Additionally, the processing unit 110 may determine the correspondence between the traffic lights detected in step 560 and the lanes present in the vicinity of the vehicle 200.

[0187] In step 564, as the vehicle 200 approaches the intersection, the processing unit 110 may update the confidence level associated with the analyzed intersection geometry and the detected traffic lights. For example, the comparison between the estimated number of traffic lights present at the intersection and the actual number of traffic lights present at the intersection may affect the confidence level. Thus, based on this confidence level, the processing unit 110 may delegate control to the driver of the vehicle 200 to improve safety conditions. By performing steps 560 to 564, the processing unit 110 may identify the traffic lights present within the set of captured images and analyze the intersection geometry information. Based on this identification and analysis, the processing unit 110 may cause one or more navigation responses in the vehicle 200, as described above in conjunction with Figure 5A that described.

[0188] Figure 5E is a flowchart showing an example process 500E for causing one or more navigation responses in a vehicle based on a vehicle path, consistent with the disclosed embodiments. In step 570, the processing unit 110 may construct an initial vehicle path associated with the vehicle 200. The vehicle path may be represented using a set of points expressed in coordinates (x, z), and the distance d between two points in the set of points i may fall within the range of 1 to 5 meters. In one embodiment, the processing unit 110 may use two polynomials, such as a left road polynomial and a right road polynomial, to construct the initial vehicle path. The processing unit 110 may calculate the geometric midpoint between the two polynomials and offset each point to be included in the resulting vehicle path by a predetermined offset (e.g., a smart lane offset), if any (a zero offset may correspond to driving in the middle of the lane). The offset may be in a direction perpendicular to the line segment between any two points in the vehicle path. In another embodiment, the processing unit 110 may use one polynomial and the estimated lane width to offset each point of the vehicle path by half of the estimated lane width plus a predetermined offset (e.g., a smart lane offset).

[0189] At step 572, the processing unit 110 may update the vehicle path constructed at step 570. The processing unit 110 may use a higher resolution to reconstruct the vehicle path constructed at step 570 such that the distance d between two points in the set of points representing the vehicle path k is less than the above distance d i . For example, the distance d k may fall within the range of 0.1 to 0.3 meters. The processing unit 110 may reconstruct the vehicle path using a parabolic spline algorithm, which may generate a cumulative distance vector S corresponding to the total length of the vehicle path (i.e., based on the set of points representing the vehicle path).

[0190] At step 574, the processing unit 110 may determine a look-ahead point (expressed in coordinates as (x l , z l )) based on the updated vehicle path constructed at step 572. The processing unit 110 may extract the look-ahead point from the cumulative distance vector S, and the look-ahead point may be associated with a look-ahead distance and a look-ahead time. The look-ahead distance may have a lower limit ranging from 10 meters to 20 meters and may be calculated as the product of the speed of the vehicle 200 and the look-ahead time. For example, as the speed of the vehicle 200 decreases, the look-ahead distance may also decrease (e.g., until it reaches the lower limit). The look-ahead time may range from 0.5 to 1.5 seconds and may be inversely proportional to the gain of one or more control loops such as a heading error tracking control loop associated with causing a navigation response in the vehicle 200. For example, the gain of the heading error tracking control loop may depend on the bandwidths of a yaw rate loop, a steering actuator loop, vehicle lateral dynamics, etc. Thus, the higher the gain of the heading error tracking control loop, the shorter the look-ahead time.

[0191] At step 576, the processing unit 110 may determine a heading error and a yaw rate command based on the look-ahead point determined at step 574. The processing unit 110 may determine the heading error by calculating the arctangent of the look-ahead point, e.g., arctan(x l , z l ). The processing unit 110 may determine the yaw rate command as the product of the heading error and a high-level control gain. If the look-ahead distance is not at the lower limit, the high-level control gain may be equal to: (2 / look-ahead time). Otherwise, the high-level control gain may be equal to: (2 × speed of the vehicle 200 / look-ahead distance).

[0192] Figure 5Fis a flowchart showing an example process 500F for determining whether a vehicle ahead is changing lanes in accordance with the disclosed embodiments. At step 580, the processing unit 110 may determine navigation information associated with the vehicle ahead (e.g., a vehicle traveling in front of vehicle 200). For example, the processing unit 110 may use the techniques described above in connection with Figure 5A and Figure 5B to determine the position, speed (e.g., direction and rate), and / or acceleration of the vehicle ahead. The processing unit 110 may also use the techniques described above in connection with Figure 5E to determine one or more road polynomials, a front view point (associated with vehicle 200), and / or a snail trail (e.g., a set of points describing the path taken by the vehicle ahead).

[0193] At step 582, the processing unit 110 may analyze the navigation information determined at step 580. In one embodiment, the processing unit 110 may calculate the distance between the snail trail and the road polynomial (e.g., along the trail). If the variance of this distance along the trail exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters on a straight road, 0.3 to 0.4 meters on a moderately curved road, and 0.5 to 0.6 meters on a sharp turn road), then the processing unit 110 may determine that the vehicle ahead is likely changing lanes. In a situation where multiple vehicles are detected traveling in front of vehicle 200, the processing unit 110 may compare the snail trails associated with each vehicle. Based on this comparison, the processing unit 110 may determine that a vehicle whose snail trail does not match the snail trails of other vehicles is likely changing lanes. The processing unit 110 may additionally compare the curvature of the snail trail (associated with the vehicle ahead) with the expected curvature of the road segment in which the vehicle ahead is traveling. The expected curvature may be extracted from map data (e.g., data from map database 160), from the road polynomial, from the snail trails of other vehicles, from existing knowledge about the road, etc. If the difference between the curvature of the snail trail and the expected curvature of the road segment exceeds a predetermined threshold, then the processing unit 110 may determine that the vehicle ahead is likely changing lanes.

[0194] In another embodiment, the processing unit 110 may compare the instantaneous position of the vehicle ahead with the front viewing point (associated with vehicle 200) over a specific time period (e.g., 0.5 to 1.5 seconds). If the distance between the instantaneous position of the vehicle ahead and the front viewing point changes during this specific time period and the cumulative sum of the changes exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a moderately curved road, and 1.3 to 1.7 meters on a sharp turn road), then the processing unit 110 may determine that the vehicle ahead is likely changing lanes. In another embodiment, the processing unit 110 may analyze the geometry of the tracking trajectory by comparing the lateral distance traveled along the tracking trajectory with the desired curvature of the tracking trajectory. The desired radius of curvature can be determined according to the formula: (δ z 2 +δ x 2 ) / 2 / (δ x ), where δ x represents the lateral distance traveled and δ z represents the longitudinal distance traveled. If the difference between the lateral distance traveled and the desired curvature exceeds a predetermined threshold (e.g., 500 to 700 meters), then the processing unit 110 may determine that the vehicle ahead is likely changing lanes. In another embodiment, the processing unit 110 may analyze the position of the vehicle ahead. If the position of the vehicle ahead obscures the road polynomial (e.g., the leading vehicle covers above the road polynomial), then the processing unit 110 may determine that the vehicle ahead is likely changing lanes. In the case where the position of the vehicle ahead is such that another vehicle is detected ahead of the vehicle ahead and the tracking trajectories of the two vehicles are not parallel, the processing unit 110 may determine that the (nearer) vehicle ahead is likely changing lanes.

[0195] In step 584, the processing unit 110 may determine whether the vehicle ahead 200 is changing lanes based on the analysis performed in step 582. For example, the processing unit 110 may make this determination based on a weighted average of the individual analyses performed in step 582. In such a scenario, for example, a determination that the vehicle ahead is likely changing lanes made by the processing unit 110 based on a particular type of analysis may be assigned a value of "1" (and "0" to represent a determination that the vehicle ahead is unlikely changing lanes). Different analyses performed in step 582 may be assigned different weights, and the disclosed embodiments are not limited to any particular combination of analyses and weights. Additionally, in some embodiments, the analysis may utilize a trained system (e.g., a machine learning or deep learning system) that may, for example, estimate the future path ahead of the vehicle's current position based on an image captured at the current position.

[0196] Figure 6FIG. 600 is a flow chart showing an example process for causing one or more navigation responses based on stereo image analysis consistent with the disclosed embodiments. At step 610, the processing unit 110 may receive the first and second pluralities of images via the data interface 128. For example, cameras included in the image acquisition unit 120 (such as the image capture devices 122 and 124 having fields of view 202 and 204) may capture the first and second pluralities of images of the area in front of the vehicle 200 and transmit them to the processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, the processing unit 110 may receive the first and second pluralities of images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.

[0197] At step 620, the processing unit 110 may execute the stereo image analysis module 404 to perform stereo image analysis on the first and second pluralities of images to create a 3D map of the road in front of the vehicle and detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. The stereo image analysis may be performed in a manner similar to the steps described above in connection with Figures 5A - 5D For example, the processing unit 110 may execute the stereo image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road markings, traffic lights, road hazards, etc.) within the first and second pluralities of images, filter out a subset of the candidate objects based on various criteria, and perform multi-frame analysis, construct measurements, and determine confidence levels on the remaining candidate objects. In performing the above steps, the processing unit 110 may consider information from both the first and second pluralities of images, rather than information from a single set of images. For example, the processing unit 110 may analyze the differences in pixel-level data (or other subsets of data from the two streams of captured images) of candidate objects that appear in both the first and second pluralities of images. As another example, the processing unit 110 may estimate the position and / or speed of a candidate object (e.g., relative to the vehicle 200) by observing that the candidate object appears in one of the pluralities of images but not the other, or other differences that may exist relative to an object that may appear in both image streams. For example, the position, speed, and / or acceleration relative to the vehicle 200 may be determined based on features such as the trajectory, position, movement characteristics, etc. associated with an object that appears in one or both of the image streams.

[0198] At step 630, the processing unit 110 may execute the navigation response module 408 to cause one or more navigation responses based on the analysis performed at step 620 and as described above in connection with Figure 4The described techniques cause one or more navigation responses in vehicle 200. Navigation responses can include, for example, steering, lane changes, changes in acceleration, changes in speed, braking, etc. In some embodiments, processing unit 110 can use data obtained from the execution of speed and acceleration module 406 to cause the one or more navigation responses. Additionally, multiple navigation responses can occur simultaneously, sequentially, or in any combination thereof.

[0199] Figure 7 FIG. 7 is a flow chart illustrating an example process 700 for causing one or more navigation responses based on an analysis of three sets of images consistent with the disclosed embodiments. In step 710, processing unit 110 can receive first, second, and third pluralities of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) can capture first, second, and third pluralities of images of regions in front of and / or to the sides of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 can receive first, second, and third pluralities of images via three or more data interfaces. For example, each of image capture devices 122, 124, 126 can have an associated data interface for transmitting data to processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.

[0200] In step 720, processing unit 110 can analyze the first, second, and third pluralities of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. The analysis can be performed in a manner similar to the steps described above in connection with Figures 5A - 5D and Figure 6 For example, processing unit 110 can perform monocular image analysis on each of the first, second, and third pluralities of images (e.g., via the execution of monocular image analysis module 402 and based on the steps described above in connection with Figures 5A - 5D Alternatively, processing unit 110 can perform stereo image analysis on the first and second pluralities of images, the second and third pluralities of images, and / or the first and third pluralities of images (e.g., via the execution of stereo image analysis module 404 and based on the steps described above in connection with Figure 6The described steps). The processed information corresponding to the analysis of the first, second, and / or third plurality of images can be combined. In some embodiments, the processing unit 110 can perform a combination of monocular and stereo image analysis. For example, the processing unit 110 can perform monocular image analysis on the first plurality of images (e.g., via the execution of the monocular image analysis module 402) and perform stereo image analysis on the second and third plurality of images (e.g., via the execution of the stereo image analysis module 404). The configurations of the image capture devices 122, 124, and 126 - including their respective positions and fields of view 202, 204, and 206 - can affect the type of analysis performed on the first, second, and third plurality of images. The disclosed embodiments are not limited to a particular configuration of the image capture devices 122, 124, and 126 or the type of analysis performed on the first, second, and third plurality of images.

[0201] In some embodiments, the processing unit 110 can perform tests on the system 100 based on the images acquired and analyzed in steps 710 and 720. Such tests can provide an indicator of the overall performance of the system 100 for certain configurations of the image acquisition devices 122, 124, and 126. For example, the processing unit 110 can determine the ratios of "false positives" (e.g., situations where the system 100 incorrectly determines the presence of a vehicle or pedestrian) and "misses".

[0202] In step 730, the processing unit 110 can cause one or more navigation responses in the vehicle 200 based on information obtained from two of the first, second, and third plurality of images. The selection of two of the first, second, and third plurality of images can depend on various factors, such as, for example, the number, type, and size of the objects detected in each of the plurality of images. The processing unit 110 can also make the selection based on image quality and resolution, the effective field of view reflected in the image, the number of frames captured, the extent to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which an object appears, the proportion of objects in each such frame), etc.

[0203] In some embodiments, the processing unit 110 may select information obtained from two of the first, second, and third plurality of images by determining the degree of consistency between information obtained from one image source and information obtained from other image sources. For example, the processing unit 110 may combine the processed information obtained from each of the image capture devices 122, 124, and 126 (whether by monocular analysis, stereoscopic analysis, or any combination of both) and determine visual indicators of consistency (e.g., lane markings, detected vehicles and their positions and / or paths, detected traffic lights, etc.) between the images captured from each of the image capture devices 122, 124, and 126. The processing unit 110 may also exclude information that is inconsistent between the captured images (e.g., a vehicle changing lanes, a lane model indicating that a vehicle is too close to vehicle 200, etc.). Thus, the processing unit 110 may select information obtained from two of the first, second, and third plurality of images based on the determination of consistent and inconsistent information.

[0204] The navigation response may include, for example, steering, lane changes, changes in acceleration, etc. The processing unit 110 may cause one or more navigation responses based on the analysis performed in step 720 and the techniques described above in connection with Figure 4 The techniques described cause one or more navigation responses. The processing unit 110 may also cause one or more navigation responses using data obtained from the execution of the speed and acceleration module 406. In some embodiments, the processing unit 110 may cause one or more navigation responses based on the relative position, relative speed, and / or relative acceleration between vehicle 200 and an object detected within any of the first, second, and third plurality of images. The plurality of navigation responses may occur simultaneously, sequentially, or in any combination thereof.

[0205] Reinforcement Learning and Trained Navigation System

[0206] The following sections discuss autonomous driving and systems and methods for achieving autonomous control of a vehicle, whether the control is fully autonomous (self-driving vehicle) or partially autonomous (e.g., one or more driver assistance systems or functions). As Figure 8As shown, the autonomous driving task can be divided into three main modules, including a sensing module 801, a driving strategy module 803, and a control module 805. In some embodiments, the modules 801, 803, and 805 can be stored in the memory unit 140 and / or the memory unit 150 of the system 100, or the modules 801, 803, and 805 (or portions thereof) can be stored remotely from the system 100 (e.g., stored in a server accessible by the system 100 via, for example, a wireless transceiver 172). Additionally, any module disclosed herein (e.g., modules 801, 803, and 805) can implement techniques associated with a trained system, such as a neural network or a deep neural network, or an untrained system.

[0207] The sensing module 801, which can be implemented using the processing unit 110, can process various tasks related to sensing the navigation state in the environment of the host vehicle. These tasks can rely on inputs from various sensors and sensing systems associated with the host vehicle. These inputs can include images or image streams from one or more on-vehicle cameras, GPS location information, accelerometer outputs, user feedback or user input to one or more user interface devices, radar, lidar, etc. The sensing, which can include data from cameras and / or any other available sensors as well as map information, can be collected, analyzed, and formulated into a "sensed state" that describes the information extracted from the scene in the environment of the host vehicle. The sensed state can include sensed information related to target vehicles, lane markings, pedestrians, traffic lights, road geometry, lane shape, obstacles, distance to other objects / vehicles, relative speed, relative acceleration, and any other potential sensed information. Supervised machine learning can be implemented to generate a sensed state output based on the sensed data provided to the sensing module 801. The output of the sensing module can represent the sensed navigation "state" of the host vehicle, which can be passed to the driving strategy module 803.

[0208] Although a sensed state may be generated based on image data received from one or more cameras or image sensors associated with the host vehicle, any suitable sensor or combination of sensors may also be used to generate a sensed state for use in navigation. In some embodiments, a sensed state may be generated without relying on captured image data. In fact, any of the navigation principles described herein may be applied to a sensed state generated based on captured image data as well as a sensed state generated using other non-image-based sensors. A sensed state may also be determined via a source external to the host vehicle. For example, a sensed state may be generated based, in whole or in part, on information received from a source remote from the host vehicle (e.g., sensor information, processed state information, etc. shared from other vehicles, shared from a central server, or from any other information source related to the navigation state of the host vehicle).

[0209] The driving strategy module 803 (discussed in more detail below and may be implemented using the processing unit 110) may implement a desired driving strategy to determine one or more navigation actions for the host vehicle to take in response to the sensed navigation state. If there are no other agents (e.g., target vehicles or pedestrians) in the environment of the host vehicle, the sensed state input to the driving strategy module 803 may be processed in a relatively straightforward manner. The task becomes more complex when the sensed state requires negotiation with one or more other agents. Techniques for generating the output of the driving strategy module 803 may include reinforcement learning (discussed in more detail below). The output of the driving strategy module 803 may include at least one navigation action for the host vehicle and may include a desired acceleration (which may translate to an updated speed of the host vehicle), a desired yaw rate of the host vehicle, a desired trajectory, and other potential desired navigation actions.

[0210] Based on the output from the driving strategy module 803, the control module 805, which may also be implemented using the processing unit 110, may generate control instructions for one or more actuators or controlled devices associated with the host vehicle. Such actuators and devices may include an accelerator, one or more steering controllers, brakes, signal transmitters, displays, or any other actuator or device that may be controlled as part of the navigation operation associated with the host vehicle. Aspects of control theory may be used to generate the output of the control module 805. The control module 805 may be responsible for generating and outputting instructions to the controllable components of the host vehicle to implement the desired navigation goals or requirements of the driving strategy module 803.

[0211] Returning to the driving policy module 803, in some embodiments, the driving policy module 803 can be implemented using a trained system trained by reinforcement learning. In other embodiments, the driving policy module 803 can be implemented by "manually" solving various scenarios that may arise during autonomous navigation by using a specified algorithm without using machine learning methods. However, while this approach is feasible, it may lead to an oversimplification of the driving policy and may lack the flexibility of a machine learning-based trained system. For example, a trained system may be better equipped to handle complex navigation states and can better determine whether a taxi is stopping or pulling over to pick up / drop off a passenger; determine whether a pedestrian intends to cross the street in front of the main vehicle; defensively balance the unexpected behavior of other drivers; negotiate in dense traffic involving target vehicles and / or pedestrians; decide when to suspend certain navigation rules or augment other rules; anticipate un-sensed but expected conditions (e.g., whether a pedestrian will emerge from behind a car or obstacle); and so on. A trained system based on reinforcement learning can also be better equipped to solve continuous and high-dimensional state spaces as well as continuous action spaces.

[0212] Training the system using reinforcement learning can involve learning a driving policy to map from sensed states to navigation actions. The driving policy is a function π: S → A, where S is a set of states, and is the action space (e.g., desired speed, acceleration, yaw command, etc.). The state space is S = S s × S p , where S s is the sensed state, and S p is additional information about the state saved by the policy. Working at discrete time intervals, at time t, the current state s t ∈ S can be observed, and the policy can be applied to obtain the desired action a t = π(s t ).

[0213] The system can be trained by exposing it to various navigation states, having it apply the policy, and providing a reward (based on a reward function designed to reward desired navigation behavior). Based on the reward feedback, the system can "learn" the policy and become trained in generating desired navigation actions. For example, the learning system can observe the current state s t ∈ SA and decide on an action a ∈ A based on the policy t . Based on the decided action (and the implementation of that action), the environment moves to the next state s t+1∈ S and is observable by the learning system. For each action generated in response to the observed state, the feedback to the learning system is a reward signal r 1 , r 2 ,...

[0214] The goal of reinforcement learning (RL) is to find a policy π. It is usually assumed that at time t, there exists a reward function r t , which measures the instantaneous quality of being in state s t and taking action a t . However, taking action a t at time t affects the environment and thus affects the value of future states. Therefore, when deciding which action to take, not only the current reward but also the future rewards should be considered. In some cases, when the system determines that if a lower-reward option is adopted now, a higher reward may be achieved in the future, then the system should take that action even if an action is associated with a reward lower than another available option. To formalize this, it is observed that the policy π and the initial state s induce a distribution over , where if the agent starts in state s 0 = s and follows the policy π from there, the probability of the vector (r 1 ,..., r T ) is the probability of observing the rewards r 1 ,..., r T . The value of the initial state s can be defined as:

[0215]

[0216] Instead of limiting the time horizon to T, the future rewards can be discounted for some fixed γ ∈ (0, ), and defined as:

[0217]

[0218] In any case, the optimal policy is the solution to the following:

[0219]

[0220] where the expectation is over the initial state s

[0221] There are several possible methods for training a driving policy system. For example, an imitation method (e.g., behavioral cloning) can be used, in which the system learns from state / action pairs, where the actions are those that a good agent (e.g., a human) would choose in response to a particular observed state. Suppose a human driver is observed. Through this observation, many pairs of the form (s t , at ) example where s t is the state and a t is the action of a human driver) can be obtained, observed, and used as the basis for training a driving policy system. For example, supervised learning can be used to learn a policy π such that π(s t ) ≈ a t . This approach has many potential advantages. First, there is no need to define a reward function. Second, the learning is supervised and occurs offline (no need to apply an actor during the learning process). The disadvantages of this approach are that different human drivers, and even the same human driver, are not deterministic in their policy choices. Therefore, learning a function for which ||π(s t ) - a t || is very small is usually infeasible. Moreover, over time, even small errors can accumulate and result in large errors.

[0222] Another technique that can be employed is policy-based learning. Here, the policy can be expressed in parametric form, and appropriate optimization techniques (e.g., stochastic gradient descent) are used to directly optimize the policy. This method directly solves the problem given by . Of course, there are many ways to solve this problem. One advantage of this method is that it directly addresses the problem and thus often leads to good practical results. A potential disadvantage is that it often requires "on-policy" training, i.e., the learning of π is an iterative process where, at iteration j, there is a non-perfect policy π j , and in order to construct the next policy π j , it is necessary to interact with the environment while acting based on π j .

[0223] The system can also be trained through value-based learning (learning the Q or V function). Suppose a good approximation can learn the optimal value function V*. The optimal policy can be constructed (e.g., relying on the Bellman equation). Some versions of value-based learning can be implemented offline (referred to as "off-policy" training). Some of the disadvantages of value-based methods may be due to their strong dependence on Markovian assumptions and the need to approximate complex functions (approximating the value function may be more difficult than directly approximating the policy).

[0224] Another technique can include model-based learning and planning (learning the probability of state transitions and solving the optimization problem of finding the optimal V). Combinations of these techniques can also be used to train the learning system. In this method, the dynamics of the learning process can be learned, i.e., taking (s t , at ) and a function that generates a distribution in the next state s t+1 . Once this function is learned, the optimal problem can be solved to find the policy π whose value is optimal. This is what is called "planning". One advantage of this method may lie in that the learning part is supervised and can be applied offline by observing the triples (s t , a t , s t+1 ). Similar to the "imitation" method, one disadvantage of this method may lie in that small errors in the learning process may accumulate and produce a policy that is not fully executed.

[0225] Another method for training the driving policy module 803 may include decomposing the driving policy function into semantically meaningful components. This allows for the manual implementation of some parts of the policy, which can ensure the safety of the policy; and reinforcement learning techniques can be used to implement other parts of the policy, which can adaptively achieve a human-like balance between many scenarios, defensive / aggressive behaviors, and human-like negotiation with other drivers. From a technical perspective, the reinforcement learning method can combine several methods and provide an easy-to-operate training program, most of which can be executed using recorded data or self-built simulators.

[0226] In some embodiments, the training of the driving policy module 803 can rely on the "option" mechanism. For illustration, consider the simple scenario of the driving policy on a two-lane highway. In the direct RL method, the policy π maps the state to , where the first component of π(s) is the desired acceleration command and the second component of π(s) is the yaw rate. In the modified method, the following policies can be constructed:

[0227] An adaptive cruise control (ACC) policy, o ACC : S → A: This policy always outputs a yaw rate of 0 and only changes the speed to achieve smooth and accident-free driving.

[0228] ACC + left policy, o L : S → A: The longitudinal command of this policy is the same as the ACC command. The yaw rate is a direct implementation that centers the vehicle in the middle of the left lane while ensuring safety (for example, if there is a vehicle on the left, do not move to the left).

[0229] ACC + right policy, o R : S → A: The same as o L , but the vehicle can be centered in the middle of the right lane.

[0230] These policies can be called "options". Relying on these "options", a policy π that selects the option can be learnedo : S → O, where O is a set of available options. In one case, O = {o ACC , o L , o R}. The option selector policy π o defines the actual policy π: S → A by setting for each s .

[0231] In practice, the policy function can be decomposed into an option graph 901, as Figure 9 shown. Figure 10 Another example option graph 1000 is shown in. The option graph can represent a hierarchical decision set organized as a directed acyclic graph (DAG). There is a special node called the root node 903 of the graph. This node has no incoming nodes. The decision-making process starts from the root node and traverses the entire graph until it reaches a "leaf" node, which is a node that has no outgoing decision lines. As Figure 9 shown, the leaf nodes can include, for example, nodes 905, 907, and 909. When a leaf node is encountered, the driving policy module 803 can output acceleration and steering commands associated with the desired navigation action associated with the leaf node.

[0232] Internal nodes (such as nodes 911, 913, and 915) can, for example, cause the implementation of a policy that selects a child among its available options. The set of available children of an internal node includes all nodes associated with a particular internal node via decision lines. For example, Figure 9 the internal node 913 designated as "merge" in includes three child nodes 909, 915, and 917 (respectively "hold", "overtake on the right", and "overtake on the left") each connected to node 913 via a decision line.

[0233] Flexibility in the decision-making system can be obtained by enabling nodes to adjust their position in the hierarchy of the option graph. For example, any node can be allowed to declare itself "critical". Each node can implement a function "is critical" that outputs "true" if the node is in a critical part of its policy implementation. For example, a node responsible for overtaking can declare itself critical midway through the maneuver. This can impose a constraint on the set of available children of node u, which can include all nodes v that are children of node u and there is a path from v to a leaf node passing through all nodes designated as critical. On the one hand, this approach can allow the declaration of a desired path on the graph at each time step, while on the other hand, it can preserve the stability of the policy, especially when the critical part of the policy is being implemented.

[0234] By defining an option graph, the problem of learning a driving policy π: S→A can be decomposed into the problem of defining a policy for each node in the graph, where the policy at the internal node should be selected from among the available child nodes. For some nodes, the corresponding policy can be implemented manually (e.g., by specifying a set of actions in response to the observed state through an if-then type algorithm), while for other policies a trained system built by reinforcement learning can be used for implementation. The choice between manual or trained / learned approaches can depend on the safety aspects associated with the task and on its relative simplicity. The option graph can be constructed in such a way that some nodes are implemented directly, while other nodes can rely on trained models. This approach can ensure the safe operation of the system.

[0235] The following discussion provides information on how to use the driving strategy module 803. Figure 9 Further details on the role of the option map in . As described above, the input to the driving policy module is a "sensed state" that summarizes the environment map, for example, obtained from available sensors. The output of the driving policy module 803 is a set of expectations (optionally together with a set of hard constraints) that define the trajectory as a solution to the optimization problem.

[0236] As described above, the option graph represents a hierarchical set of decisions organized as a DAG. There is a special node called the "root" of the graph. The root node is the only node with no incoming edges (e.g., decision lines). The decision process traverses the graph starting from the root node until it reaches a "leaf" node, i.e., a node with no outgoing edges. Each internal node should implement a strategy that selects a child from its available children. Each leaf node should implement a strategy that defines a set of expectations (e.g., a set of navigation goals for the host vehicle) based on the entire path from the root to the leaf. This set of expectations, which is defined directly based on the sensed states, together with a set of hard constraints, establishes an optimization problem, the solution of which is the trajectory of the vehicle. Hard constraints can be employed to further improve the safety of the system, and the expectations can be used to provide driving comfort and human-like driving behavior of the system. The trajectory provided as a solution to the optimization problem in turn defines the commands that should be provided to the steering, braking, and / or engine actuators in order to complete the trajectory.

[0237] Back Figure 9 ​, Option diagram 901 represents an option diagram for a two-lane highway, including merge lanes (meaning that at certain points, a third lane merges into the right or left lane of the highway). Root node 903 first decides whether the host vehicle is in a normal road scenario or approaching a lane change scenario. This is an example of a decision that can be implemented based on the sensed state. The normal road node 911 includes three child nodes: the maintain node 909, the left passing node 917, and the right passing node 915. Maintaining means the situation where the host vehicle wants to continue driving in the same lane. The maintain node is a leaf node (without outgoing edges / lines). Therefore, the maintain node defines a set of expectations. The first expectation it defines can include the desired lateral position - for example, as close as possible to the center of the currently traveled lane. Smooth navigation can also be expected (e.g., within a predetermined or allowed maximum acceleration). The maintain node can also define how the host vehicle responds to other vehicles. For example, the maintain node can look at the sensed target vehicles and assign a semantic meaning to each vehicle, which can be converted into components of a trajectory.

[0238] Various semantic meanings can be assigned to target vehicles in the environment of the host vehicle. For example, in some embodiments, the semantic meaning can include any of the following designations: 1) Irrelevant: Indicates that the vehicle sensed in the scene is currently irrelevant; 2) Next lane: Indicates that the sensed vehicle is in an adjacent lane and an appropriate offset should be maintained relative to that vehicle (the exact offset can be calculated in an optimization problem that constructs a trajectory given the expectations and hard constraints, and the exact offset can potentially be vehicle-dependent - the maintain leaf of the option diagram sets the semantic type of the target vehicle, which defines the expectations relative to the target vehicle); 3) Yield: The host vehicle will try to yield to the sensed target vehicle, for example, by reducing speed (especially in the case where the host vehicle determines that the target vehicle may cut into the host vehicle's lane); 4) Takeway: The host vehicle will try to take the right of way, for example, by increasing speed; 5) Follow: The host vehicle expects to follow the target vehicle and drive smoothly; 6) Left / right passing: This means that the host vehicle wants to initiate a lane change to the left or right lane. The left passing node 917 and the left passing node 915 are internal nodes where the expectations have not been defined yet.

[0239] The next node in option graph 901 is the Select Gap node 919. This node can be responsible for selecting a gap between two target vehicles in a specific target lane that the host vehicle desires to enter. By selecting a node of form IDj, for some value of j, the host vehicle reaching the leaf specifies the desired leaf for the trajectory optimization problem - for example, the host vehicle wishes to perform a maneuver to reach the selected gap. Such a maneuver may involve first accelerating / braking in the current lane and then going to the target lane at the appropriate time to enter the selected gap. If the Select Gap node 919 cannot find an appropriate gap, it moves to the Abort node 921, which defines the desire to return to the center of the current lane and cancel the overtaking.

[0240] Returning to the merge node 913, when the host vehicle approaches a merge, it has several options that can depend on the specific situation. For example, as Figure 11A shown, the host vehicle 1105 is traveling along a two-lane road where no other target vehicles are detected in the host lane or the merge lane 1111 of the two-lane road. In this case, the driving strategy module 803 can select the Hold node 909 when reaching the merge node 913. That is, in the absence of target vehicles being sensed as merging onto the road, it can be expected to remain in its current lane.

[0241] In Figure 11B , the situation is slightly different. Here, the host vehicle 1105 senses one or more target vehicles 1107 entering the main road 1112 from the merge lane 1111. In this case, once the driving strategy module 803 encounters the merge node 913, it can select to initiate a left overtaking maneuver to avoid the merge situation.

[0242] In Figure 11C , the host vehicle 1105 encounters one or more target vehicles 1107 entering the main road 1112 from the merge lane 1111. The host vehicle 1105 also detects a target vehicle 1109 traveling in the lane adjacent to the host vehicle's lane. The host vehicle also detects one or more target vehicles 1110 traveling in the same lane as the host vehicle 1105. In this case, the driving strategy module 803 can decide to adjust the speed of the host vehicle 1105 to yield to the target vehicle 1107 and proceed in front of the target vehicle 1115. This can be achieved, for example, by proceeding to the Select Gap node 919, which in turn will select a gap between ID0 (vehicle 1107) and ID1 (vehicle 1115) as the appropriate merge gap. In this case, the appropriate gap for the merge situation defines the objective of the trajectory planner optimization problem.

[0243] As described above, a node of the option graph can declare itself as "critical", which can ensure that the selected option passes through critical nodes. Formally, each node can implement the IsCritical function. After performing a forward pass from the root to the leaves on the option graph and solving the optimization problem of the trajectory planner, a backward pass can be performed from the leaves back to the root. Along this backward pass, the IsCritical function of all nodes in the pass can be called, and a list of all critical nodes can be saved. In the forward path corresponding to the next time frame, the driving strategy module 803 may need to select a path passing through all critical nodes from the root node to the leaves.

[0244] Figures 11A to 11C This can be used to show the potential benefits of the method. For example, in the case where an overtaking action is initiated and the driving strategy module 803 reaches a leaf corresponding to IDk, such as when the host vehicle is in the middle of an overtaking maneuver, it would be undesirable to select the hold node 909. To avoid such a jump, the IDj node can specify itself as critical. During the maneuver, the success of the trajectory planner can be monitored, and if the overtaking maneuver proceeds as expected, the function IsCritical will return a "true" value. This method can ensure that within the next time frame, the overtaking maneuver will continue (instead of jumping to another potentially inconsistent maneuver before completing the initially selected maneuver). On the other hand, if the monitoring of the maneuver indicates that the selected maneuver is not proceeding as expected, or if the maneuver becomes unnecessary or impossible, the IsCritical function can return a "false" value. This can allow the selection gap node to select a different gap in the next time frame, or abort the overtaking maneuver completely. This method can, on the one hand, allow the desired path to be declared on the option graph at each time step, and on the other hand, help improve the stability of the strategy during the critical part of the execution.

[0245] Hard constraints can be different from navigation expectations, which will be discussed in more detail below. For example, hard constraints can ensure safe driving by applying an additional filtering layer to the planned navigation actions. The involved hard constraints can be determined based on the sensed state, and the hard constraints can be manually programmed and defined rather than through the use of a trained system established based on reinforcement learning. However, in some embodiments, the trained system can learn the applicable hard constraints to be applied and followed. This method can prompt the driving strategy module 803 to reach the selected actions that already comply with the applicable hard constraints, which can reduce or eliminate the selected actions that may need to be modified later to comply with the applicable hard constraints. Nevertheless, as a redundant safety measure, even in the case where the driving strategy module 803 has been trained to handle the predetermined hard constraints, hard constraints can be applied to the output of the driving strategy module 803.

[0246] There are many examples of potential hard constraints. For example, a hard constraint can be defined in conjunction with a guardrail on the edge of a road. Under any circumstances, the host vehicle is not allowed to pass through the guardrail. Such a rule creates a hard lateral constraint on the trajectory of the host vehicle. Another example of a hard constraint can include a road bump (e.g., a speed control bump), which can cause a hard constraint on the driving speed before and while passing through the bump. Hard constraints can be considered safety-critical and thus can be defined manually rather than relying solely on a trained system that learns the constraints during training.

[0247] Compared to hard constraints, the desired goal can be to achieve or attain a comfortable drive. As discussed above, examples of desires can include the goal of placing the host vehicle at a lateral position corresponding to the center of the host vehicle's lane. Another desire may include an ID for a suitable entry gap. Note that the host vehicle does not need to be precisely at the center of the lane. Instead, the desire to be as close to it as possible can ensure that even when deviating from the lane center, the host vehicle tends to migrate to the center of the lane. Desires may not be safety-critical. In some embodiments, desires may need to be negotiated with other drivers and pedestrians. One way to construct desires can rely on an option graph, and the policies implemented in at least some of the nodes of the graph can be based on reinforcement learning.

[0248] For a node of the option graph 901 or 1000 implemented as a learning-based training node, the training process can include decomposing the problem into a supervised learning phase and a reinforcement learning phase. In the supervised learning phase, a differentiable mapping from (s t , a t ) to can be learned such that This can be similar to "model-based" reinforcement learning. However, in the forward loop of the network, the actual value of s t+1 can be used to replace thereby eliminating the problem of error accumulation. The role of prediction is to propagate messages from the future back to past actions. In this sense, the algorithm can be a combination of "model-based" reinforcement learning and "policy-based learning".

[0249] An important element that can be provided in certain scenarios is a differentiable path from future loss / reward back to the decision of an action. Through the option graph structure, the implementation of options involving safety constraints is typically non-differentiable. To overcome this problem, the selection of children in a learned policy node can be random. That is, the node can output a probability vector p, which assigns the probability for selecting each child of a particular node. Suppose the node has k children and let a (1) ,..., a (k)Actions for the paths from each child to the leaf. The resulting predicted actions are thus This may result in a differentiable path from this action to p. In practice, the action a can be chosen as a for i~p (i) , and the difference between a and can be referred to as additive noise.

[0250] For a given s t , a t under training, supervised learning can be used with real data. For the training of the node policy, a simulator can be used. Subsequently, fine-tuning of the policy can be done using real data. Two concepts can make the simulation more realistic. First, using imitation, the "behavior cloning" paradigm can be used with large real-world datasets to construct an initial policy. In some cases, the resulting actor can be suitable. In other cases, the resulting actor forms at least a very good initial policy for other actors on the road. Second, using self-play, our own policies can be used to enhance the training. For example, given a preliminary implementation that can be other experienced actors (cars / pedestrians), the policy can be trained based on the simulator. Some other actors can be replaced by the new policy, and the process can be repeated. Thus, the policy can continue to improve as it should respond to more types of other actors with different levels of complexity.

[0251] In addition, in some embodiments, the system can implement a multi-actor approach. For example, the system can consider data from various sources and / or images captured from multiple perspectives. Additionally, some disclosed embodiments can provide energy economy as the anticipation of events that may not be directly involved with the host vehicle but can have an impact on the host vehicle, or even the anticipation of events that may lead to unpredictable situations involving other vehicles can be considered (e.g., a radar may "see through" a vehicle ahead and the inevitable anticipation, or even the high likelihood of an event that will affect the host vehicle).

[0252] Trained System with Applied Navigation Constraints

[0253] In the context of autonomous driving, a key concern is how to ensure that the learned policies of a trained navigation network are safe. In some embodiments, constraints can be used to train a driving policy system such that the actions selected by the trained system already account for applicable safety constraints. Additionally, in some embodiments, an additional layer of safety can be provided by passing the selected actions of the trained system through one or more hard constraints involved in a particular sensed scenario in the environment of the host vehicle. This approach can ensure that the actions taken by the host vehicle are limited to those that are confirmed to satisfy the applicable safety constraints.

[0254] At its core, a navigation system can include a learning algorithm based on a policy function that maps observed states to one or more desired actions. In some implementations, the learning algorithm is a deep learning algorithm. The desired actions can include at least one action expected to maximize the expected reward of the vehicle. While in some cases, the actual action taken by the vehicle can correspond to one of the desired actions, in other cases, the actual action taken can be determined based on the observed state, one or more desired actions, and non-learned hard constraints (e.g., safety constraints) imposed on the learned navigation engine. These constraints can include no-driving zones around various types of detected objects (e.g., target vehicles, pedestrians, stationary objects on the road side or on the road, moving objects on the road side or on the road, guardrails, etc.). In some cases, the size of the zone can vary based on the detected motion (e.g., speed and / or direction) of the detected object. Other constraints can include a maximum driving speed when passing through the influence zone of a pedestrian, a maximum deceleration rate (to handle the distance to a target vehicle behind the host vehicle), a mandatory stop at a sensed crosswalk or railroad crossing, etc.

[0255] Hard constraints used in conjunction with a system trained by machine learning can provide a level of safety in autonomous driving that can exceed the level of safety achievable based solely on the output of the trained system. For example, a desired set of constraints can be used as training guidance to train a machine learning system, and thus, the trained system can select actions in response to sensed navigation states that account for and comply with the limitations of the applicable navigation constraints. However, the trained system still has some flexibility in selecting navigation actions, and thus, at least in some cases, the actions selected by the trained system may not strictly comply with the relevant navigation constraints. Therefore, in order to require the selected actions to strictly comply with the relevant navigation constraints, non-machine learning components outside the learning / trained framework that ensure the strict application of the relevant navigation constraints can be used to combine, compare, filter, regulate, modify, etc. the output of the trained system.

[0256] The following discussion provides additional details regarding the trained system and the potential benefits (especially from a safety perspective) collected from combining the trained system with algorithm components outside the trained / learning framework. As previously mentioned, the reinforcement learning objective for optimizing the policy can be achieved through stochastic gradient ascent. The objective (e.g., expected reward) can be defined as

[0257] The objective involving expectations can be used in machine learning scenarios. However, such an objective, without being restricted by navigation constraints, may not return actions that are strictly limited by these constraints. For example, consider a reward function where, represents a trajectory of a rare "corner" event (e.g., such as an accident) to be avoided, while represents the remaining trajectories. One objective of the learning system can be to learn to perform overtaking maneuvers. Generally, in accident-free trajectories, a successful and smooth overtaking will be rewarded, and staying in the lane without completing the overtaking will be penalized - thus, the range is [-1, 1]. If the sequence represents an accident, then the reward -r should provide a high enough penalty to prevent such an event from occurring. The problem is to determine what value of r should be to ensure accident-free driving.

[0258] It is observed that the impact of an accident on is an additional term -pr, where p is the probability mass of the trajectory with the accident event. If this term is negligible, i.e., p << 1 / r, then the learning system may more often prefer a policy that has executed an accident (or generally adopts a reckless driving policy) in order to successfully perform the overtaking maneuver, rather than a more defensive policy at the cost of some overtaking maneuvers not being successfully completed. In other words, if the probability of an accident is at most p, then r must be set such that r >> 1 / p. It can be expected that p is made very small (e.g., on the order of p = 10 -9 ). Therefore, r should be large. In the policy gradient, the gradient of can be estimated. The following lemma shows that the variance of the random variable increases with pr 2 and this variance is greater than r for r >> 1 / p. Therefore, estimating the objective may be difficult, and estimating its gradient may be even more difficult.

[0259] Lemma: Let π o be a policy, and let p and r be scalars such that with probability p, is obtained and with probability 1 - p, is obtained. Then,

[0260]

[0261] Among them, the final approximation applies to the case where r ≥ 1 / p.

[0262] This discussion shows that an objective of the form may not ensure functional safety without causing variance problems. The baseline subtraction method for reducing variance may not provide sufficient remedies for this problem, because the problem will transfer from with high variance to the equally high variance of the baseline constant, and the estimation of this baseline constant will also be affected by numerical instability. And, if the probability of an accident is p, then before obtaining the accident event, at least 1 / p sequences should be sampled on average. This means that the lower bound of the samples of the sequences for the learning algorithm aiming to minimize is 1 / p. The solution to this problem can be found in the architecture design described in this article, rather than through digital adjustment techniques. The method here is based on the concept that hard constraints should be injected outside the learning framework. In other words, the policy function can be decomposed into a learnable part and a non-learnable part. Formally, the policy function can be constructed as where maps the (agnostic) state space to a set of expectations (e.g., expected navigation goals, etc.), while π (T) maps this expectation to a trajectory (which can determine how the vehicle should move within a short distance). The function π (T) is responsible for the comfort of driving and making strategic decisions, such as which other cars should be overtaken or given way, and what is the expected position of the host vehicle within its lane, etc. The mapping from the sensed navigation state to the expectation is the policy which can be learned from experience by maximizing the expected reward. The expectations generated by can be transformed into a cost function for the driving trajectory. The function π (T) is not a learned function, and it can be achieved by finding a trajectory that minimizes the cost subject to hard constraints in terms of functional safety. This decomposition can ensure functional safety while providing comfortable driving.

[0263] As Figure 11D shows, the double lane change navigation scenario provides an example that further illustrates these concepts. In a double lane change, vehicles approach the lane change area 1130 from both the left and right sides. And, vehicles from each side (such as vehicle 1133 or vehicle 1135) can decide whether to change lanes into the lane on the other side of the lane change area 1130. Successfully performing a double lane change in heavy traffic may require significant negotiation skills and experience, and it may be difficult to execute by heuristic or brute force methods by enumerating all possible trajectories that all actors in the scenario may take. In this double lane change example, a set of expectations suitable for the double lane change maneuver can be defined can be the Cartesian product of the following sets:

[0264] where [0, v max is the desired target speed of the host vehicle, L = {1, 1.5, 2, 2.5, 3, 3.5, 4} is the desired lateral position in lane units, where integers represent the lane center and fractions represent the lane boundaries, and {g, t, o} are the classification labels assigned to each of the n other vehicles. If the host vehicle is to yield to another vehicle, the other vehicle can be assigned "g", if the host vehicle is to encroach on the lane of another vehicle, the other vehicle can be assigned "t", or if the host vehicle is to maintain an offset distance relative to another vehicle, the other vehicle can be assigned "o".

[0265] A set of expectations is described below how it can be transformed into a cost function for the driving trajectory. The driving trajectory can be represented by (x 1 , y 1 ),..., (x k , y k ), where (x i , y i ) is the (lateral, longitudinal) position (in ego-centric units) of the host vehicle at time τ·i. In some experiments, τ = 0.1 s and k = 10. Of course, other values can also be chosen. The cost assigned to the trajectory can include a weighted sum of the individual costs assigned to the desired speed, lateral position, and the label assigned to each of the n other vehicles.

[0266] Given the desired speed v ∈ [0, υ max , the cost of the trajectory associated with the speed is

[0267]

[0268] Given the desired lateral position l ∈ L, the cost associated with the desired lateral position is

[0269]

[0270] where dist(x, y, l) is the distance from the point (x, y) to the lane position l. Regarding the cost due to other vehicles, for any other vehicle, (x′ 1 , y′ 1 ),..., (x′ k , y′ k ) can represent the other vehicle in ego-centric units of the host vehicle, and i can be the earliest point for which there exists a j such that (x i , yi ) and (x' j , y' j ) is very small. If there is no such point, then i can be set to i = ∞. If another vehicle is classified as "yield", it can be expected that τi > τj + 0.5, which means that the host vehicle will reach the trajectory intersection point at least 0.5 seconds after the other vehicle reaches the same point. The formula for converting the above constraints into cost can be [τ(j - i) + 0.5] + .

[0271] Similarly, if another vehicle is classified as "occupying the road", it can be expected that τj > τi + 0.5, which can be converted into a cost [τ(i - j) + 0.5] + . If another vehicle is classified as "offset", it can be expected that i = ∞, which means that the trajectories of the host vehicle and the offset vehicle do not intersect. This situation can be converted into a cost by penalizing the distance between the trajectories.

[0272] Assigning weights to each of these costs can provide a single objective function π for the trajectory planner T) . The cost that encourages smooth driving can be added to this objective. Moreover, in order to ensure the functional safety of the trajectory, hard constraints can be added to this objective. For example, it can be prohibited that (x i , y i ) leaves the road, and if |i - j| is small, then for any trajectory point (x' j , y' j ) of any other vehicle, it can be prohibited that (x i , y i ) approaches (x' j , y' j ).

[0273] In summary, the policy π θ can be decomposed into a mapping from an agnostic state to a set of expectations and a mapping from the expectations to an actual trajectory. The latter mapping is not learning-based and can be achieved by solving an optimization problem whose cost depends on the expectations and whose hard constraints can guarantee the functional safety of the policy.

[0274] The following discussion describes the mapping from an agnostic state to the set of expectations. As mentioned above, in order to comply with functional safety, a system that relies solely on reinforcement learning may suffer from high and inconvenient variance regarding the rewards . By using policy gradient iteration to decompose the problem into a mapping from the (agnostic) state space to a set of expectations, and then mapping to an actual trajectory without involving a system based on machine learning training, this result can be avoided.

[0275] For various reasons, decision-making can be further decomposed into semantically meaningful components. For example, the size of may be large or even continuous. In the double merge scenario described above regarding Figure 11D Figure 11D , In addition, the gradient estimator may involve the term In such an expression, the variance can grow with the time horizon T. In some cases, the value of T can be approximately 250, which may be sufficient to produce significant variance. Assuming a sampling rate in the range of 10 Hz and a merge region 1130 of 100 meters, the preparation for the merge can start approximately 300 meters before the merge region. If the host vehicle is traveling at 16 meters per second (about 60 kilometers per hour), then the value of T for an episode can be approximately 250.

[0276] Returning to the concept of the option graph, Figure 11E shown in Figure 11E is an option graph that can represent the double merge scenario depicted in Figure 11D Figure 11D . As mentioned before, the option graph can represent a hierarchical decision set organized as a directed acyclic graph (DAG). In this graph, there may be a special node called the "root" node 1140, which can be the only node without incoming edges (e.g., decision lines). The decision-making process can start from the root node and traverse the graph until it reaches a "leaf" node, i.e., a node without outgoing edges. Each internal node can implement a policy function to select one of its available children. There can be a predefined mapping from a set of traversals on the option graph to a set of desired Given a node v in the graph, the parameter vector θ . In other words, a traversal on the option graph can be automatically transformed into a v v can specify the policy for selecting the children of v. If θ is the concatenation of all θ v v , then, a can be defined by traversing from the root of the graph to the leaf while using the policy defined by θ v v at each node v to select the child node.

[0277] In the Figure 11E double merge option graph 1139 in Figure 11E , the root node 1140 can first determine whether the host vehicle is within the merge region (e.g., Figure 11Dthe area 1130) in it, or whether the host vehicle is approaching a merging area and needs to prepare for a possible merge. In both cases, the host vehicle may need to decide whether to change lanes (e.g., left or right) or whether to stay in the current lane. If the host vehicle has decided to change lanes, the host vehicle may need to decide whether the conditions are suitable to continue and perform a lane change maneuver (e.g., at the "continue" node 1142). If it is not possible to change lanes, the host vehicle can attempt to "advance" towards the desired lane by aiming to be on the lane markings (e.g., at node 1144, as part of the negotiation with the vehicles in the desired lane). Alternatively, the host vehicle can choose to "stay" in the same lane (e.g., at node 1146). This process can determine the lateral position of the host vehicle in a natural way. For example,

[0278] This can achieve determining the desired lateral position in a natural way. For example, if the host vehicle changes lanes from lane 2 to lane 3, the "continue" node can set the desired lateral position to 3, the "stay" node can set the desired lateral position to 2, and the "advance" node can set the desired lateral position to 2.5. Next, the host vehicle can decide whether to maintain the "same" speed (node 1148), "accelerate" (node 1150), or "decelerate" (node 1152). Next, the host vehicle can enter a "chain-like" structure 1154 that overtakes other vehicles and sets their semantic meanings to values in the set {g, t, o}. This process can set the expectations relative to other vehicles. The parameters of all nodes in the chain can be shared (similar to a recurrent neural network).

[0279] A potential benefit of the option is the interpretability of the results. Another potential benefit is that one can rely on the decomposable structure of the set, and thus, can choose the strategy at each node from a small number of probabilities. Additionally, the structure can allow reducing the variance of the policy gradient estimator.

[0280] As mentioned above, the length of the segment in the double merge scenario can be approximately T = 250 time steps. Such a value (or any other suitable value depending on the specific navigation scenario) can provide enough time to see the consequences of the host vehicle's actions (e.g., if the host vehicle decides to change lanes as a preparation for merging, the host vehicle can only see the benefits after successfully completing the merge). On the other hand, due to the dynamics of driving, the host vehicle must make decisions at a high enough frequency (e.g., 10 Hz in the above case).

[0281] The option graph can achieve a reduction in the effective value of T in at least two ways. First, given a higher-level decision, rewards can be defined for lower-level decisions while considering shorter segments. For example, when the host vehicle has selected the "lane change" and "continue" nodes, by watching segments of 2 to 3 seconds (meaning T becomes 20 - 30 instead of 250), a strategy for assigning semantic meanings to the vehicle can be learned. Second, for high-level decisions such as whether to change lanes or stay in the same lane, the host vehicle may not need to make a decision every 0.1 seconds. Instead, the host vehicle can make decisions at a lower frequency (e.g., once per second), or implement an "option termination" function, and then the gradient can be calculated only after each option termination. In both cases, the effective value of T may be an order of magnitude smaller than its original value. In summary, the estimator for each node can depend on the value of T, which is an order of magnitude smaller than the original 250 steps, which may immediately lead to a smaller variance.

[0282] As described above, hard constraints can promote safer driving, and there can be several different types of constraints. For example, static hard constraints can be defined directly based on the sensed state. These hard constraints can include speed bumps, speed limits, road bends, intersections, etc. within the environment of the host vehicle, which may involve one or more constraints on vehicle speed, heading, acceleration, braking (deceleration), etc. Static hard constraints may also include semantically free spaces, where, for example, the host vehicle is prohibited from driving outside the free space and from driving too close to physical obstacles. Static hard constraints can also limit (e.g., prohibit) maneuvers that do not conform to aspects of the vehicle's kinematic motion. For example, static hard constraints can be used to prohibit maneuvers that may cause the host vehicle to flip, slide, or otherwise lose control.

[0283] Hard constraints can also be associated with the vehicle. For example, a constraint can be adopted that requires the vehicle to maintain a longitudinal distance of at least one meter from other vehicles and a lateral distance of at least 0.5 meters from other vehicles. Constraints can also be applied such that the host vehicle will avoid maintaining a collision course with one or more other vehicles. For example, the time τ can be a time metric based on a specific scenario. The predicted trajectories of the host vehicle and one or more other vehicles can be considered from the current time to time τ. In the case where the two trajectories intersect, can represent the time when vehicle i arrives at and leaves the intersection point. That is, each vehicle will reach the intersection point when the first part of the vehicle passes through the intersection point, and a certain amount of time is required before the last part of the vehicle passes through the intersection point. This amount of time separates the arrival time from the departure time. Assuming (i.e., the arrival time of vehicle 1 is less than the arrival time of vehicle 2), then we will want to ensure that vehicle 1 has left the intersection point before vehicle 2 arrives. Otherwise, a collision will occur. Therefore, it can be implemented such that Hard constraints. Also, to ensure that Vehicle 1 and Vehicle 2 are staggered from each other by a minimum amount, an additional safety margin can be obtained by including a buffer time (e.g., 0.5 seconds or another appropriate value) in the constraints. The hard constraints related to the predicted intersection trajectories of the two vehicles can be expressed as

[0284] The amount of time τ for which the trajectories of the host vehicle and one or more other vehicles are tracked can vary. However, in an intersection scenario, the rate may be lower, τ may be longer, and τ can be defined such that the host vehicle will enter and leave the intersection in less than τ seconds.

[0285] Of course, applying hard constraints to vehicle trajectories requires predicting the trajectories of those vehicles. For the host vehicle, trajectory prediction may be relatively straightforward because the host vehicle typically already understands and is in fact planning its expected trajectory at any given time. Relative to other vehicles, predicting their trajectories may not be as straightforward. For other vehicles, the baseline calculations for determining the predicted trajectories may depend on the current speed and heading of the other vehicles, e.g., as determined based on an analysis of the image stream captured by one or more cameras and / or other sensors (radar, lidar, acoustic, etc.) on the host vehicle.

[0286] However, there may be some exceptions that can simplify the problem or at least provide increased confidence in the predicted trajectory for another vehicle. For example, for structured roads with lane markings and where there may be right-of-way rules, the trajectories of other vehicles can be at least partially based on the position of the other vehicles relative to the lanes and based on the applicable right-of-way rules. Thus, in some cases, when the lane structure is observed, it can be assumed that the vehicle in the next lane will comply with the lane boundaries. That is, the host vehicle can assume that the vehicle in the next lane will stay in its lane unless evidence is observed indicating that the vehicle in the next lane will cut into the host vehicle's lane (e.g., signal lights, strong lateral movement, movement across the lane boundary).

[0287] Other situations can also provide clues about the expected trajectories of other vehicles. For example, at a stop sign, traffic light, roundabout, etc., where the host vehicle may have the right-of-way, it can be assumed that other vehicles will comply with that right-of-way. Thus, unless evidence of a rule violation is observed, it can be assumed that other vehicles will continue along a trajectory that complies with the right-of-way held by the host vehicle.

[0288] Hard constraints can also be applied with respect to pedestrians in the host vehicle's environment. For example, a buffer distance can be established with respect to pedestrians such that the host vehicle is prohibited from driving closer to any observed pedestrian than the specified buffer distance. The pedestrian buffer distance can be any suitable distance. In some embodiments, the buffer distance can be at least one meter with respect to the observed pedestrian.

[0289] Similar to the case with a vehicle, hard constraints can also be applied relative to the relative movement between a pedestrian and the host vehicle. For example, the trajectory of a pedestrian can be monitored relative to the predicted trajectory of the host vehicle (based on heading and speed). Given a particular pedestrian trajectory, where for each point p on the trajectory, t(p) can represent the time it takes for the pedestrian to reach point p. To maintain the required buffer distance of at least 1 meter from the pedestrian, t(p) must be greater than the time the host vehicle will reach point p (with a sufficient time difference such that the host vehicle passes in front of the pedestrian at a distance of at least one meter), or t(p) must be less than the time the host vehicle will reach point p (e.g., if the host vehicle brakes to yield to the pedestrian). Nevertheless, in the latter example, the hard constraint may require the host vehicle to reach point p sufficiently later than the pedestrian such that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter. Of course, there may be exceptions to the hard constraints on pedestrians. For example, in the case where the host vehicle has the right of way or is moving very slowly and there is no evidence that the pedestrian will refuse to yield to the host vehicle or will otherwise navigate towards the host vehicle, the pedestrian hard constraint can be relaxed (e.g., to a smaller buffer of at least 0.75 meters or 0.50 meters).

[0290] In some examples, if it is determined that not all constraints can be satisfied, the constraints can be relaxed. For example, in the case where the road is too narrow to leave the required spacing (e.g., 0.5 meters) between two curbs or between a curb and a parked vehicle, one or more constraints can be relaxed if there are mitigating circumstances. For example, if there are no pedestrians (or other objects) on the sidewalk, the vehicle can drive slowly at 0.1 meter from the curb. In some embodiments, a constraint can be relaxed if relaxing the constraint improves the user experience. For example, to avoid potholes, a constraint can be relaxed to allow the vehicle to be closer to the lane edge, curb, or pedestrian than might normally be allowed. Additionally, when determining which constraints to relax, in some embodiments, the one or more constraints selected to be relaxed are those that are considered to have the least available negative impact on safety. For example, before relaxing a constraint related to proximity to other vehicles, a constraint regarding how close the vehicle can drive to a curb or concrete barrier can be relaxed. In some embodiments, pedestrian constraints can be the last to be relaxed or may never be relaxed in some cases.

[0291] Figure 12 Examples of scenarios that can be captured and analyzed during the navigation of the host vehicle are shown. For example, the host vehicle can include a navigation system as described above (e.g., system 100), which can receive a plurality of images representing the environment of the host vehicle from a camera associated with the host vehicle (e.g., at least one of image capture device 122, image capture device 124, and image capture device 126). Figure 12The scenario shown is an example of one of the images that can be captured of the environment of a host vehicle traveling in lane 1210 along predicted trajectory 1212 at time t. The navigation system can include at least one processing device (e.g., including any of the EyeQ processors or other devices described above), which is specifically programmed to receive multiple images and analyze the images to determine actions in response to the scenario. Specifically, as Figure 8 shown, at least one processing device can implement a sensing module 801, a driving strategy module 803, and a control module 805. The sensing module 801 can be responsible for collecting and outputting image information collected from a camera and providing this information to the driving strategy module 803 in the form of an identified navigation state, which can constitute a trained navigation system that has been trained by machine learning techniques such as supervised learning, reinforcement learning, etc. Based on the navigation state information provided by the sensing module 801 to the driving strategy module 803, the driving strategy module 803 (e.g., by implementing the option graph method described above) can generate desired navigation actions for the host vehicle to perform in response to the identified navigation state.

[0292] In some embodiments, at least one processing device can use, for example, the control module 805 to directly convert the desired navigation action into a navigation command. However, in other embodiments, hard constraints can be applied such that the desired navigation action provided by the driving strategy module 803 is tested against various predefined navigation constraints that the scenario and the desired navigation action may involve. For example, in the case where the driving strategy module 803 outputs a desired navigation action that would cause the host vehicle to follow trajectory 1212, this navigation action can be tested against one or more hard constraints associated with various aspects of the host vehicle's environment. For example, the captured image 1201 can show a curb 1213, a pedestrian 1215, a target vehicle 1217, and stationary objects (e.g., a tipped-over box) present in the scenario. Each of these can be associated with one or more hard constraints. For example, the curb 1213 can be associated with a static constraint that prohibits the host vehicle from navigating onto or past the curb and onto the sidewalk 1214. The curb 1213 can also be associated with a roadblock envelope that defines a distance (e.g., a buffer zone) away from the curb (e.g., 0.1 meter, 0.25 meter, 0.5 meter, 1 meter, etc.) and extending along the curb, which defines a prohibited navigation zone for the host vehicle. Of course, static constraints can also be associated with other types of roadside boundaries (e.g., guardrails, concrete posts, traffic cones, bridge towers, or any other type of roadside obstacle).

[0293] Note that distances and ranges can be determined by any suitable method. For example, in some embodiments, distance information may be provided by on-vehicle radar and / or lidar systems. Alternatively or additionally, distance information may be derived from the analysis of one or more images captured of the environment of the host vehicle. For example, the number of pixels of an identified object represented in the image may be determined and compared with the known field of view and focal length geometry of the image capture device to determine scale and distance. For example, speed and acceleration can be determined by observing the change in scale between objects from image to image over a known time interval. This analysis can indicate the direction of movement of the object towards or away from the host vehicle, and how fast the object is moving away from or towards the host vehicle. The cross-over speed can be determined by analyzing the change in the X coordinate position of an object from one image to another over a known time period.

[0294] Pedestrian 1215 may be associated with a pedestrian envelope that defines a buffer 1216. In some cases, a hard constraint may be imposed to prohibit the host vehicle from navigating within a distance of one meter from pedestrian 1215 (in any direction relative to pedestrian 1215). Pedestrian 1215 may also define the location of a pedestrian influence zone 1220. Such an influence zone may be associated with a constraint that limits the speed of the host vehicle within the influence zone. The influence zone may extend 5 meters, 10 meters, 20 meters, etc. from pedestrian 1215. Each graduation of the influence zone may be associated with a different speed limit. For example, within the area from one meter to five meters from pedestrian 1215, the host vehicle may be restricted to a first speed (e.g., 10 mph (miles per hour), 20 mph, etc.), which may be less than the speed limit in the pedestrian influence zone that extends from 5 meters to 10 meters. Any graduation may be used for the levels of the influence zone. In some embodiments, the first level may be narrower than from one meter to five meters and may only extend from one meter to two meters. In other embodiments, the first level of the influence zone may extend from one meter (the boundary of the no-navigation zone around the pedestrian) to a distance of at least 10 meters. The second level may then extend from ten meters to at least about twenty meters. The second level may be associated with the maximum driving speed of the host vehicle, which is greater than the maximum driving speed associated with the first level of the pedestrian influence zone.

[0295] One or more stationary object constraints may also be related to a scene detected in the environment of the host vehicle. For example, in Image 1201, at least one processing device may detect a stationary object, such as a box 1219 present in the road. The detected stationary object may include various objects, such as at least one of a tree, a pole, a road sign, or an object in the road. One or more predetermined navigation constraints may be associated with the detected stationary object. For example, such a constraint may include a stationary object envelope, where the stationary object envelope defines a buffer zone with respect to the object, and the host vehicle may be prohibited from navigating into this buffer zone. At least a portion of the buffer zone may extend a predetermined distance from the edge of the detected stationary object. For example, in the scene represented by Image 1201, a buffer zone of at least 0.1 meter, 0.25 meter, 0.5 meter, or more may be associated with the box 1219, such that the host vehicle will pass by the right or left side of the box at least at a distance (e.g., the buffer zone distance) to avoid a collision with the detected stationary object.

[0296] Predetermined hard constraints may also include one or more target vehicle constraints. For example, a target vehicle 1217 may be detected in Image 1201. To ensure that the host vehicle does not collide with the target vehicle 1217, one or more hard constraints may be employed. In some cases, a target vehicle envelope may be associated with a single buffer zone distance. For example, a buffer zone may be defined by a distance of 1 meter around the target vehicle in all directions. The buffer zone may define an area extending at least one meter from the target vehicle, and the host vehicle may be prohibited from navigating into this area.

[0297] However, the envelope around the target vehicle 1217 does not need to be defined by a fixed buffer distance. In some cases, the predetermined hard constraints associated with a target vehicle (or any other movable object detected in the environment of the host vehicle) may depend on the orientation of the host vehicle relative to the detected target vehicle. For example, in some cases, the longitudinal buffer zone distance (e.g., the distance extending from the target vehicle towards the front or rear of the host vehicle - such as when the host vehicle is moving towards the target vehicle) may be at least one meter. The lateral buffer zone distance (e.g., the distance extending from the target vehicle towards either side of the host vehicle - such as when the host vehicle is moving in the same or opposite direction as the target vehicle, such that one side of the host vehicle will pass adjacent to one side of the target vehicle) may be at least 0.5 meter.

[0298] As described above, other constraints may also be involved through the detection of target vehicles or pedestrians in the environment of the host vehicle. For example, the predicted trajectories of the host vehicle and the target vehicle 1217 may be considered, and in the case where the two trajectories intersect (e.g., at the intersection point 1230), hard constraints may be required Or Where the host vehicle is Vehicle 1 and the target vehicle 1217 is Vehicle 2. Similarly, the trajectory of the pedestrian 1215 can be monitored relative to the predicted trajectory of the host vehicle (based on the forward direction and speed). Given a specific pedestrian trajectory, for each point p on the trajectory, t(p) will represent the time required for the pedestrian to reach point p (i.e., Figure 12 the point 1231 in). To maintain the required buffer distance of at least 1 meter from the pedestrian, t(p) must be greater than the time at which the host vehicle will reach point p (with a sufficient time difference such that the host vehicle passes in front of the pedestrian at a distance of at least one meter), or t(p) must be less than the time at which the host vehicle will reach point p (e.g., if the host vehicle brakes to make way for the pedestrian). Nevertheless, in the latter example, the hard constraint will require the host vehicle to reach point p sufficiently later than the pedestrian so that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter.

[0299] Other hard constraints can also be adopted. For example, at least in some cases, the maximum deceleration rate of the host vehicle can be adopted. This maximum deceleration rate can be determined based on the detected distance to the target vehicle following the host vehicle (e.g., using images collected from a rear camera). Hard constraints can include a forced stop at a sensed crosswalk or railroad crossing, or other applicable constraints.

[0300] In cases where the analysis of the scene in the environment of the host vehicle indicates that one or more predetermined navigation constraints may be involved, these constraints can be imposed relative to one or more planned navigation actions of the host vehicle. For example, in cases where the analysis of the scene causes the driving strategy module 803 to return a desired navigation action, the desired navigation action can be tested against one or more of the involved constraints. If the desired navigation action is determined to violate any aspect of the involved constraints (e.g., if the desired navigation action will drive the host vehicle within 0.7 meters of the pedestrian 1215, where the predetermined hard constraint requires the host vehicle to maintain a distance of at least 1.0 meter from the pedestrian 1215), then at least one modification can be made to the desired navigation action based on one or more predetermined navigation constraints. Adjusting the desired navigation action in this way can follow the constraints involved in the specific scene detected in the environment of the host vehicle and provide an actual navigation action for the host vehicle.

[0301] After determining the actual navigation action of the host vehicle, the navigation action can be implemented by causing at least one adjustment to the navigation actuator of the host vehicle in response to the determined actual navigation action of the host vehicle. Such a navigation actuator can include at least one of the steering mechanism, brakes, or accelerator of the host vehicle.

[0302] Priority Constraints

[0303] As described above, a navigation system may employ various hard constraints to ensure the safe operation of a host vehicle. The constraints may include a minimum safe driving distance relative to a pedestrian, a target vehicle, a roadblock, or a detected object, a maximum driving speed when passing through the influence zone of a detected pedestrian, or a maximum deceleration rate of the host vehicle, among others. These constraints may be imposed on a trained system trained based on machine learning (supervised, reinforcement, or a combination), but they may also be useful for untrained systems (e.g., those that use algorithms to directly handle expected situations that occur in scenarios from the host vehicle's environment).

[0304] In either case, there may be a hierarchy of constraints. In other words, some navigation constraints take precedence over others. Thus, if a situation arises where there is no available navigation action that would satisfy all relevant constraints, the navigation system may first determine the available navigation actions that achieve the highest-priority constraints. For example, the system may also cause the vehicle to first avoid the pedestrian, even if navigating to avoid the pedestrian would result in a collision with another vehicle or object detected in the road. In another example, the system may cause the vehicle to drive up onto a curb to avoid the pedestrian.

[0305] Figure 13 A flowchart is provided that illustrates an algorithm for implementing a hierarchy of relevant constraints determined based on an analysis of a scenario in the environment of a host vehicle. For example, at step 1301, at least one processing device associated with the navigation system (e.g., an EyeQ processor, etc.) may receive a plurality of images representing the environment of the host vehicle from a camera mounted on the host vehicle. By analyzing the images of the scenario representing the environment of the host vehicle at step 1303, a navigation state associated with the host vehicle may be identified. For example, the navigation state may indicate that the host vehicle is traveling along a two-lane road 1210, as Figure 12 shown, where a target vehicle 1217 is moving through an intersection in front of the host vehicle, a pedestrian 1215 is waiting to cross the road on which the host vehicle is traveling, an object 1219 is present in front of the host vehicle's lane, and various other attributes of the scenario.

[0306] At step 1305, one or more navigation constraints involved in the navigation state of the host vehicle may be determined. For example, after analyzing the scenario in the environment of the host vehicle represented by one or more captured images, at least one processing device may determine one or more navigation constraints involved with the objects, vehicles, pedestrians, etc. identified through image analysis of the captured images. In some embodiments, at least one processing device may determine at least a first predetermined navigation constraint and a second predetermined navigation constraint involved in the navigation state, and the first predetermined navigation constraint may be different from the second predetermined navigation constraint. For example, the first navigation constraint may relate to one or more target vehicles detected in the environment of the host vehicle, while the second navigation constraint may relate to pedestrians detected in the environment of the host vehicle.

[0307] In step 1307, at least one processing device may determine a priority associated with the constraints identified in step 1305. In the example described, a second predetermined navigation constraint involving a pedestrian may have a higher priority than a first predetermined navigation constraint involving a target vehicle. Although the priority associated with a navigation constraint may be determined or assigned based on various factors, in some embodiments, the priority of a navigation constraint may be related to its relative importance from a safety perspective. For example, while it may be important to comply with or satisfy all implemented navigation constraints in as many cases as possible, some constraints may be associated with a higher safety risk compared to other constraints, and thus, these constraints may be assigned a higher priority. For example, a navigation constraint that requires the host vehicle to maintain a distance of at least 1 meter from a pedestrian may have a higher priority than a constraint that requires the host vehicle to maintain a distance of at least 1 meter from a target vehicle. This may be because a collision with a pedestrian can have more severe consequences than a collision with another vehicle. Similarly, maintaining a distance between the host vehicle and the target vehicle may have a higher priority than the following constraints: requiring the host vehicle to avoid a box on the road, drive below a specific speed over a speed bump, or expose the occupant of the host vehicle to an acceleration not exceeding a maximum acceleration level.

[0308] Although the driving strategy module 803 is designed to maximize safety by satisfying the navigation constraints involved in a particular scenario or navigation state, in some cases, it may not be possible to satisfy every involved constraint. In such a case, as shown in step 1309, the priority of each involved constraint may be used to determine which involved constraint should be satisfied first. Continuing with the above example, in a situation where it is not possible to satisfy both the pedestrian gap constraint and the target vehicle gap constraint simultaneously and only one of the constraints can be satisfied, the higher-priority pedestrian gap constraint may cause that constraint to be satisfied before attempting to maintain a gap to the target vehicle. Thus, under normal circumstances, as shown in step 1311, in a situation where both the first predetermined navigation constraint and the second predetermined navigation constraint can be satisfied, at least one processing device may determine a first navigation action of the host vehicle that satisfies both the first predetermined navigation constraint and the second predetermined navigation constraint based on the identified navigation state of the host vehicle. However, in other cases, in a situation where not all involved constraints can be satisfied, as shown in step 1313, in a situation where the first predetermined navigation constraint and the second predetermined navigation constraint cannot be satisfied simultaneously, at least one processing device may determine a second navigation action of the host vehicle that satisfies the second predetermined navigation constraint (i.e., the higher-priority constraint) but does not satisfy the first predetermined navigation constraint (having a lower priority than the second navigation constraint) based on the identified navigation state.

[0309] Next, at step 1315, to implement the determined navigation action of the host vehicle, at least one processing device may cause at least one adjustment of a navigation actuator of the host vehicle in response to the determined first navigation action of the host vehicle or the determined second navigation action of the host vehicle. As described in the previous example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.

[0310] Constraint Relaxation

[0311] As described above, navigation constraints may be imposed for safety purposes. The constraints may include a minimum safe driving distance relative to a pedestrian, a target vehicle, a roadblock, or a detected object, a maximum driving speed when passing within the influence area of a detected pedestrian, or a maximum deceleration rate of the host vehicle, etc. These constraints may be imposed in a learned or non-learned navigation system. In some cases, these constraints may be relaxed. For example, in a case where the host vehicle decelerates or stops near a pedestrian and then moves forward slowly to convey the intention to drive past the pedestrian, the response of the pedestrian may be detected from the acquired image. If the response of the pedestrian is to remain stationary or stop moving (and / or if eye contact with the pedestrian is sensed), it may be understood that the pedestrian recognizes the intention of the navigation system to drive past the pedestrian. In such a case, the system may relax one or more predetermined constraints and implement less stringent constraints (e.g., allowing the vehicle to navigate within 0.5 meters of the pedestrian instead of within a more stringent 1-meter boundary).

[0312] Figure 14 A flowchart for implementing control of a host vehicle based on relaxation of one or more navigation constraints is provided. At step 1401, at least one processing device may receive a plurality of images representing the environment of the host vehicle from a camera associated with the host vehicle. Analysis of the images at step 1403 may enable identification of a navigation state associated with the host vehicle. At step 1405, at least one processor may determine a navigation constraint associated with the navigation state of the host vehicle. The navigation constraint may include a first predetermined navigation constraint involved in at least one aspect of the navigation state. At step 1407, analysis of the plurality of images may reveal the existence of at least one navigation constraint relaxation factor.

[0313] A navigation constraint relaxation factor can include any suitable indicator that one or more navigation constraints can be suspended, altered, or otherwise relaxed in at least one aspect. In some embodiments, at least one navigation constraint relaxation factor can include a determination (based on image analysis) that a pedestrian's eyes are looking in the direction of the host vehicle. In such a case, it can be more safely assumed that the pedestrian is aware of the host vehicle. Thus, the confidence that the pedestrian will not engage in an unexpected action that causes the pedestrian to move into the path of the host vehicle may be higher. Other constraint relaxation factors can also be used. For example, at least one navigation constraint relaxation factor can include a pedestrian determined to not be moving (e.g., a pedestrian assumed to be unlikely to enter the path of the host vehicle); or a pedestrian whose movement is determined to be slowing down. A navigation constraint relaxation factor can also include more complex actions, such as a pedestrian determined to not be moving after the host vehicle has stopped and then resumed movement. In such a case, it can be assumed that the pedestrian knows that the host vehicle has the right of way, and the pedestrian's stopping can indicate an intention to yield to the host vehicle. Other situations that can cause one or more constraints to be relaxed include the type of curb (e.g., a low curb or a curb with a gradual slope may allow a relaxed distance constraint), the absence of pedestrians or other objects on the sidewalk, a vehicle with its engine not running may have a relaxed distance, or a situation where a pedestrian is moving away from and / or with their back to the area the host vehicle is heading towards.

[0314] In the case where a navigation constraint relaxation factor is identified (e.g., at step 1407), a second navigation constraint can be determined or generated in response to the detection of the constraint relaxation factor. The second navigation constraint can be different from the first navigation constraint, and the second navigation constraint can include at least one characteristic that is relaxed relative to the first navigation constraint. The second navigation constraint can include a newly generated constraint based on the first constraint, where the newly generated constraint includes at least one modification that relaxes the first constraint in at least one aspect. Alternatively, the second constraint can constitute a predetermined constraint that is less strict than the first navigation constraint in at least one aspect. In some embodiments, such a second constraint can be retained for use only in cases where a navigation constraint relaxation factor is identified in the environment of the host vehicle. Whether the second constraint is newly generated or selected from a set of fully or partially available predetermined constraints, applying the second navigation constraint in place of the more strict first navigation constraint (which can be applied in the absence of detection of a relevant navigation constraint relaxation factor) can be referred to as constraint relaxation and can be done at step 1409.

[0315] If at least one constraint relaxation factor is detected in step 1407 and at least one constraint has been relaxed in step 1409, the navigation action of the host vehicle can be determined in step 1411. The navigation action of the host vehicle can be based on the identified navigation state and can satisfy the second navigation constraint. The navigation action can be implemented in step 1413 by causing at least one adjustment to the navigation actuator of the host vehicle in response to the determined navigation action.

[0316] As described above, the use of navigation constraints and relaxed navigation constraints can be employed for trained (e.g., via machine learning) or untrained (e.g., programmed to respond to a particular navigation state with a predetermined action) navigation systems. In the case of using a trained navigation system, the availability of relaxed navigation constraints for certain navigation situations can represent a mode switch from a trained system response to an untrained system response. For example, a trained navigation network can determine an original navigation action of the host vehicle based on a first navigation constraint. However, the action taken by the vehicle can be an action different from the navigation action that satisfies the first navigation constraint. Instead, the action taken can satisfy a more relaxed second navigation constraint and can be an action generated by an untrained system (e.g., in response to detecting a particular condition in the environment of the host vehicle, such as the presence of a constraint relaxation factor).

[0317] There are many examples of navigation constraints that can be relaxed in response to detecting a constraint relaxation factor in the environment of the host vehicle. For example, in the case where a predetermined navigation constraint includes a buffer associated with a detected pedestrian and at least a portion of the buffer extends a certain distance from the detected pedestrian, the relaxed navigation constraint (newly generated, retrieved from memory from a predetermined set, or generated as a relaxed version of a previously existing constraint) can include a different or modified buffer. For example, the different or modified buffer can have a smaller distance relative to the pedestrian than the original or unmodified buffer relative to the detected pedestrian. Thus, in the case where an appropriate constraint relaxation factor is detected in the environment of the host vehicle, in view of the relaxed constraint, the host vehicle can be allowed to navigate closer to the detected pedestrian.

[0318] As described above, the relaxed nature of the navigation constraint can include a reduction in the width of the buffer associated with at least one pedestrian. However, the relaxed nature can also include a reduction in the width of the buffer associated with a target vehicle, a detected object, a roadside obstacle, or any other object detected in the environment of the host vehicle.

[0319] At least one relaxed characteristic may also include other types of modifications to the navigation constraint characteristics. For example, the relaxed characteristic may include an increase in the rate associated with at least one predetermined navigation constraint. The relaxed characteristic may also include an increase in the maximum allowable deceleration / acceleration associated with at least one predetermined navigation constraint.

[0320] Although the constraints may be relaxed in some cases as described above, in other cases, the navigation constraints may be augmented. For example, in some cases, the navigation system may determine conditions that warrant an augmentation of a set of normal navigation constraints. Such augmentation may include adding new constraints to a set of predetermined constraints or adjusting one or more aspects of the predetermined constraints. The addition or adjustment may result in a more conservative navigation relative to the set of predetermined constraints applicable under normal driving conditions. Conditions that may warrant constraint augmentation may include sensor failures, adverse environmental conditions (rain, snow, fog, or other situations associated with reduced visibility or reduced vehicle traction), etc.

[0321] Figure 15 A flowchart for implementing control of a host vehicle based on an augmentation of one or more navigation constraints is provided. At step 1501, at least one processing device may receive a plurality of images representing the environment of the host vehicle from a camera associated with the host vehicle. Analysis of the images at step 1503 may enable identification of a navigation state associated with the host vehicle. At step 1505, at least one processor may determine navigation constraints associated with the navigation state of the host vehicle. The navigation constraints may include a first predetermined navigation constraint involved in at least one aspect of the navigation state. At step 1507, analysis of the plurality of images may reveal the presence of at least one navigation constraint augmentation factor.

[0322] The navigation constraints involved may include those described above (e.g., regarding Figure 12)Any navigation constraint or any other suitable navigation constraint. A navigation constraint enhancer can include any indicator that one or more navigation constraints can be supplemented / enhanced in at least one aspect. The supplementation or enhancement of the navigation constraint can be performed on a per-group basis (e.g., by adding a new navigation constraint to a predetermined set of constraints) or on a per-constraint basis (e.g., modifying a specific constraint such that the modified constraint is more restrictive than the original constraint, or adding a new constraint corresponding to a predetermined constraint, where the new constraint is more restrictive than the corresponding constraint in at least one aspect). Additionally or alternatively, the supplementation or enhancement of the navigation constraint can refer to a selection from a predetermined set of constraints based on a hierarchy. For example, a set of enhanced constraints can be used for selection based on whether a navigation enhancer is detected in the environment of the host vehicle or relative to the host vehicle. In normal situations where no enhancer is detected, the navigation constraint involved can be drawn from the constraints applicable to normal situations. On the other hand, in the case where one or more constraint enhancers are detected, the constraint involved can be drawn from the enhanced constraints that are generated or predetermined relative to the one or more enhancers. The enhanced constraints may be more restrictive than the corresponding constraints applicable in normal situations in at least one aspect.

[0323] In some embodiments, at least one navigation constraint enhancer can include the detection (e.g., based on image analysis) of the presence of ice, snow, or water on the road surface in the environment of the host vehicle. For example, such determination can be based on the detection of: areas with reflectivity higher than expected for a dry road (e.g., indicating ice or water on the road); white areas on the road indicating the presence of snow; shadows on the road consistent with the presence of longitudinal grooves on the road (e.g., tire tracks in snow); water droplets or ice / snow particles on the windshield of the host vehicle; or any other suitable indicator of the presence of water or ice / snow on the road surface.

[0324] At least one navigation constraint enhancer can also include the detection of particles on the outer surface of the windshield of the host vehicle. Such particles may degrade the image quality of one or more image capture devices associated with the host vehicle. Although described with respect to the windshield of the host vehicle, which is related to a camera mounted behind the windshield of the host vehicle, the detection of particles on other surfaces (e.g., the lens or lens cap of the camera, the headlight lens, the rear windshield, the taillight lens, or any other surface of the host vehicle visible (or detected by a sensor) to the image capture device associated with the host vehicle) can also indicate the presence of a navigation constraint enhancer.

[0325] Navigation constraint enhancement factors can also be detected as attributes of one or more image acquisition devices. For example, a detected degradation in aspects of the image quality of one or more images captured by an image capture device (e.g., a camera) associated with the host vehicle can also constitute a navigation constraint enhancement factor. The degradation in image quality may be associated with a hardware failure or a partial hardware failure that is associated with the image capture device or a component associated with the image capture device. Such a degradation in image quality may also be caused by environmental conditions. For example, the presence of smoke, fog, rain, snow, etc. in the air around the host vehicle may also result in a degradation of the image quality with respect to roads, pedestrians, target vehicles, etc. that may be present in the environment of the host vehicle.

[0326] Navigation constraint enhancement factors can also relate to other aspects of the host vehicle. For example, in some cases, a navigation constraint enhancement factor can include a detected failure or partial failure of a system or sensor associated with the host vehicle. Such enhancement factors can include, for example, a detected failure or partial failure of a speed sensor, GPS receiver, accelerometer, camera, radar, lidar, brakes, tires, or any other system associated with the host vehicle, which failure can affect the ability of the host vehicle to navigate with respect to navigation constraints associated with the navigation state of the host vehicle.

[0327] In the case where a navigation constraint enhancement factor is identified (e.g., at step 1507), a second navigation constraint can be determined or generated in response to the detected constraint enhancement factor. The second navigation constraint can be different from the first navigation constraint, and the second navigation constraint can include at least one characteristic that is enhanced relative to the first navigation constraint. The second navigation constraint can be more restrictive than the first navigation constraint because the detected constraint enhancement factor in the environment of the host vehicle or associated with the host vehicle can indicate that the host vehicle can have at least one navigation ability that is degraded relative to normal operating conditions. Such degraded ability can include lower road traction (e.g., ice, snow, or water on the road; reduced tire pressure, etc.); impaired vision (e.g., rain, snow, dust, sand, smoke, fog, etc. that reduce the captured image quality); impaired detection ability (e.g., sensor failure or partial failure, reduced sensor performance, etc.), or any other degradation in the ability of the host vehicle to navigate in response to the detected navigation state.

[0328] In the case where at least one constraint enhancement factor is detected in step 1507 and at least one constraint is enhanced in step 1509, the navigation action of the host vehicle can be determined in step 1511. The navigation action of the host vehicle can be based on the identified navigation state and can satisfy the second (i.e., enhanced) navigation constraint. The navigation action can be implemented in step 1513 by causing at least one adjustment to the navigation actuators of the host vehicle in response to the determined navigation action.

[0329] As discussed, the use of navigation constraints and enhanced navigation constraints can be employed for trained (e.g., via machine learning) or untrained (e.g., systems programmed to respond with a predetermined action in response to a particular navigation state) navigation systems. In the case of using a trained navigation system, the availability of enhanced navigation constraints for certain navigation situations can represent a mode switch from a trained system response to an untrained system response. For example, a trained navigation network can determine an original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle can be an action different from the navigation action that satisfies the first navigation constraint. Instead, the action taken can satisfy an enhanced second navigation constraint and can be an action generated by an untrained system (e.g., in response to detecting a particular condition in the environment of the host vehicle, such as the presence of a navigation constraint enhancer).

[0330] There are many examples of navigation constraints that can be generated, supplemented, or enhanced in response to detecting a constraint enhancer in the environment of the host vehicle. For example, in the case where a predetermined navigation constraint includes a buffer associated with a detected pedestrian, object, vehicle, etc., and at least a portion of the buffer extends a certain distance from the detected pedestrian / object / vehicle, the enhanced navigation constraint (newly generated, recalled from memory from a predetermined set, or generated as an enhanced version of a pre-existing constraint) can include a different or modified buffer. For example, the different or modified buffer can have a greater distance relative to the pedestrian / object / vehicle than the original or unmodified buffer relative to the detected pedestrian / object / vehicle. Thus, in the case where an appropriate constraint enhancer is detected in the environment of the host vehicle or relative to the host vehicle, in view of the enhanced constraint, the host vehicle can be forced to navigate further away from the detected pedestrian / object / vehicle.

[0331] At least one enhanced characteristic can also include other types of modifications to navigation constraint characteristics. For example, the enhanced characteristic can include a reduction in speed associated with at least one predetermined navigation constraint. The enhanced characteristic can also include a reduction in the maximum allowable deceleration / acceleration associated with at least one predetermined navigation constraint.

[0332] Navigation Based on Long - Term Planning

[0333] In some embodiments, the disclosed navigation system can not only respond to the detected navigation state in the environment of the host vehicle, but also determine one or more navigation actions based on long range planning. For example, the system can consider the potential impact of one or more navigation actions, which are available as options for navigating relative to the detected navigation state, on future navigation states. Considering the impact of the available actions on future states can enable the navigation system to determine navigation actions based not only on the currently detected navigation state but also on long range planning. Navigation using long range planning techniques may be particularly applicable in cases where the navigation system employs one or more reward functions as a technique for selecting navigation actions from available options. Potential rewards can be analyzed relative to the available navigation actions that can be taken in response to the currently detected navigation state of the host vehicle. However, further, potential rewards can also be analyzed relative to actions that can be taken in response to future navigation states that are expected to result from the available actions on the current navigation state. Thus, the disclosed navigation system can, in some cases, select a navigation action from among the available actions that can be taken in response to the detected navigation state, even when the selected navigation action may not result in the highest reward. This is particularly true when the system determines that the selected action may lead to a future navigation state that gives rise to one or more potential navigation actions, where the potential navigation action provides a higher reward than the selected action or, in some cases, any action available relative to the current navigation state. The principle can be expressed more simply as: taking a less favorable action now in order to generate a higher reward option in the future. Thus, the disclosed navigation system capable of long range planning can select a short-term sub-optimal action where long-term prediction indicates that a short-term loss of reward will result in an increase in long-term reward.

[0334] Typically, autonomous driving applications can involve a series of planning problems, where the navigation system can determine immediate actions to optimize longer-term goals. For example, when a vehicle encounters a merging situation at a roundabout, the navigation system can determine immediate acceleration or braking commands to initiate navigation to the roundabout. Although the immediate actions for the detected navigation state at the roundabout may involve acceleration or braking commands in response to the detected state, the long-term goal is a successful merge, and the long-term impact of the selected commands is the success / failure of the merge. The planning problem can be solved by decomposing the problem into two phases. First, supervised learning can be applied to predict the near future based on the present (assuming the representation of the predictor relative to the present will be differentiable). Second, a recurrent neural network can be used to model the complete trajectory of the actor, where the unexplained factors are modeled as (additive) input nodes. This can allow the use of supervised learning techniques and direct optimization of the recurrent neural network to determine the solution to the long-term planning problem. This approach can also enable the learning of robust strategies by incorporating adversarial elements into the environment.

[0335] The two most fundamental elements of an autonomous driving system are sensing and planning. Sensing deals with finding a compact representation of the present state of the environment, while planning deals with deciding what actions to take to optimize future goals. Supervised machine learning techniques are useful for solving sensing problems. Machine learning algorithm frameworks can also be used for the planning part, especially the reinforcement learning (RL) framework, such as those described above.

[0336] RL can be executed in a sequence of consecutive rounds. At the t-th round, the planner (also known as the actor or driving policy module 803) can observe the state s t ∈ S, which represents the actor and the environment. Then it should decide the action a t ∈ A. After executing the action, the actor receives an immediate reward and moves to a new state s t+1 . As an example, the host vehicle can include an adaptive cruise control (ACC) system, where the vehicle should autonomously perform acceleration / braking to maintain an appropriate distance from the vehicle in front while maintaining smooth driving. The state can be modeled as a pair where x t is the distance to the vehicle in front, and v t is the speed of the host vehicle relative to the speed of the vehicle in front. The action will be an acceleration command (if a t < 0, the host vehicle decelerates). The reward can depend on |a t | (reflecting the smoothness of driving) and s tA function that (reflects the safety distance maintained between the host vehicle and the vehicle ahead). The goal of the planner is to maximize the cumulative reward (which can be up to a time horizon of future rewards or a discounted sum). To do this, the planner can rely on a policy π: S → A, which maps states to actions.

[0337] Supervised learning (SL) can be regarded as a special case of RL, where s t is sampled from some distribution over S, and the reward function can have the form r t = -l(a t , y t ), where l is a loss function, and the learner observes the value of y t , where y t is the value of the best action (possibly noisy) to take when viewing the state s t . There can be several differences between the general RL model and the special case SL, and these differences make the general RL problem more challenging.

[0338] In some SL cases, the actions (or predictions) taken by the learner may have no impact on the environment. In other words, s t+1 and a t are independent. This has two important implications. First, in SL, the samples (s 1 , y 1 ),..., (s m , y m ) can be collected in advance before starting to search for a policy (or predictor) that has good accuracy relative to the samples. In contrast, in RL, the state s t+1 usually depends on the actions taken (and previous states), which in turn depends on the policy used to generate that action. This ties the data generation process to the policy learning process. Second, since actions do not affect the environment in SL, the contribution of the choice of a t to the performance of π is local. Specifically, a t only affects the value of the immediate reward. In contrast, in RL, the action taken at round t may have a long-term impact on the reward values in future rounds.

[0339] In SL, the knowledge of the "correct" answer y t , together with the shape of the reward r t = -l(a t , y t ), can provide complete knowledge of the rewards for all possible choices of a t , which can enable the calculation of the reward with respect to a tThe derivative. In contrast, in RL, the "one-time" value of the reward may be all that can be observed for a specific choice of the action taken. This can be referred to as "Bandit" feedback. This is one of the most important reasons why "exploration" is needed as part of long-term navigation planning, because in an RL-based system, if only "Bandit" feedback is available, the system may not always know whether the action taken is the best action to take.

[0340] Many RL algorithms rely at least in part on the mathematical optimization model of the Markov Decision Process (MDP, Markov Decision Process). The Markov assumption is that given s t and a t , the distribution of s t+1 is completely determined. Based on the stationary distribution over the MDP states, it gives a closed-form expression for the cumulative reward of a given policy. The stationary distribution of the policy can be expressed as the solution of a linear programming problem. This gives rise to two classes of algorithms: 1) Optimization with respect to the original problem, which is called policy search; 2) Optimization with respect to the dual problem, whose variables are called the value function V π . If the MDP starts from the initial state s and actions are chosen from there according to π, the value function determines the expected cumulative reward. The relevant quantity is the state-action value function Q π (s, a), assuming that starting from the state s, the action a is chosen immediately, and the actions are chosen from there according to π, the state-action value function determines the cumulative reward. The Q function may lead to the characterization of the optimal policy (using the Bellman equation). In particular, the Q function may indicate that the optimal policy is a deterministic function from S to A (in fact, it may be characterized as a "greedy" policy with respect to the optimal Q function).

[0341] A potential advantage of the MDP model is that it allows the future to be coupled to the present using the Q function. For example, assume that the host vehicle is now in the state s, and the value of Q π (s, a) can indicate the impact of performing the action a on the future. Therefore, the Q function can provide a local measurement of the quality of the action a, making the RL problem more similar to the SL scenario.

[0342] Many RL algorithms approximate the V function or the Q function in some way. Value iteration algorithms (such as Q-learning algorithms) can rely on the fact that the V and Q functions of the optimal policy can be fixed points of some operators derived from the Bellman equation. The Actor-critic policy iteration algorithm aims to learn the policy in an iterative manner, where at iteration t, the "critic" estimates and based on this estimate, the "actor" improves the policy.

[0343] Despite the mathematical advantages of MDP and the convenience of switching to the Q-function representation, the method may have several limitations. For example, the approximate concept of a Markovian behavioral state may be all that can be found in some cases. Additionally, state transitions depend not only on the agent's actions but also on the actions of other participants in the environment. For example, in the ACC example mentioned above, although the dynamics of the autonomous vehicle can be Markovian, the next state can depend on the actions of the driver of another vehicle, which may not be Markovian. A possible solution to this problem is to use a partially observable MDP, where it is assumed that there are Markovian states, but the observations are distributed according to hidden states.

[0344] A more direct approach can consider game-theoretic generalizations of MDP (e.g., the stochastic game framework). In fact, algorithms for MDP can be generalized to multi-agent games (e.g., Minimax-Q learning or Nash-Q learning). Other methods can include explicit modeling of other participants and vanishing regret learning algorithms. Learning in multi-agent scenarios may be more complex than in single-agent scenarios.

[0345] A second limitation of the Q-function representation can arise from deviating from the tabular setting. The tabular setting is when the number of states and actions is small, so Q can be represented as a table with |S| rows and |A| columns. However, if the natural representation of S and A includes Euclidean space and the state and action spaces are discrete, the number of state / action pairs can be exponential in the dimension. In this case, using the tabular setting may not be practical. Instead, the Q-function can be approximated by some function from a parametric hypothesis class (e.g., a neural network of a certain architecture). For example, the deep Q-network (DQN) learning algorithm can be used. In DQN, the state space can be continuous while the action space can still be a small discrete set. There are methods to handle continuous action spaces, but they may rely on approximating the Q-function. In any case, the Q-function can be complex and sensitive to noise, thus posing challenges to learning.

[0346] A different approach can be to use a recurrent neural network (RNN) to solve the RL problem. In some cases, the RNN can be combined with the concept of multi-agent games and the robustness of adversarial environments from game theory. Additionally, this approach may not explicitly rely on any Markov assumption.

[0347] The method of navigating by planning based on prediction is described in more detail below. In this method, it can be assumed that the state space S is a subset of, and the action space A is a subset of. This may be a natural manifestation in many applications. As mentioned above, there may be two key differences between RL and SL: (1) Since past actions affect future rewards, information from the future may need to be propagated back to the past; (2) The "bandit" nature of the rewards may obscure the dependence between (state, action) and rewards, which can complicate the learning process.

[0348] As the first step of this method, it can be observed that there are interesting problems where the bandit nature of the rewards is not an issue. For example, the reward value in the ACC application (as will be discussed in more detail below) may be differentiable with respect to the current state and action. In fact, even if the rewards are given in a "bandit" manner, learning a differentiable function such that the problem of can be a relatively straightforward SL problem (e.g., a one-dimensional regression problem). Therefore, the first step of this method can be to define the reward as a function that is differentiable with respect to s and a, or the first step of this method can be to use a regression learning algorithm in order to learn a differentiable function that at least minimizes at least some regression loss over samples, where the instance vector is and the target scalar is r t . In some cases, elements of exploration can be used to create the training set.

[0349] To address the connection between the past and the future, a similar idea can be used. For example, assume that a differentiable function can be learned such that Learning such a function can be characterized as an SL problem. can be considered a predictor of the near future. Next, a parametric function π θ : S → A can be used to describe the policy that maps from S to A. Expressing π θ as a neural network can enable the use of a recurrent neural network (RNN) to express a segment of the running actor for T steps, where the next state is defined as Here, can be defined by the environment and can represent the unpredictable aspects of the near future. s t+1 depends on s t and a t in a differentiable manner, which can enable the connection between future reward values and past behavior. The parameter vector of the policy function π θ can be learned by backpropagation on the resulting RNN. Note that there is no need to be in v tExplicit probabilistic assumptions are imposed. In particular, there is no requirement for Markov relationships. Instead, recurrent networks can be relied upon to propagate "sufficient" information between the past and the future. Intuitively, can describe the predictable part of the near future, while v t can express the unpredictable aspects that may arise due to the actions of other participants in the environment. The learning system should learn strategies that are robust to the actions of other participants. If ||v t || is large, the connection between past actions and future reward values may be too noisy to learn meaningful strategies. Explicitly expressing the dynamics of the system in a transparent manner can make it easier to incorporate prior knowledge. For example, prior knowledge can simplify the problem of defining .

[0350] As described above, the learning system can benefit from robustness with respect to adversarial environments (such as the environment of the host vehicle), which may include multiple other drivers who may act in unexpected ways. In a model that does not impose probabilistic assumptions on v t , an environment can be considered in which v t is selected in an adversarial manner. In some cases, restrictions can be imposed on μ t , otherwise the adversary may make the planning problem difficult or even impossible. A natural constraint may be to require ||μ t || to be bounded by a constant.

[0351] Robustness to adversarial environments may be useful in autonomous driving applications. Selecting μ t in an adversarial manner can even speed up the learning process, as it can focus the learning system on robust optimal strategies. This concept can be illustrated with a simple game. The state is the action is the immediate loss function is 0.1|a t | + [|s t | - 2] + , where [x] + = max{x, 0} is the ReLU (rectified linear unit) function. The next state is s t+1 = s t + a t + v t , where v t ∈[-0.5, 0.5] is selected in an adversarial manner for the environment. Here, the optimal strategy can be written as a two-layer network with ReLU: a t = [s t - 1.5] + + [-s t - 1.5] + . Observe that when |st When ∈(1.5, 2], the optimal action may have a greater immediate loss than action a = 0. Therefore, the system can plan for the future and can rely not only on the immediate loss. It is observed that the derivative of the loss with respect to a t is 0.1sign(a t ), and the derivative with respect to s t is 1[|s t | > 2]sign(s t ). In the case where s t ∈(1.5, 2], the adversarial choice of v t will set v t = 0.5. Therefore, whenever a t > 1.5 - s t , there can be a non-zero loss in the (t + 1)-th round. In this case, the derivative of the loss can be directly backpropagated to a t . Therefore, in the case where the choice of a t is non-optimal, the adversarial choice of v t can help the navigation system obtain a non-zero backpropagation message. This relationship can help the navigation system select the current action based on the expectation that this current behavior (even if this behavior leads to a non-optimal reward or even a loss) will provide an opportunity for a more optimized action that leads to a higher reward in the future.

[0352] This method can be applied to any navigation situation that may actually occur. The following description applies to a sample method: Adaptive Cruise Control (ACC). In the ACC problem, the host vehicle may attempt to maintain an appropriate distance from the target vehicle ahead (e.g., 1.5 seconds to the target vehicle). Another goal may be to drive as smoothly as possible while maintaining the desired gap. The model representing this situation can be defined as follows. The state space is The action space is The first coordinate of the state is the speed of the target vehicle, the second coordinate is the speed of the host vehicle, and the last coordinate is the distance between the host vehicle and the target vehicle (e.g., the position of the host vehicle along the road curve minus the position of the target vehicle). The action taken by the host vehicle is to accelerate and can be represented by a t . The quantity τ can represent the time difference between consecutive rounds. Although τ can be set to any appropriate quantity, in one example, τ can be 0.1 seconds. The position s t can be expressed as and the (unknown) acceleration of the target vehicle can be expressed as

[0353] The complete dynamics of the system can be described by the following description:

[0354]

[0355] This can be described as the sum of two vectors:

[0356]

[0357] The first vector is the predictable part, while the second vector is the unpredictable part. The reward for the t-th round is defined as follows:

[0358] where

[0359] The first term can result in a penalty for non-zero acceleration, thus encouraging smooth driving. The second term depends on the ratio between the distance x t to the target vehicle and the desired distance where the desired distance is defined as the maximum between a distance of 1 meter and the braking distance in 1.5 seconds. In some cases, this ratio may be exactly 1, but as long as the ratio is within the range [0.7, 1.3], the policy may waive any penalty, which may relax the main vehicle in terms of navigation - a property that may be important in achieving smooth driving.

[0360] Implementing the method outlined above, the navigation system of the main vehicle (e.g., through the operation of the driving policy module 803 within the processing unit 110 of the navigation system) can select an action in response to the observed state. The selected action can be based not only on the analysis of the reward associated with the response actions available relative to the sensed navigation state, but also on the consideration and analysis of future states, potential actions in response to future states, and the rewards associated with the potential actions.

[0361] Figure 16 An algorithmic approach to navigation based on detection and long-term planning is shown. For example, at step 1601, at least one processing device 110 of the navigation system of the main vehicle can receive a plurality of images. These images can capture scenes representing the environment of the main vehicle and can be provided by any of the above-mentioned image capture devices (e.g., cameras, sensors, etc.). Analyzing one or more of these images at step 1603 can enable at least one processing device 110 to identify the current navigation state associated with the main vehicle (as described above).

[0362] At steps 1605, 1607, and 1609, various potential navigation actions in response to the sensed navigation state can be determined. These potential navigation actions (e.g., the first navigation action to the Nth available navigation action) can be determined based on the sensed state and the long-term goals of the navigation system (e.g., completing a merge, smoothly following a leading vehicle, overtaking a target vehicle, avoiding an object in the road, decelerating for a detected stop sign, avoiding a cutting-in target vehicle, or any other navigation action that can further the navigation goals of the system).

[0363] For each of the determined potential navigation actions, the system can determine an expected reward. The expected reward can be determined according to any of the techniques described above and can include an analysis of one or more reward functions for a particular potential action. For each of the potential navigation actions determined at steps 1605, 1607, and 1609 (e.g., the first, second, and Nth), the expected rewards 1606, 1608, and 1610 can be determined, respectively.

[0364] In some cases, the navigation system of the host vehicle can select from the available potential actions based on values associated with the expected rewards 1606, 1608, and 1610 (or any other type of indicator of expected reward). For example, in certain cases, the action that produces the highest expected reward can be selected.

[0365] In other cases, especially where the navigation system is performing long-term planning to determine the navigation actions of the host vehicle, the system may not select the potential action that produces the highest expected reward. Instead, the system can look ahead to analyze whether there is an opportunity to achieve a higher reward later if a lower-reward action is selected in response to the current navigation state. For example, for any or all of the potential actions determined at steps 1605, 1607, and 1609, future states can be determined. Each of the future states determined at steps 1613, 1615, and 1617 can represent a future navigation state expected to be modified by the corresponding potential action (e.g., the potential actions determined at steps 1605, 1607, and 1609) based on the current navigation state.

[0366] For each of the future states predicted at steps 1613, 1615, and 1617, one or more future actions (as navigation options available in response to the determined future state) can be determined and evaluated. At steps 1619, 1621, and 1623, for example, values of expected rewards or any other type of indicator associated with one or more future actions can be generated (e.g., based on one or more reward functions). The expected rewards associated with one or more future actions can be evaluated by comparing the values of the reward functions associated with each future action or by comparing any other indicator associated with the expected reward.

[0367] At step 1625, the navigation system of the host vehicle can select a navigation action for the host vehicle based on a comparison of expected rewards, based not only on potential actions identified relative to the current navigation state (e.g., those in steps 1605, 1607, and 1609), but also on expected rewards determined as a result of potential future actions available in response to predicted future states (e.g., those determined at steps 1613, 1615, and 1617). The selection at step 1625 can be based on the option and reward analysis performed at steps 1619, 1621, and 1623.

[0368] The selection of a navigation action at step 1625 can be based solely on a comparison of expected rewards associated with future action options. In this case, the navigation system can select an action for the current state based solely on a comparison of expected rewards resulting from actions on potential future navigation states. For example, the system can select the potential actions identified at steps 1650, 1610, and 1609 that are associated with the highest future reward values determined by the analysis at steps 1619, 1621, and 1623.

[0369] The selection of a navigation action at step 1625 can also be based solely on a comparison of current action options (as described above). In this case, the navigation system can select the potential action identified at step 1605, 1607, or 1609 that is associated with the highest expected reward 1606, 1608, or 1610. This selection can be performed with little or no consideration of future navigation states or future expected rewards for available navigation actions in response to expected future navigation states.

[0370] On the other hand, in some cases, the selection of the navigation action at step 1625 can be based on a comparison of the expected rewards associated with both future action options and current action options. In fact, this may be one of the navigation principles based on long-term planning. For example, the expected rewards for future actions can be analyzed to determine whether any of the expected rewards can warrant the selection of a lower-reward action in response to the current navigation state, in order to achieve a potentially higher reward in response to a subsequent navigation action that is expected to be available in response to a future navigation state. As an example, the value of the expected reward 1606 or other indicator can indicate the highest expected reward among the rewards 1606, 1608, and 1610. On the other hand, the expected reward 1608 can indicate the lowest expected reward among the rewards 1606, 1608, and 1610. Instead of simply selecting the potential action determined at step 1605 (i.e., the action that results in the highest expected reward 1606), the analysis of future states, potential future actions, and future rewards can be used for the navigation action selection at step 1625. In one example, it can be determined that the reward identified at step 1621 (in response to at least one future action for the future state determined at step 1615, and the future state determined at step 1615 is based on the second potential action determined at step 1607) can be higher than the expected reward 1606. Based on this comparison, the second potential action determined at step 1607 can be selected instead of the first potential action determined at step 1605, even though the expected reward 1606 is higher than the expected reward 1608. In one example, the potential navigation action determined at step 1605 can include a lane change in front of the detected target vehicle, while the potential navigation action determined at step 1607 can include a lane change behind the target vehicle. Although the expected reward 1606 for changing lanes in front of the target vehicle may be higher than the expected reward 1608 associated with changing lanes behind the target vehicle, it can be determined that changing lanes behind the target vehicle may lead to a future state for which there may be action options that result in potential rewards even higher than the expected rewards 1606, 1608, or other rewards based on the available actions in response to the current, sensed navigation state.

[0371] The selection from among the potential actions at step 1625 can be based on any suitable comparison of expected rewards (or any other measure or indicator of the benefits associated with one potential action compared to another). In some cases, as described above, if a second potential action is expected to provide at least one future action associated with an expected reward higher than the reward associated with the first potential action, the first potential action can be bypassed and the second potential action can be selected. In other cases, more complex comparisons can be employed. For example, the rewards associated with action options in response to an expected future state can be compared to more than one expected reward associated with the determined potential action.

[0372] In some cases, if at least one future action is expected to yield a reward that is higher than any reward expected as a result of a potential action for the current state (e.g., expected rewards 1606, 1608, 1610, etc.), then the action and expected reward based on the projected future state can influence the selection of the potential action for the current state. In certain cases, the future action option that yields the highest expected reward (e.g., from among the expected rewards associated with potential actions for the sensed current state and the expected rewards associated with potential future action options relative to potential future navigation states) can be used as a guide for selecting the potential action for the current navigation state. That is, after identifying the future action option that yields the highest expected reward (or a reward above a predetermined threshold, etc.), at step 1625, the potential action that will lead to the future state associated with the identified future action that yields the highest expected reward can be selected.

[0373] In other cases, the selection of available actions can be made based on the determined difference between expected rewards. For example, if the difference between the expected reward associated with the future action determined at step 1621 and expected reward 1606 is greater than the difference between expected reward 1608 and expected reward 1606 (assuming a positive-signed difference), then the second potential action determined at step 1607 can be selected. In another example, if the difference between the expected reward associated with the future action determined at step 1621 and the expected reward associated with the future action determined at step 1619 is greater than the difference between expected reward 1608 and expected reward 1606, then the second potential action determined at step 1607 can be selected.

[0374] Several examples for selecting from among potential actions for the current navigation state have been described. However, any other suitable comparison technique or criterion can be used for selecting available actions through long-term planning based on the analysis of actions and rewards extending to the projected future state. Additionally, while Figure 16 two layers in the long-term planning analysis are represented (e.g., the first layer considers the rewards generated by potential actions for the current state, and the second layer considers the rewards generated by future action options in response to the projected future state), however, an analysis based on more layers can be possible. For example, rather than basing the long-term planning analysis on one or two layers, a three-layer, four-layer, or more-layer analysis can be used to select from among the available potential actions in response to the current navigation state.

[0375] After selecting among potential actions responsive to a sensed navigation state, at step 1627, at least one processor may cause at least one adjustment action of a navigation actuator of the host vehicle responsive to the selected potential navigation. The navigation actuator may include any suitable device for controlling at least one aspect of the host vehicle. For example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.

[0376] Navigation Based on Inferred Aggression by Other Parties

[0377] The target vehicle may be monitored by analyzing an acquired image stream to determine indicators of driving aggression. Aggression is described herein as a qualitative or quantitative parameter, but other characteristics may be used: the perceived level of attention (potential impairment of the driver, distraction caused by a mobile phone, asleep, etc.). In some cases, the target vehicle may be considered to have a defensive posture, and in some cases, it may be determined that the target vehicle has a more aggressive posture. Navigation actions may be selected or generated based on the aggression indicators. For example, in some cases, the relative speed, relative acceleration, increase in relative acceleration, following distance, etc. with respect to the host vehicle may be tracked to determine whether the target vehicle is aggressive or defensive. For example, if the target vehicle is determined to have an aggression level exceeding a threshold, the host vehicle may tend to yield to the target vehicle. The aggression level of the target vehicle may also be discerned based on the determined behavior of the target vehicle with respect to one or more obstacles (e.g., a vehicle ahead, an obstacle on the road, a traffic light, etc.) in or near the path of the target vehicle.

[0378] As an introduction to this concept, an exemplary experiment will be described of merging into a roundabout relative to the host vehicle, where the navigation goal is to pass through the roundabout and exit the roundabout. The scenario may start with the host vehicle approaching the entrance of the roundabout and may end with the host vehicle reaching the exit of the roundabout (e.g., the second exit). Success may be measured based on whether the host vehicle always maintains a safe distance from all other vehicles, whether the host vehicle completes the route as soon as possible, and whether the host vehicle follows a smooth acceleration strategy. In this example, N T target vehicles may be randomly placed on the roundabout. To model a mixture of adversarial and typical behaviors, with probability p, the target vehicle may be modeled by an "aggressive" driving strategy such that the aggressive target vehicle accelerates when the host vehicle attempts to merge in front of the target vehicle. With probability 1 - p, the target vehicle may be modeled by a "defensive" driving strategy such that the target vehicle decelerates and allows the host vehicle to merge. In this experiment, p = 0.5, and no information about the type of other drivers may be provided to the navigation system of the host vehicle. The type of other drivers may be randomly selected at the start of the segment.

[0379] The navigation state can be represented by the speed and position of the host vehicle (the agent), as well as the position, speed, and acceleration of the target vehicle. It may be important to maintain an observation of the target acceleration in order to distinguish between aggressive and defensive drivers based on the current state. All target vehicles may move along a one-dimensional curve outlined by a circular path. The host vehicle may move along its own one-dimensional curve that intersects the curve of the target vehicle at a merge point, and this point is the starting point of both curves. To model reasonable driving, the absolute value of the acceleration of all vehicles may be bounded by a constant. Since reverse driving is not allowed, the speed may also be passed through a ReLU. Note that by not allowing reverse driving, long-term planning may become necessary because the agent will not regret its past actions.

[0380] As described above, the next state s t+1 can be decomposed into a predictable part and an unpredictable part v t in sum. The expression can represent the dynamics of the vehicle position and speed (which can be well-defined in a differentiable manner), while v t can represent the acceleration of the target vehicle. It can be verified that can be expressed as a composition of ReLU functions on an affine transformation, and thus it is differentiable with respect to s t and a t . The vector v t can be defined by the simulator in a non-differentiable manner and can exhibit aggressive behavior towards some targets and defensive behavior towards other targets. Two frames from such a simulator are shown in Figure 17A and 17B . In this exemplary experiment, the host vehicle 1701 learns to decelerate as it approaches the entrance of the roundabout. It also learns to yield to aggressive vehicles (e.g., vehicles 1703 and 1705) and to safely continue driving when merging in front of defensive vehicles (e.g., vehicles 1706, 1708, and 1710). In the example represented by Figure 17A and 17B , the type of the target vehicle is not provided to the navigation system of the host vehicle 1701. Instead, based on the inference of the observed (e.g., target vehicle's) position and acceleration, it is determined whether a particular vehicle is identified as aggressive or defensive. In Figure 17A , based on the position, speed, and / or relative acceleration, the host vehicle 1701 can determine that vehicle 1703 has an aggressive tendency, and thus the host vehicle 1701 can stop and wait for the target vehicle 1703 to pass rather than attempting to merge in front of the target vehicle 1703. However, in Figure 17BIn this case, the target vehicle 1701 identifies that the target vehicle 1710 traveling behind the vehicle 1703 exhibits a defensive tendency (also based on the observed position, speed, and / or relative acceleration of the vehicle 1710), and thus successfully changes lanes in front of the target vehicle 1710 and behind the target vehicle 1703.

[0381] Figure 18 A flowchart representing an exemplary algorithm for navigating a host vehicle based on the aggression of predicted other vehicles is provided. In Figure 18 the example, the aggression level associated with at least one target vehicle can be inferred based on the observed behavior of the target vehicle with respect to objects in the environment relative to the target vehicle. For example, in step 1801, at least one processing device (e.g., the processing device 110) of the host vehicle navigation system can receive a plurality of images representing the environment of the host vehicle from a camera associated with the host vehicle. In step 1803, the analysis of one or more of the received images can enable at least one processor to identify a target vehicle (e.g., the vehicle 1703) in the environment of the host vehicle 1701. In step 1805, the analysis of one or more of the received images can enable at least one processing device to identify at least one obstacle of the target vehicle in the environment of the host vehicle. The objects can include debris in the road, stop lights / traffic lights, pedestrians, another vehicle (e.g., a vehicle traveling in front of the target vehicle, a parked vehicle, etc.), boxes in the road, roadblocks, curbs, or any other type of object that may be encountered in the environment of the host vehicle. In step 1807, the analysis of one or more of the received images can enable at least one processing device to determine at least one navigation characteristic of the target vehicle with respect to at least one identified obstacle of the target vehicle.

[0382] Various navigation characteristics can be used to infer the aggression level of the detected target vehicle in order to generate an appropriate navigation response to the target vehicle. For example, such navigation characteristics can include the relative acceleration between the target vehicle and at least one identified obstacle, the distance of the target vehicle from the obstacle (e.g., the following distance of the target vehicle behind another vehicle), and / or the relative speed between the target vehicle and the obstacle, etc.

[0383] In some embodiments, the navigation characteristics of a target vehicle can be determined based on the output from sensors associated with the host vehicle (e.g., radar, speed sensors, GPS, etc.). However, in some cases, the navigation characteristics of the target vehicle can be determined, at least in part or entirely, based on an analysis of an image of the environment of the host vehicle. For example, the image analysis techniques described above and in, e.g., U.S. Patent No. 9,168,868, which is incorporated herein by reference, can be used to identify a target vehicle within the environment of the host vehicle. And monitoring the position of the target vehicle in the captured images over time and / or monitoring the position of one or more features associated with the target vehicle (e.g., taillights, headlights, bumpers, wheels, etc.) in the captured images can enable determination of the relative distance, speed, and / or acceleration between the target vehicle and the host vehicle or between the target vehicle and one or more other objects in the environment of the host vehicle.

[0384] The aggression level of the identified target vehicle can be inferred from any suitable observed navigation characteristics of the target vehicle or any combination of the observed navigation characteristics. For example, a determination of aggression can be made based on any observed characteristics and one or more predetermined threshold levels or any other suitable qualitative or quantitative analysis. In some embodiments, if the target vehicle is observed to follow the host vehicle or another vehicle at a distance less than a predetermined aggression distance threshold, the target vehicle can be considered aggressive. On the other hand, a target vehicle observed to follow the host vehicle or another vehicle at a distance greater than a predetermined defensive distance threshold can be considered defensive. The predetermined aggression distance threshold need not be the same as the predetermined defensive distance threshold. Additionally, either or both of the predetermined aggression distance threshold and the predetermined defensive distance threshold can include a range of values rather than a bright line value. Further, neither the predetermined aggression distance threshold nor the predetermined defensive distance threshold need be fixed. Instead, these values or ranges of values can change over time and different thresholds / ranges of thresholds can be applied based on the characteristics of the observed target vehicle. For example, the thresholds applied can depend on one or more other characteristics of the target vehicle. A higher observed relative speed and / or acceleration can warrant the application of a larger threshold / range. Conversely, a lower relative speed and / or acceleration (including zero relative speed and / or acceleration) can warrant the application of a smaller distance threshold / range when making aggression / defense inferences.

[0385] Aggressive / defensive inference can also be based on relative speed and / or relative acceleration thresholds. If the observed relative speed and / or relative acceleration of a target vehicle with respect to another vehicle exceeds a predetermined level or range, the target vehicle can be considered aggressive. If the observed relative speed and / or relative acceleration of a target vehicle with respect to another vehicle is below a predetermined level or range, the target vehicle can be considered defensive.

[0386] Although the aggressive / defensive determination can be made based on any observed navigation characteristic alone, the determination can also depend on any combination of the observed characteristics. For example, as described above, in some cases, a target vehicle can be considered aggressive based solely on observing that it is following another vehicle at a distance below a certain threshold or range. However, in other cases, if a target vehicle follows another vehicle at a distance less than a predetermined amount (which can be the same or different from the threshold applied when making a determination based solely on distance) and has a relative speed and / or relative acceleration greater than a predetermined amount or range, the target vehicle can be considered aggressive. Similarly, a target vehicle can be considered defensive based solely on observing that it is following another vehicle at a distance greater than a certain threshold or range. However, in other cases, if a target vehicle follows another vehicle at a distance greater than a predetermined amount (which can be the same or different from the threshold applied when making a determination based solely on distance) and has a relative speed and / or relative acceleration less than a predetermined amount or range, the target vehicle can be considered defensive. If, for example, a vehicle exceeds 0.5G of acceleration or deceleration (e.g., a jerk of 5 meters per cubic second (m / s3)), a vehicle has a lateral acceleration of 0.5G during a lane change or on a curve, one vehicle causes another vehicle to perform any of the above, a vehicle changes lanes and causes another vehicle to yield with a deceleration of 0.3G or an acceleration rate of more than 3 m / s3, and / or a vehicle changes two lanes without stopping, system 100 can make an aggressive / defensive determination.

[0387] It should be understood that a reference to a quantity exceeding a range can indicate that the quantity exceeds all values associated with the range or falls within the range. Similarly, a reference to a quantity below a range can indicate that the quantity is below all values associated with the range or falls within the range. Additionally, although examples for making aggressive / defensive inferences have been described with respect to distance, relative acceleration, and relative speed, any other suitable quantity can be used. For example, a calculation of the time until a collision is to occur, or any indirect indicator of the distance, acceleration, and / or speed of the target vehicle can be used. It should also be noted that although the above examples focus on a target vehicle with respect to other vehicles, aggressive / defensive inferences can be made by observing the navigation characteristics of a target vehicle with respect to any other type of obstacle (e.g., pedestrians, roadblocks, traffic lights, debris, etc.).

[0388] Return to Figure 17A and 17B In the example shown, when the host vehicle 1701 approaches the roundabout, the navigation system (including at least one of its processing devices) may receive a stream of images from a camera associated with the host vehicle. Based on the analysis of one or more of the received images, any one of the target vehicles 1703, 1705, 1706, 1708, and 1710 may be identified. Additionally, the navigation system may analyze the navigation characteristics of one or more of the identified target vehicles. The navigation system may recognize that the gap between target vehicles 1703 and 1705 represents a first opportunity for a potential merge onto the roundabout. The navigation system may analyze target vehicle 1703 to determine an aggression indicator associated with target vehicle 1703. If target vehicle 1703 is considered aggressive, the host vehicle navigation system may choose to yield to vehicle 1703 rather than merge in front of vehicle 1703. On the other hand, if target vehicle 1703 is considered defensive, the host vehicle navigation system may attempt to complete a merge maneuver in front of vehicle 1703.

[0389] When the host vehicle 1701 approaches the roundabout, at least one processing device of the navigation system may analyze the captured image to determine the navigation characteristics associated with the target vehicle 1703. For example, based on the image, it may determine that vehicle 1703 is following vehicle 1705 at a distance that provides sufficient clearance for the safe entry of the host vehicle 1701. In fact, it may be determined that vehicle 1703 is following vehicle 1705 at a distance that exceeds the aggressive distance threshold, and based on this information, the host vehicle navigation system may tend to identify the target vehicle 1703 as defensive. However, in some cases, as described above, more than one navigation characteristic of the target vehicle may be analyzed when making the aggressive / defensive determination. Further analysis may show that the host vehicle navigation system determines that when the target vehicle 1703 is following behind the target vehicle 1705 at a non-aggressive distance, vehicle 1703 has a relative speed and / or relative acceleration that exceeds one or more thresholds associated with aggressive behavior with respect to vehicle 1705. In fact, the host vehicle 1701 may determine that the target vehicle 1703 is accelerating with respect to vehicle 1705 and approaching the gap that exists between vehicles 1703 and 1705. Based on further analysis of the relative speed, acceleration, and distance (and even the rate at which the gap between vehicles 1703 and 1705 is closing), the host vehicle 1701 may determine that the target vehicle 1703 is behaving aggressively. Thus, although there may be sufficient clearance for the host vehicle to navigate safely, the host vehicle 1701 may anticipate that a lane change in front of the target vehicle 1703 will result in an aggressively navigating vehicle immediately behind the host vehicle. Additionally, based on the behavior observed through image analysis or other sensor outputs, the target vehicle 1703 may be expected to: if the host vehicle 1701 were to change lanes in front of vehicle 1703, the target vehicle 1703 would continue to accelerate towards the host vehicle 1701 or continue to travel towards the host vehicle 1701 at a non-zero relative speed. From a safety perspective, this situation may be undesirable and may also cause discomfort to the passengers of the host vehicle. For this reason, as Figure 17B shown, the host vehicle 1701 may choose to yield to vehicle 1703 and change lanes to the roundabout behind vehicle 1703 and in front of vehicle 1710, which is considered defensive based on an analysis of one or more of its navigation characteristics.

[0390] Back to Figure 18, at step 1809, at least one processing device of the navigation system of the host vehicle may determine a navigation action of the host vehicle (e.g., changing lanes in front of vehicle 1710 and behind vehicle 1703) based on at least one navigation characteristic of the identified target vehicle relative to the identified obstacle. To implement this navigation action (at step 1811), at least one processing device may cause at least one adjustment of the navigation actuator of the host vehicle in response to the determined navigation action. For example, braking may be applied to yield to Figure 17A vehicle 1703 in, and the accelerator may be applied in conjunction with the steering of the wheels of the host vehicle to cause the host vehicle to enter the roundabout behind vehicle 1703, as Figure 17B shown.

[0391] As described in the above example, the navigation of the host vehicle may be based on the navigation characteristics of the target vehicle relative to another vehicle or object. Additionally, the navigation of the host vehicle may be based only on the navigation characteristics of the target vehicle without specifically referring to another vehicle or object. For example, at Figure 18 step 1807, the analysis of multiple images captured from the environment of the host vehicle may enable the determination of at least one navigation characteristic of the identified target vehicle that indicates an aggression level associated with the target vehicle. The navigation characteristics may include speed, acceleration, etc., which do not require reference to another object or the target vehicle in order to make an aggressiveness / defensiveness determination. For example, observed acceleration and / or speed associated with the target vehicle that exceed a predetermined threshold or fall within or exceed a numerical range may indicate aggressive behavior. Conversely, observed acceleration and / or speed associated with the target vehicle that are below a predetermined threshold or fall within or exceed a numerical range may indicate defensive behavior.

[0392] Of course, in some cases, in order to make an aggressiveness / defensiveness determination, the observed navigation characteristics (e.g., position, distance, acceleration, etc.) may be referenced relative to the host vehicle. For example, the observed navigation characteristics of the target vehicle that indicate an aggression level associated with the target vehicle may include an increase in the relative acceleration between the target vehicle and the host vehicle, the following distance of the target vehicle behind the host vehicle, the relative speed between the target vehicle and the host vehicle, etc.

[0393] Navigation Based on Accident Liability Constraints

[0394] As described in the above sections, planned navigation actions can be tested against predefined constraints to ensure compliance with certain rules. In some embodiments, this concept can be extended to the consideration of potential accident liability. As described below, the primary goal of autonomous navigation is safety. Since absolute safety may be impossible (e.g., at least because a particular host vehicle under autonomous control cannot control other vehicles in its vicinity - it can only control its own actions), using potential accident liability as a consideration and indeed as a constraint on planned actions in autonomous navigation can help ensure that a particular autonomous vehicle does not take any actions that are considered unsafe - for example, potential accident liability may be attached to those actions of the host vehicle. If the host vehicle only takes actions that are safe and determined not to result in an accident for which the host vehicle itself is at fault or liable, then a desired level of accident avoidance can be achieved (e.g., driving less than 10 -9 per hour).

[0395] Challenges posed by most current autonomous driving methods include a lack of safety guarantees (or at least the inability to provide a desired level of safety), as well as a lack of scalability. Consider the problem of ensuring safe driving for multi-agent scenarios. Since society is unlikely to tolerate road accident deaths caused by machines, an acceptable level of safety is crucial for the acceptance of autonomous vehicles. While the goal may be to provide zero accidents, this may be impossible because accidents typically involve multiple agents, and situations can be envisioned where an accident occurs entirely due to the liability of other agents. For example, as Figure 19 shown, the host vehicle 1901 is driving on a multi-lane highway. While the host vehicle 1901 can control its own actions relative to the target vehicles 1903, 1905, 1907, and 1909, it cannot control the actions of the target vehicles in its vicinity. As a result, if vehicle 1905 suddenly cuts into the host vehicle's lane, for example, on a collision course with the host vehicle, the host vehicle 1901 may not be able to avoid an accident with at least one of the target vehicles. To address this conundrum, the typical response of autonomous vehicle practitioners has been to adopt a statistics-driven approach, in which safety verification becomes more stringent as more data is collected over more miles.

[0396] However, to understand the nature of the problems with data-driven approaches to safety, the first thing to consider is that the probability of death caused by accidents per hour of (human) driving is known to be 10 -6 . A reasonable assumption is that in order for society to accept machines replacing humans in the driving task, the mortality rate should be reduced by three orders of magnitude, i.e., reduced to a probability of 10 -9 per hour. This estimate is similar to the assumed mortality rate of airbags and is derived from aviation standards. For example, 10 -9is the probability that a wing spontaneously detaches from an aircraft in mid-air. However, it is not practical to attempt to guarantee safety using data-driven statistical methods that provide additional confidence by accumulating driving mileage. The amount of data required to guarantee a death probability of 1 in 10 -9 is proportional to its reciprocal (i.e., 10 9 hours of data), which is on the order of about 30 billion miles. Additionally, multi-agent systems interact with their environment and may not be verifiable offline (unless there is a realistic simulator available that mimics real human driving in all its richness and complexity, such as reckless driving—but the problem of validating the simulator is even more difficult than creating a safe autonomous vehicle agent). Any change to the planning and control software will require a similar amount of new data collection, which is clearly impractical and unrealistic. Moreover, systems developed through data always suffer from a lack of interpretability and explainability of the actions taken—if an autonomous vehicle (AV) is involved in an accident resulting in death, we need to know why. Therefore, a model-based approach to safety is needed, but the existing "functional safety" and ASIL requirements in the automotive industry are not designed to address multi-agent environments.

[0397] The second major challenge in developing a safe driving model for autonomous vehicles is the need for scalability. The premise of AVs is not just "building a better world," but rather is based on the premise that mobility without a driver can be maintained at a lower cost than mobility with a driver. This premise is always coupled with the concept of scalability—that is, supporting the mass production of AVs (in the millions), and more importantly, supporting negligible incremental costs in order to be able to drive in new cities. Therefore, the cost of computing and sensing does matter, and the cost of verification and the ability to drive "everywhere" rather than in a few selected cities are also necessary requirements for maintaining the business if AVs are to be manufactured on a large scale.

[0398] The problem with most current methods lies in the "brute force" mindset along three axes: (i) the required "computational density", (ii) the way in which high-definition maps are defined and created, and (iii) the required specifications of sensors. The brute force approach runs counter to scalability and shifts the focus to a future where unconstrained in-vehicle computing is ubiquitous, where the cost of building and maintaining HD maps becomes negligible and scalable, and where exotic, super-advanced sensors will be developed, produced to automotive grade, and cost negligible. A future where any one of these scenarios is realized is indeed possible, but it is very likely a low-probability event that all of the above hold. Therefore, it is necessary to provide a formal model that combines safety and scalability into an AV program that is socially acceptable and scalable in the sense of supporting millions of vehicles driving anywhere in developed countries.

[0399] The disclosed embodiments represent a solution that can provide a target safety level (or even exceed safety goals) and can also scale to systems that include millions (or more) of autonomous vehicles. In terms of safety, a model called "Responsibility Sensitive Safety (RSS)" is introduced, which formalizes the concept of "accident blameworthiness", is interpretable and explainable, and incorporates the notion of "responsibility" into the actions of robotic agents. The definition of RSS is agnostic to the way it is implemented - this is a key feature for the goal of facilitating the creation of a compelling global safety model. The motivation for RSS is the observation that (as Figure 19 shown) agents play an asymmetric role in accidents, in which typically only one of the agents is responsible for the accident and is thus to blame for it. The RSS model also includes a formal treatment of "prudent driving" under limited sensing conditions, where not all agents are always visible (e.g., due to occlusion). A major goal of the RSS model is to ensure that an agent never has an accident "blamed" on it or for which it is responsible. It can only be useful if the model comes with an effective policy that complies with RSS (e.g., a function that maps "sensing state" to actions). For example, an action that seems innocent at the current moment may lead to a catastrophic event in the distant future ("butterfly effect"). RSS can help build a set of local constraints in the short-term future that can guarantee (or at least virtually guarantee) that no accident will occur due to the actions of the host vehicle in the future.

[0400] Another contribution progresses around the introduction of a "semantic" language that includes units, measurements, and action spaces, as well as specifications on how to incorporate them into the planning, sensing, and actuation of AVs. To gain a semantic perspective, consider how someone taking a driving course is instructed to think about "driving strategies" in this context. These instructions are not geometric—they do not take the form of "drive 13.7 meters at the current speed and then accelerate at a rate of 0.8 m / s 2 ." Instead, these instructions are semantic in nature—"follow the car in front of you" or "pass that car on your left." The typical language of human driving strategies is about longitudinal and lateral goals, rather than geometric units via acceleration vectors. A formalized semantic language can be useful in several ways that relate to the computational complexity of planning, which does not increase exponentially with time and the number of agents, the ways of safe and comfortable interaction, the ways of defining the computations for sensing, the specifications of sensor modalities, and how they interact in fusion methods. The fusion method (based on the semantic language) can ensure that the RSS model can achieve the required probability of death of 1 in 10 5 for every hour of driving with only an offline validation on a dataset of driving data of magnitude 10 -9 hours.

[0401] For example, in a reinforcement learning setting, a Q-function can be defined over a semantic space (e.g., a function that evaluates the long-term quality of performing an action a ∈ A when the agent is in state s ∈ S; given such a Q-function, the natural choice of action can be to pick the action with the highest quality π(s) = argmax a Q(s,a)), in which the number of trajectories to examine at any given time is bounded by 10 4 , regardless of the time horizon used for planning. The signal-to-noise ratio in this space can be high, allowing effective machine learning methods to successfully model the Q-function. In the case of computations for sensing, semantics can allow for a distinction between errors that affect safety and those that affect driving comfort. We define a PAC model for sensing (Probably Approximate Correct, borrowing the PAC learning terminology of Valiant), couple this model with the Q-function, and show how measurement errors can be incorporated into planning in a way that is consistent with RSS while still allowing for the optimization of driving comfort. The semantic language may be important for the success of certain aspects of this model because other standard measurements of errors (such as errors relative to a global coordinate system) may not conform to the PAC sensing model. Additionally, the semantic language may be an important enabler for defining HD maps, which can be constructed using low-bandwidth sensing data, thus enabling construction via crowdsourcing and supporting scalability.

[0402] In summary, the disclosed embodiments may include a formal model covering the following important components of AV: sensing, planning, and acting. This model may help ensure that, from a planning perspective, accidents for which the AV is itself responsible do not occur. And through the PAC sensing model, even in the presence of sensing errors, the described fusion method may only require a very reasonable amount of offline data collection to conform to the described safety model. In addition, this model can combine safety and scalability through semantic language, thus providing a complete approach for safe and scalable AVs. Finally, it is worth noting that developing an acceptable safety model that will be adopted by the industry and regulatory agencies may be a prerequisite for the success of AVs.

[0403] The RSS model generally follows the classical sensing-planning-acting robot control approach. The sensing system can be responsible for understanding the current state of the host vehicle's environment. The planning part, which can be referred to as the "driving strategy" and can be implemented by a set of hard-coded instructions, a trained system (e.g., a neural network), or a combination, can be responsible for determining what the best next move is based on the available options for achieving driving goals (e.g., how to move from the left lane to the right lane to exit the highway). The acting part is responsible for implementing the plan (e.g., a system of actuators and one or more controllers for steering, accelerating, and / or braking the vehicle, etc. to implement the selected navigation actions). The embodiments described below mainly focus on the sensing and planning parts.

[0404] Accidents can stem from sensing errors or planning errors. Planning is a multi-agent effort because there are other road users (humans and machines) who react to the actions of the AV. The described RSS model is designed to address safety issues in the planning part. This can be called multi-agent safety. In statistical methods, the estimation of the probability of planning errors can be done "online". That is, after each software update, the vehicle must be driven billions of miles with the new version to provide an acceptable level of estimate of the frequency of planning errors. This is clearly infeasible. As an alternative, the RSS model can provide a 100% guarantee (or almost 100% guarantee) that the planning module will not make errors attributable to the AV (formally defining the concept of "attributable"). The RSS model can also provide an effective method for its verification that does not rely on online testing.

[0405] Errors in the sensing system may be easier to verify because sensing can be independent of vehicle actions, and thus we can use "offline" data to verify the probability of severe sensing errors. However, even collecting offline data for more than 10 9 hours of driving is challenging. As part of the description of the disclosed sensing system, a fusion method is described that can use significantly less data for verification.

[0406] The described RSS system can also scale to millions of vehicles. For example, the described semantic driving strategies and the applied safety constraints can be made consistent with sensing and mapping requirements that can scale to millions of vehicles even in today's technology.

[0407] The fundamental building block of such a system is a thorough safety definition, which is the minimum standard that an AV system may need to comply with. In the following technical lemma, the statistical methods used to verify AV systems are shown to be infeasible, even for verifying simple statements such as "the system has N accidents per hour". This means that a model-based safety definition is the only viable tool for verifying AV systems.

[0408] Lemma 1 Let X be a probability space, and A be an event with Pr(A) = p 1 < 0.1. Suppose we sample (i.i.d.) samples from X, and let Then

[0409] Pr(Z = 0) ≥ e -2 .

[0410] Proof We use the inequality 1 - x ≥ e -2x (proved in Appendix A.1 for completeness), and obtain

[0411]

[0412] Corollary 1 Suppose an AV system AV 1 has accidents with a small yet insufficient probability p 1 . Any deterministic verification procedure given 1 / p 1 samples will fail to distinguish AV 1 from a different AV system AV 0 that never has accidents, with a constant probability.

[0413] To gain a perspective on typical values of these probabilities, assume we desire an accident probability of 10 -9 per hour, while a certain AV system offers only a probability of 10 -8 . Even if the system has 10 8 hours of driving, there is a constant probability that the verification process will fail to indicate that the system is dangerous.

[0414] Finally, note that the difficulty lies in invalidating a single, specific,...

Claims

1. A navigation system for navigating a host vehicle according to at least one navigation goal of an autonomous host vehicle, the navigation system comprising: At least one processor programmed to: receiving a sensor output from one or more sensors indicative of at least one aspect of a motion of a host vehicle relative to a host vehicle environment, wherein the sensor output is generated at a first time that is later than a data acquisition time when a measurement or data acquisition upon which the sensor output is based was acquired and earlier than a second time when the sensor output is received by the at least one processor; for a motion prediction time, generating a prediction of at least one aspect of host vehicle motion based at least in part on the received sensor output and an estimate of how at least one aspect of host vehicle motion has changed over a time interval between the data acquisition time and the motion prediction time; determining a planned navigation maneuver for the host vehicle based at least in part on at least one navigation goal for the host vehicle and based on the generated prediction of at least one aspect of the host vehicle's motion; generating a navigation command for implementing at least a portion of the planned navigation action; as well as providing the navigation command to at least one actuation system of the host vehicle such that the at least one actuation system receives the navigation command at a third time that is later than the second time and earlier than or substantially equal to an actuation time for a component of the at least one actuation system to respond to the received command; wherein the motion prediction time is after the data acquisition time and earlier than or equal to the actuation time; and Wherein the prediction of at least one aspect of host vehicle motion at the motion prediction time addresses a mismatch between: a data acquisition rate associated with one or more sensors and a control rate associated with a rate at which the at least one processor generates the navigation commands.

2. The navigation system according to claim 1, wherein: The motion prediction time substantially corresponds to the third time.

3. The navigation system according to claim 1, wherein: The motion prediction time substantially corresponds to the second time.

4. The navigation system according to claim 1, wherein: The motion prediction time substantially corresponds to the actuation time. 5 . The navigation system of claim 1 , wherein the one or more sensors include a velocity sensor, an accelerometer, a camera, a lidar system, or a radar system.

6. The navigation system according to claim 1, wherein: The prediction of at least one aspect of the host vehicle's motion includes a prediction of at least one of a speed or an acceleration of the host vehicle at the time of the motion prediction.

7. The navigation system according to claim 1, wherein: The prediction of at least one aspect of the host vehicle's motion includes a prediction of a path of the host vehicle at the time of the motion prediction.

8. The navigation system according to claim 7, wherein: The prediction of the path of the host vehicle at the motion prediction time includes a target heading of the host vehicle.

9. The navigation system according to claim 7, wherein: The one or more sensors include a camera, and the prediction of the path of the host vehicle at the motion prediction time is based on at least one image captured by the camera.

10. The navigation system according to claim 7, wherein: The prediction of the path of the host vehicle at the motion prediction time is based at least on the determined speed of the host vehicle and a target trajectory of the host vehicle contained in a map of the road segment on which the host vehicle is traveling.

11. The navigation system according to claim 10, wherein: The target trajectory includes a predetermined three-dimensional spline representing a preferred path along at least one lane of the road segment.

12. The navigation system according to claim 1, wherein: The prediction of at least one aspect of the host vehicle's motion is based on at least one of a determined brake pedal position, a determined throttle position, a determined air resistance opposing the host vehicle's motion, a friction force, or a grade of a road segment on which the host vehicle is traveling.

13. The navigation system according to claim 1, wherein: The planned navigation maneuver includes at least one of a speed change or a heading change of the host vehicle.

14. The navigation system according to claim 1, wherein: The at least one actuation system includes one or more of a throttle actuation system, a brake actuation system, or a steering actuation system.

15. The navigation system according to claim 1, wherein: The navigation goal of the host vehicle includes a transition from a first location to a second location.

16. The navigation system according to claim 1, wherein: The navigation goal of the host vehicle includes a lane change from a current lane occupied by the host vehicle to an adjacent lane.

17. The navigation system according to claim 1, wherein: The navigation goal of the host vehicle includes maintaining a proximity buffer zone between the host vehicle and a detected target vehicle, wherein the proximity buffer zone is determined based on a detected current speed of the host vehicle, a maximum braking rate capability of the host vehicle, a determined current speed of the target vehicle, and an assumed maximum braking rate capability of the target vehicle, and wherein the proximity buffer zone relative to the target vehicle is further determined based on a maximum acceleration capability of the host vehicle such that the proximity buffer zone includes at least a sum of a host vehicle acceleration distance, a host vehicle stopping distance, and a target vehicle stopping distance, wherein the host vehicle acceleration distance is determined as a distance that the host vehicle will travel if it accelerates at the maximum acceleration capability of the host vehicle within a reaction time associated with the host vehicle, the host vehicle stopping distance is determined as a distance required to reduce the current speed of the host vehicle to zero at the maximum braking rate capability of the host vehicle, and the target vehicle stopping distance is determined as a distance required to reduce the current speed of the target vehicle to zero at the assumed maximum braking rate capability of the target vehicle.

18. The navigation system according to claim 1, wherein: The prediction of at least one aspect of the host vehicle's motion is based on a predetermined function associated with the host vehicle, wherein the predetermined function enables prediction of the host vehicle's future speed and acceleration based on a determined current speed of the host vehicle and a determined brake pedal position or a determined throttle position of the host vehicle.

19. The navigation system of claim 1, wherein the motion prediction time is at least 100 milliseconds after the data acquisition time.

20. The navigation system of claim 1, wherein the motion prediction time is at least 200 milliseconds after the data acquisition time.

21. The navigation system of claim 1, wherein: The navigation command includes at least one of a pedal command for controlling a speed of the host vehicle or a yaw rate command for controlling a heading of the host vehicle.

22. A method for navigating an autonomous host vehicle, the method comprising: receiving a sensor output from one or more sensors indicative of at least one aspect of a motion of a host vehicle relative to a host vehicle environment, wherein the sensor output is generated at a first time that is later than a data acquisition time when a measurement or data acquisition upon which the sensor output is based was acquired and earlier than a second time when the sensor output is received by the at least one processor; for a motion prediction time, generating a prediction of at least one aspect of host vehicle motion based at least in part on the received sensor output and an estimate of how at least one aspect of host vehicle motion has changed over a time interval between the data acquisition time and the motion prediction time; determining a planned navigation maneuver for the host vehicle based at least in part on at least one navigation goal for the host vehicle and based on the generated prediction of at least one aspect of the host vehicle's motion; generating a navigation command for implementing at least a portion of the planned navigation action; as well as providing the navigation command to at least one actuation system of the host vehicle such that the at least one actuation system receives the navigation command at a third time that is later than the second time and earlier than or substantially equal to an actuation time for a component of the at least one actuation system to respond to the received command; wherein the motion prediction time is after the data acquisition time and earlier than or equal to the actuation time; and Wherein the prediction of at least one aspect of host vehicle motion at the motion prediction time addresses a mismatch between: a data acquisition rate associated with one or more sensors and a control rate associated with a rate at which the at least one processor generates the navigation commands.

23. The method according to claim 22, wherein: The motion prediction time substantially corresponds to the third time.

24. The method according to claim 22, wherein: The motion prediction time substantially corresponds to the second time.

25. The method according to claim 22, wherein: The motion prediction time substantially corresponds to the actuation time.

26. A navigation system for navigating a host vehicle according to at least one navigation goal of an autonomous host vehicle, the navigation system comprising a computer-readable non-transitory memory storing instructions that when executed by a processor cause operations comprising: receiving a sensor output from one or more sensors indicative of at least one aspect of motion of the host vehicle relative to a host vehicle environment, wherein the sensor output is generated at a first time that is later than a data acquisition time when a measurement or data acquisition upon which the sensor output is based was acquired and earlier than a second time when the sensor output is received by the at least one processor; for a motion prediction time, generating a prediction of at least one aspect of host vehicle motion based at least in part on the received sensor output and an estimate of how at least one aspect of host vehicle motion has changed over a time interval between the data acquisition time and the motion prediction time; determining a planned navigation maneuver for the host vehicle based at least in part on at least one navigation goal for the host vehicle and based on the generated prediction of at least one aspect of the host vehicle's motion; generating a navigation command for implementing at least a portion of the planned navigation action; as well as providing the navigation command to at least one actuation system of the host vehicle such that the at least one actuation system receives the navigation command at a third time that is later than the second time and earlier than or substantially equal to an actuation time for a component of the at least one actuation system to respond to the received command; wherein the motion prediction time is after the data acquisition time and earlier than or equal to the actuation time; and Wherein the prediction of at least one aspect of host vehicle motion at the motion prediction time addresses a mismatch between: a data acquisition rate associated with one or more sensors and a control rate associated with a rate at which the at least one processor generates the navigation commands.

Citation Information

Patent Citations

  • Collision Warning System

    US9168868B2

  • Driver assistance system, and method for operating said driver assistance system

    CN104837707A

  • robust dead time and dynamic compensation for trajectory tracking control

    DE102014215243A1