System and method for navigation with safe distance
Through image capture and processing equipment combined with GPS data, the autonomous vehicle system analyzes the environment and plans navigation actions, solving the security navigation problems of autonomous vehicles in complex environments and achieving safety scalability.
Patent Information
- Application Number
- CN201980043159.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-11
- Filing Date
- 2019-08-14
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2039-08-14
AI Technical Summary
Existing autonomous vehicle navigation systems are difficult to perform navigation decisions safely and accurately in complex environments, especially when taking into account potential accident liability constraints and are difficult to scale to millions of vehicles.
Image capture equipment and processing equipment are used to analyze the vehicle environment, combine GPS data and sensor information, determine navigation targets and plan navigation actions, consider vehicle braking rate, acceleration capability and speed, and implement safe navigation strategies to avoid collisions.
It realizes safe and accurate navigation decisions for autonomous vehicles in complex environments, meets safety assurance requirements, and can be expanded to a large number of vehicles.
Smart Images

Figure CN112601686B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 718,554, filed on August 14, 2018, U.S. Provisional Patent Application No. 62 / 724,355, filed on August 29, 2018, U.S. Provisional Patent Application No. 62 / 772,366, filed on November 28, 2018, and U.S. Provisional Patent Application No. 62 / 777,914, filed on December 11, 2018. All of the above applications are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure generally relates to autonomous vehicle navigation. In addition, the present disclosure relates to systems and methods for navigating within potential accident liability constraints. Background Art
[0004] As technology continues to advance, the goal of fully autonomous vehicles capable of navigating roadways is imminent. Autonomous vehicles may need to consider a variety of factors and make appropriate decisions based on those factors to safely and accurately reach a desired destination. For example, autonomous vehicles may need to process and interpret visual information (e.g., captured by cameras), information from radar or lidar, and may also use information obtained from other sources (e.g., GPS devices, speed sensors, accelerometers, suspension sensors, etc.). Furthermore, to navigate to a destination, autonomous vehicles may also need to identify their position within a specific roadway (e.g., a specific lane on a multi-lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, observe traffic signals and signs, cross from one road to another at appropriate intersections or junctions, and respond to any other circumstances that occur or develop during vehicle operation. Furthermore, navigation systems may need to adhere to certain imposed constraints. In some cases, those constraints may involve interactions between the host vehicle and one or more other objects (such as other vehicles, pedestrians, etc.). In other cases, the constraints may involve responsibility rules to be followed when performing one or more navigation actions on behalf of the host vehicle.
[0005] In the field of autonomous driving, there are two key considerations for viable autonomous vehicle systems. The first is standardization of safety assurance, including the requirements that each autonomous vehicle must meet to guarantee safety, and how to verify those requirements. The second is scalability, as engineering solutions that incur significant costs will not scale to millions of vehicles and could prevent widespread adoption of autonomous vehicles, or even prevent widespread adoption. Therefore, there is a need for an interpretable mathematical model for safety assurance and a system design that adheres to safety assurance requirements while being scalable to millions of vehicles. Summary of the Invention
[0006] Embodiments consistent with the present disclosure provide systems and methods for autonomous vehicle navigation. The disclosed embodiments may use cameras to provide autonomous vehicle navigation features. For example, consistent with the disclosed embodiments, the disclosed system may include one, two, or more cameras that monitor the vehicle environment. The disclosed system may provide a navigation response based on, for example, analysis of images captured by one or more cameras. The navigation response may also take into account other data, including, for example, global positioning system (GPS) data, sensor data (e.g., from accelerometers, velocity sensors, suspension sensors, etc.), and / or other map data.
[0007] In one embodiment, a system for navigating a host vehicle is disclosed. The system may include at least one processing device programmed to receive at least one image representing an environment of the host vehicle from an image capture device; determine a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy; analyze the at least one image to identify a target vehicle in the environment of the host vehicle, wherein the target vehicle is traveling toward the host vehicle; determine a next-state distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determine a host vehicle braking rate, a host vehicle maximum acceleration capability, and a current speed of the host vehicle; determine a stopping distance for the host vehicle based on the host vehicle braking rate, the host vehicle maximum acceleration capability, and the current speed of the host vehicle; determine a target vehicle current speed, the target vehicle maximum acceleration capability, and a target vehicle braking rate; determine a stopping distance for the target vehicle based on the target vehicle braking rate, the target vehicle maximum acceleration capability, and the current speed of the target vehicle; and implement the planned navigation action if the determined next-state distance is greater than the sum of the stopping distance of the host vehicle and the stopping distance of the target vehicle.
[0008] In one embodiment, a method for navigating a host vehicle is disclosed. The method may include: receiving at least one image representing an environment of the host vehicle from an image capture device; determining a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy; analyzing the at least one image to identify a target vehicle in the environment of the host vehicle, wherein the target vehicle is traveling toward the host vehicle; determining a next-state distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determining a host vehicle braking rate, a host vehicle maximum acceleration capability, and a current speed of the host vehicle; determining a stopping distance for the host vehicle based on the host vehicle braking rate, the host vehicle maximum acceleration capability, and the current speed of the host vehicle; determining the current speed of the target vehicle, the target vehicle maximum acceleration capability, and the target vehicle braking rate; determining a stopping distance for the target vehicle based on the target vehicle braking rate, the target vehicle maximum acceleration capability, and the current speed of the target vehicle; and implementing the planned navigation action if the determined next-state distance is greater than the sum of the stopping distance of the host vehicle and the stopping distance of the target vehicle.
[0009] In one embodiment, a system for navigating a host vehicle is disclosed. The system may include at least one processing device programmed to receive at least one image representing an environment of the host vehicle from an image capture device; determine a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy; analyze the at least one image to identify a target vehicle in the environment of the host vehicle; determine a next-state lateral distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determine a maximum yaw rate capability of the host vehicle, a maximum variable turning radius capability of the host vehicle, and a current lateral speed of the host vehicle; determine a lateral braking distance of the host vehicle based on the maximum yaw rate capability of the host vehicle, the maximum variable turning radius capability of the host vehicle, and the current lateral speed of the host vehicle; determine a current lateral speed of a target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum variable turning radius capability of the target vehicle; determine a lateral braking distance of the target vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum variable turning radius capability of the target vehicle; and implement the planned navigation action if the determined next-state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle.
[0010] In one embodiment, a method for navigating a host vehicle is disclosed. The method may include: receiving at least one image representing an environment of the host vehicle from an image capture device; determining a planned navigation action for achieving a navigation goal of the host vehicle based on at least one driving strategy; analyzing the at least one image to identify a target vehicle in the environment of the host vehicle; determining a next-state lateral distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; determining a maximum yaw rate capability of the host vehicle, a maximum variable turning radius capability of the host vehicle, and a current lateral speed of the host vehicle; determining a lateral braking distance of the host vehicle based on the maximum yaw rate capability of the host vehicle, the maximum variable turning radius capability of the host vehicle, and the current lateral speed of the host vehicle; determining a current lateral speed of a target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum variable turning radius capability of the target vehicle; determining a lateral braking distance of the target vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum variable turning radius capability of the target vehicle; and implementing the planned navigation action if the determined next-state lateral distance is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle.
[0011] In one embodiment, a system for navigating a host vehicle near a crosswalk is disclosed. The system may include at least one processing device programmed to: receive at least one image representing an environment of the host vehicle from an image capture device; detect a representation of a crosswalk in the at least one image based on an analysis of the at least one image; determine whether a representation of a pedestrian appears in the at least one image based on the analysis of the at least one image; detect the presence of a traffic light in the environment of the host vehicle; determine whether the detected traffic light is associated with the host vehicle and the crosswalk; determine a state of the detected traffic light; when a representation of a pedestrian appears in the at least one image, determine a proximity of the pedestrian to the detected crosswalk; determine a planned navigation action based on at least one driving strategy for navigating the host vehicle relative to the detected crosswalk, wherein the determination of the planned navigation action is further based on the determined state of the detected traffic light and the determined proximity of the pedestrian to the detected crosswalk; and cause one or more actuator systems of the host vehicle to implement the planned navigation action.
[0012] In one embodiment, a method for navigating a host vehicle near a crosswalk is disclosed. The method may include: receiving at least one image representing an environment of the host vehicle from an image capture device; detecting a representation of a crosswalk in the at least one image based on an analysis of the at least one image; determining whether a representation of a pedestrian appears in the at least one image based on the analysis of the at least one image; detecting the presence of a traffic light in the environment of the host vehicle; determining whether the detected traffic light is associated with the host vehicle and the crosswalk; determining a state of the detected traffic light; determining a proximity of the pedestrian relative to the detected crosswalk when the representation of the pedestrian appears in the at least one image; determining a planned navigation action based on at least one driving strategy for navigating the host vehicle relative to the detected crosswalk, wherein the determination of the planned navigation action is further based on the determined state of the detected traffic light and the determined proximity of the pedestrian relative to the detected crosswalk; and causing one or more actuator systems of the host vehicle to implement the planned navigation action.
[0013] In one embodiment, a method for navigating a host vehicle near a crosswalk is disclosed. The method may include: receiving at least one image representing an environment of the host vehicle from an image capture device; detecting a starting location and an ending location of the crosswalk; determining, based on an analysis of the at least one image, whether a pedestrian is near the crosswalk; detecting the presence of a traffic light in the environment of the host vehicle; determining whether the traffic light is relevant to the host vehicle and the crosswalk; determining a state of the traffic light; determining a navigation action for the host vehicle near the detected crosswalk based on: the relevance of the traffic light; the determined state of the traffic light; the presence of a pedestrian near the crosswalk; a selected shortest distance between the pedestrian and one of the starting location or the ending location of the crosswalk; and a motion vector of the pedestrian; and causing one or more actuator systems of the host vehicle to implement the navigation action.
[0014] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions that are executable by at least one processing device to perform any of the steps and / or methods described herein.
[0015] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments.
[0017] In the attached figure:
[0018] Figure 1 is a pictorial representation of an example system consistent with the disclosed embodiments.
[0019] Figure 2Ais a diagrammatic side view representation of an example vehicle including systems consistent with the disclosed embodiments.
[0020] Figure 2B is consistent with the disclosed embodiments Figure 2A Diagrammatic top view representation of the vehicle and systems shown in .
[0021] Figure 2C is a diagrammatic top view representation of another embodiment of a vehicle including a system consistent with the disclosed embodiments.
[0022] Figure 2D is a diagrammatic top view representation of yet another embodiment of a vehicle including a system consistent with the disclosed embodiments.
[0023] Figure 2E is a diagrammatic top view representation of yet another embodiment of a vehicle including a system consistent with the disclosed embodiments.
[0024] Figure 2F is a pictorial representation of an example vehicle control system consistent with the disclosed embodiments.
[0025] Figure 3A is a pictorial representation of the interior of a vehicle including a rearview mirror and a user interface for a vehicle imaging system, consistent with the disclosed embodiments.
[0026] Figure 3B is an illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0027] Figure 3C is consistent with the disclosed embodiments Figure 3B Illustrations of the camera installation from different viewing angles are shown in FIG.
[0028] Figure 3D is an illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0029] Figure 4 is an exemplary block diagram of a memory configured to store instructions for performing one or more operations, consistent with the disclosed embodiments.
[0030] Figure 5A is a flow chart illustrating an exemplary process for eliciting one or more navigation responses based on monocular image analysis, consistent with the disclosed embodiments.
[0031] Figure 5B is a flow chart illustrating an exemplary process for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments.
[0032] Figure 5C is a flow chart illustrating an exemplary process for detecting road markings and / or lane geometry information in a set of images, consistent with the disclosed embodiments.
[0033] Figure 5D is a flow chart illustrating an exemplary process for detecting traffic lights in a set of images, consistent with the disclosed embodiments.
[0034] Figure 5E is a flow chart illustrating an example process for eliciting one or more navigation responses based on a vehicle path, consistent with the disclosed embodiments.
[0035] Figure 5F is a flow chart illustrating an exemplary process for determining whether a leading vehicle is changing lanes, consistent with the disclosed embodiments.
[0036] Figure 6 is a flow chart illustrating an exemplary process for eliciting one or more navigation responses based on stereo image analysis, consistent with the disclosed embodiments.
[0037] Figure 7 is a flow chart illustrating an exemplary process for eliciting one or more navigational responses based on analysis of three sets of images, consistent with the disclosed embodiments.
[0038] Figure 8 is a block diagram representation of modules that may be implemented by one or more specially programmed processing devices of a navigation system of an autonomous vehicle, consistent with the disclosed embodiments.
[0039] Figure 9 is a diagram of navigation options consistent with the disclosed embodiments.
[0040] Figure 10 is a diagram of navigation options consistent with the disclosed embodiments.
[0041] Figure 11A 、 Figure 11B and Figure 11C A schematic representation of navigation options for a host vehicle in a merge zone consistent with the disclosed embodiments is provided.
[0042] Figure 11D An illustrative description of a dual-lane merging scenario consistent with the disclosed embodiments is provided.
[0043] Figure 11E A diagram is provided of options that are potentially useful in a double-merge scenario consistent with the disclosed embodiments.
[0044] Figure 12A diagram is provided that captures a representative image of a host vehicle's environment and potential navigation constraints consistent with the disclosed embodiments.
[0045] Figure 13 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.
[0046] Figure 14 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.
[0047] Figure 15 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.
[0048] Figure 16 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.
[0049] Figure 17A and 17B An illustrative illustration of a host vehicle navigating into a roundabout is provided, consistent with the disclosed embodiments.
[0050] Figure 18 A flow chart of an algorithm for navigating a vehicle consistent with the disclosed embodiments is provided.
[0051] Figure 19 An example of a host vehicle traveling on a multi-lane highway is shown, consistent with the disclosed embodiments.
[0052] Figure 20A and 20B An example of a vehicle cutting in front of another vehicle is shown, consistent with the disclosed embodiments.
[0053] Figure 21 An example of a vehicle following another vehicle is shown, consistent with the disclosed embodiments.
[0054] Figure 22 An example of a vehicle exiting a parking lot and merging onto a potentially busy road is shown, consistent with the disclosed embodiments.
[0055] Figure 23 A vehicle is shown traveling on a road, consistent with the disclosed embodiments.
[0056] Figures 24A-24D Four example scenarios consistent with the disclosed embodiments are shown.
[0057] Figure 25 An example scenario consistent with the disclosed embodiments is shown.
[0058] Figure 26An example scenario consistent with the disclosed embodiments is shown.
[0059] Figure 27 An example scenario consistent with the disclosed embodiments is shown.
[0060] Figure 28A and 28B An example of a scenario in which a vehicle is following another vehicle is shown, consistent with the disclosed embodiments.
[0061] Figure 29A and 29B Example attribution of blame in a cut-through scenario consistent with disclosed embodiments is shown.
[0062] Figure 30A and 30B Example attribution in a cut-through scenario consistent with the disclosed embodiments is shown.
[0063] Figures 31A-31D Example attribution in a drift scenario consistent with the disclosed embodiments is shown.
[0064] Figure 32A and 32B Example attribution in a two-way traffic scenario consistent with the disclosed embodiments is shown.
[0065] Figure 33A and 33B Example attribution in a two-way traffic scenario consistent with the disclosed embodiments is shown.
[0066] Figure 34A and 34B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0067] Figure 35A and 35B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0068] Figure 36A and 36B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0069] Figure 37A and 37B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0070] Figure 38A and 38B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0071] Figure 39A and39B Example attribution in a route priority scenario consistent with the disclosed embodiments is shown.
[0072] Figure 40A and 40B Example attribution in a traffic light scenario consistent with disclosed embodiments is shown.
[0073] Figure 41A and 41B Example attribution in a traffic light scenario consistent with disclosed embodiments is shown.
[0074] Figure 42A and 42B Example attribution in a traffic light scenario consistent with disclosed embodiments is shown.
[0075] Figures 43A-43C An example vulnerable road user (VRU) scenario consistent with the disclosed embodiments is shown.
[0076] Figures 44A-44C An example vulnerable road user (VRU) scenario consistent with the disclosed embodiments is shown.
[0077] Figures 45A-45C An example vulnerable road user (VRU) scenario consistent with the disclosed embodiments is shown.
[0078] Figures 46A-46D An example vulnerable road user (VRU) scenario consistent with the disclosed embodiments is shown.
[0079] Figure 47A is a graphical illustration of blame time and appropriate responses consistent with disclosed embodiments.
[0080] Figure 47B is a graphical representation of route priorities for routes of different geometries consistent with the disclosed embodiments.
[0081] Figure 47C is an illustration of the longitudinal ordering of routes of different geometries, consistent with the disclosed embodiments.
[0082] Figure 47D is a graphical representation of a safe longitudinal distance between vehicles consistent with the disclosed embodiments.
[0083] Figure 47E is an illustration of a situation in which a vehicle cannot predict the path of another vehicle, consistent with the disclosed embodiments.
[0084] Figure 47F is an illustration of route priorities at a traffic light consistent with the disclosed embodiments.
[0085] Figure 47G is a diagram of an exemplary unstructured route consistent with the disclosed embodiments.
[0086] Figure 47H is an illustration of exemplary lateral behavior in an unstructured situation consistent with the disclosed embodiments.
[0087] Figure 47I is a graphical representation of exposure time and attribution time for an occluded region consistent with disclosed embodiments.
[0088] Figure 48A An example scenario of two vehicles traveling in opposite directions is shown, consistent with the disclosed embodiments.
[0089] Figure 48B An example is shown of a target vehicle traveling toward a host vehicle, consistent with the disclosed embodiments.
[0090] Figure 49 An example of a host vehicle maintaining a safe longitudinal distance is shown, consistent with the disclosed embodiments.
[0091] Figure 50A and 50B A flowchart depicting an exemplary process for maintaining a safe longitudinal distance consistent with the disclosed embodiments is provided.
[0092] Figure 51A An example scenario is shown of two vehicles laterally spaced apart from each other, consistent with the disclosed embodiments.
[0093] Figure 51B An example of a host vehicle maintaining a safe lateral distance is shown consistent with the disclosed embodiments.
[0094] Figure 52A An example of a host vehicle performing a planned navigation maneuver is shown consistent with the disclosed embodiments.
[0095] Figure 52B An example is shown in which a host vehicle determines whether to perform a navigation action.
[0096] Figure 53A and 53B A flowchart depicting an example process for maintaining a safe lateral distance consistent with the disclosed embodiments is provided.
[0097] Figure 54 is a schematic illustration of a roadway including a crosswalk, consistent with disclosed embodiments.
[0098] Figure 55A and 55Bis a schematic illustration of possible navigation actions performed by a vehicle traveling along a roadway, consistent with the disclosed embodiments.
[0099] Figure 55C and 55D An example of estimating the distance from a pedestrian to a crosswalk is shown, consistent with the disclosed embodiments.
[0100] Figure 56 is a flow chart describing a process for navigating a vehicle near a crosswalk, consistent with the disclosed embodiments. DETAILED DESCRIPTION
[0101] The following detailed description refers to the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. Although several illustrative embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, replacements, additions, or modifications may be made to the components shown in the drawings, and the illustrative methods described herein may be modified by replacing, reordering, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the proper scope is defined by the appended claims.
[0102] Autonomous Vehicle Overview
[0103] As used throughout this disclosure, the term "autonomous vehicle" refers to a vehicle that is capable of implementing at least one navigation change without driver input. A "navigation change" refers to one or more of the vehicle's steering, braking, or acceleration / deceleration. By autonomous, it is meant that the vehicle need not be fully automatic (e.g., fully operable without a driver or without driver input). Instead, autonomous vehicles include those that are capable of operating under driver control during certain time periods and are capable of operating without driver control during other time periods. Autonomous vehicles may also include vehicles that only control certain aspects of vehicle navigation, such as steering (e.g., maintaining the vehicle's line between vehicle lane constraints) or certain steering operations in certain circumstances (but not all circumstances), but may leave other aspects to the driver (e.g., braking or braking in certain circumstances). In some cases, an autonomous vehicle may handle some or all aspects of the vehicle's braking, speed control, and / or steering.
[0104] Because human drivers often rely on visual cues and observation to control vehicles, transportation infrastructure is built with lane markings, traffic signs, and traffic lights designed to provide visual information to drivers. Given these design characteristics of transportation infrastructure, autonomous vehicles can include cameras and processing units that analyze visual information captured from the vehicle's environment. This visual information can include, for example, images representing components of the transportation infrastructure observable by the driver (e.g., lane markings, traffic signs, traffic lights, etc.) as well as other obstacles (e.g., other vehicles, pedestrians, debris, etc.). Autonomous vehicles can also use stored information, such as information that provides a model of the vehicle's environment when navigating. For example, a vehicle can use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data to provide information about its environment while the vehicle is traveling, and the vehicle (and other vehicles) can use this information to locate itself on the model. Some vehicles are also capable of communicating with each other, sharing information, and alerting peer vehicles to hazards or changes in the vehicle's surroundings.
[0105] System Overview
[0106] Figure 1 is a block diagram representation of a system 100 consistent with the disclosed example embodiments. Depending on the requirements of a particular implementation, system 100 may include various components. In some embodiments, system 100 may include a processing unit 110, an image acquisition unit 120, a location sensor 130, one or more memory units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. Processing unit 110 may include one or more processing devices. In some embodiments, processing unit 110 may include an application processor 180, an image processor 190, or any other suitable processing device. Similarly, image acquisition unit 120 may include any number of image acquisition devices and components, depending on the requirements of a particular application. In some embodiments, image acquisition unit 120 may include one or more image capture devices (e.g., a camera, a CCD, or any other type of image sensor), such as image capture device 122, image capture device 124, and image capture device 126. System 100 may also include a data interface 128 that communicatively connects processing unit 110 to image acquisition unit 120. For example, the data interface 128 may include one or more any wired and / or wireless links for transmitting image data acquired by the image acquisition unit 120 to the processing unit 110 .
[0107] The wireless transceiver 172 may include one or more devices configured to exchange transmissions over an air interface to one or more networks (e.g., cellular, the Internet, etc.) using radio frequencies, infrared frequencies, magnetic fields, or electric fields. The wireless transceiver 172 may use any known standard to transmit and / or receive data (e.g., Wi-Fi, Bluetooth®, Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions may include communications from the host vehicle to one or more remotely located servers. Such transmissions may also include (one-way or two-way) communications between the host vehicle and one or more target vehicles in the host vehicle's environment (e.g., to facilitate navigation of the host vehicle with or in conjunction with target vehicles in the host vehicle's environment), or even broadcast transmissions to unspecified recipients in the vicinity of the transmitting vehicle.
[0108] Both application processor 180 and image processor 190 may include various types of hardware-based processing devices. For example, either or both application processor 180 and image processor 190 may include a microprocessor, a preprocessor (such as an image preprocessor), a graphics processor, a central processing unit (CPU), support circuits, a digital signal processor, an integrated circuit, memory, or any other type of device suitable for running applications and for image processing and analysis. In some embodiments, application processor 180 and / or image processor 190 may include any type of single-core or multi-core processor, a mobile device microcontroller, a central processing unit, or the like. A variety of processing devices may be used, including, for example, processors available from manufacturers such as Intel®, AMD®, and the like, and may include various architectures (e.g., x86 processors, ARM®, etc.).
[0109] In some embodiments, the application processor 180 and / or the image processor 190 may include any of the EyeQ series processor chips available from Mobileye®. These processor designs all include multiple processing units with local memory and an instruction set. Such a processor may include a video input for receiving image data from multiple image sensors and may also include video output capabilities. In one example, the EyeQ2® uses 90nm-micron technology operating at 332 MHz. The EyeQ2® architecture consists of two floating-point hyperthreaded 32-bit RISC CPUs (MIPS32®34K® cores), five Vision Compute Engines (VCEs), three Vector Microcode Processors (VMP®), a Denali 64-bit mobile DDR controller, a 128-bit internal Sonics Interconnect, dual 16-bit video input and 18-bit video output controllers, a 16-channel DMA, and several peripherals. The MIPS34K CPU manages the five VCEs, three VMPsTM and DMA, a second MIPS34K CPU and multi-channel DMA, and other peripherals. Five VCEs, three VMPs®, and a MIPS34K CPU can perform the intensive visual computations required for multi-function bundled applications. In another example, EyeQ3®, a third-generation processor six times more powerful than EyeQ2®, can be used in the disclosed embodiments. In other examples, EyeQ4® and / or EyeQ5® can be used in the disclosed embodiments. Of course, any newer or future EyeQ processing device can also be used with the disclosed embodiments.
[0110] Any of the processing devices disclosed herein can be configured to perform certain functions. Configuring a processing device (such as any described EyeQ processor or other controller or microprocessor) to perform certain functions can include programming computer-executable instructions and making those instructions available to the processing device for execution during operation of the processing device. In some embodiments, configuring the processing device can include programming the processing device directly with architectural instructions. In other embodiments, configuring the processing device can include storing executable instructions on a memory accessible to the processing device during operation. For example, the processing device can access the memory during operation to obtain and execute the stored instructions. In either case, a processing device configured to perform the sensing, image analysis, and / or navigation functions disclosed herein represents a dedicated hardware-based system in the control of multiple hardware-based components of a host vehicle.
[0111] although Figure 1 Two separate processing devices are depicted as being included in processing unit 110, but more or fewer processing devices may be used. For example, in some embodiments, a single processing device may be used to perform the tasks of application processor 180 and image processor 190. In other embodiments, these tasks may be performed by more than two processing devices. Additionally, in some embodiments, system 100 may include one or more processing units 110 without including other components such as image acquisition unit 120.
[0112] The processing unit 110 may include various types of devices. For example, the processing unit 110 may include various devices such as a controller, an image preprocessor, a central processing unit (CPU), support circuits, a digital signal processor, an integrated circuit, memory, or any other type of device used for image processing and analysis. The image preprocessor may include a video processor for capturing, digitizing, and processing images from the image sensor. The CPU may include any number of microcontrollers or microprocessors. The support circuits may include any number of circuits known in the art, including caches, power supplies, clocks, and input / output circuits. The memory may store software that, when executed by the processor, controls the operation of the system. The memory may include a database and image processing software. The memory may include any number of random access memories, read-only memories, flash memories, disk drives, optical storage devices, tape storage devices, removable storage devices, and other types of storage devices. In one embodiment, the memory may be separate from the processing unit 110. In another embodiment, the memory may be integrated into the processing unit 110.
[0113] Each memory 140, 150 may include software instructions that, when executed by a processor (e.g., application processor 180 and / or image processor 190), may control the operation of various aspects of system 100. For example, these memory units may include various databases and image processing software, as well as trained systems such as neural networks and deep neural networks. The memory units may include random access memory, read-only memory, flash memory, disk drives, optical storage devices, tape storage devices, removable storage devices, and / or any other type of storage device. In some embodiments, the memory units 140, 150 may be separate from the application processor 180 and / or image processor 190. In other embodiments, these memory units may be integrated into the application processor 180 and / or image processor 190.
[0114] Position sensor 130 may include any type of device suitable for determining a location associated with at least one component of system 100. In some embodiments, position sensor 130 may include a GPS receiver. Such a receiver may determine user location and velocity by processing signals broadcast by global positioning system satellites. Position information from position sensor 130 may be made available to application processor 180 and / or image processor 190.
[0115] In some embodiments, system 100 may include components such as a speed sensor (e.g., a speedometer) for measuring the speed of vehicle 200. System 100 may also include one or more accelerometers (single-axis or multi-axis) for measuring the acceleration of vehicle 200 along one or more axes.
[0116] The memory units 140, 150 may include a database, or data organized in any other form, indicating the locations of known landmarks. Sensor information of the environment (such as images from lidar or stereo processing of two or more images, radar signals, depth information) may be processed together with position information (such as GPS coordinates, the vehicle's ego motion, etc.) to determine the vehicle's current position relative to known landmarks and refine the vehicle's position. Certain aspects of this technology are included in what is known as REM. TM The positioning technology is sold by the assignee of this application.
[0117] The user interface 170 may include any device suitable for providing information to or receiving input from one or more users of the system 100. In some embodiments, the user interface 170 may include a user input device including, for example, a touch screen, a microphone, a keyboard, a pointing device, a tracking wheel, a camera, knobs, buttons, etc. Using such input devices, a user can provide information input or commands to the system 100 by typing instructions or information, providing voice commands, selecting on-screen menu options using buttons, a pointer, or eye tracking capabilities, or by any other suitable technique for communicating information to the system 100.
[0118] The user interface 170 may be equipped with one or more processing devices configured to provide and receive information to and from a user and to process that information for use by, for example, the application processor 180. In some embodiments, such processing devices may execute instructions to recognize and track eye movements, receive and interpret voice commands, recognize and interpret touches and / or gestures made on a touch screen, respond to keyboard input or menu selections, etc. In some embodiments, the user interface 170 may include a display, a speaker, a haptic device, and / or any other device for providing output information to a user.
[0119] Map database 160 may comprise any type of database for storing map data useful to system 100. In some embodiments, map database 160 may include data relating to the locations of various items within a reference coordinate system, including roads, water features, geographic features, businesses, points of interest, restaurants, gas stations, and the like. Map database 160 may store not only the locations of these items but also descriptors associated with these items, including, for example, names associated with any stored features. In some embodiments, map database 160 may be physically located with other components of system 100. Alternatively or additionally, map database 160 or portions thereof may be remotely located relative to other components of system 100 (e.g., processing unit 110). In such embodiments, information from map database 160 may be downloaded to a network via a wired or wireless data connection (e.g., via a cellular network and / or the Internet, etc.). In some cases, map database 160 may store a sparse data model comprising a polynomial representation of certain road features (e.g., lane markings) or a target trajectory of the host vehicle. The map database 160 may also include stored representations of various recognized landmarks that may be used to determine or update the known position of the host vehicle relative to the target trajectory. The landmark representation may include data fields such as landmark type, landmark location, among other potential identifiers.
[0120] Image capture devices 122, 124, and 126 may each comprise any type of device suitable for capturing at least one image from an environment. Furthermore, any number of image capture devices may be used to acquire images for input to an image processor. Some embodiments may include only a single image capture device, while other embodiments may include two, three, or even four, or more image capture devices. Figures 2B to 2E Image capture devices 122, 124, and 126 are further described.
[0121] One or more cameras (e.g., image capture devices 122, 124, and 126) may be part of a sensing block included on the vehicle. The sensing block may include various other sensors, and any or all of these sensors may be relied upon to develop a sensed navigational state for the vehicle. In addition to cameras (front, side, rear, etc.), other sensors (such as RADAR, LIDAR, and acoustic sensors) may be included in the sensing block. Furthermore, the sensing block may include one or more components configured to transmit and / or receive information related to the vehicle's environment. For example, such components may include wireless transceivers (RF, etc.) that can receive sensor-based information or any other type of information related to the host vehicle's environment from a source remotely located relative to the host vehicle. This information may include sensor output information or related information received from vehicle systems other than the host vehicle. In some embodiments, this information may include information received from a remote computing device, a central server, or the like. Furthermore, cameras may be configured in many different ways: a single camera unit, multiple cameras, a camera cluster, with a long field of view (FOV), a short FOV, a wide angle, a fisheye, and the like.
[0122] The system 100 or its various components may be incorporated into a variety of different platforms. In some embodiments, the system 100 may be included on a vehicle 200, such as Figure 2A For example, the vehicle 200 may be equipped with Figure 1 While in some embodiments the vehicle 200 may be equipped with only a single image capture device (e.g., a camera), in other embodiments, such as a combination of Figures 2B to 2E For those embodiments discussed, multiple image capture devices may be used. For example, Figure 2A Either of the image capture devices 122 and 124 of the vehicle 200 shown in FIG. 2 may be part of an ADAS (Advanced Driver Assistance System) imaging set.
[0123] The image capture device included on the vehicle 200 as part of the image acquisition unit 120 may be located in any suitable location. Figures 2A to 2E ,as well as Figures 3A to 3C As shown in , the image capture device 122 can be located near the rearview mirror. This location can provide a line of sight similar to the line of sight of the driver of the vehicle 200, which can assist in determining what is visible and invisible to the driver. The image capture device 122 can be placed anywhere near the rearview mirror, and placing the image capture device 122 on the driver's side of the mirror can also assist in obtaining an image representative of the driver's field of view and / or line of sight.
[0124] Other locations of the image capture device of image acquisition unit 120 can also be used. For example, image capture device 124 can be located on or in the bumper of vehicle 200. This location is particularly suitable for image capture devices with a wide field of view. The line of sight of the image capture device located in the bumper may be different from the driver's line of sight, and therefore, the bumper image capture device and the driver may not always see the same object. Image capture devices (e.g., image capture devices 122, 124, and 126) can also be located in other locations. For example, the image capture device can be located on or in one or both of the side mirrors of vehicle 200, on the roof of vehicle 200, on the hood of vehicle 200, on the trunk of vehicle 200, on the side of vehicle 200, mounted on any window of vehicle 200, located behind any window of vehicle 200, located in front of any window of vehicle 200, and in or near lighting devices installed on the front and / or rear of vehicle 200, etc.
[0125] In addition to the image capture device, the vehicle 200 may also include various other components of the system 100. For example, the processing unit 110 may be included on the vehicle 200, either integrated with or separate from the vehicle's engine control unit (ECU). The vehicle 200 may also be equipped with a location sensor 130, such as a GPS receiver, and may also include a map database 160 and memory units 140 and 150.
[0126] As discussed earlier, the wireless transceiver 172 can transmit and / or receive data over one or more networks (e.g., a cellular network, the Internet, etc.). For example, the wireless transceiver 172 can upload data collected by the system 100 to one or more servers and download data from one or more servers. Via the wireless transceiver 172, the system 100 can receive, for example, periodic or on-demand updates to data stored in the map database 160, the memory 140, and / or the storage 150. Similarly, the wireless transceiver 172 can upload any data from the system 100 (e.g., images captured by the image acquisition unit 120, data received by the position sensor 130 or other sensors, the vehicle control system, etc.) and / or any data processed by the processing unit 110 to one or more servers.
[0127] The system 100 can upload data to a server (e.g., to the cloud) based on a privacy level setting. For example, the system 100 can implement a privacy level setting to specify or limit the type of data (including metadata) sent to the server that can uniquely identify the vehicle and / or the driver / owner of the vehicle. Such settings can be set by a user via, for example, the wireless transceiver 172, can be set by a factory default setting, or can be initialized by data received by the wireless transceiver 172.
[0128] In some embodiments, the system 100 may upload data according to a "high" privacy level, and if the setting is set, the system 100 may transmit data (e.g., location information related to the route, captured images, etc.) without any details about the specific vehicle and / or driver / owner. For example, when uploading data according to the "high" privacy setting, the system 100 may not include the vehicle identification number (VIN) or the name of the driver or owner of the vehicle, and may instead transmit data (such as captured images and / or limited location information related to the route).
[0129] Other privacy levels may also be considered. For example, the system 100 may transmit data to the server according to a "medium" privacy level and may include additional information that is not included at a "high" privacy level, such as the make and / or model of the vehicle and / or the type of vehicle (e.g., passenger vehicle, sport utility vehicle, truck, etc.). In some embodiments, the system 100 may upload data according to a "low" privacy level. At the "low" privacy level setting, the system 100 may upload data and include information sufficient to uniquely identify a particular vehicle, owner / driver, and / or a portion or the entire route traveled by the vehicle. For example, such "low" privacy level data may include one or more of the following: VIN, driver / owner name, the vehicle's point of origin prior to departure, the vehicle's desired destination, the vehicle's make and / or model, the vehicle type, etc.
[0130] Figure 2A is a diagrammatic side view representation of an exemplary vehicle imaging system consistent with the disclosed embodiments. Figure 2B yes Figure 2A A diagrammatic top view illustration of the embodiment shown in FIG. Figure 2B As shown, the disclosed embodiments may include a vehicle 200 including a system 100 in its body with a first image capture device 122 located near a rearview mirror of the vehicle 200 and / or near a driver, a second image capture device 124 located on or in a bumper area (e.g., one of the bumper areas 210) of the vehicle 200, and a processing unit 110.
[0131] like Figure 2C As shown, both image capture devices 122 and 124 may be located near the rearview mirror of vehicle 200 and / or near the driver. Figure 2B and Figure 2C Two image capture devices 122 and 124 are shown, it being understood that other embodiments may include more than two image capture devices. Figure 2D and Figure 2E In the embodiment shown in , a first image capture device 122 , a second image capture device 124 , and a third image capture device 126 are included in the system 100 of a vehicle 200 .
[0132] like Figure 2D As shown, image capture device 122 may be located near a rearview mirror of vehicle 200 and / or near the driver, and image capture devices 124 and 126 may be located on or in a bumper area (e.g., one of bumper areas 210) of vehicle 200. Figure 2E As shown, image capture devices 122, 124, and 126 may be located near rearview mirrors and / or near the driver's seat of vehicle 200. The disclosed embodiments are not limited to any particular number and configuration of image capture devices, and the image capture devices may be located in any suitable location within and / or on vehicle 200.
[0133] It should be understood that the disclosed embodiments are not limited to vehicles and can be applied in other scenarios. It should also be understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and can be applicable to all types of vehicles, including cars, trucks, trailers, and other types of vehicles.
[0134] The first image capture device 122 may include any suitable type of image capture device. The image capture device 122 may include an optical axis. In one example, the image capture device 122 may include an Aptina M9V024W VGA sensor with a global shutter. In other embodiments, the image capture device 122 may provide a resolution of 1280×960 pixels and may include a rolling shutter. The image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included, for example, to provide a desired focal length and field of view for the image capture device. In some embodiments, the image capture device 122 may be associated with a 6 mm lens or a 12 mm lens. In some embodiments, as Figure 2DAs shown, image capture device 122 can be configured to capture images with a desired field of view (FOV) 202. For example, image capture device 122 can be configured to have a conventional FOV, such as in the range of 40 to 56 degrees, including a 46-degree FOV, a 50-degree FOV, a 52-degree FOV, or a larger FOV. Alternatively, image capture device 122 can be configured to have a narrow FOV in the range of 23 to 40 degrees, such as a 28-degree FOV or a 36-degree FOV. Furthermore, image capture device 122 can be configured to have a wide FOV in the range of 100 to 180 degrees. In some embodiments, image capture device 122 can include a wide-angle bumper camera or a camera with a FOV of up to 180 degrees. In some embodiments, image capture device 122 can be a 7.2 megapixel image capture device with an aspect ratio of approximately 2:1 (e.g., H×V=3800×1900 pixels) and a horizontal FOV of approximately 100 degrees. Such an image capture device can be used in place of a three-image capture device configuration. Due to significant lens distortion, in embodiments where the image capture device uses a radially symmetric lens, the vertical FOV of such an image capture device may be significantly less than 50 degrees. For example, such a lens may not be radially symmetric, which would allow a vertical FOV greater than 50 degrees with a 100 degree horizontal FOV.
[0135] The first image capture device 122 can acquire a plurality of first images of a scene associated with the vehicle 200. Each of the plurality of first images can be acquired as a series of image scan lines, which can be captured using a rolling shutter. Each scan line can include a plurality of pixels.
[0136] The first image capture device 122 may have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate may refer to the rate at which the image sensor may acquire image data associated with each pixel included in a particular scan line.
[0137] Image capture devices 122, 124, and 126 may include any suitable type and number of image sensors, including, for example, CCD sensors or CMOS sensors. In one embodiment, a CMOS image sensor and a rolling shutter may be employed such that each pixel in a row is read one at a time, and scanning of the rows continues on a row-by-row basis until the entire image frame has been captured. In some embodiments, the rows may be captured sequentially from top to bottom relative to the frame.
[0138] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124, and 126) may constitute a high-resolution imager and may have a resolution greater than 5 M pixels, 7 M pixels, 10 M pixels, or more.
[0139] The use of a rolling shutter may cause pixels in different rows to be exposed and captured at different times, which may cause distortions and other image artifacts in the captured image frame. On the other hand, when image capture device 122 is configured to operate with a global or synchronized shutter, all pixels can be exposed for the same amount of time and during a common exposure period. As a result, the image data in a frame collected from a system employing a global shutter represents a snapshot of the entire FOV (such as FOV 202) at a particular time. In contrast, in a rolling shutter application, each row in the frame is exposed and data is captured at a different time. As a result, moving objects may appear distorted in an image capture device with a rolling shutter. This phenomenon will be described in more detail below.
[0140] Second image capture device 124 and third image capture device 126 can be any type of image capture device. Similar to first image capture device 122, each of image capture devices 124 and 126 can include an optical axis. In one embodiment, each of image capture devices 124 and 126 can include an Aptina M9V024 WVGA sensor with a global shutter. Alternatively, each of image capture devices 124 and 126 can include a rolling shutter. Similar to image capture device 122, image capture devices 124 and 126 can be configured to include various lenses and optical elements. In some embodiments, the lenses associated with image capture devices 124 and 126 can provide a FOV (such as FOVs 204 and 206) that is equal to or narrower than the FOV associated with image capture device 122 (such as FOV 202). For example, image capture devices 124 and 126 can have a FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less.
[0141] Image capture devices 124 and 126 can capture a plurality of second and third images of a scene associated with vehicle 200. Each of the plurality of second and third images can be captured as a second series of image scan lines and a third series of image scan lines, which can be captured using a rolling shutter. Each scan line or row can have a plurality of pixels. Image capture devices 124 and 126 can have a second scan rate and a third scan rate associated with the capture of each image scan line included in the second and third series.
[0142] Each image capture device 122, 124, and 126 may be placed at any suitable location and orientation relative to vehicle 200. The relative positions of image capture devices 122, 124, and 126 may be selected to facilitate fusing information obtained from the image capture devices. For example, in some embodiments, the FOV associated with image capture device 124 (such as FOV 204) may partially or completely overlap with the FOV associated with image capture device 122 (such as FOV 202) and the FOV associated with image capture device 126 (such as FOV 206).
[0143] Image capture devices 122, 124, and 126 may be located at any suitable relative heights on vehicle 200. In one example, there may be height differences between image capture devices 122, 124, and 126 that may provide sufficient parallax information to enable stereo analysis. Figure 2A As shown, the two image capture devices 122 and 124 are at different heights. For example, there may also be lateral displacement differences between the image capture devices 122, 124, and 126, providing additional parallax information for the stereo analysis of the processing unit 110. Figure 2C and Figure 2D As shown, the difference in lateral displacement can be expressed by d x In some embodiments, there may be a forward or backward displacement (e.g., a range displacement) between image capture devices 122, 124, and 126. For example, image capture device 122 may be positioned 0.5 to 2 meters or more behind image capture device 124 and / or image capture device 126. This type of displacement may enable one of the image capture devices to cover a potential blind spot of the other(s).
[0144] Image capture device 122 may have any suitable resolution capability (e.g., the number of pixels associated with the image sensor), and the resolution of the image sensor(s) associated with image capture device 122 may be higher, lower, or the same as the resolution of the image sensor(s) associated with image capture devices 124 and 126. In some embodiments, the image sensor(s) associated with image capture device 122 and / or image capture devices 124 and 126 may have a resolution of 640×480, 1024×768, 1280×960, or any other suitable resolution.
[0145] The frame rate (e.g., the rate at which an image capture device acquires a set of pixel data for one image frame before continuing to capture pixel data associated with the next image frame) can be controllable. The frame rate associated with image capture device 122 can be higher, lower, or the same as the frame rates associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 can depend on various factors that may affect the timing of the frame rates. For example, one or more of image capture devices 122, 124, and 126 can include a selectable pixel delay period that is applied before or after acquiring image data associated with one or more pixels of the image sensors in image capture devices 122, 124, and / or 126. Typically, image data corresponding to each pixel can be acquired based on the clock rate used for the device (e.g., one pixel per clock cycle). Additionally, in embodiments including a rolling shutter, one or more of image capture devices 122, 124, and 126 may include a selectable horizontal blanking period that is applied before or after acquiring image data associated with a row of pixels of an image sensor in image capture devices 122, 124, and / or 126. Additionally, one or more of image capture devices 122, 124, and 126 may include a selectable vertical blanking period that is applied before or after acquiring image data associated with an image frame of image capture devices 122, 124, and 126.
[0146] These timing controls may enable synchronization of the frame rates associated with image capture devices 122, 124, and 126, even if the line scan rate of each is different. Furthermore, as will be discussed in greater detail below, these selectable timing controls may enable synchronization of image capture from areas where the FOV of image capture device 122 overlaps with one or more of the FOVs of image capture devices 124 and 126, even if the field of view of image capture device 122 is different from the FOVs of image capture devices 124 and 126, in addition to other factors (e.g., image sensor resolution, maximum line scan rate, etc.).
[0147] The frame rate timing in image capture devices 122, 124, and 126 may depend on the resolution of the associated image sensors. For example, assuming similar line scan rates for both devices, if one device includes an image sensor with a resolution of 640×480 and the other device includes an image sensor with a resolution of 1280×960, more time will be required to acquire a frame of image data from the sensor with the higher resolution.
[0148] Another factor that may affect the timing of image data acquisition in image capture devices 122, 124, and 126 is the maximum line scan rate. For example, it will take a certain minimum amount of time to acquire a line of image data from the image sensors included in image capture devices 122, 124, and 126. Assuming no pixel delay period is added, this minimum amount of time for acquiring a line of image data will be related to the maximum line scan rate for the particular device. Devices that offer higher maximum line scan rates have the potential to provide higher frame rates than devices with lower maximum line scan rates. In some embodiments, one or more of image capture devices 124 and 126 may have a maximum line scan rate that is higher than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture devices 124 and / or 126 may be 1.25, 1.5, 1.75, or 2 times, or more, the maximum line scan rate of image capture device 122.
[0149] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may operate at a scan rate less than or equal to its maximum scan rate. The system may be configured such that one or more of image capture devices 124 and 126 operate at a line scan rate equal to the line scan rate of image capture device 122. In other examples, the system may be configured such that the line scan rate of image capture device 124 and / or image capture device 126 may be 1.25, 1.5, 1.75, or 2 or more times the line scan rate of image capture device 122.
[0150] In some embodiments, image capture devices 122, 124, and 126 may be asymmetrical. That is, they may include cameras with different fields of view (FOVs) and focal lengths. For example, the fields of view of image capture devices 122, 124, and 126 may include any desired area of the environment surrounding vehicle 200. In some embodiments, one or more of image capture devices 122, 124, and 126 may be configured to acquire image data from the environment in front of vehicle 200, behind vehicle 200, to the sides of vehicle 200, or a combination thereof.
[0151] Furthermore, the focal length associated with each image capture device 122, 124, and / or 126 can be selectable (e.g., by including an appropriate lens, etc.) so that each device captures images of objects at a desired range of distances relative to the vehicle 200. For example, in some embodiments, the image capture devices 122, 124, and 126 can capture images of objects that are within a few meters of the vehicle. The image capture devices 122, 124, and 126 can also be configured to capture images of objects at greater ranges from the vehicle (e.g., 25 m, 50 m, 100 m, 150 m, or more). In addition, the focal lengths of image capture devices 122, 124, and 126 may be selected so that one image capture device (e.g., image capture device 122) can capture images of objects that are relatively close to the vehicle (e.g., within 10 m or within 20 m), while other image capture devices (e.g., image capture devices 124 and 126) can capture images of objects that are farther away from vehicle 200 (e.g., greater than 20 m, 50 m, 100 m, 150 m, etc.).
[0152] According to some embodiments, the FOV of one or more image capture devices 122, 124, and 126 may have a wide angle. For example, having a FOV of 140 degrees may be advantageous, particularly for image capture devices 122, 124, and 126 that may be used to capture images of areas near vehicle 200. For example, image capture device 122 may be used to capture images of areas to the right or left of vehicle 200, and in these embodiments, it may be desirable for image capture device 122 to have a wide FOV (e.g., at least 140 degrees).
[0153] The field of view associated with each of image capture devices 122, 124, and 126 may depend on the respective focal lengths. For example, as the focal length increases, the corresponding field of view decreases.
[0154] Image capture devices 122, 124, and 126 can be configured to have any suitable field of view. In one specific example, image capture device 122 can have a horizontal FOV of 46 degrees, image capture device 124 can have a horizontal FOV of 23 degrees, and image capture device 126 can have a horizontal FOV between 23 and 46 degrees. In another example, image capture device 122 can have a horizontal FOV of 52 degrees, image capture device 124 can have a horizontal FOV of 26 degrees, and image capture device 126 can have a horizontal FOV between 26 and 52 degrees. In some embodiments, the ratio of the FOV of image capture device 122 to the FOV of image capture device 124 and / or image capture device 126 can vary from 1.5 to 2.0. In other embodiments, the ratio can vary between 1.25 and 2.25.
[0155] System 100 can be configured such that the field of view of image capture device 122 at least partially or completely overlaps the field of view of image capture device 124 and / or image capture device 126. In some embodiments, system 100 can be configured such that the fields of view of image capture devices 124 and 126, for example, fall within (e.g., are narrower than) the field of view of image capture device 122 and share a common center with the field of view of image capture device 122. In other embodiments, image capture devices 122, 124, and 126 can capture adjacent FOVs or can have partial overlap in their FOVs. In some embodiments, the fields of view of image capture devices 122, 124, and 126 can be aligned such that the center of the narrower FOV image capture devices 124 and / or 126 can be located in the lower half of the field of view of the wider FOV device 122.
[0156] Figure 2F is a diagrammatic representation of an exemplary vehicle control system consistent with the disclosed embodiments. Figure 2F As indicated, the vehicle 200 may include a throttle regulation (throtting) system 220, a braking system 230, and a steering system 240. The system 100 may provide input (e.g., control signals) to one or more of the throttle regulation system 220, the braking system 230, and the steering system 240 via one or more data links (e.g., any wired and / or wireless links for transmitting data). For example, based on analysis of images acquired by the image capture devices 122, 124, and / or 126, the system 100 may provide control signals to one or more of the throttle regulation system 220, the braking system 230, and the steering system 240 to navigate the vehicle 200 (e.g., by causing acceleration, turning, lane changes, etc.). In addition, the system 100 may receive input from one or more of the throttle regulation system 220, the braking system 230, and the steering system 240 indicating an operating condition of the vehicle 200 (e.g., speed, whether the vehicle 200 is braking and / or turning, etc.). The following is in conjunction with Figures 4 to 7 Provide further details.
[0157] like Figure 3AAs shown, the vehicle 200 may also include a user interface 170 for interacting with the driver or passengers of the vehicle 200. For example, the user interface 170 in a vehicle application may include a touch screen 320, a knob 330, a button 340, and a microphone 350. The driver or passenger of the vehicle 200 may also interact with the system 100 using a handle (e.g., located on or near the steering column of the vehicle 200, including, for example, a turn signal handle), a button (e.g., located on the steering wheel of the vehicle 200), or the like. In some embodiments, the microphone 350 may be located adjacent to the rearview mirror 310. Similarly, in some embodiments, the image capture device 122 may be located near the rearview mirror 310. In some embodiments, the user interface 170 may also include one or more speakers 360 (e.g., speakers of the vehicle's audio system). For example, the system 100 may provide various notifications (e.g., alarms) via the speakers 360.
[0158] Figures 3B to 3D is an illustration of an example camera mount 370 configured to be located behind a rearview mirror (e.g., rearview mirror 310) and opposite a vehicle windshield, consistent with the disclosed embodiments. Figure 3B As shown, camera mount 370 may include image capture devices 122, 124, and 126. Image capture devices 124 and 126 may be located behind a sun visor 380, wherein sun visor 380 may be flush with the vehicle windshield and include a combination of film and / or anti-reflective material. For example, sun visor 380 may be positioned so that it is aligned with the vehicle windshield having a matching bevel. In some embodiments, each of image capture devices 122, 124, and 126 may be located behind sun visor 380, such as at Figure 3D The disclosed embodiments are not limited to any particular configuration of image capture devices 122 , 124 , and 126 , camera mount 370 , and light shield 380 . Figure 3C yes Figure 3B An illustration of camera mount 370 is shown from a front perspective.
[0159] As will be appreciated by those skilled in the art having the benefit of this disclosure, numerous variations and / or modifications may be made to the aforementioned disclosed embodiments. For example, not all components are necessary for the operation of system 100. Furthermore, any component may be located in any suitable part of system 100 and the components may be rearranged into various configurations while still providing the functionality of the disclosed embodiments. Thus, the aforementioned configurations are exemplary, and regardless of the configurations discussed above, system 100 may provide a wide range of functionality to analyze the surrounding environment of vehicle 200 and navigate vehicle 200 in response to that analysis.
[0160] As discussed in greater detail below and consistent with the various disclosed embodiments, system 100 can provide various features related to autonomous driving and / or driver assistance technologies. For example, system 100 can analyze image data, location data (e.g., GPS location information), map data, speed data, and / or data from sensors included in vehicle 200. System 100 can collect data for analysis from, for example, image acquisition unit 120, location sensor 130, and other sensors. Furthermore, system 100 can analyze the collected data to determine whether vehicle 200 should take a certain action, and then automatically take the determined action without human intervention. For example, while vehicle 200 is navigating without human intervention, system 100 can automatically control braking, acceleration, and / or steering of vehicle 200 (e.g., by sending control signals to one or more of throttle control system 220, braking system 230, and steering system 240). Furthermore, system 100 can analyze the collected data and, based on the analysis of the collected data, issue warnings and / or alerts to vehicle occupants. Additional details regarding various embodiments provided by system 100 are provided below.
[0161] Forward multi-imaging system
[0162] As discussed above, system 100 can provide driver assistance functionality using a multi-camera system. A multi-camera system can use one or more cameras facing the front of the vehicle. In other embodiments, the multi-camera system can include one or more cameras facing the sides or rear of the vehicle. In one embodiment, for example, system 100 can use a dual-camera imaging system, where the first and second cameras (e.g., image capture devices 122 and 124) can be located in front of and / or on the sides of a vehicle (e.g., vehicle 200). Other camera configurations are consistent with the disclosed embodiments, and the configurations disclosed herein are examples. For example, system 100 can include a configuration with any number of cameras (e.g., one, two, three, four, five, six, seven, eight, etc.). Furthermore, system 100 can include camera "clusters." For example, a camera cluster (including any appropriate number of cameras, such as one, four, eight, etc.) can be forward-facing relative to the vehicle, or can face any other direction (e.g., rearward, sideways, angled, etc.). Thus, system 100 can include multiple camera clusters, each oriented in a specific direction to capture images from a specific area of the vehicle's environment.
[0163] The first camera may have a field of view that is larger than, smaller than, or partially overlaps with the field of view of the second camera. Furthermore, the first camera may be connected to a first image processor to perform monocular image analysis on the images provided by the first camera, and the second camera may be connected to a second image processor to perform monocular image analysis on the images provided by the second camera. The outputs of the first and second image processors (e.g., processed information) may be combined. In some embodiments, the second image processor may receive images from both the first and second cameras to perform stereo analysis. In another embodiment, system 100 may utilize a three-camera imaging system, each camera having a different field of view. Thus, such a system may make decisions based on information derived from objects located at varying distances in front of and to the sides of the vehicle. References to monocular image analysis may refer to instances in which image analysis is performed based on images captured from a single viewpoint (e.g., from a single camera). Stereo image analysis may refer to instances in which image analysis is performed based on two or more images captured using one or more variations in image capture parameters. For example, captured images suitable for performing stereo image analysis may include images captured from two or more different positions, from different fields of view, using different focal lengths, with disparity information, and the like.
[0164] For example, in one embodiment, system 100 may implement a three-camera configuration using image capture devices 122-126. In this configuration, image capture device 122 may provide a narrow field of view (e.g., 34 degrees or another value selected from a range of approximately 20 degrees to 45 degrees, etc.), image capture device 124 may provide a wide field of view (e.g., 150 degrees or another value selected from a range of approximately 100 degrees to approximately 180 degrees), and image capture device 126 may provide an intermediate field of view (e.g., 46 degrees or another value selected from a range of approximately 35 degrees to approximately 60 degrees). In some embodiments, image capture device 126 may serve as the primary or base camera. Image capture devices 122-126 may be located behind rearview mirror 310 and substantially side-by-side (e.g., 6 cm apart). Furthermore, in some embodiments, as discussed above, one or more of image capture devices 122-126 may be mounted behind sun visor 380 flush with the windshield of vehicle 200. This shielding may act to minimize the effect of any reflections from within the vehicle on the image capture devices 122 - 126 .
[0165] In another embodiment, as above combined Figure 3B and 3CAs discussed, the wide field of view camera (e.g., image capture device 124 in the above example) can be mounted lower than the narrow field of view camera and the main field of view camera (e.g., image capture devices 122 and 126 in the above example). This configuration provides a clear line of sight from the wide field of view camera. To reduce reflections, the camera can be mounted closer to the windshield of vehicle 200, and a polarizer can be included on the camera to dampen reflected light.
[0166] A three-camera system can provide certain performance characteristics. For example, some embodiments may include the ability to verify the detection of an object by one camera based on the detection results from another camera. In the three-camera configuration discussed above, the processing unit 110 may include, for example, three processing devices (e.g., three EyeQ series processor chips as discussed above), where each processing device is dedicated to processing images captured by one or more of the image capture devices 122-126.
[0167] In a three-camera system, a first processing device can receive images from both the main camera and the narrow field of view camera and perform vision processing on the narrow FOV camera to, for example, detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the first processing device can calculate the disparity of pixels between the images from the main camera and the narrow camera and create a 3D reconstruction of the vehicle 200's environment. The first processing device can then combine the 3D reconstruction with 3D map data, or with 3D information calculated based on information from another camera.
[0168] The second processing device can receive images from the primary camera and perform visual processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the second processing device can calculate camera displacement and, based on this displacement, calculate the disparity of pixels between consecutive images and create a 3D reconstruction of the scene (e.g., structure from motion). The second processing device can send the structure from motion based on the 3D reconstruction to the first processing device for combination with the stereoscopic 3D image.
[0169] The third processing device can receive images from the wide FOV camera and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing device can also execute additional processing instructions to analyze the images to identify moving objects in the images, such as vehicles changing lanes, pedestrians, etc.
[0170] In some embodiments, having streams of image-based information captured and processed independently can provide an opportunity to provide redundancy in the system. Such redundancy can include, for example, using a first image capture device and images processed from that device to verify and / or supplement information obtained by capturing and processing image information from at least a second image capture device.
[0171] In some embodiments, system 100 utilizes two image capture devices (e.g., image capture devices 122 and 124) to provide navigation assistance for vehicle 200, and utilizes a third image capture device (e.g., image capture device 126) to provide redundancy and validate the analysis of data received from the other two image capture devices. For example, in this configuration, image capture devices 122 and 124 may provide images for stereo analysis performed by system 100 to navigate vehicle 200, while image capture device 126 may provide images for monocular analysis performed by system 100 to provide redundancy and validation of information obtained based on images captured from image capture device 122 and / or image capture device 124. That is, image capture device 126 (and a corresponding processing device) may be considered to provide a redundant subsystem for providing a check on the analysis obtained from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Additionally, in some embodiments, redundancy and verification of the received data may be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors, information received from one or more transceivers outside the vehicle, etc.).
[0172] Those skilled in the art will recognize that the above-described camera configurations, camera placements, number of cameras, camera positions, etc. are merely examples. These components and other components described with respect to the overall system can be assembled and used in a variety of different configurations without departing from the scope of the disclosed embodiments. Further details regarding the use of a multi-camera system to provide driver assistance and / or autonomous vehicle functionality are as follows.
[0173] Figure 4 1 is an exemplary functional block diagram of memory 140 and / or storage 150 that may be stored / programmed with instructions for performing one or more operations consistent with the disclosed embodiments. Although reference is made below to memory 140, those skilled in the art will recognize that instructions may be stored in memory 140 and / or storage 150.
[0174] like Figure 4As shown, memory 140 may store a monocular image analysis module 402, a stereo image analysis module 404, a velocity and acceleration module 406, and a navigation response module 408. The disclosed embodiments are not limited to any particular configuration of memory 140. Furthermore, application processor 180 and / or image processor 190 may execute instructions stored in any of modules 402 to 408 included in memory 140. Those skilled in the art will understand that in the following discussion, references to processing unit 110 may refer individually or collectively to application processor 180 and image processor 190. Therefore, the steps of any of the following processes may be performed by one or more processing devices.
[0175] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the set of images with additional sensory information (e.g., information from radar) to perform monocular image analysis. 5A to 5D As described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, the system 100 (e.g., via the processing unit 110) may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in conjunction with the navigation response module 408.
[0176] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform monocular image analysis. 5A to 5D As described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, the system 100 (e.g., via the processing unit 110) may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in conjunction with determining navigation responses.
[0177] In one embodiment, stereo image analysis module 404 may store instructions (such as computer vision software) that, when executed by processing unit 110, perform stereo image analysis on a first set of images and a second set of images acquired by a combination of image capture devices selected from any one of image capture devices 122, 124, and 126. In some embodiments, processing unit 110 may combine information from the first set of images and the second set of images with additional sensory information (e.g., information from radar) to perform stereo image analysis. For example, stereo image analysis module 404 may include instructions for performing stereo image analysis based on a first set of images acquired by image capture device 124 and a second set of images acquired by image capture device 126. As described below in conjunction with Figure 6 As described, the stereo image analysis module 404 may include instructions for detecting a set of features within the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, etc. Based on this analysis, the processing unit 110 may cause one or more navigation responses in the vehicle 200, such as steering, lane changes, changes in acceleration, etc., as discussed below in conjunction with the navigation response module 408. Furthermore, in some embodiments, the stereo image analysis module 404 may implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.
[0178] In one embodiment, the speed and acceleration module 406 may store software configured to analyze data received from one or more computing and electromechanical devices in the vehicle 200, which are configured to cause changes in the speed and / or acceleration of the vehicle 200. For example, the processing unit 110 may execute instructions associated with the speed and acceleration module 406 to calculate a target velocity for the vehicle 200 based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include, for example, target location, velocity, and / or acceleration, the location and / or velocity of the vehicle 200 relative to nearby vehicles, pedestrians, or road objects, position information of the vehicle 200 relative to lane markings on the road, etc. In addition, the processing unit 110 may calculate the target velocity for the vehicle 200 based on sensory input (e.g., information from radar) and input from other systems of the vehicle 200, such as the throttle control system 220, the braking system 230, and / or the steering system 240 of the vehicle 200. Based on the calculated target rate, the processing unit 110 can transmit electronic signals to the throttle regulation system 220, the braking system 230 and / or the steering system 240 of the vehicle 200, for example by physically pressing the brake or releasing the accelerator of the vehicle 200, to trigger a change in speed and / or acceleration.
[0179] In one embodiment, the navigation response module 408 may store software that is executable by the processing unit 110 to determine a desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include position and velocity information associated with nearby vehicles, pedestrians, and road objects, target position information for the vehicle 200, and the like. Additionally, in some embodiments, the navigation response may be based (in part or in whole) on map data, a predetermined position of the vehicle 200, and / or a relative velocity or acceleration between the vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 may also determine the desired navigation response based on sensory input (e.g., information from radar) and input from other systems of the vehicle 200, such as the throttle control system 220, the braking system 230, and the steering system 240 of the vehicle 200. Based on the desired navigation response, the processing unit 110 may transmit electronic signals to the throttle adjustment system 220, the braking system 230, and the steering system 240 of the vehicle 200 to trigger the desired navigation response, such as by turning the steering wheel of the vehicle 200 to achieve a predetermined angle of rotation. In some embodiments, the processing unit 110 may use the output of the navigation response module 408 (e.g., the desired navigation response) as input to the execution of the speed and acceleration module 406 for calculating the change in the velocity of the vehicle 200.
[0180] Furthermore, any modules disclosed herein (e.g., modules 402, 404, and 406) may implement techniques associated with trained systems (such as neural networks or deep neural networks) or untrained systems.
[0181] Figure 5A is a flow chart illustrating an exemplary process 500A for causing one or more navigation responses based on monocular image analysis consistent with the disclosed embodiments. At step 510, the processing unit 110 may receive a plurality of images via the data interface 128 between the processing unit 110 and the image acquisition unit 120. For example, a camera (such as the image capture device 122 having the field of view 202) included in the image acquisition unit 120 may capture a plurality of images of an area in front of the vehicle 200 (e.g., or to the side or rear of the vehicle) and transmit them to the processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). At step 520, the processing unit 110 may execute the monocular image analysis module 402 to analyze the plurality of images, as described below in conjunction with Figures 5B to 5D By performing this analysis, processing unit 110 may detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, etc.
[0182] At step 520, the processing unit 110 may also execute the monocular image analysis module 402 to detect various road hazards, such as, for example, parts of a truck tire, fallen road signs, loose cargo, small animals, etc. Road hazards may vary in structure, shape, size, and color, which may make the detection of such hazards more challenging. In some embodiments, the processing unit 110 may execute the monocular image analysis module 402 to perform multi-frame analysis on the multiple images to detect road hazards. For example, the processing unit 110 may estimate the camera motion between consecutive image frames and calculate the disparity in pixels between frames to construct a 3D map of the road. The processing unit 110 may then use the 3D map to detect the road surface and the hazards present on the road surface.
[0183] At step 530, the processing unit 110 may execute the navigation response module 408 to perform the navigation response based on the analysis performed at step 520 and the above combination. Figure 4 The described techniques cause one or more navigation responses. The navigation responses may include, for example, steering, lane changes, acceleration changes, etc. In some embodiments, the processing unit 110 may use data obtained from the execution of the speed and acceleration module 406 to cause one or more navigation responses. In addition, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof. For example, the processing unit 110 may cause the vehicle 200 to change lanes and then accelerate by, for example, sequentially transmitting control signals to the steering system 240 and the throttle adjustment system 220 of the vehicle 200. Alternatively, the processing unit 110 may cause the vehicle 200 to brake while changing lanes by, for example, simultaneously transmitting control signals to the braking system 230 and the steering system 240 of the vehicle 200.
[0184] Figure 5B is a flow chart illustrating an exemplary process 500B for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments. Processing unit 110 may execute monocular image analysis module 402 to implement process 500B. At step 540, processing unit 110 may determine a set of candidate objects representing possible vehicles and / or pedestrians. For example, processing unit 110 may scan one or more images, compare the images to one or more predetermined patterns, and identify possible locations within each image that may contain objects of interest (e.g., vehicles, pedestrians, or portions thereof). The predetermined patterns may be designed to achieve a high "false hit" rate and a low "miss" rate. For example, processing unit 110 may apply a low similarity threshold to the predetermined patterns to identify candidate objects as possible vehicles or pedestrians. Doing so may allow processing unit 110 to reduce the likelihood of missing (e.g., failing to identify) candidate objects representing vehicles or pedestrians.
[0185] At step 542, processing unit 110 may filter the set of candidate objects based on classification criteria to exclude certain candidates (e.g., irrelevant or less relevant objects). Such criteria may be derived from various attributes associated with object types stored in a database (e.g., a database stored in memory 140). Attributes may include object shape, size, texture, location (e.g., relative to vehicle 200), etc. Thus, processing unit 110 may use one or more sets of criteria to reject false candidates from the set of candidate objects.
[0186] At step 544, processing unit 110 may analyze multiple frames of imagery to determine whether an object in the set of candidate objects represents a vehicle and / or a pedestrian. For example, processing unit 110 may track detected candidate objects across consecutive frames and accumulate frame-by-frame data associated with the detected objects (e.g., size, position relative to vehicle 200, etc.). In addition, processing unit 110 may estimate parameters of the detected objects and compare the frame-by-frame position data of the objects with the predicted position.
[0187] In step 546, the processing unit 110 may construct a set of measurements for the detected object. Such measurements may include, for example, position, velocity, and acceleration values (relative to the vehicle 200) associated with the detected object. In some embodiments, the processing unit 110 may construct the measurements based on estimation techniques such as a Kalman filter or linear quadratic estimation (LQE) using a series of time-based observations and / or based on modeling data available for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). The Kalman filter may be based on a measure of the scale of the object, where the measure of scale is proportional to the time to collision (e.g., the amount of time it takes for the vehicle 200 to reach the object). Thus, by executing steps 540 to 546, the processing unit 110 may identify vehicles and pedestrians that appear within the set of captured images and obtain information associated with the vehicles and pedestrians (e.g., position, velocity, size). Based on this identification and the obtained information, the processing unit 110 may cause one or more navigation responses in the vehicle 200, as described above in conjunction with Figure 5A described.
[0188] At step 548, the processing unit 110 may perform an optical flow analysis on the one or more images to reduce the likelihood of detecting "false hits" and missing candidate objects representing vehicles or pedestrians. Optical flow analysis may refer to, for example, analyzing motion patterns relative to the vehicle 200, associated with other vehicles and pedestrians, and distinct from road motion in one or more images. The processing unit 110 may calculate the motion of the candidate objects by observing different positions of the objects across multiple image frames captured at different times. The processing unit 110 may use the position and time values as inputs to a mathematical model for calculating the motion of the candidate objects. Thus, optical flow analysis may provide another method for detecting vehicles and pedestrians near the vehicle 200. The processing unit 110 may perform the optical flow analysis in conjunction with steps 540 to 546 to provide redundancy in detecting vehicles and pedestrians and improve the reliability of the system 100.
[0189] Figure 5C is a flow chart illustrating an exemplary process 500C for detecting road markings and / or lane geometry in a set of images consistent with the disclosed embodiments. Processing unit 110 may execute monocular image analysis module 402 to implement process 500C. At step 550, processing unit 110 may detect a set of objects by scanning one or more images. In order to detect segments of lane markings, lane geometry, and other relevant road markings, processing unit 110 may filter the set of objects to exclude those determined to be irrelevant (e.g., small potholes, small rocks, etc.). At step 552, processing unit 110 may group together the segments detected in step 550 that belong to the same road marking or lane marking. Based on the grouping, processing unit 110 may develop a model, such as a mathematical model, that represents the detected segments.
[0190] In step 554, the processing unit 110 may construct a set of measurements associated with the detected segment. In some embodiments, the processing unit 110 may create a projection of the detected segment from the image plane onto a real-world plane. The projection may be characterized using a cubic polynomial having coefficients corresponding to physical properties such as the position, slope, curvature, and curvature derivatives of the detected road. In generating the projection, the processing unit 110 may take into account variations in the road surface, as well as the pitch and roll rates associated with the vehicle 200. In addition, the processing unit 110 may model the road elevation by analyzing position and motion cues present on the road surface. Furthermore, the processing unit 110 may estimate the pitch and roll rates associated with the vehicle 200 by tracking a set of feature points in one or more images.
[0191] At step 556, processing unit 110 may perform a multi-frame analysis by, for example, tracking the detected segments across consecutive image frames and accumulating frame-by-frame data associated with the detected segments. As processing unit 110 performs multi-frame analysis, the set of measurements constructed in step 554 may become more reliable and associated with increasingly higher confidence levels. Thus, by executing steps 550 through 556, processing unit 110 may identify road markings present in the set of captured images and derive lane geometry information. Based on this identification and the derived information, processing unit 110 may cause one or more navigation responses in vehicle 200, as described above in conjunction with Figure 5A described.
[0192] At step 558, processing unit 110 may consider additional information sources to further develop a safety model of vehicle 200 in its surrounding environment. Processing unit 110 may use this safety model to define an environment in which system 100 can safely perform autonomous control of vehicle 200. To develop this safety model, in some embodiments, processing unit 110 may consider the position and motion of other vehicles, detected curbs and guardrails, and / or a general road shape description extracted from map data (such as data from map database 160). By considering additional information sources, processing unit 110 may provide redundancy for detecting road markings and lane geometry and increase the reliability of system 100.
[0193] Figure 5D is a flow chart illustrating an exemplary process 500D for detecting traffic lights in a set of images consistent with the disclosed embodiments. Processing unit 110 may execute monocular image analysis module 402 to implement process 500D. At step 560, processing unit 110 may scan the set of images and identify objects appearing at locations in the images that may contain traffic lights. For example, processing unit 110 may filter the identified objects to construct a set of candidate objects, excluding those that are unlikely to correspond to traffic lights. Filtering may be based on various attributes associated with traffic lights, such as shape, size, texture, and position (e.g., relative to vehicle 200). Such attributes may be based on multiple examples of traffic lights and traffic control signals and stored in a database. In some embodiments, processing unit 110 may perform multi-frame analysis on the set of candidate objects reflecting possible traffic lights. For example, processing unit 110 may track candidate objects across consecutive image frames, estimate the real-world positions of the candidate objects, and filter out objects that are moving (thus unlikely to be traffic lights). In some embodiments, processing unit 110 may perform color analysis on the candidate objects and identify the relative positions of the detected colors appearing within the possible traffic lights.
[0194] At step 562, processing unit 110 may analyze the geometry of the intersection. This analysis may be based on any combination of: (i) the number of lanes detected on either side of vehicle 200, (ii) detected markings on the road (e.g., arrows), and (iii) a description of the intersection extracted from map data (e.g., data from map database 160). Processing unit 110 may use information derived from execution of monocular analysis module 402 to perform the analysis. Furthermore, processing unit 110 may determine the correspondence between the traffic lights detected at step 560 and the lanes present near vehicle 200.
[0195] At step 564, as the vehicle 200 approaches the intersection, the processing unit 110 may update the confidence level associated with the analyzed intersection geometry and detected traffic lights. For example, the number of traffic lights estimated to be present at the intersection compared to the number of traffic lights actually present at the intersection may affect the confidence level. Therefore, based on the confidence level, the processing unit 110 may delegate control to the driver of the vehicle 200 in order to improve safety conditions. By executing steps 560 to 564, the processing unit 110 may identify the traffic lights that appear within the set of captured images and analyze the intersection geometry information. Based on this identification and analysis, the processing unit 110 may cause one or more navigation responses in the vehicle 200, as described above in conjunction with Figure 5A described.
[0196] Figure 5E is a flow chart illustrating an exemplary process 500E for inducing one or more navigation responses in vehicle 200 based on a vehicle path, consistent with the disclosed embodiments. At step 570, processing unit 110 may construct an initial vehicle path associated with vehicle 200. The vehicle path may be represented by a set of points expressed in coordinates (x, z), and the distance d between two points in the set of points may be i May fall within the range of 1 to 5 meters. In one embodiment, the processing unit 110 may construct an initial vehicle path using two polynomials, such as a left road polynomial and a right road polynomial. The processing unit 110 may calculate the geometric midpoint between the two polynomials and offset each point included in the resulting vehicle path by a predetermined offset (e.g., a smart lane offset), if any (a zero offset may correspond to traveling in the middle of the lane). The offset may be in a direction perpendicular to the line segment between any two points in the vehicle path. In another embodiment, the processing unit 110 may use a polynomial and an estimated lane width to offset each point of the vehicle path by half the estimated lane width plus a predetermined offset (e.g., a smart lane offset).
[0197] At step 572, the processing unit 110 may update the vehicle path constructed at step 570. The processing unit 110 may reconstruct the vehicle path constructed at step 570 using a higher resolution so that the distance d between two points in the set of points representing the vehicle path is less than d. k Smaller than the above distance d i For example, the distance d k The processing unit 110 may reconstruct the vehicle path using a parabolic spline algorithm, which may produce a cumulative distance vector S corresponding to the total length of the vehicle path (ie, based on the set of points representing the vehicle path).
[0198] At step 574, the processing unit 110 may determine a look-ahead point (expressed in coordinates (x l ,z l )). The processing unit 110 can extract a look-ahead point from the accumulated distance vector S, and the look-ahead point can be associated with a look-ahead distance and a look-ahead time. The look-ahead distance, which can have a lower bound ranging from 10 meters to 20 meters, can be calculated as the product of the velocity of the vehicle 200 and the look-ahead time. For example, as the velocity of the vehicle 200 decreases, the look-ahead distance can also decrease (e.g., until it reaches the lower bound). The look-ahead time, which can range from 0.5 to 1.5 seconds, can be inversely proportional to the gain of one or more control loops associated with causing a navigation response in the vehicle 200, such as a heading error tracking control loop. For example, the gain of the heading error tracking control loop can depend on the bandwidth of the yaw rate loop, the steering actuator loop, the lateral dynamics of the vehicle, etc. Therefore, the higher the gain of the heading error tracking control loop, the shorter the look-ahead time.
[0199] At step 576, the processing unit 110 may determine the heading error and yaw rate command based on the foresight point determined at step 574. The processing unit 110 may calculate the arc tangent of the foresight point, e.g. The yaw rate command may be determined by the processing unit 110 as the product of the heading error and the high-level control gain. If the look-ahead distance is not at the lower bound, the high-level control gain may be equal to: (2 / look-ahead time). Otherwise, the high-level control gain may be equal to: (2×velocity of vehicle 200 / look-ahead distance).
[0200] Figure 5FFIGURE 5 is a flow chart illustrating an exemplary process 500F for determining whether a leading vehicle is changing lanes consistent with the disclosed embodiments. In step 580, the processing unit 110 may determine navigation information associated with a leading vehicle (e.g., a vehicle traveling in front of the vehicle 200). For example, the processing unit 110 may use the above combined Figure 5A and Figure 5B The described techniques can be used to determine the position, velocity (e.g., direction and speed), and / or acceleration of the vehicle ahead. The processing unit 110 can also use the above combined Figure 5E The described techniques determine one or more road polynomials, look-ahead points (associated with vehicle 200 ), and / or snail trails (eg, a set of points describing a path taken by a leading vehicle).
[0201] In step 582, processing unit 110 may analyze the navigation information determined in step 580. In one embodiment, processing unit 110 may calculate the distance between the tracking trajectory and the road polynomial (e.g., along the trajectory). If the variance of this distance along the trajectory exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters on a straight road, 0.3 to 0.4 meters on a moderately curved road, and 0.5 to 0.6 meters on a sharply curved road), processing unit 110 may determine that the leading vehicle is likely changing lanes. If multiple vehicles are detected traveling ahead of vehicle 200, processing unit 110 may compare the tracking trajectories associated with each vehicle. Based on this comparison, processing unit 110 may determine that a vehicle whose tracking trajectory does not match the tracking trajectories of the other vehicles is likely changing lanes. Processing unit 110 may additionally compare the curvature of the tracking trajectory (associated with the leading vehicle) with the expected curvature of the road segment in which the leading vehicle is traveling. The expected curvature may be extracted from map data (e.g., data from map database 160), from a road polynomial, from tracked trajectories of other vehicles, from prior knowledge about the road, etc. If the difference between the curvature of the tracked trajectory and the expected curvature of the road segment exceeds a predetermined threshold, processing unit 110 may determine that the leading vehicle is likely changing lanes.
[0202] In another embodiment, the processing unit 110 may compare the instantaneous position of the preceding vehicle with the forward view point (associated with the vehicle 200) over a specific time period (e.g., 0.5 to 1.5 seconds). If the distance between the instantaneous position of the preceding vehicle and the forward view point changes during the specific time period, and the cumulative sum of the changes exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a moderately curved road, and 1.3 to 1.7 meters on a sharp curve), the processing unit 110 may determine that the preceding vehicle is likely changing lanes. In another embodiment, the processing unit 110 may analyze the geometry of the tracking trajectory by comparing the lateral distance traveled along the trajectory with the expected curvature of the tracking trajectory. The expected radius of curvature may be determined according to the following calculation: ,in represents the lateral distance traveled and denoted by the longitudinal distance traveled. If the difference between the lateral distance traveled and the expected curvature exceeds a predetermined threshold (e.g., 500 to 700 meters), the processing unit 110 may determine that the leading vehicle is likely to be changing lanes. In another embodiment, the processing unit 110 may analyze the position of the leading vehicle. If the position of the leading vehicle occludes the road polynomial (e.g., the leading vehicle is overlaid on top of the road polynomial), the processing unit 110 may determine that the leading vehicle is likely to be changing lanes. In the event that the position of the leading vehicle is such that another vehicle is detected in front of the leading vehicle and the tracking trajectories of the two vehicles are not parallel, the processing unit 110 may determine that the (closer) leading vehicle is likely to be changing lanes.
[0203] At step 584, the processing unit 110 may determine whether the leading vehicle 200 is changing lanes based on the analysis performed at step 582. For example, the processing unit 110 may make this determination based on a weighted average of the individual analyses performed at step 582. Under such an approach, for example, a determination by the processing unit 110 that the leading vehicle is likely changing lanes based on a particular type of analysis may be assigned a value of "1" (and "0" used to represent a determination that the leading vehicle is unlikely to be changing lanes). The different analyses performed at step 582 may be assigned different weights, and the disclosed embodiments are not limited to any particular combination of analyses and weights. Furthermore, in some embodiments, the analysis may utilize a trained system (e.g., a machine learning or deep learning system) that may, for example, estimate a future path ahead of the vehicle's current location based on images captured at the current location.
[0204] Figure 6is a flow chart illustrating an exemplary process 600 for eliciting one or more navigation responses based on stereo image analysis consistent with the disclosed embodiments. At step 610, processing unit 110 may receive a first and second plurality of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122 and 124 having fields of view 202 and 204) may capture a first and second plurality of images of an area in front of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first and second plurality of images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0205] At step 620, the processing unit 110 may execute the stereo image analysis module 404 to perform stereo image analysis on the first and second plurality of images to create a 3D map of the road ahead of the vehicle and detect features within the image, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. The stereo image analysis may be performed in a manner similar to the above combined Figures 5A-5D The steps described above may be performed in the manner described above. For example, processing unit 110 may execute stereo image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road signs, traffic lights, road hazards, etc.) within the first and second pluralities of images, filter out a subset of candidate objects based on various criteria, and perform multi-frame analysis, construct measurements, and determine confidence levels for the remaining candidate objects. In performing the above steps, processing unit 110 may consider information from both the first and second pluralities of images, rather than from a single set of images. For example, processing unit 110 may analyze differences in pixel-level data (or other subsets of data from the two streams of captured images) for candidate objects appearing in both the first and second pluralities of images. As another example, processing unit 110 may estimate the position and / or velocity of a candidate object (e.g., relative to vehicle 200) by observing whether an object appears in one of the multiple images but not in the other, or other differences that may exist relative to objects appearing in both image streams. For example, the position, velocity, and / or acceleration relative to vehicle 200 may be determined based on the trajectory, position, movement characteristics, etc. of features associated with the object appearing in one or both image streams.
[0206] At step 630, the processing unit 110 may execute the navigation response module 408 to perform the navigation response based on the analysis performed at step 620 and the above combination. Figure 4The techniques described herein may be used to cause one or more navigation responses in the vehicle 200. The navigation responses may include, for example, steering, lane changes, changes in acceleration, changes in speed, braking, etc. In some embodiments, the processing unit 110 may use data obtained from the execution of the speed and acceleration module 406 to cause the one or more navigation responses. Furthermore, multiple navigation responses may occur simultaneously, sequentially, or any combination thereof.
[0207] Figure 7 is a flow chart illustrating an exemplary process 700 for inducing one or more navigation responses based on analysis of three sets of images, consistent with the disclosed embodiments. At step 710, processing unit 110 may receive first, second, and third pluralities of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) may capture first, second, and third pluralities of images of areas in front of and / or to the sides of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first, second, and third pluralities of images via three or more data interfaces. For example, each of image capture devices 122, 124, 126 may have an associated data interface for transmitting data to processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0208] At step 720, the processing unit 110 may analyze the first, second, and third plurality of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. This analysis may be performed similarly to the above combined Figures 5A-5D and Figure 6 For example, the processing unit 110 may perform monocular image analysis on each of the first, second, and third plurality of images (e.g., via execution of the monocular image analysis module 402 and based on the above combination). Figures 5A-5D Alternatively, the processing unit 110 may perform stereoscopic image analysis on the first and second pluralities of images, the second and third pluralities of images, and / or the first and third pluralities of images (e.g., via execution of the stereoscopic image analysis module 404 and based on the above combination). Figure 64 and 3 pluralities of images). The processed information corresponding to the analysis of the first, second, and / or third pluralities of images may be combined. In some embodiments, the processing unit 110 may perform a combination of monocular and stereo image analysis. For example, the processing unit 110 may perform monocular image analysis on the first pluralities of images (e.g., via execution by the monocular image analysis module 402) and perform stereo image analysis on the second and third pluralities of images (e.g., via execution by the stereo image analysis module 404). The configuration of the image capture devices 122, 124, and 126—including their respective positions and fields of view 202, 204, and 206—may affect the type of analysis performed on the first, second, and third pluralities of images. The disclosed embodiments are not limited to a particular configuration of the image capture devices 122, 124, and 126 or the type of analysis performed on the first, second, and third pluralities of images.
[0209] In some embodiments, processing unit 110 may perform testing on system 100 based on the images acquired and analyzed in steps 710 and 720. Such testing may provide an indicator of the overall performance of system 100 for certain configurations of image acquisition devices 122, 124, and 126. For example, processing unit 110 may determine the ratio of "false hits" (e.g., instances where system 100 incorrectly determines the presence of a vehicle or pedestrian) to "misses."
[0210] At step 730, the processing unit 110 may cause one or more navigation responses in the vehicle 200 based on information obtained from two of the first, second, and third pluralities of images. The selection of two of the first, second, and third pluralities of images may depend on various factors, such as, for example, the number, type, and size of objects detected in each of the plurality of images. The processing unit 110 may also make a selection based on image quality and resolution, the effective field of view reflected in the image, the number of frames captured, the degree to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which the object appears, the proportion of the object appearing in each such frame), etc.
[0211] In some embodiments, processing unit 110 may select information derived from two of the first, second, and third pluralities of images by determining how consistent the information derived from one image source is with information derived from other image sources. For example, processing unit 110 may combine processed information derived from each of image capture devices 122, 124, and 126 (whether through monocular analysis, stereo analysis, or any combination thereof) and determine visual indicators (e.g., lane markings, detected vehicles and their positions and / or paths, detected traffic lights, etc.) that are consistent across each captured image from image capture devices 122, 124, and 126. Processing unit 110 may also exclude information that is inconsistent across the captured images (e.g., vehicles changing lanes, lane models indicating that a vehicle is too close to vehicle 200, etc.). Thus, processing unit 110 may select information derived from two of the first, second, and third pluralities of images based on the determination of consistent and inconsistent information.
[0212] The navigation response may include, for example, steering, lane change, acceleration change, etc. The processing unit 110 may perform the following actions based on the analysis performed in step 720 and the above combination. Figure 4 The described techniques cause one or more navigation responses. Processing unit 110 may also cause one or more navigation responses using data derived from execution of velocity and acceleration module 406. In some embodiments, processing unit 110 may cause one or more navigation responses based on the relative position, relative velocity, and / or relative acceleration between vehicle 200 and an object detected within any of the first, second, and third pluralities of images. Multiple navigation responses may occur simultaneously, sequentially, or any combination thereof.
[0213] Reinforcement learning and trained navigation systems
[0214] The following section discusses autonomous driving and systems and methods for achieving autonomous control of a vehicle, whether that control is fully autonomous (a self-driving vehicle) or partially autonomous (e.g., one or more driver assistance systems or functions). Figure 8As shown, the autonomous driving task can be divided into three main modules, including a sensing module 801, a driving strategy module 803, and a control module 805. In some embodiments, modules 801, 803, and 805 can be stored in the memory unit 140 and / or the memory unit 150 of the system 100, or modules 801, 803, and 805 (or portions thereof) can be stored remotely from the system 100 (e.g., stored in a server accessible to the system 100 via, for example, the wireless transceiver 172). In addition, any module disclosed herein (e.g., modules 801, 803, and 805) can implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.
[0215] Sensing module 801, which may be implemented using processing unit 110, can handle various tasks related to sensing the navigational state of the host vehicle's environment. Such tasks may rely on input from various sensors and sensing systems associated with the host vehicle. These inputs may include images or image streams from one or more onboard cameras, GPS location information, accelerometer output, user feedback or user input to one or more user interface devices, radar, lidar, and the like. Sensing, which may include data from cameras and / or any other available sensors, as well as map information, can be collected, analyzed, and formulated into a "sensed state," which describes information extracted from the scene in the host vehicle's environment. This sensed state may include, among other potential sensed information, sensed information related to target vehicles, lane markings, pedestrians, traffic lights, road geometry, lane shape, obstacles, distance to other objects / vehicles, relative velocity, and relative acceleration. Supervised machine learning may be implemented to generate a sensed state output based on the sensed data provided to sensing module 801. The output of the sensing module may represent the sensed navigational "state" of the host vehicle, which may be passed to driving strategy module 803.
[0216] While the sensed state can be developed based on image data received from one or more cameras or image sensors associated with the host vehicle, the sensed state for use in navigation can also be developed using any suitable sensor or combination of sensors. In some embodiments, the sensed state can be developed without relying on captured image data. Indeed, any of the navigation principles described herein can be applied to sensed states developed based on captured image data as well as sensed states developed using other non-image based sensors. The sensed state can also be determined via a source external to the host vehicle. For example, the sensed state can be developed based in whole or in part on information received from a source remote from the host vehicle (e.g., based on sensor information, processed state information, etc. shared from other vehicles, shared from a central server, or from any other information source relevant to the navigation state of the host vehicle).
[0217] The driving policy module 803 (discussed in more detail below and may be implemented using the processing unit 110) may implement a desired driving policy to determine one or more navigation actions to be taken by the host vehicle in response to the sensed navigation state. If there are no other agents (e.g., target vehicles or pedestrians) in the host vehicle's environment, the sensed state input to the driving policy module 803 may be processed in a relatively straightforward manner. The task becomes more complex when the sensed state requires negotiation with one or more other agents. Techniques for generating the output of the driving policy module 803 may include reinforcement learning (discussed in more detail below). The output of the driving policy module 803 may include at least one navigation action for the host vehicle and may include a desired acceleration (which may be converted into an updated velocity for the host vehicle), a desired yaw rate for the host vehicle, and a desired trajectory, in addition to other potential desired navigation actions.
[0218] Based on the output from the driving strategy module 803, a control module 805, which may also be implemented using the processing unit 110, may develop control instructions for one or more actuators or controlled devices associated with the host vehicle. Such actuators and devices may include an accelerator, one or more steering controls, brakes, signal transmitters, displays, or any other actuators or devices that may be controlled as part of navigation operations associated with the host vehicle. Aspects of control theory may be used to generate the output of the control module 805. The control module 805 may be responsible for developing and outputting instructions to the controllable components of the host vehicle to implement the desired navigation goals or requirements of the driving strategy module 803.
[0219] Returning to the driving policy module 803, in some embodiments, the driving policy module 803 can be implemented using a trained system trained using reinforcement learning. In other embodiments, the driving policy module 803 can be implemented without machine learning methods by using a specified algorithm to "manually" solve various scenarios that may arise during autonomous navigation. However, while feasible, this approach may result in overly simplistic driving policies and may lack the flexibility of a trained system based on machine learning. For example, a trained system may be better equipped to handle complex navigation states and may be better able to determine whether a taxi is stopping or stopping to pick up or drop off a passenger; determine whether a pedestrian intends to cross the street in front of the host vehicle; defensively balance unexpected actions of other drivers; negotiate dense traffic involving target vehicles and / or pedestrians; decide when to suspend certain navigation rules or augment others; anticipate unsensed but expected conditions (e.g., whether a pedestrian will emerge from behind a car or obstacle); and so on. A trained system based on reinforcement learning may also be better equipped to solve continuous and high-dimensional state spaces and continuous action spaces.
[0220] Training a system using reinforcement learning can involve learning a driving policy to map from sensed states to navigation actions. A driving policy is a function ,in is a set of states, and is the action space (e.g., desired velocity, acceleration, yaw command, etc.). The state space is ,in is the sensing state, and is additional information about the state saved by the strategy. Working in discrete time intervals, at time t, the current state can be observed , and the strategy can be applied to obtain the desired action .
[0221] The system can be trained by exposing it to various navigation states, causing it to apply the policy, and providing rewards (based on a reward function designed to reward the desired navigation behavior). Based on the reward feedback, the system can "learn" the policy and become trained in producing the desired navigation actions. For example, the learning system can observe the current state , and based on strategy To decide the action Based on the action decided (and the execution of that action), the environment moves to the next state , to be observed by the learning system. For each action developed in response to the observed state, the feedback to the learning system is a reward signal .
[0222] The goal of reinforcement learning (RL) is to find a policy It is usually assumed that at time t, there is a reward function , which is measured in the state and take action However, taking action at time t affects the environment and, therefore, the value of future states. Therefore, when deciding which action to take, not only the current reward but also future rewards should be considered. In some cases, when the system determines that a higher reward may be achieved in the future if a lower reward option is taken now, then the system should take an action even if that action is associated with a lower reward than another available option. To formalize this, observe that the policy And the initial state s is summarized in distribution over , where if the agent is in state Start at the beginning and follow the strategy from there , then the vector The probability of observing the reward The probability of the initial state s can be defined as:
[0223] .
[0224] Instead of defining the time horizon as T, we can define it by discounting future rewards, for some fixed :
[0225] .
[0226] In any case, the best strategy is the solution to the following equation:
[0227]
[0228] Here, the expectation is above the initial state s.
[0229] There are several possible approaches for training a driving policy system. For example, one could use imitation methods (e.g., behavioral cloning), in which the system learns from state / action pairs, where the actions are those that a good agent (e.g., a human) would choose in response to a particular observed state. Suppose a human driver is observed. From this observation, the form Many examples (including is the state, and are human driver actions) can be obtained, observed, and used as a basis for training a driving policy system. For example, supervised learning can be used to learn a policy , making This approach has many potential advantages. First, it does not require the definition of a reward function. Second, the learning is supervised and occurs offline (no agent is required during the learning process). A disadvantage of this approach is that different human drivers, and even the same human driver, are not deterministic in their strategy choices. Therefore, learning is not very reliable. Very small functions are usually not feasible. Moreover, even small errors can accumulate over time to produce larger errors.
[0230] Another technique that can be used is policy-based learning. Here, the policy can be expressed in parameter form and optimized directly using a suitable optimization technique (e.g., stochastic gradient descent). There are many other ways to solve this problem. One advantage of this approach is that it directly solves the problem, and therefore often leads to good practical results. A potential disadvantage is that it often requires "on-policy" training, i.e. The learning process is an iterative process, where at iteration j, there is a non-perfect policy , and in order to construct the next strategy , must be based on Interact with the environment while performing actions.
[0231] The system can also be trained by value-based learning (learning Q or V functions). Assume that the optimal value function can be learned A good approximation of . An optimal policy can be constructed (e.g., by relying on the Bellman equation). Some versions of value-based learning can be implemented offline (called "off-policy" training). Some disadvantages of value-based methods may be due to their strong reliance on Markov assumptions and the need to approximate complex functions (which can be more difficult to approximate than directly approximating the policy).
[0232] Another technique may include model-based learning and planning (learning the probabilities of state transitions and solving the optimization problem of finding the optimal V). A combination of these techniques can also be used to train a learning system. In this approach, the dynamics of the process can be learned, i.e., taking And in the next state Once this function is learned, we can solve the optimization problem to find the policy whose value is the best. This is called “planning”. One advantage of this approach may be that the learning is partially supervised and can be done by observing the triples Similar to the "imitation" approach, a drawback of this approach may be that small errors in the learning process may accumulate and produce an underperforming policy.
[0233] Another approach for training the driving policy module 803 can involve decomposing the driving policy function into semantically meaningful components. This allows for manual implementation of certain parts of the policy, which can ensure policy safety, while other parts can be implemented using reinforcement learning techniques, which can enable adaptability to a wide range of scenarios, human-like balance between defensive / aggressive behavior, and human-like negotiation with other drivers. From a technical perspective, reinforcement learning methods can combine several approaches and provide an easy-to-use training procedure, where much of the training can be performed using recorded data or a custom simulator.
[0234] In some embodiments, the training of the driving policy module 803 can rely on an "option" mechanism. To illustrate, consider a simple scenario of a driving policy on a two-lane highway. In a direct RL approach, the policy Mapping the state to Among them The first component of is the desired acceleration command, and The second component of is the yaw rate. In the modified approach, the following strategy can be constructed:
[0235] Automatic Cruise Control (ACC) strategy, : This policy always outputs a yaw rate of 0 and only varies the rate to achieve smooth and accident-free driving.
[0236] ACC+left strategy, : The longitudinal commands for this strategy are the same as for ACC. The yaw rate is a straightforward implementation to center the vehicle toward the middle of the left lane while ensuring safe lateral movement (e.g., not moving left if there is a car to the left).
[0237] ACC+right strategy, :and Same, but the vehicle may center toward the center of the right lane.
[0238] These strategies can be called "options". Based on these "options", we can learn the strategy of choosing the option. ,in is a set of available options. In one case, . Option Selector Strategy By setting To define the actual strategy .
[0239] In practice, the policy function can be decomposed into an option graph 901, such as Figure 9 shown. Figure 10 Another example option graph 1000 is shown in FIG. An option graph can represent a hierarchical decision set organized as a directed acyclic graph (DAG). There is a special node called the root node 903 of the graph. This node has no incoming nodes. The decision process starts at the root node and traverses the entire graph until it reaches a "leaf" node, which is a node with no outgoing decision lines. Figure 9 As shown, the leaf nodes may include, for example, nodes 905, 907, and 909. Upon encountering a leaf node, the driving strategy module 803 may output acceleration and steering commands associated with the desired navigation action associated with the leaf node.
[0240] Internal nodes, such as nodes 911, 913, and 915, for example, may result in the implementation of a policy that selects a child among its available options. The set of available children of an internal node includes all nodes associated with the particular internal node via a decision line. For example, Figure 9 The internal node 913 designated as "merge" includes three child nodes 909, 915 and 917 ("keep", "pass right" and "pass left"), each connected to the node 913 by a decision line.
[0241] Flexibility in the decision making system can be achieved by enabling nodes to adjust their position in the hierarchy of the options graph. For example, any node can be allowed to declare itself as "critical". Each node can implement an "is critical" function that outputs "true" if the node is in a critical part of its policy implementation. For example, a node responsible for overtaking can declare itself as critical in the middle of a maneuver. This can impose a constraint on the set of available children of a node u, which can include all nodes v that are children of node u, and there is a path from v to a leaf node that passes through all nodes designated as critical. On the one hand, this approach can allow the desired path to be declared on the graph at each time step, while on the other hand, the stability of the policy can be preserved, especially when the critical part of the policy is being implemented.
[0242] Learning driving strategies by defining option maps The problem of can be decomposed into the problem of defining a policy for each node of the graph, where the policy at the internal nodes should be selected from among the available child nodes. For some nodes, the corresponding policy can be implemented manually (for example, by an if-then type algorithm to specify a set of actions in response to the observed state), while for other policies it can be implemented using a trained system built by reinforcement learning. The choice between manual or trained / learned approaches can depend on the safety aspects associated with the task and on its relative simplicity. The option graph can be constructed in such a way that some nodes are implemented directly, while other nodes can rely on trained models. This approach can ensure safe operation of the system.
[0243] The following discussion provides information on how to Figure 9 Further details on the role of the option map in [ 803 ] are provided. As described above, the input to the driving policy module is a "sensed state," which summarizes the environment map, e.g., obtained from available sensors. The output of the driving policy module 803 is a set of expectations (optionally along with a set of hard constraints) that define the trajectory as a solution to the optimization problem.
[0244] As described above, an option graph represents a hierarchical set of decisions organized as a DAG. There is a special node called the "root" of the graph. The root node is the only node with no incoming edges (e.g., decision lines). The decision process traverses the graph starting from the root node until it reaches a "leaf" node, which is a node with no outgoing edges. Each internal node should implement a strategy for selecting a child from its available children. Each leaf node should implement a strategy that defines a set of expectations (e.g., a set of navigation goals for the host vehicle) based on the entire path from root to leaf. This set of expectations, defined directly based on the sensed states, together with a set of hard constraints, establishes an optimization problem whose solution is the vehicle's trajectory. Hard constraints can be employed to further improve the safety of the system, and the expectations can be used to provide the system with driving comfort and human-like driving behavior. The trajectory provided as the solution to the optimization problem, in turn, defines the commands that should be provided to the steering, braking, and / or engine actuators in order to complete the trajectory.
[0245] return Figure 9, option graph 901 represents an option graph for a two-lane highway, including a merging lane (meaning that at some point, the third lane merges into the right or left lane of the highway). Root node 903 first determines whether the host vehicle is in a normal road scenario or an approaching merging scenario. This is an example of a decision that can be made based on the sensed state. Normal road node 911 includes three child nodes: hold node 909, left overtake node 917, and right overtake node 915. Hold refers to a situation where the host vehicle wants to continue traveling in the same lane. The hold node is a leaf node (no outgoing edges / lines). Therefore, the hold node defines a set of expectations. The first expectation it defines can include a desired lateral position - for example, as close to the center of the current lane as possible. It can also be expected to navigate smoothly (for example, within a predetermined or allowed acceleration maximum). The hold node can also define how the host vehicle reacts to other vehicles. For example, the hold node can view the sensed target vehicles and assign semantic meaning to each vehicle, which can be converted into components of the trajectory.
[0246] Various semantic meanings can be assigned to target vehicles in the host vehicle's environment. For example, in some embodiments, semantic meanings can include any of the following designations: 1) Not relevant: indicates that the sensed vehicle in the scene is currently irrelevant; 2) Next lane: indicates that the sensed vehicle is in an adjacent lane and should maintain an appropriate offset relative to the vehicle (the exact offset can be calculated in an optimization problem that constructs a trajectory given a desired and hard constraints, and the exact offset can potentially be vehicle-dependent—the hold leaf of the option graph sets the semantic type of the target vehicle, which defines the desired relative to the target vehicle); 3) Yield: The host vehicle will attempt to yield to the sensed target vehicle by, for example, reducing speed (particularly if the host vehicle determines that the target vehicle may cut into the host vehicle's lane); 4) Takeway: The host vehicle will attempt to take the right of way by, for example, increasing speed; 5) Follow: The host vehicle desires to follow the target vehicle and maintain a steady course; 6) Left / right pass: This means that the host vehicle intends to initiate a lane change to the left or right lane. Left pass node 917 and left pass node 915 are internal nodes where the desired behavior is not yet defined.
[0247] The next node in option graph 901 is select gap node 919. This node may be responsible for selecting a gap between two target vehicles in a particular target lane that the subject vehicle desires to enter. By selecting a node of the form IDj, for some value of j, the subject vehicle reaches a leaf that specifies a desire for the trajectory optimization problem—for example, the subject vehicle desires to maneuver in order to reach the selected gap. This maneuver may involve first accelerating / braking in the current lane, then moving to the target lane at an appropriate time to enter the selected gap. If select gap node 919 cannot find a suitable gap, it moves to abort node 921, which defines the desire to return to the center of the current lane and cancel the overtaking.
[0248] Returning to the merge node 913, when the host vehicle approaches the merge, it has several options that may depend on the specific situation. Figure 11A As shown, the host vehicle 1105 is traveling along a two-lane road, wherein no other target vehicles are detected in the main lane or merging lane 1111 of the two-lane road. In this case, the driving strategy module 803 may select the hold node 909 upon reaching the merge node 913. That is, it may be desirable to remain in its current lane if no target vehicles are sensed merging onto the roadway.
[0249] exist Figure 11B In , the situation is slightly different. Here, the host vehicle 1105 senses one or more target vehicles 1107 entering the main roadway 1112 from the merging lane 1111. In this case, once the driving strategy module 803 encounters the merging node 913, it can choose to initiate a left passing maneuver to avoid the merging situation.
[0250] exist Figure 11C In Figure 1, host vehicle 1105 encounters one or more target vehicles 1107 entering main roadway 1112 from merging lane 1111. Host vehicle 1105 also detects target vehicle 1109 traveling in a lane adjacent to the host vehicle's lane. Host vehicle also detects one or more target vehicles 1110 traveling in the same lane as host vehicle 1105. In this case, driving policy module 803 may decide to adjust the speed of host vehicle 1105 to yield to target vehicle 1107 and travel ahead of target vehicle 1115. This can be achieved, for example, by proceeding to select gap node 919, which in turn selects the gap between ID0 (vehicle 1107) and ID1 (vehicle 1115) as the appropriate merging gap. In this case, the appropriate gap for the merging situation defines the objective of the trajectory planner optimization problem.
[0251] As described above, nodes of the option graph can declare themselves as "critical", which ensures that the selected option passes through critical nodes. Formally, each node can implement an IsCritical function. After performing a forward pass from the root to the leaf on the option graph and solving the trajectory planner's optimization problem, a backward pass can be performed from the leaf back to the root. Along this backward pass, the IsCritical function of all nodes in the pass can be called, and a list of all critical nodes can be saved. In the forward path corresponding to the next time frame, the driving policy module 803 may be required to select a path from the root node to the leaf that passes through all critical nodes.
[0252] Figures 11A to 11C This can be used to illustrate the potential benefits of this approach. For example, when an overtaking maneuver is initiated and the driving strategy module 803 reaches the leaf corresponding to IDk, such as when the host vehicle is in the middle of an overtaking maneuver, it would be undesirable to select the hold node 909. To avoid this jump, the IDj node can designate itself as critical. During the maneuver, the success of the trajectory planner can be monitored, and if the overtaking maneuver proceeds as expected, the function IsCritical will return a "true" value. This approach can ensure that the overtaking maneuver will continue in the next time frame (rather than jumping to another potentially inconsistent maneuver before completing the originally selected maneuver). On the other hand, if monitoring of the maneuver indicates that the selected maneuver did not proceed as expected, or if the maneuver becomes unnecessary or impossible, the IsCritical function can return a "false" value. This can allow the selection gap node to select a different gap in the next time frame, or to abort the overtaking maneuver entirely. This approach can, on the one hand, allow the desired path to be declared on the option graph at each time step, while on the other hand, it can help improve the stability of the strategy during the critical portion of execution.
[0253] Hard constraints can be different from navigation expectations, which will be discussed in more detail below. For example, hard constraints can ensure safe driving by applying an additional layer of filtering to the planned navigation actions. The hard constraints involved can be determined based on the sensed state, and the hard constraints can be manually programmed and defined, rather than by using a trained system built on reinforcement learning. However, in some embodiments, the trained system can learn the applicable hard constraints to apply and follow. This approach can prompt the driving policy module 803 to arrive at a selected action that already complies with the applicable hard constraints, which can reduce or eliminate the selected action that may need to be modified later to comply with the applicable hard constraints. Nevertheless, as a redundant safety measure, hard constraints can be applied to the output of the driving policy module 803 even if the driving policy module 803 has been trained to take into account predetermined hard constraints.
[0254] There are many examples of potential hard constraints. For example, a hard constraint can be defined in conjunction with a guardrail at the edge of the road. The host vehicle is not allowed to pass the guardrail under any circumstances. Such a rule creates a hard lateral constraint on the trajectory of the host vehicle. Another example of a hard constraint can include a bump in the road (e.g., a rate control bump), which can cause a hard constraint on the driving speed before and while traversing the bump. Hard constraints can be considered safety critical and therefore can be defined manually rather than relying solely on a trained system that learns the constraints during training.
[0255] In contrast to hard constraints, a desired goal may be to achieve or attain comfortable driving. As discussed above, an example of a desire may include a goal to place the host vehicle within a lane at a lateral position corresponding to the center of the host vehicle's lane. Another desire may include the ID of a gap that is suitable to enter. Note that the host vehicle does not need to be exactly in the center of the lane, but rather, a desire to be as close to it as possible can ensure that the host vehicle tends to migrate to the center of the lane even when drifting off the lane center. The desire may not be safety critical. In some embodiments, the desire may require negotiation with other drivers and pedestrians. One approach to constructing the desire may rely on an option graph, and the policies implemented in at least some nodes of the graph may be based on reinforcement learning.
[0256] For nodes of option graph 901 or 1000 implemented as learning-based training nodes, the training process may include decomposing the problem into a supervised learning phase and a reinforcement learning phase. In the supervised learning phase, the learning arrive A differentiable map of , which can be similar to "model-based" reinforcement learning. However, in the forward loop of the network, we can use The actual value is replaced by This eliminates the problem of error accumulation. The role of prediction is to propagate messages from the future back to past actions. In this sense, the algorithm can be a combination of "model-based" reinforcement learning and "policy-based learning".
[0257] An important element that can be provided in some scenarios is a differentiable path from future loss / reward back to the decision on the action. With the option graph structure, the implementation of options involving safety constraints is usually not differentiable. To overcome this problem, the selection of children in the learned policy node can be random. That is, the node can output a probability vector p that assigns the probability of selecting each child of a particular node. Suppose the node has k children, and let is the action of the path from each child to the leaf. The resulting predicted action is thus , which may lead to a differentiable path from that action to p. In practice, the action a can be chosen to be of , and a and The difference between them can be called additive noise.
[0258] For a given Down For training of node policies, supervised learning can be used with real data. For training of node policies, a simulator can be used. Afterwards, fine-tuning of the policy can be done using real data. Two concepts can make the simulation more realistic. First, using imitation, an initial policy can be built using a "behavioral cloning" paradigm using large real-world datasets. In some cases, the resulting agent can be suitable. In other cases, the resulting agent at least forms a very good initial policy for other agents on the road. Second, using self-play, our own policy can be used to reinforce the training. For example, given an initial implementation of other agents (cars / pedestrians), which can be experienced, the policy can be trained based on a simulator. Some other agents can be replaced by the new policy and the process can be repeated. Therefore, the policy can continue to improve because it should respond to a wider variety of other agents with different levels of complexity.
[0259] Furthermore, in some embodiments, the system can implement a multi-agent approach. For example, the system can consider data from various sources and / or images captured from multiple angles. Furthermore, some disclosed embodiments can provide energy economy because the anticipation of events that do not directly involve the host vehicle but may have an impact on the host vehicle can be taken into account, or even the anticipation of events that may lead to unpredictable situations involving other vehicles can be taken into account (e.g., radar may "see through" the vehicle ahead and the high probability of an unavoidable anticipated or even impacting event of the host vehicle).
[0260] Trained system with imposed navigation constraints
[0261] In the context of autonomous driving, a key concern is ensuring that the learned policies of a trained navigation network are safe. In some embodiments, constraints can be used to train the driving policy system so that the actions selected by the trained system already account for applicable safety constraints. Furthermore, in some embodiments, an additional layer of safety can be provided by passing the selected actions of the trained system through one or more hard constraints related to a specific sensed scenario in the host vehicle's environment. This approach can ensure that the actions taken by the host vehicle are limited to those actions that are confirmed to satisfy the applicable safety constraints.
[0262] At its core, the navigation system may include a learning algorithm based on a policy function that maps observed states to one or more desired actions. In some embodiments, the learning algorithm is a deep learning algorithm. The desired action may include at least one action that is expected to maximize the vehicle's expected reward. While in some cases, the actual action taken by the vehicle may correspond to one of the desired actions, in other cases, the actual action taken may be determined based on the observed state, one or more desired actions, and non-learned hard constraints (e.g., safety constraints) imposed on the learning navigation engine. These constraints may include no-drive zones around various types of detected objects (e.g., target vehicles, pedestrians, stationary objects on the side of the road or in the roadway, moving objects on the side of the road or in the roadway, guardrails, etc.). In some cases, the size of this zone may vary based on the detected motion (e.g., rate and / or direction) of the detected object. Other constraints may include a maximum travel speed when traveling within the influence zone of a pedestrian, a maximum deceleration (to account for the distance between target vehicles behind the host vehicle), mandatory stops at sensed crosswalks or railroad crossings, and the like.
[0263] Hard constraints used in conjunction with systems trained via machine learning can provide a degree of safety in autonomous driving that may exceed that achievable based solely on the outputs of the trained system. For example, a machine learning system may be trained using a desired set of constraints as training guidance, and thus, the trained system may select an action in response to a sensed navigation state that accounts for and adheres to the limitations of the applicable navigation constraints. However, the trained system still has some flexibility in selecting navigation actions, and thus, there are at least some situations in which the action selected by the trained system may not strictly adhere to the relevant navigation constraints. Therefore, in order to require that the selected action strictly adhere to the relevant navigation constraints, the outputs of the trained system may be combined, compared, filtered, adjusted, modified, etc. using non-machine learning components outside of the learning / trained framework that ensure the strict application of the relevant navigation constraints.
[0264] The following discussion provides additional details about the trained system and the potential benefits (particularly from a safety perspective) gleaned from combining the trained system with algorithmic components outside the trained / learning framework. As previously mentioned, the reinforcement learning objective of the policy can be optimized via stochastic gradient ascent. The objective (e.g., expected reward) can be defined as .
[0265] Goals involving expectations can be used in machine learning contexts. However, such goals, when not constrained by navigation constraints, may not return actions that are strictly constrained by those constraints. For example, consider a reward function where is used to represent trajectories of rare "corner" events (e.g., such as accidents) that are to be avoided, while For the remaining trajectories, one goal of the learning system can be to learn to perform overtaking maneuvers. Typically, in accident-free trajectories, will reward successful, smooth overtakes and penalize staying in the lane without completing the overtake - hence the range [-1,1]. Represents an accident, then the reward The penalty should be high enough to prevent this from happening. The question is the value of ensuring accident-free driving. What it should be.
[0266] Observed that the accident The impact is additional ,in is the probability mass of the trajectory with the accident event. If this term is negligible, that is, , then the learning system can more often prefer a strategy that results in accidents (or generally adopts a reckless driving strategy) to successfully perform overtaking maneuvers, compared to a more defensive strategy at the expense of some overtaking maneuvers not being successfully completed. In other words, if the probability of an accident is at most p, then one must set Make It is desirable to minimize p (e.g., of the order of magnitude). Therefore, Should be large. In the policy gradient, it can be estimated The following lemma shows that the random variable The variance of And becomes larger, the variance is is greater than Therefore, estimating the target can be difficult, and estimating its gradient can be even more difficult.
[0267] Lemma: Let Be the strategy, and let p and is a scalar such that with probability p and obtain with probability 1-p .So,
[0268]
[0269] where the final approximation applies to situation.
[0270] This discussion shows that the form The goal may not ensure functional safety without causing variance problems. Baseline subtraction methods for variance reduction may not provide an adequate remedy for this problem because the problem will be The high variance of is transferred to the equally high variance of the baseline constant, the estimate of which is also subject to numerical instability. Moreover, if the probability of an accident is p, then on average at least 1 / p sequences should be sampled before an accident event is obtained. This means that for the Minimize the lower bound of 1 / p samples of the sequence of learning algorithms. The solution to this problem can be found in the architectural design described in this paper, rather than through numerical tuning techniques. The approach here is based on the idea that hard constraints should be injected outside the learning framework. In other words, the policy function can be decomposed into a learnable part and a non-learnable part. Formally, the policy function can be constructed as ,in Mapping the (unknowable) state space to a set of expectations (e.g., desired navigation goals, etc.), while Map this expectation to a trajectory (which determines how the car should move over short distances). Function Responsible for driving comfort and making strategic decisions, such as which other cars should be overtaken or given way, and what is the desired position of the host vehicle within its lane. The mapping from the sensed navigation state to the desired state is the policy , which can learn from experience by maximizing the expected reward. The generated expectation can be transformed into a cost function for the driving trajectory. Function Rather than learning a function, it can be achieved by finding a trajectory that minimizes the cost subject to hard constraints on functional safety. This decomposition can ensure functional safety while providing a comfortable ride.
[0271] like Figure 11D As shown, a double merging navigation scenario provides an example that further illustrates these concepts. In a double merging, vehicles approach the merging area 1130 from both the left and right sides. Also, vehicles from each side (such as vehicle 1133 or vehicle 1135) can decide whether to merge into the lane on the other side of the merging area 1130. Successfully executing a double merging in busy traffic may require significant negotiation skills and experience, and may be difficult to perform with a heuristic or brute force approach by enumerating all possible trajectories that all agents in the scenario may take. In this double merging example, a desired set of trajectories suitable for a double merging maneuver can be defined. . Can be the Cartesian product of the following sets:
[0272] in, is the desired target velocity of the host vehicle, is the desired lateral position in lane units, where integers represent lane centers and fractions represent lane boundaries, and is a classification label assigned to each of the n other vehicles. The other vehicle may be assigned “g” if the host vehicle is to yield to the other vehicle, “t” if the host vehicle is to hold the lane relative to the other vehicle, or “o” if the host vehicle is to maintain an offset distance relative to the other vehicle.
[0273] The following describes a set of expectations How can it be transformed into a cost function for the driving trajectory? The driving trajectory can be represented by Indicates that The main vehicle at time The (horizontal, vertical) position (in egocentric units) of the And k = 10. Of course, other values may also be chosen. The cost assigned to the trajectory may comprise a weighted sum of the individual costs assigned to the desired velocity, lateral position, and label assigned to each of the other n vehicles.
[0274] Given the desired rate , the cost of the trajectory associated with the rate is
[0275] .
[0276] Given the desired horizontal position , the cost associated with the desired lateral position is
[0277]
[0278] in It is from the point The distance to lane position l. Regarding the cost due to other vehicles, for any other vehicle, can represent other vehicles in the egocentric unit of the main vehicle, and i can be the earliest point for which there exists j such that and If there is no such point, then i can be set to If the other vehicle is classified as "give way", you can expect , which means that the host vehicle will arrive at the trajectory intersection point at least 0.5 seconds after the other vehicle arrives at the same point. A possible formula for converting the above constraints into costs is .
[0279] Likewise, if another car is classified as "blocking the road," one can expect , which can be converted into cost If another car is classified as "offset", then you can expect , which means that the trajectory of the host vehicle and the trajectory of the offset car do not intersect. This situation can be converted into a cost by penalizing them relative to the distance between the trajectories.
[0280] Assigning weights to each of these costs provides a single objective function for the trajectory planner. A cost that encourages smooth driving can be added to the objective. Furthermore, to ensure the functional safety of the trajectory, hard constraints can be added to the objective. For example, Leave the roadway, and then for any trajectory point of any other vehicle , you can prohibit near ,if Small words.
[0281] In short, strategy It can be decomposed into a mapping from unknowable states to desired sets and a mapping from desired to actual trajectories. The latter mapping is not based on learning and can be achieved by solving an optimization problem whose cost depends on the desired set and whose hard constraints can guarantee the functional safety of the policy.
[0282] The following discussion describes the mapping from unknowable states to desired sets. As mentioned above, in order to comply with functional safety, a system that relies solely on reinforcement learning may suffer from This result can be avoided by using policy gradients to iteratively decompose the problem into a mapping from an (unknowable) state space to a set of desired trajectories, and subsequently to actual trajectories without involving a machine learning-based training system.
[0283] Decision making can be further broken down into semantically meaningful components for various reasons. For example, The size of may be large or even continuous. Figure 11D In the double lane merging scenario described above, . In addition, the gradient estimator may involve the term In such an expression, the variance can grow over the time range T. In some cases, the value of T can be approximately 250, which may be high enough to produce significant variance. Assuming the sampling rate is in the range of 10 Hz and the merging area 1130 is 100 meters, then the preparation for the merge can begin approximately 300 meters before the merging area. If the host vehicle is traveling at a speed of 16 meters per second (about 60 kilometers per hour), the value of T for the episode can be approximately 250.
[0284] Back to the concept of option maps, Figure 11E It is shown that it can be expressed Figure 11D As previously described, an option graph can represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There may be a special node in the graph called a "root" node 1140, which may be the only node with no incoming edges (e.g., decision lines). The decision process can traverse the graph starting from the root node until it reaches a "leaf" node, i.e., a node with no outgoing edges. Each internal node can implement a policy function that selects a child from its available children. There may be a set of traversals on the option graph to a desired set In other words, traversals on the option graph can be automatically converted to Given a node v in the graph, the parameter vector A strategy for selecting the descendants of v can be specified. If θ is all The concatenation of , then, can be defined by traversing from the root of the graph to the leaves , and at each node v, using Defines the strategy for selecting child nodes.
[0285] exist Figure 11E In the double merging option diagram 1139, the root node 1140 may first determine whether the host vehicle is within the merging area (e.g., Figure 11D 1130), or whether the host vehicle is approaching a merging area and needs to prepare for a possible merge. In both cases, the host vehicle may need to decide whether to change lanes (e.g., to the left or right) or whether to remain in the current lane. If the host vehicle has decided to change lanes, the host vehicle may need to decide whether conditions are suitable to continue and execute the lane change maneuver (e.g., at a “Continue” node 1142). If a lane change is not possible, the host vehicle may attempt to “advance” toward the desired lane by aiming to be on the lane markings (e.g., at a node 1144 as part of a negotiation with the vehicle in the desired lane). Alternatively, the host vehicle may choose to “remain” in the same lane (e.g., at a node 1146). This process may determine the lateral position of the host vehicle in a natural manner. For example,
[0286] This allows for a natural way to determine the desired lateral position. For example, if the host vehicle changes lanes from lane 2 to lane 3, the "continue" node may set the desired lateral position to 3, the "hold" node may set the desired lateral position to 2, and the "advance" node may set the desired lateral position to 2.5. Next, the host vehicle may decide whether to maintain the "same" rate (node 1148), "accelerate" (node 1150), or "slow down" (node 1152). Next, the host vehicle may enter a "chain-like" structure 1154 of overtaking other vehicles and setting their semantic meaning to values in the set {g, t, o}. This process may set expectations relative to other vehicles. The parameters of all nodes in the chain may be shared (similar to a recurrent neural network).
[0287] A potential benefit of this option is the interpretability of the results. Another potential benefit is the ability to rely on the set The decomposable structure of , and therefore, the policy at each node can be selected from a small number of possibilities. In addition, this structure can allow reducing the variance of the policy gradient estimator.
[0288] As described above, the length of a segment in a double-merge scenario can be approximately T = 250 steps. This value (or any other suitable value depending on the specific navigation scenario) provides sufficient time to see the consequences of the host vehicle's actions (e.g., if the host vehicle decides to change lanes in preparation for a merge, the host vehicle will only see the benefits after successfully completing the merge). On the other hand, due to the dynamic nature of driving, the host vehicle must make decisions at a sufficiently high frequency (e.g., 10 Hz in the above case).
[0289] The option graph can achieve a reduction in the effective value of T in at least two ways. First, given higher-level decisions, rewards can be defined for lower-level decisions while considering shorter segments. For example, when the host vehicle has selected the "Lane Change" and "Continue" nodes, a policy for assigning semantic meaning to the vehicle can be learned by observing 2- to 3-second segments (meaning T becomes 20-30 instead of 250). Second, for high-level decisions (such as whether to change lanes or stay in the same lane), the host vehicle may not need to make a decision every 0.1 seconds. Alternatively, the host vehicle can make decisions less frequently (e.g., every second) or implement an "option expiration" function, and then the gradient can be calculated only after each option expiration. In both cases, the effective value of T can be an order of magnitude smaller than its original value. In summary, the estimator for each node can depend on a value of T that is an order of magnitude smaller than the original 250 steps, which may immediately lead to lower variance.
[0290] As described above, hard constraints can promote safer driving, and there can be several different types of constraints. For example, static hard constraints can be defined directly based on the sensed state. These hard constraints can include speed bumps, rate limits, road bends, intersections, etc. within the environment of the host vehicle, which may involve one or more constraints on vehicle speed, heading, acceleration, braking (deceleration), etc. Static hard constraints may also include semantic free space, where, for example, the host vehicle is prohibited from driving outside the free space and is prohibited from driving too close to physical obstacles. Static hard constraints can also restrict (e.g., prohibit) maneuvers that are inconsistent with various aspects of the vehicle's kinematic motion. For example, static hard constraints can be used to prohibit maneuvers that may cause the host vehicle to flip, slide, or otherwise lose control.
[0291] Hard constraints may also be associated with the vehicle. For example, a constraint may be employed that requires the vehicle to maintain a longitudinal distance of at least one meter to other vehicles and a lateral distance of at least 0.5 meters to other vehicles. Constraints may also be applied such that the host vehicle will avoid maintaining a collision course with one or more other vehicles. For example, time τ may be a time metric based on a particular scenario. The predicted trajectories of the host vehicle and one or more other vehicles may be considered from the current time to time τ. In the event that the two trajectories intersect, Let be the time for vehicle i to arrive at and leave the intersection. That is, each car will arrive at the intersection when the first part of the car passes through it, and it will take a certain amount of time before the last part of the car passes through it. This amount of time separates the arrival time from the departure time. Assume (i.e., the arrival time of vehicle 1 is less than the arrival time of vehicle 2), then we will want to ensure that vehicle 1 has left the intersection before vehicle 2 arrives. Otherwise, a collision will result. Therefore, a hard constraint can be implemented so that Furthermore, to ensure that vehicle 1 and vehicle 2 do not miss each other by a minimum amount, an additional safety margin can be obtained by including a buffer time (e.g., 0.5 seconds or another appropriate value) in the constraint. The hard constraint related to the predicted crossing trajectories of the two vehicles can be expressed as .
[0292] The amount of time τ that the trajectories of the host vehicle and one or more other vehicles are tracked can vary. However, in intersection scenarios, where the velocity may be lower, May be longer, and can be defined so that the host vehicle will be less than Enter and exit intersections in seconds.
[0293] Of course, applying hard constraints to vehicle trajectories requires predicting the trajectories of those vehicles. For the host vehicle, trajectory prediction may be relatively straightforward, as the host vehicle typically already understands and is in fact planning the desired trajectory at any given time. Relative to other vehicles, predicting their trajectories may be less straightforward. For other vehicles, the baseline calculation used to determine the predicted trajectory may rely on the current velocity and heading of the other vehicles, for example, as determined based on analysis of image streams captured by one or more cameras and / or other sensors (radar, lidar, acoustics, etc.) on the host vehicle.
[0294] However, there may be some exceptions that may simplify the problem or at least provide increased confidence in the trajectory predicted for the other vehicle. For example, for structured roads where lane indications are present and where give-way rules may be present, the trajectory of the other vehicle may be based at least in part on the position of the other vehicle relative to the lane and on the applicable give-way rules. Thus, in some cases, when the lane structure is observed, it may be assumed that the vehicle in the next lane will respect the lane boundaries. That is, the host vehicle may assume that the vehicle in the next lane will remain in its lane unless evidence is observed (e.g., a signal light, strong lateral motion, movement across a lane boundary) indicating that the vehicle in the next lane will cut into the host vehicle's lane.
[0295] Other situations can also provide clues about the intended trajectory of other vehicles. For example, in situations where the host vehicle may have the right of way, such as at stop signs, traffic lights, roundabouts, etc., it can be assumed that other vehicles will respect that right of way. Therefore, unless evidence of a violation is observed, it can be assumed that other vehicles will continue along a trajectory that respects the right of way held by the host vehicle.
[0296] Hard constraints can also be applied to pedestrians in the host vehicle's environment. For example, a buffer distance can be established with respect to pedestrians, prohibiting the host vehicle from traveling closer than a specified buffer distance to any observed pedestrian. The pedestrian buffer distance can be any suitable distance. In some embodiments, the buffer distance can be at least one meter relative to an observed pedestrian.
[0297] Similar to the case with vehicles, hard constraints can also be applied with respect to the relative motion between pedestrians and the host vehicle. For example, the trajectory of a pedestrian can be monitored relative to the projected trajectory of the host vehicle (based on heading direction and velocity). Given a particular pedestrian trajectory, where for each point p on the trajectory, t(p) can represent the time required for the pedestrian to reach point p. In order to maintain the required buffer distance of at least 1 meter from the pedestrian, t(p) must be greater than the time the host vehicle will reach point p (with enough time difference so that the host vehicle passes at least one meter in front of the pedestrian), or t(p) must be less than the time the host vehicle will reach point p (for example, if the host vehicle brakes to make way for the pedestrian). Nevertheless, in the latter example, the hard constraint may require the host vehicle to arrive at point p enough time after the pedestrian so that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter. Of course, there may be exceptions to the hard constraints for pedestrians. For example, in situations where the host vehicle has the right of way or is moving very slowly, and there is no observed evidence that the pedestrian will refuse to yield to the host vehicle or will otherwise navigate toward the host vehicle, the hard pedestrian constraint can be relaxed (e.g., to a smaller buffer of at least 0.75 meters or 0.50 meters).
[0298] In some examples, constraints may be relaxed if it is determined that not all constraints can be satisfied. For example, if the road is too narrow to allow for two curbs or the required spacing between a curb and a parked vehicle (e.g., 0.5 meters), one or more constraints may be relaxed if there are mitigating circumstances. For example, if there are no pedestrians (or other objects) on the sidewalk, one may slowly advance at 0.1 meters from the curb. In some embodiments, constraints may be relaxed if relaxing the constraints improves the user experience. For example, to avoid potholes, constraints may be relaxed to allow vehicles to get closer to the edge of the lane, curb, or pedestrians than would normally be allowed. In addition, when determining which constraints to relax, in some embodiments, the one or more constraints selected to be relaxed are those that are considered to have the smallest available negative impact on safety. For example, constraints on how close a vehicle can get to a curb or concrete barrier may be relaxed before constraints on proximity to other vehicles are relaxed. In some embodiments, pedestrian constraints may be the last to be relaxed, or in some cases may never be relaxed.
[0299] Figure 12 Examples of scenes that may be captured and analyzed during navigation of a host vehicle are shown. For example, the host vehicle may include a navigation system as described above (e.g., system 100), which may receive a plurality of images representing the host vehicle's environment from a camera associated with the host vehicle (e.g., at least one of image capture device 122, image capture device 124, and image capture device 126). Figure 12The scene shown in is an example of one of the images that may be captured at time t from the environment of a host vehicle traveling in lane 1210 along a predicted trajectory 1212. The navigation system may include at least one processing device (e.g., including any EyeQ processor described above or other device) that is specifically programmed to receive multiple images and analyze the images to determine an action in response to the scene. Specifically, as Figure 8 As shown, at least one processing device may implement a sensing module 801, a driving strategy module 803, and a control module 805. The sensing module 801 may be responsible for collecting and outputting image information collected from a camera, and providing this information in the form of a recognized navigation state to the driving strategy module 803, which may constitute a trained navigation system that has been trained through machine learning techniques (such as supervised learning, reinforcement learning, etc.). Based on the navigation state information provided by the sensing module 801 to the driving strategy module 803, the driving strategy module 803 (e.g., by implementing the option map method described above) may generate a desired navigation action for the host vehicle to perform in response to the recognized navigation state.
[0300] In some embodiments, at least one processing device may convert the desired navigation action directly into a navigation command using, for example, a control module 805. However, in other embodiments, hard constraints may be applied such that the desired navigation action provided by the driving strategy module 803 is tested against various predetermined navigation constraints that may be involved in the scenario and the desired navigation action. For example, where the driving strategy module 803 outputs a desired navigation action that will cause the host vehicle to follow trajectory 1212, the navigation action may be tested against one or more hard constraints associated with various aspects of the host vehicle's environment. For example, the captured image 1201 may show a curb 1213, a pedestrian 1215, a target vehicle 1217, and a stationary object (e.g., an overturned box) present in the scene. Each of these may be associated with one or more hard constraints. For example, the curb 1213 may be associated with a static constraint that prohibits the host vehicle from navigating into or past the curb and onto the sidewalk 1214. The curb 1213 may also be associated with a barrier envelope that defines a distance (e.g., a buffer zone) extending away from and along the curb (e.g., 0.1 m, 0.25 m, 0.5 m, 1 m, etc.) that defines a no-navigation zone for the host vehicle. Of course, static constraints may also be associated with other types of roadside boundaries (e.g., guardrails, concrete bollards, traffic cones, bridge pylons, or any other type of roadside obstacle).
[0301] It should be noted that distance and range can be determined by any suitable method. For example, in some embodiments, distance information can be provided by an onboard radar and / or lidar system. Alternatively or additionally, distance information can be derived from an analysis of one or more images captured from the environment of the host vehicle. For example, a number of pixels of an identified object represented in an image can be determined and compared to the known field of view and focal length geometry of the image capture device to determine scale and distance. For example, speed and acceleration can be determined by observing the change in scale between objects from image to image over a known time interval. The analysis can indicate the direction of motion toward or away from the host vehicle, and how fast the object is moving away from or toward the host vehicle. The crossing speed can be determined by analyzing the change in the X-coordinate position of the object from one image to another over a known time period.
[0302] Pedestrian 1215 may be associated with a pedestrian envelope that defines a buffer zone 1216. In some cases, a hard constraint may be imposed that prohibits the host vehicle from navigating within a distance of one meter from pedestrian 1215 (in any direction relative to pedestrian 1215). Pedestrian 1215 may also define the location of a pedestrian influence zone 1220. This influence zone may be associated with a constraint that limits the host vehicle's speed within the influence zone. The influence zone may extend from pedestrian 1215 to 5 meters, 10 meters, 20 meters, and so on. Each graduation of the influence zone may be associated with a different speed limit. For example, within an area from one to five meters from pedestrian 1215, the host vehicle may be limited to a first speed (e.g., 10 mph, 20 mph, etc.), which may be less than the speed limit in the pedestrian influence zone extending from 5 to 10 meters. Any graduation may be used for each level of the influence zone. In some embodiments, the first graduation may be narrower than from one to five meters and may extend only from one to two meters. In other embodiments, a first level of the influence zone may extend from one meter (the boundary of the no-navigation zone surrounding the pedestrian) to a distance of at least 10 meters. A second level may then extend from ten meters to at least about twenty meters. The second level may be associated with a maximum travel speed of the host vehicle that is greater than the maximum travel speed associated with the first level of the pedestrian influence zone.
[0303] One or more stationary object constraints may also be associated with a scene detected in the host vehicle's environment. For example, in image 1201, at least one processing device may detect a stationary object, such as a box 1219, located in the roadway. The detected stationary object may include various objects, such as at least one of a tree, a pole, a road sign, or an object in the roadway. One or more predetermined navigation constraints may be associated with the detected stationary object. For example, such a constraint may include a stationary object envelope, wherein the stationary object envelope defines a buffer zone about the object within which navigation of the host vehicle may be prohibited. At least a portion of the buffer zone may extend a predetermined distance from the edge of the detected stationary object. For example, in the scene represented by image 1201, a buffer zone of at least 0.1 meters, 0.25 meters, 0.5 meters, or more may be associated with box 1219, such that the host vehicle will pass at least a distance (e.g., a buffer zone distance) to the right or left of the box to avoid collision with the detected stationary object.
[0304] Predefined hard constraints may also include one or more target vehicle constraints. For example, a target vehicle 1217 may be detected in image 1201. To ensure that the host vehicle does not collide with target vehicle 1217, one or more hard constraints may be employed. In some cases, the target vehicle envelope may be associated with a single buffer zone distance. For example, a buffer zone may be defined by a distance of 1 meter around the target vehicle in all directions. The buffer zone may define an area extending at least one meter from the target vehicle, into which the host vehicle is prohibited from navigating.
[0305] However, the envelope around target vehicle 1217 need not be defined by a fixed buffer distance. In some cases, the predefined hard constraints associated with the target vehicle (or any other movable object detected in the host vehicle's environment) may depend on the orientation of the host vehicle relative to the detected target vehicle. For example, in some cases, the longitudinal buffer distance (e.g., the distance extending from the target vehicle to the front or rear of the host vehicle—such as when the host vehicle is traveling toward the target vehicle) may be at least one meter. The lateral buffer distance (e.g., the distance extending from the target vehicle to either side of the host vehicle—such as when the host vehicle is traveling in the same or opposite direction as the target vehicle so that one side of the host vehicle will pass adjacent to one side of the target vehicle) may be at least 0.5 meters.
[0306] As described above, other constraints may also be involved by detecting target vehicles or pedestrians in the host vehicle's environment. For example, the predicted trajectories of the host vehicle and the target vehicle 1217 may be considered, and in the event that the two trajectories intersect (e.g., at the intersection point 1230), a hard constraint may require that or , where the host vehicle is vehicle 1 and the target vehicle 1217 is vehicle 2. Similarly, the trajectory of pedestrian 1215 (based on heading direction and velocity) can be monitored relative to the projected trajectory of the host vehicle. Given a particular pedestrian trajectory, for each point p on the trajectory, t(p) will represent the pedestrian's arrival at point p (i.e., Figure 12 In order to maintain the required buffer distance of at least 1 meter from the pedestrian, t(p) must be greater than the time the host vehicle will arrive at point p (with enough time difference so that the host vehicle passes at least one meter ahead of the pedestrian), or t(p) must be less than the time the host vehicle will arrive at point p (for example, if the host vehicle brakes to make way for the pedestrian). Nevertheless, in the latter example, the hard constraint will require the host vehicle to arrive at point p enough after the pedestrian so that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter.
[0307] Other hard constraints may also be employed. For example, at least in some cases, a maximum deceleration rate for the host vehicle may be employed. This maximum deceleration rate may be determined based on the detected distance to a target vehicle following the host vehicle (e.g., using images collected from a rear-facing camera). Hard constraints may include mandatory stops at sensed road crossings or railroad crossings, or other applicable constraints.
[0308] In the event that the analysis of the scene in the host vehicle's environment indicates that one or more predefined navigation constraints may be involved, these constraints can be imposed relative to one or more planned navigation actions of the host vehicle. For example, in the event that the analysis of the scene causes the driving strategy module 803 to return a desired navigation action, the desired navigation action can be tested against one or more involved constraints. If the desired navigation action is determined to violate any aspect of the involved constraints (e.g., if the desired navigation action will drive the host vehicle within 0.7 meters of a pedestrian 1215, where a predefined hard constraint requires the host vehicle to maintain a distance of at least 1.0 meters from the pedestrian 1215), at least one modification can be made to the desired navigation action based on the one or more predefined navigation constraints. Adjusting the desired navigation action in this manner can provide an actual navigation action for the host vehicle in accordance with the constraints involved in the particular scene detected in the host vehicle's environment.
[0309] After determining the actual navigational maneuver of the host vehicle, the navigational maneuver may be implemented by causing at least one adjustment of a navigational actuator of the host vehicle in response to the determined actual navigational maneuver of the host vehicle. Such a navigational actuator may include at least one of a steering mechanism, a brake, or an accelerator of the host vehicle.
[0310] Precedence Constraints
[0311] As described above, the navigation system may employ various hard constraints to ensure safe operation of the host vehicle. Constraints may include a minimum safe driving distance relative to pedestrians, target vehicles, roadblocks, or detected objects, a maximum travel speed when passing within the influence zone of a detected pedestrian, or a maximum deceleration rate for the host vehicle, among others. These constraints may be imposed on trained systems based on machine learning (supervised, reinforcement, or a combination thereof), but they may also be useful for untrained systems (e.g., those that employ algorithms to directly process expected situations occurring in scenarios from the host vehicle's environment).
[0312] In either case, there may be a hierarchy of constraints. In other words, some navigation constraints take precedence over others. Therefore, if a situation arises where no navigation action is available that would satisfy all involved constraints, the navigation system may first determine the available navigation action that implements the highest-priority constraint. For example, the system may cause the vehicle to avoid a pedestrian first, even if navigating to avoid the pedestrian would result in a collision with another vehicle or object detected in the road. In another example, the system may cause the vehicle to ride on a curb to avoid the pedestrian.
[0313] Figure 13 A flow chart illustrating an algorithm for implementing a hierarchy of associated constraints determined based on an analysis of a scene in the host vehicle's environment is provided. For example, at step 1301, at least one processing device associated with a navigation system (e.g., an EyeQ processor, etc.) may receive a plurality of images representing the host vehicle's environment from a camera mounted on the host vehicle. By analyzing the one or more images representing the scene of the host vehicle's environment at step 1303, a navigation state associated with the host vehicle may be identified. In addition to various other attributes of the scene, the navigation state may indicate, for example, that the host vehicle is traveling along a two-lane road 1210, such as Figure 12 As shown, a target vehicle 1217 is moving through an intersection in front of the host vehicle, a pedestrian 1215 is waiting to cross the path of the host vehicle, and an object 1219 exists in front of the host vehicle's lane.
[0314] In step 1305, one or more navigation constraints related to the navigation state of the host vehicle may be determined. For example, after analyzing a scene in the environment of the host vehicle represented by one or more captured images, at least one processing device may determine one or more navigation constraints related to objects, vehicles, pedestrians, etc. identified through image analysis of the captured images. In some embodiments, the at least one processing device may determine at least a first predefined navigation constraint and a second predefined navigation constraint related to the navigation state, and the first predefined navigation constraint may be different from the second predefined navigation constraint. For example, the first navigation constraint may relate to one or more target vehicles detected in the environment of the host vehicle, and the second navigation constraint may relate to pedestrians detected in the environment of the host vehicle.
[0315] In step 1307, at least one processing device may determine the priority associated with the constraints identified in step 1305. In the depicted example, a second predefined navigation constraint involving pedestrians may have a higher priority than a first predefined navigation constraint involving a target vehicle. While the priority associated with navigation constraints can be determined or assigned based on a variety of factors, in some embodiments, the priority of a navigation constraint may be related to its relative importance from a safety perspective. For example, while it may be important to comply with or satisfy all implemented navigation constraints in as many situations as possible, some constraints may be associated with a higher safety risk than others and, therefore, may be assigned a higher priority. For example, a navigation constraint requiring the host vehicle to maintain a spacing of at least 1 meter from a pedestrian may have a higher priority than a constraint requiring the host vehicle to maintain a spacing of at least 1 meter from a target vehicle. This may be because a collision with a pedestrian may have more severe consequences than a collision with another vehicle. Similarly, maintaining space between the host vehicle and the target vehicle may have a higher priority than a constraint requiring the host vehicle to avoid boxes in the road, drive below a certain speed over speed bumps, or expose the host vehicle's occupants to no more than a maximum acceleration level.
[0316] While the driving strategy module 803 is designed to maximize safety by satisfying navigation constraints associated with a particular scenario or navigation state, in some cases, it may not be practical to satisfy each of the associated constraints. In such cases, as shown in step 1309, the priority of each associated constraint may be used to determine which associated constraint should be satisfied first. Continuing with the above example, in a situation where it is not possible to satisfy both the pedestrian gap constraint and the target vehicle gap constraint, and only one of the constraints can be satisfied, the higher-priority pedestrian gap constraint may cause that constraint to be satisfied before attempting to maintain a gap to the target vehicle. Therefore, under normal circumstances, as shown in step 1311, in a situation where both the first predefined navigation constraint and the second predefined navigation constraint can be satisfied, at least one processing device may determine, based on the identified navigation state of the host vehicle, a first navigation action for the host vehicle that satisfies both the first predefined navigation constraint and the second predefined navigation constraint. However, in other cases, when not all involved constraints can be satisfied, as shown in step 1313, when the first predefined navigation constraint and the second predefined navigation constraint cannot be satisfied at the same time, at least one processing device can determine, based on the identified navigation state, a second navigation action of the host vehicle that satisfies the second predefined navigation constraint (i.e., a higher priority constraint) but does not satisfy the first predefined navigation constraint (having a lower priority than the second navigation constraint).
[0317] Next, in step 1315, to implement the determined navigation maneuver of the host vehicle, the at least one processing device may cause at least one adjustment of a navigation actuator of the host vehicle in response to the determined first navigation maneuver or the determined second navigation maneuver of the host vehicle. As described in the previous example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.
[0318] Constraint relaxation
[0319] As described above, navigation constraints can be imposed for safety purposes. Constraints can include a minimum safe driving distance relative to a pedestrian, target vehicle, roadblock, or detected object, a maximum driving speed when passing within the influence zone of a detected pedestrian, or a maximum deceleration rate for the host vehicle, among others. These constraints can be imposed on either a learning or non-learning navigation system. In some cases, these constraints can be relaxed. For example, in a scenario where the host vehicle slows down or stops near a pedestrian and then slowly advances to communicate its intention to pass the pedestrian, the pedestrian's response can be detected from the acquired imagery. If the pedestrian responds by remaining still or stopping (and / or if eye contact with the pedestrian is sensed), it can be understood that the pedestrian recognizes the navigation system's intention to pass the pedestrian. In this case, the system can relax one or more predefined constraints and implement less stringent constraints (e.g., allowing the vehicle to navigate within 0.5 meters of the pedestrian, rather than within the more stringent 1-meter boundary).
[0320] Figure 14 A flow chart for implementing control of a host vehicle based on the relaxation of one or more navigation constraints is provided. In step 1401, at least one processing device may receive a plurality of images representing an environment of the host vehicle from a camera associated with the host vehicle. In step 1403, analysis of the images may enable identification of a navigation state associated with the host vehicle. In step 1405, at least one processor may determine a navigation constraint associated with the navigation state of the host vehicle. The navigation constraint may include a first predefined navigation constraint related to at least one aspect of the navigation state. In step 1407, analysis of the plurality of images may reveal the presence of at least one navigation constraint relaxation factor.
[0321] A navigation constraint relaxation factor may include any suitable indicator that one or more navigation constraints may be suspended, changed, or otherwise relaxed in at least one aspect. In some embodiments, at least one navigation constraint relaxation factor may include a determination (based on image analysis) that the pedestrian's eyes are looking in the direction of the host vehicle. In this case, it can be more safely assumed that the pedestrian is aware of the host vehicle. Therefore, the confidence that the pedestrian will not engage in unexpected maneuvers that result in the pedestrian moving into the path of the host vehicle may be higher. Other constraint relaxation factors may also be used. For example, at least one navigation constraint relaxation factor may include: a pedestrian determined to be not moving (e.g., a pedestrian assumed to be unlikely to enter the path of the host vehicle); or a pedestrian whose movement is determined to be slowing down. Navigation constraint relaxation factors may also include more complex actions, such as a pedestrian determined to be not moving after the host vehicle has come to a stop and then resumed movement. In this case, it can be assumed that the pedestrian understands that the host vehicle has the right of way, and that the pedestrian coming to a stop may indicate the pedestrian's intention to give way to the host vehicle. Other situations that may cause one or more constraints to be relaxed include the type of curb (e.g., a low curb or a curb with a gradual slope may allow a relaxed distance constraint), the absence of pedestrians or other objects on the sidewalk, vehicles with their engines not running may have a relaxed distance, or situations where pedestrians are moving towards and / or away from the area toward which the host vehicle is traveling.
[0322] In the event that a navigation constraint relaxation factor is identified (e.g., in step 1407), a second navigation constraint may be determined or developed in response to the detection of the constraint relaxation factor. The second navigation constraint may be different from the first navigation constraint and may include at least one characteristic that is relaxed relative to the first navigation constraint. The second navigation constraint may include a newly generated constraint based on the first constraint, wherein the newly generated constraint includes at least one modification that relaxes the first constraint in at least one aspect. Alternatively, the second constraint may constitute a predetermined constraint that is less stringent than the first navigation constraint in at least one aspect. In some embodiments, such a second constraint may be retained for use only in situations where a constraint relaxation factor is identified in the environment of the host vehicle. Regardless of whether the second constraint is newly generated or selected from a set of fully or partially available predetermined constraints, applying the second navigation constraint in place of the more stringent first navigation constraint (which may be applied in situations where no relevant navigation constraint relaxation factor is detected) may be referred to as constraint relaxation and may be completed in step 1409.
[0323] If at least one constraint relaxation factor is detected in step 1407 and at least one constraint has been relaxed in step 1409, a navigation action of the host vehicle may be determined in step 1411. The navigation action of the host vehicle may be based on the identified navigation state and may satisfy the second navigation constraint. The navigation action may be implemented in step 1413 by causing at least one adjustment of a navigation actuator of the host vehicle in response to the determined navigation action.
[0324] As described above, the use of navigation constraints and relaxed navigation constraints can be employed by navigation systems that are trained (e.g., through machine learning) or untrained (e.g., a navigation system that is programmed to respond with a predetermined action in response to a specific navigation state). In the case of using a trained navigation system, the availability of relaxed navigation constraints for certain navigation situations can represent a mode switch from a trained system response to an untrained system response. For example, a trained navigation network can determine an original navigation action for the host vehicle based on a first navigation constraint. However, the action taken by the vehicle can be an action that is different from the navigation action that satisfies the first navigation constraint. Instead, the action taken can satisfy a second, more relaxed navigation constraint and can be an action developed by an untrained system (e.g., as a response to detecting a specific condition in the host vehicle's environment, such as the presence of a navigation constraint relaxation factor).
[0325] There are many examples of navigation constraints that can be relaxed in response to a constraint relaxation factor being detected in the host vehicle's environment. For example, where a predefined navigation constraint includes a buffer zone associated with a detected pedestrian, and at least a portion of the buffer zone extends a certain distance from the detected pedestrian, the relaxed navigation constraint (newly generated, recalled from memory from a predetermined set, or generated as a relaxed version of a pre-existing constraint) can include a different or modified buffer zone. For example, the different or modified buffer zone can have a smaller distance relative to the pedestrian than the original or unmodified buffer zone relative to the detected pedestrian. Thus, in view of the relaxed constraint, the host vehicle can be allowed to navigate closer to the detected pedestrian, where an appropriate constraint relaxation factor is detected in the host vehicle's environment.
[0326] As described above, the relaxed characteristic of the navigation constraints may include a reduced width of a buffer zone associated with at least one pedestrian. However, the relaxed characteristic may also include a reduced width of a buffer zone associated with a target vehicle, a detected object, a roadside obstacle, or any other object detected in the host vehicle's environment.
[0327] At least one relaxed characteristic may also include other types of modifications to navigation constraint characteristics. For example, a relaxed characteristic may include an increase in the velocity associated with at least one predefined navigation constraint. A relaxed characteristic may also include an increase in the maximum allowable deceleration / acceleration associated with at least one predefined navigation constraint.
[0328] While constraints may be relaxed in some cases as described above, in other cases, navigation constraints may be strengthened. For example, in some cases, the navigation system may determine that conditions warrant an enhancement to a set of normal navigation constraints. Such enhancements may include adding new constraints to a predefined set of constraints or adjusting one or more aspects of the predefined constraints. The additions or adjustments may result in more conservative navigation relative to the predefined set of constraints applicable under normal driving conditions. Conditions that may warrant a constraint enhancement may include sensor failure, adverse environmental conditions (rain, snow, fog, or other conditions associated with reduced visibility or reduced vehicle traction), etc.
[0329] Figure 15 A flow chart is provided for implementing control of a host vehicle based on augmentation of one or more navigation constraints. In step 1501, at least one processing device may receive a plurality of images representing an environment of the host vehicle from a camera associated with the host vehicle. In step 1503, analysis of the images may enable identification of a navigation state associated with the host vehicle. In step 1505, at least one processor may determine a navigation constraint associated with the navigation state of the host vehicle. The navigation constraint may include a first predefined navigation constraint related to at least one aspect of the navigation state. In step 1507, analysis of the plurality of images may reveal the presence of at least one navigation constraint augmentation factor.
[0330] The navigation constraints involved may include those mentioned above (e.g., Figure 12) any navigation constraint or any other suitable navigation constraint. A navigation constraint enhancement factor may include any indicator that one or more navigation constraints may be supplemented / enhanced in at least one aspect. Supplementation or enhancement of navigation constraints may be performed on a per-set basis (e.g., by adding a new navigation constraint to a predetermined set of constraints) or on a per-constraint basis (e.g., by modifying a particular constraint so that the modified constraint is more restrictive than the original constraint, or adding a new constraint corresponding to a predetermined constraint, wherein the new constraint is more restrictive in at least one aspect than the corresponding constraint). Additionally or alternatively, supplementation or enhancement of navigation constraints may involve selection from a predetermined set of constraints on a hierarchical basis. For example, a set of enhanced constraints may be selected based on whether a navigation enhancement factor is detected in or relative to the host vehicle's environment. In a normal situation where no enhancement factor is detected, the navigation constraints involved may be extracted from the constraints applicable to the normal situation. On the other hand, in the event that one or more constraint enhancement factors are detected, the constraints involved may be extracted from the enhanced constraints that are generated or predefined relative to the one or more enhancement factors. The enhanced constraints may be more restrictive in at least one aspect than the corresponding constraints applicable under normal conditions.
[0331] In some embodiments, at least one navigation constraint enhancing factor may include detecting (e.g., based on image analysis) the presence of ice, snow, or water on a road surface in the host vehicle's environment. For example, such a determination may be based on detecting: areas of higher reflectivity than would be expected for a dry roadway (e.g., indicating ice or water on the roadway); white areas on the roadway indicating the presence of snow; shadows on the roadway consistent with the presence of longitudinal grooves on the roadway (e.g., tire tracks in the snow); water droplets or ice / snow particles on the host vehicle's windshield; or any other suitable indicator of the presence of water or ice / snow on the road surface.
[0332] At least one navigation constraint enhancing factor may also include detection of particles on an exterior surface of a windshield of the host vehicle. Such particles may impair image quality of one or more image capture devices associated with the host vehicle. Although described with respect to the windshield of the host vehicle, in relation to a camera mounted behind the windshield of the host vehicle, detection of particles on other surfaces (e.g., a lens or lens cover of a camera, a headlight lens, a rear windshield, a taillight lens, or any other surface of the host vehicle that is visible to (or detected by) an image capture device associated with the host vehicle) may also indicate the presence of a navigation constraint enhancing factor.
[0333] A navigation constraint enhancing factor may also be detected as a property of one or more image acquisition devices. For example, a detected degradation in the image quality of one or more images captured by an image capture device (e.g., a camera) associated with the host vehicle may also constitute a navigation constraint enhancing factor. The degradation in image quality may be associated with a hardware failure or partial hardware failure associated with the image capture device or a component associated with the image capture device. Such degradation in image quality may also be caused by environmental conditions. For example, the presence of smoke, fog, rain, snow, etc. in the air surrounding the host vehicle may also result in reduced image quality relative to roads, pedestrians, target vehicles, etc. that may be present in the environment of the host vehicle.
[0334] Navigation constraint enhancing factors may also relate to other aspects of the host vehicle. For example, in some cases, a navigation constraint enhancing factor may include a detected failure or partial failure of a system or sensor associated with the host vehicle. Such enhancing factors may include, for example, a detected failure or partial failure of a speed sensor, GPS receiver, accelerometer, camera, radar, lidar, brakes, tires, or any other system associated with the host vehicle that may affect the host vehicle's ability to navigate relative to the navigation constraints associated with the host vehicle's navigational state.
[0335] In the event that the presence of a navigation constraint enhancing factor is identified (e.g., in step 1507), a second navigation constraint may be determined or developed in response to the detection of the constraint enhancing factor. The second navigation constraint may be different from the first navigation constraint, and the second navigation constraint may include at least one characteristic that is enhanced relative to the first navigation constraint. The second navigation constraint may be more restrictive than the first navigation constraint because the detection of the constraint enhancing factor in the environment of the host vehicle or associated with the host vehicle may indicate that the host vehicle may have at least one navigation capability that is reduced relative to normal operating conditions. Such reduced capability may include lower road traction (e.g., ice, snow, or water on the roadway; reduced tire pressure, etc.); impaired vision (e.g., rain, snow, dust, smoke, fog, etc. that reduces the quality of captured images); impaired detection capability (e.g., sensor failure or partial failure, reduced sensor performance, etc.), or any other reduction in the ability of the host vehicle to navigate in response to the detected navigation condition.
[0336] If at least one constraint-enhancing factor is detected in step 1507 and at least one constraint has been enhanced in step 1509, a navigation action for the host vehicle may be determined in step 1511. The navigation action for the host vehicle may be based on the identified navigation state and may satisfy the second navigation (i.e., enhanced) constraint. The navigation action may be implemented in step 1513 by causing at least one adjustment to be made to a navigation actuator of the host vehicle in response to the determined navigation action.
[0337] As discussed, the use of navigation constraints and augmented navigation constraints can be employed by either trained (e.g., through machine learning) or untrained (e.g., systems programmed to respond with predetermined actions in response to specific navigation conditions) navigation systems. In the case of a trained navigation system, the availability of augmented navigation constraints for certain navigation situations may represent a paradigm shift from a trained system response to an untrained system response. For example, a trained navigation network may determine an original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle may be a different action from the navigation action that satisfies the first navigation constraint. Instead, the action taken may satisfy a second augmented navigation constraint and may be an action developed by an untrained system (e.g., as a response to detecting a specific condition in the host vehicle's environment, such as the presence of a navigation constraint augmenting factor).
[0338] There are many examples of navigation constraints that can be generated, supplemented, or enhanced in response to the detection of a constraint-enhancing factor in the host vehicle's environment. For example, where a predefined navigation constraint includes a buffer zone associated with a detected pedestrian, object, vehicle, etc., and at least a portion of the buffer zone extends a certain distance from the detected pedestrian / object / vehicle, the enhanced navigation constraint (newly developed, recalled from memory from a predetermined set, or generated as an enhanced version of a pre-existing constraint) can include a different or modified buffer zone. For example, the different or modified buffer zone can have a greater distance relative to the detected pedestrian / object / vehicle than the original or unmodified buffer zone relative to the detected pedestrian / object / vehicle. Thus, in view of the enhanced constraint, the host vehicle can be forced to navigate further away from the detected pedestrian / object / vehicle, where an appropriate constraint-enhancing factor is detected in or relative to the host vehicle's environment.
[0339] At least one enhanced characteristic may also include other types of modifications to navigation constraint characteristics. For example, the enhanced characteristic may include a reduction in speed associated with at least one predefined navigation constraint. The enhanced characteristic may also include a reduction in the maximum allowable deceleration / acceleration associated with at least one predefined navigation constraint.
[0340] Navigation based on long-term planning
[0341] In some embodiments, the disclosed navigation system can not only respond to a navigation state detected in the host vehicle's environment but also determine one or more navigation actions based on long-term planning. For example, the system can consider the potential impact of one or more navigation actions available as options for navigating relative to the detected navigation state on future navigation states. Considering the impact of available actions on future states allows the navigation system to determine navigation actions based not only on the currently detected navigation state but also on long-term planning. Navigation using long-term planning techniques may be particularly useful in situations where the navigation system employs one or more reward functions as a technique for selecting navigation actions from available options. Potential rewards can be analyzed relative to available navigation actions that can be taken in response to the detected current navigation state of the host vehicle. However, further, potential rewards can also be analyzed relative to actions that can be taken in response to future navigation states that are expected to result from available actions in response to the current navigation state. Thus, the disclosed navigation system can, in some cases, select a navigation action from available actions that can be taken in response to the detected navigation state, even when the selected navigation action may not yield the highest reward. This is particularly true when the system determines that a selected action may result in a future navigation state that gives rise to one or more potential navigation actions that offer a higher reward than the selected action or, in some cases, than any action available relative to the current navigation state. This principle can be more simply expressed as taking a less favorable action now in order to yield a higher reward option in the future. Thus, the disclosed navigation system capable of long-term planning can select suboptimal short-term actions where the long-term prediction indicates that a short-term loss of reward will result in an increase in long-term reward.
[0342] Typically, autonomous driving applications may involve a series of planning problems in which the navigation system may decide on immediate actions in order to optimize a longer-term goal. For example, when a vehicle encounters a merging situation at a roundabout, the navigation system may decide on immediate acceleration or braking commands in order to initiate navigation to the roundabout. While the immediate action on the detected navigation state at the roundabout may involve acceleration or braking commands in response to the detected state, the long-term goal is a successful merge, and the long-term impact of the selected command is the success / failure of the merge. The planning problem can be solved by breaking the problem into two stages. First, supervised learning can be applied to predict the near future based on the present (assuming that the predictor will be differentiable with respect to the present representation). Second, a recurrent neural network can be used to model the complete trajectory of the agent, where unexplained factors are modeled as (additive) input nodes. This can allow the use of supervised learning techniques and direct optimization of the recurrent neural network to determine the solution to the long-term planning problem. This approach can also enable robust policies to be learned by incorporating adversarial elements into the environment.
[0343] The two most fundamental elements of autonomous driving systems are sensing and planning. Sensing seeks a compact representation of the current state of the environment, while planning determines which actions to take to optimize future goals. Supervised machine learning techniques are useful for solving sensing problems. Machine learning algorithmic frameworks, particularly reinforcement learning (RL) frameworks such as those mentioned above, can also be used for planning.
[0344] RL can be executed in a sequence of consecutive rounds. In round t, the planner (also known as the agent or driving policy module 803) can observe the state , which represents the agent and the environment. Then it should decide the action After performing the action, the agent receives an immediate reward and move to a new state As an example, the host vehicle may include an adaptive cruise control (ACC) system, where the vehicle should autonomously implement acceleration / braking to maintain a proper distance from the vehicle ahead while maintaining a smooth drive. This state can be modeled as a pair of , where x t is the distance to the vehicle ahead, and v t is the speed of the host vehicle relative to the speed of the preceding vehicle. Will be the acceleration command (where a t <0 then the host vehicle slows down). The reward can be determined by (reflects the smoothness of driving) and s t(reflecting the function of maintaining a safe distance between the main vehicle and the preceding vehicle). The goal of the planner is to maximize the cumulative reward (which can be up to the time horizon or the discounted sum of future rewards). To do this, the planner can rely on the strategy , which maps states to actions.
[0345] Supervised learning (SL) can be viewed as a special case of RL, where s t is sampled from some distribution on S, and the reward function can have in the form of is the loss function, and the learner observes y t The value of y t To view the status t There are several differences between general RL models and special case SL, and these differences make the general RL problem more challenging.
[0346] In some SL situations, the actions (or predictions) taken by the learner may have no effect on the environment. In other words, s t+1 and a t are independent. This has two important implications. First, in SL, the sample can be collected in advance before we can start searching for a policy (or predictor) that has good accuracy with respect to the sample. In contrast, in RL, the state s t+1 Typically depends on the action taken (and the previous state), which in turn depends on the policy used to generate that action. This will make the data generation process bound to the policy learning process. Second, because in SL actions do not affect the environment, choosing a t The contribution to the performance of π is local. Specifically, a t Only affects the value of the immediate reward. In contrast, in RL, the action taken in round t may have a long-term impact on the reward value in future rounds.
[0347] In SL, knowledge of the “correct” answer t , together with rewards The shape together can provide for a t Full knowledge of the rewards of all possible choices of , which can enable the calculation of rewards relative to a t In contrast, in RL, the "one-time" value of the reward may be all that can be observed for a specific choice of action taken. This can be called "bandit" feedback. This is one of the most important reasons for the need for "exploration" as part of long-term navigation planning, because in RL-based systems, if only "bandit" feedback is available, the system may not always know whether the action taken is the best action to take.
[0348] Many RL algorithms rely at least in part on a mathematical optimization model called a Markov Decision Process (MDP). The Markov assumption is that given s t and a t , s t+1 The distribution of is fully determined. This gives a closed-form expression for the cumulative reward of a given policy in terms of the stable distribution over the states of the MDP. The stable distribution of a policy can be expressed as a solution to a linear programming problem. This gives rise to two classes of algorithms: 1) optimization with respect to the original problem, which is called policy search; and 2) optimization with respect to the dual problem, whose variables are called value functions. If the MDP starts from the initial state s and from there follows To choose an action, the value function determines the expected cumulative reward. The related quantity is the state-action value function , assuming that we start from state s, select action a immediately, and from there according to The state-action-value function determines the cumulative reward for the action chosen. The Q-function may lead to a characterization of the optimal policy (using the Bellman equation). In particular, the Q-function may indicate that the optimal policy is a deterministic function from S to A (in fact, it may be characterized as a "greedy" policy with respect to the optimal Q-function).
[0349] One potential advantage of the MDP model is that it allows the future to be coupled to the present using Q functions. For example, suppose the host vehicle is now in state s, The value of can indicate the impact of performing action a on the future. Therefore, the Q function can provide a local measure of the quality of action a, making the RL problem more similar to the SL scenario.
[0350] Many RL algorithms approximate the V-function or the Q-function in one way or another. Value iteration algorithms (e.g., Q-learning) can rely on the fact that the V- and Q-functions of the optimal policy can be fixed points of some operators derived from the Bellman equation. Actor-critic policy iteration algorithms aim to learn a policy in an iterative manner, where at iteration t, the “critic” estimates And based on that estimate, the “actor” refines the strategy.
[0351] Despite the mathematical advantages of MDPs and the convenience of switching to a Q-function representation, the approach may have several limitations. For example, an approximate notion of a Markovian behavior state may be all that can be found in some cases. Furthermore, transitions between states depend not only on the actions of the agent, but also on the actions of other actors in the environment. For example, in the ACC example mentioned above, while the dynamics of the autonomous vehicle may be Markovian, the next state may depend on the behavior of the driver of another vehicle, which is not necessarily Markovian. One possible solution to this problem is to use a partially observed MDP, where it is assumed that a Markovian state exists, but that what can be observed are observations distributed according to the hidden state.
[0352] A more direct approach might consider game-theoretic generalizations of MDPs (e.g., the stochastic game framework). Indeed, algorithms for MDPs can be generalized to multi-agent games (e.g., min-max-Q learning or Nash-Q learning). Other approaches might include explicit modeling of other players and vanishing regret learning algorithms. Learning in multi-agent scenarios can be more complex than in single-agent scenarios.
[0353] The second limitation of the Q function representation can be introduced by deviating from the tabular setting. The tabular setting is when the number of states and actions is small, so Q can be represented as Rows and columns. However, if the natural representation of S and A involves Euclidean space, and the state and action spaces are discrete, the number of states / actions may be exponential in the dimensions. In this case, it may not be practical to adopt a tabular scenario. Instead, the Q function can be approximated by some function from a class of parameter hypotheses (e.g., a neural network of a certain architecture). For example, a deep Q-network (DQN) learning algorithm can be used. In DQN, the state space can be continuous, while the action space can still be a small discrete set. There may be ways to deal with continuous action spaces, but they may rely on approximating the Q function. In any case, the Q function may be complex and sensitive to noise, and therefore may pose challenges to learning.
[0354] A different approach could be to use recurrent neural networks (RNNs) to solve RL problems. In some cases, RNNs can be combined with concepts from multi-agent games and robustness to adversarial environments from game theory. Furthermore, this approach can be one that does not explicitly rely on any Markov assumptions.
[0355] The following describes in more detail the method of navigating by planning based on prediction. In this method, it can be assumed that the state space S is A subset of , and the action space A is This may be a natural representation in many applications. As mentioned above, there may be two key differences between RL and SL: (1) because past actions affect future rewards, information from the future may need to be propagated back to the past; (2) the “bandit” nature of rewards may obscure the dependencies between (state, action) and rewards, which can complicate the learning process.
[0356] As a first step in this approach, it can be observed that there are interesting problems where the bandit nature of the reward is not a problem. For example, the reward value of an ACC application (as discussed in more detail below) may be differentiable with respect to the current state and action. In fact, even if the reward is given in a "bandit" manner, learning a differentiable function , making The problem of can be a relatively straightforward SL problem (e.g., a one-dimensional regression problem). Therefore, the first step of the method can be to define the reward as a function , the function is differentiable with respect to s and a, or the first step of the method can be to use a regression learning algorithm in order to learn a differentiable function , which minimizes at least some regression loss on samples where the instance vector is And the target scalar is In some cases, an element of exploration can be used in order to create a training set.
[0357] To solve the connection between the past and the future, similar ideas can be used. For example, suppose that we can learn a differentiable function , making Learning such a function can be formulated as a SL problem. can be considered as a near-future predictor. Next, we can use the parameter function To describe the strategy of mapping from S to A. Expressing it as a neural network enables using a recurrent neural network (RNN) to express a fragment that runs the agent for T rounds, where the next state is defined as Here, Can be defined by the environment and can express unpredictable aspects of the near future. t+1 Depends on s in a differentiable way t and a t The fact that can make the connection between future reward value and past action. The parameter vector of can be learned by backpropagation on the resulting RNN. Note that there is no need to Explicit probabilistic assumptions are imposed on . In particular, no requirement for a Markov relation is required. Instead, the recurrent network can be relied upon to propagate “enough” information between the past and the future. Intuitively, can describe the predictable part of the near future, and It can express the unpredictable aspects that may arise due to the behavior of other actors in the environment. The learning system should learn a policy that is robust to the behavior of other actors. If If the value is large, the connection between past actions and future rewards may be too noisy for learning meaningful policies. Explicitly expressing the dynamics of the system in a transparent way can make it easier to incorporate prior knowledge. For example, prior knowledge can simplify the definition of problem.
[0358] As described above, the learning system may benefit from robustness against an adversarial environment (such as the host vehicle's environment), which may include multiple other drivers who may behave in unexpected ways. In models that impose probabilistic assumptions, it is possible to consider the selection in an adversarial manner. In some cases, it is possible to impose constraints, otherwise the adversary may make the planning problem difficult or even impossible. A natural constraint might be to require that the .
[0359] Robustness against adversarial environments may be useful in autonomous driving applications. It can even speed up the learning process because it allows the learning system to focus on the best robust strategy. This concept can be illustrated with a simple game. , the action is , and the instantaneous loss function is ,in is the ReLU (rectified linear unit) function. The next state is ,in is chosen in an adversarial way with respect to the environment. Here, the optimal strategy can be written as a two-layer network with ReLU: . It is observed that when When a=0, the optimal action may have a larger immediate loss than action a=0. Therefore, the system can plan for the future and can not only rely on the immediate loss. Observe that the loss is t The derivative of , and about s t The derivative of is .exist In the case of The adversarial choice will set , so whenever , there can be non-zero loss at round t+1. In this case, the derivative of the loss can be directly back-propagated to a t Therefore, in a t When the choice is non-optimal, Adversarial selection of can help the navigation system obtain non-zero backpropagation messages. This relationship can help the navigation system choose the current action based on the expectation that this current action (even if it results in a non-optimal reward or even a loss) will provide the opportunity for a more optimal action that results in a higher reward in the future.
[0360] This approach can be applied to virtually any navigation situation that may arise. The following description applies to the approach for one example: Adaptive Cruise Control (ACC). In the ACC problem, the host vehicle may attempt to maintain an appropriate distance from a target vehicle ahead (e.g., 1.5 seconds to the target vehicle). Another goal may be to drive as smoothly as possible while maintaining a desired gap. A model representing this situation can be defined as follows. The state space is , and the action space is The first coordinate of the state is the speed of the target car, the second coordinate is the speed of the host vehicle, and the last coordinate is the distance between the host vehicle and the target vehicle (e.g., the position of the host vehicle minus the position of the target vehicle along the curve of the road). The action taken by the host vehicle is to accelerate and can be expressed as a t To express. It can represent the time difference between consecutive rounds. can be set to any suitable amount, but in one example, It can be 0.1 seconds. Position s t It can be expressed as , and the (unknown) acceleration of the target vehicle can be expressed as .
[0361] The complete dynamics of the system can be described by the following description:
[0362] .
[0363] This can be described as the sum of two vectors:
[0364] .
[0365] The first vector is the predictable part, while the second vector is the unpredictable part. The reward for round t is defined as follows:
[0366] ,in .
[0367] The first term results in a penalty for non-zero acceleration, thus encouraging smooth driving. The second term depends on the distance x to the target car. t The distance from expectations The ratio between the desired distance is defined as the maximum between a distance of 1 meter and a braking distance of 1.5 seconds. In some cases, this ratio may be exactly 1, but as long as the ratio is in the range [0.7, 1.3], the policy may forgo any penalties, which may allow the host vehicle to relax a bit in navigation - a characteristic that may be important in achieving smooth driving.
[0368] Implementing the method outlined above, the navigation system of the host vehicle (e.g., through operation of the driving strategy module 803 within the processing unit 110 of the navigation system) can select an action in response to the observed state. The selected action can be based not only on an analysis of rewards associated with available responsive actions relative to the sensed navigation state, but also on consideration and analysis of future states, potential actions in response to the future states, and rewards associated with the potential actions.
[0369] Figure 16 An algorithmic approach to navigation based on detection and long-term planning is shown. For example, at step 1601, at least one processing device 110 of a navigation system of a host vehicle may receive a plurality of images. These images may capture a scene representing the host vehicle's environment and may be provided by any of the image capture devices described above (e.g., a camera, a sensor, etc.). Analyzing one or more of these images at step 1603 may enable the at least one processing device 110 to identify a current navigation state associated with the host vehicle (as described above).
[0370] Various potential navigation actions in response to the sensed navigation state may be determined in steps 1605, 1607, and 1609. These potential navigation actions (e.g., a first navigation action through an Nth available navigation action) may be determined based on the sensed state and a long-term goal of the navigation system (e.g., completing a merge, smoothly following a preceding vehicle, passing a target vehicle, avoiding an object in the roadway, slowing down for a detected stop sign, avoiding an oncoming target vehicle, or any other navigation action that may advance the navigation goal of the system).
[0371] For each of the potential navigation actions determined, the system can determine an expected reward. The expected reward can be determined according to any of the above techniques and can include an analysis of the specific potential action with respect to one or more reward functions. For each potential navigation action determined in steps 1605, 1607, and 1609 (e.g., the first, second, and Nth), an expected reward 1606, 1608, and 1610 can be determined, respectively.
[0372] In some cases, the navigation system of the host vehicle may select from the available potential actions based on the values associated with the expected rewards 1606, 1608, and 1610 (or any other type of indicator of expected reward). For example, in some cases, the action that yields the highest expected reward may be selected.
[0373] In other cases, particularly where the navigation system performs long-term planning to determine the navigation action of the host vehicle, the system may not select the potential action that yields the highest expected reward. Instead, the system may look into the future to analyze whether there is an opportunity to achieve a higher reward later if a lower reward action is selected in response to the current navigation state. For example, a future state may be determined for any or all of the potential actions determined at steps 1605, 1607, and 1609. Each future state determined at steps 1613, 1615, and 1617 may represent a future navigation state that is expected to be modified by the corresponding potential action (e.g., the potential action determined at steps 1605, 1607, and 1609) based on the current navigation state.
[0374] For each of the future states predicted in steps 1613, 1615, and 1617, one or more future actions (as navigation options available in response to the determined future state) may be determined and evaluated. At steps 1619, 1621, and 1623, for example, a value or any other type of indicator of an expected reward associated with the one or more future actions may be developed (e.g., based on one or more reward functions). The expected reward associated with the one or more future actions may be evaluated by comparing the value of the reward function associated with each future action or by comparing any other indicator associated with the expected reward.
[0375] In step 1625, the navigation system of the host vehicle may select a navigation action for the host vehicle based on a comparison of expected rewards based not only on the potential actions identified relative to the current navigation state (e.g., in steps 1605, 1607, and 1609), but also on the expected rewards determined as a result of potential future actions available in response to the predicted future state (e.g., determined in steps 1613, 1615, and 1617). The selection in step 1625 may be based on the option and reward analysis performed in steps 1619, 1621, and 1623.
[0376] The selection of a navigation action at step 1625 may be based solely on a comparison of the expected rewards associated with future action options. In this case, the navigation system may select an action for the current state based solely on a comparison of the expected rewards that would result from actions for potential future navigation states. For example, the system may select the potential action identified at steps 1650, 1610, and 1609 that is associated with the highest future reward value determined by the analysis at steps 1619, 1621, and 1623.
[0377] The selection of a navigation action at step 1625 may also be based solely on a comparison of current action options (as described above). In this case, the navigation system may select the potential action associated with the highest expected reward 1606, 1608, or 1610 identified at steps 1605, 1607, or 1609. This selection may be performed with little or no consideration of future navigation states or future expected rewards for navigation actions available in response to the expected future navigation states.
[0378] On the other hand, in some cases, the selection of the navigation action in step 1625 can be based on a comparison of the expected rewards associated with both the future action options and the current action options. In fact, this may be one of the navigation principles based on long-term planning. For example, the expected rewards for future actions can be analyzed to determine whether any expected rewards can authorize the selection of lower reward actions in response to the current navigation state, so as to achieve potential higher rewards in response to subsequent navigation actions that are expected to be available in response to the future navigation state. As an example, the value of expected reward 1606 or other indicators can indicate the highest expected reward among rewards 1606, 1608 and 1610. On the other hand, expected reward 1608 can indicate the lowest expected reward among rewards 1606, 1608 and 1610. Rather than simply selecting the potential action determined in step 1605 (i.e., the action that causes the highest expected reward 1606), the analysis of future state, potential future actions, and future rewards can be used to select the navigation action in step 1625. In one example, it may be determined that the reward identified at step 1621 (in response to at least one future action for the future state determined at step 1615, which is based on the second potential action determined at step 1607) may be higher than the expected reward 1606. Based on this comparison, the second potential action determined at step 1607 may be selected over the first potential action determined at step 1605, even though the expected reward 1606 is higher than the expected reward 1608. In one example, the potential navigation actions determined at step 1605 may include merging in front of a detected target vehicle, while the potential navigation actions determined at step 1607 may include merging behind the target vehicle. While the expected reward 1606 for merging in front of the target vehicle may be higher than the expected reward 1608 associated with merging behind the target vehicle, it may be determined that merging behind the target vehicle may result in a future state for which action options may exist that yield an even higher potential reward than the expected rewards 1606, 1608, or other rewards based on available actions responsive to the current, sensed navigation state.
[0379] The selection from among the potential actions at step 1625 can be based on any suitable comparison of expected rewards (or any other measure or indicator of the benefit associated with one potential action relative to another potential action). In some cases, as described above, if the second potential action is expected to provide at least one future action associated with an expected reward that is higher than the reward associated with the first potential action, then the second potential action can be selected over the first potential action. In other cases, more complex comparisons can be employed. For example, a reward associated with an action option responsive to a projected future state can be compared to more than one expected reward associated with a determined potential action.
[0380] In some scenarios, if at least one future action is expected to produce a reward that is higher than any reward expected as a result of a potential action for the current state (e.g., expected rewards 1606, 1608, 1610, etc.), then the action and expected reward based on the predicted future state can influence the selection of the potential action for the current state. In some cases, the future action option that produces the highest expected reward (e.g., from among the expected rewards associated with the potential actions for the sensed current state and from among the expected rewards associated with the potential future action options relative to the potential future navigation state) can be used as a guide for selecting the potential action for the current navigation state. That is, after identifying the future action option that produces the highest expected reward (or a reward above a predetermined threshold, etc.), in step 1625, the potential action that will result in the future state associated with the identified future action that produces the highest expected reward can be selected.
[0381] In other cases, selection of available actions may be made based on the difference determined between the expected rewards. For example, if the difference between the expected reward associated with the future action determined at step 1621 and the expected reward 1606 is greater than the difference between the expected reward 1608 and the expected reward 1606 (assuming a positive-signed difference), then the second potential action determined at step 1607 may be selected. In another example, if the difference between the expected reward associated with the future action determined at step 1621 and the expected reward associated with the future action determined at step 1619 is greater than the difference between the expected reward 1608 and the expected reward 1606, then the second potential action determined at step 1607 may be selected.
[0382] Several examples have been described for selecting from among potential actions for the current navigation state. However, any other suitable comparison technique or criteria may be used for selecting available actions by long-term planning based on action and reward analysis extending into projected future states. Figure 16 While two layers of long-term planning analysis are shown (e.g., a first layer considers rewards resulting from potential actions for the current state, and a second layer considers rewards resulting from future action options in response to projected future states), analysis based on more layers may be possible. For example, rather than basing long-term planning analysis on one or two layers, three, four, or more layers of analysis may be used to select from among available potential actions in response to the current navigational state.
[0383] After selecting from among the potential actions responsive to the sensed navigation state, at least one processor may cause at least one adjustment of a navigation actuator of the host vehicle in response to the selected potential navigation action, in step 1627. The navigation actuator may include any suitable device for controlling at least one aspect of the host vehicle. For example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.
[0384] Navigation based on inferred aggression from other parties
[0385] A target vehicle can be monitored by analyzing captured image streams to determine indicators of driving aggression. While aggression is described herein as a qualitative or quantitative parameter, other characteristics may be used, such as the perceived level of driver attention (potential impairment, distraction caused by a cell phone, falling asleep, etc.). In some cases, the target vehicle may be considered to have a defensive posture, while in other cases, the target vehicle may be determined to have a more aggressive posture. Navigation actions may be selected or developed based on indicators of aggression. For example, in some cases, relative speed, relative acceleration, increase in relative acceleration, following distance, etc. relative to the host vehicle may be tracked to determine whether the target vehicle is aggressive or defensive. For example, if the target vehicle is determined to have an aggression level exceeding a threshold, the host vehicle may be inclined to yield to the target vehicle. The target vehicle's aggression level may also be discerned based on the target vehicle's determined behavior relative to one or more obstacles in or near the target vehicle's path (e.g., a preceding vehicle, an obstacle in the road, a traffic light, etc.).
[0386] As an introduction to this concept, an exemplary experiment will be described with respect to a host vehicle merging onto a roundabout, where the navigation goal is to pass through and exit the roundabout. This scenario may begin with the host vehicle approaching the entrance to the roundabout and may end with the host vehicle reaching the exit (e.g., the second exit) of the roundabout. Success may be measured based on whether the host vehicle consistently maintains a safe distance from all other vehicles, whether the host vehicle completes the route as quickly as possible, and whether the host vehicle follows a smooth acceleration strategy. In this example, N T Target vehicles can be randomly placed on the roundabout. To model a mix of adversarial and typical behavior, the target vehicle can be modeled with probability p using an "aggressive" driving strategy, causing the aggressive target vehicle to accelerate when the host vehicle attempts to merge in front of the target vehicle. With probability 1-p, the target vehicle can be modeled with a "defensive" driving strategy, causing the target vehicle to slow down and allow the host vehicle to merge in. In this experiment, p = 0.5, and the host vehicle's navigation system may not be provided with information about the other driver's type. The other driver's type can be randomly selected at the beginning of the segment.
[0387] The navigation state can be represented as the velocity and position of the host vehicle (agent), and the position, velocity and acceleration of the target vehicles. It may be important to keep an observation of the target acceleration in order to distinguish between aggressive and defensive drivers based on the current state. All target vehicles may move along a one-dimensional curve with a circular path as the outline. The host vehicle can move on its own one-dimensional curve, which intersects the target vehicle's curve at the merge point, and this point is the starting point of both curves. In order to model reasonable driving, the absolute value of the acceleration of all vehicles can be upper bounded by a constant. Because reverse driving is not allowed, the speed can also be passed through ReLU. Note that by not allowing reverse driving, long-term planning can become necessary because the agent will not regret its past actions.
[0388] As mentioned above, the next state Can be broken down into predictable parts and the unpredictable part v t The sum of . The expression can represent the dynamics of the vehicle position and velocity (which can be well defined in a differentiable way), while v t It can represent the acceleration of the target vehicle. It can be verified that It can be expressed as a combination of ReLU functions on affine transformations, so it is relatively t and a t is differentiable. The vector v t can be defined by the simulator in a non-differentiable way and can behave aggressively towards some targets and defensively towards others. Figure 17A and 17B Two frames from such a simulator are shown in FIG. In this exemplary experiment, the host vehicle 1701 learns to slow down as it approaches the entrance to the roundabout. It also learns to yield to aggressive vehicles (e.g., vehicles 1703 and 1705) and to safely continue driving when merging in front of defensive vehicles (e.g., vehicles 1706, 1708, and 1710). Figure 17A and 17B In the example shown, the type of target vehicle is not provided to the navigation system of the host vehicle 1701. Instead, whether a particular vehicle is determined to be aggressive or defensive is determined by inference based on the observed position and acceleration (e.g., of the target vehicle). Figure 17A In , based on position, velocity, and / or relative acceleration, host vehicle 1701 may determine that vehicle 1703 has an aggressive tendency, and therefore host vehicle 1701 may stop and wait for target vehicle 1703 to pass rather than attempting to merge in front of target vehicle 1703. However, in Figure 17B, target vehicle 1701 recognizes that target vehicle 1710 traveling behind vehicle 1703 is exhibiting a defensive tendency (also based on the observed position, velocity, and / or relative acceleration of vehicle 1710), and thus completes a successful lane merge in front of target vehicle 1710 and behind target vehicle 1703.
[0389] Figure 18 A flow chart representing an exemplary algorithm for navigating a host vehicle based on predicted aggression from other vehicles is provided. Figure 18 In an example, a level of aggression associated with at least one target vehicle can be inferred based on observed behavior of the target vehicle relative to objects in the target vehicle's environment. For example, in step 1801, at least one processing device (e.g., processing device 110) of a host vehicle navigation system can receive a plurality of images representing the host vehicle's environment from a camera associated with the host vehicle. In step 1803, analysis of the one or more received images can enable the at least one processor to identify a target vehicle (e.g., vehicle 1703) in the host vehicle's environment. In step 1805, analysis of the one or more received images can enable the at least one processing device to identify at least one obstacle to the target vehicle in the host vehicle's environment. The object can include debris in the roadway, a stop / traffic light, a pedestrian, another vehicle (e.g., a vehicle traveling in front of the target vehicle, a parked vehicle, etc.), a box in the road, a roadblock, a curb, or any other type of object that may be encountered in the host vehicle's environment. In step 1807, analysis of the one or more received images can enable the at least one processing device to determine at least one navigation characteristic of the target vehicle relative to the at least one identified obstacle to the target vehicle.
[0390] Various navigational characteristics may be used to infer the aggression level of a detected target vehicle in order to develop an appropriate navigational response to the target vehicle. For example, such navigational characteristics may include the relative acceleration between the target vehicle and at least one identified obstacle, the distance of the target vehicle from the obstacle (e.g., the following distance of the target vehicle behind another vehicle), and / or the relative speed between the target vehicle and the obstacle.
[0391] In some embodiments, the navigation characteristics of the target vehicle can be determined based on outputs from sensors associated with the host vehicle (e.g., radar, speed sensors, GPS, etc.). However, in some cases, the navigation characteristics of the target vehicle can be determined based in part or in whole based on an analysis of images of the host vehicle's environment. For example, the image analysis techniques described above and in, for example, U.S. Patent No. 9,168,868, which is incorporated herein by reference, can be used to identify a target vehicle within the host vehicle's environment. Furthermore, monitoring the position of the target vehicle in captured images over time and / or monitoring the position of one or more features associated with the target vehicle (e.g., taillights, headlights, bumpers, wheels, etc.) in captured images can enable determination of the relative distance, speed, and / or acceleration between the target vehicle and the host vehicle or between the target vehicle and one or more other objects in the host vehicle's environment.
[0392] The level of aggression of an identified target vehicle can be inferred from any suitable observed navigation characteristic of the target vehicle, or any combination of observed navigation characteristics. For example, a determination of aggressiveness can be made based on any observed characteristic and one or more predetermined threshold levels, or any other suitable qualitative or quantitative analysis. In some embodiments, a target vehicle may be considered aggressive if it is observed following the host vehicle or another vehicle at a distance less than a predetermined aggressive distance threshold. On the other hand, a target vehicle observed following the host vehicle or another vehicle at a distance greater than a predetermined defensive distance threshold may be considered defensive. The predetermined aggressive distance threshold need not be the same as the predetermined defensive distance threshold. Furthermore, either or both the predetermined aggressive distance threshold and the predetermined defensive distance threshold may include a range of values, rather than a bright line value. Furthermore, neither the predetermined aggressive distance threshold nor the predetermined defensive distance threshold need be fixed. Rather, these values or ranges of values may change over time, and different thresholds / ranges of thresholds may be applied based on the observed characteristics of the target vehicle. For example, the applied threshold may depend on one or more other characteristics of the target vehicle. Higher observed relative speeds and / or accelerations may warrant application of larger thresholds / ranges. Conversely, lower relative speeds and / or accelerations (including zero relative speed and / or acceleration) may authorize the application of smaller distance thresholds / ranges when making aggressive / defensive inferences.
[0393] Aggressive / defensive inferences can also be based on relative speed and / or relative acceleration thresholds. If the observed relative speed and / or relative acceleration of a target vehicle relative to another vehicle exceeds a predetermined level or range, the target vehicle can be considered aggressive. If the observed relative speed and / or relative acceleration of a target vehicle relative to another vehicle is below a predetermined level or range, the target vehicle can be considered defensive.
[0394] While an aggressive / defensive determination can be made based solely on any observed navigation characteristic, the determination can also depend on any combination of observed characteristics. For example, as described above, in some cases, a target vehicle may be deemed aggressive based solely on the observation that it is following another vehicle at a distance below a certain threshold or range. However, in other cases, the target vehicle may be deemed aggressive if it is following another vehicle at a distance less than a predetermined amount (which may be the same or different than the threshold applied when determining based solely on distance) and with a relative speed and / or relative acceleration greater than a predetermined amount or range. Similarly, a target vehicle may be deemed defensive based solely on the observation that it is following another vehicle at a distance greater than a certain threshold or range. However, in other cases, the target vehicle may be deemed defensive if it is following another vehicle at a distance greater than a predetermined amount (which may be the same or different than the threshold applied when determining based solely on distance) and with a relative speed and / or relative acceleration less than a predetermined amount or range. System 100 may deduce aggressive / defensive if, for example, the vehicle exceeds 0.5G acceleration or deceleration (e.g., a jerk of 5 meters per second cubed (m / s3)), the vehicle has a lateral acceleration of 0.5G in a lane change or on a curve, a vehicle causes another vehicle to do any of the above, a vehicle changes lanes and causes another vehicle to yield with a deceleration of 0.3G or a jerk of 3m / s3 or more, and / or the vehicle changes lanes twice without stopping.
[0395] It should be understood that a reference to a quantity exceeding a range may indicate that the quantity exceeds all values associated with the range or falls within the range. Similarly, a reference to a quantity below a range may indicate that the quantity is below all values associated with the range or falls within the range. Additionally, while examples for making aggressive / defensive inferences are described with respect to distance, relative acceleration, and relative speed, any other suitable quantity may be used. For example, a calculation of the time to collision may be used, or any indirect indicator of the distance, acceleration, and / or speed of the target vehicle. It should also be noted that while the above examples focus on a target vehicle relative to other vehicles, aggressive / defensive inferences may be made by observing the navigation characteristics of the target vehicle relative to any other type of obstacle (e.g., pedestrians, roadblocks, traffic lights, debris, etc.).
[0396] Back to Figure 17A and 17B In the illustrated example, as host vehicle 1701 approaches a roundabout, a navigation system (including at least one of its processing devices) may receive an image stream from a camera associated with the host vehicle. Based on analysis of one or more received images, any of target vehicles 1703, 1705, 1706, 1708, and 1710 may be identified. Furthermore, the navigation system may analyze navigational characteristics of one or more identified target vehicles. The navigation system may recognize that the gap between target vehicles 1703 and 1705 represents a potential first opportunity to merge onto the roundabout. The navigation system may analyze target vehicle 1703 to determine an aggression indicator associated with target vehicle 1703. If target vehicle 1703 is deemed aggressive, the host vehicle's navigation system may choose to yield to vehicle 1703 rather than merge ahead of vehicle 1703. On the other hand, if target vehicle 1703 is deemed defensive, the host vehicle's navigation system may attempt to complete the merge ahead of vehicle 1703.
[0397] As host vehicle 1701 approaches the roundabout, at least one processing device of the navigation system may analyze the captured images to determine navigational characteristics associated with target vehicle 1703. For example, based on the images, it may be determined that vehicle 1703 is following vehicle 1705 at a distance that provides sufficient clearance for host vehicle 1701 to safely enter. In fact, it may be determined that vehicle 1703 is following vehicle 1705 at a distance exceeding an aggressive distance threshold, and therefore, based on this information, the host vehicle navigation system may be inclined to identify target vehicle 1703 as defensive. However, in some cases, as described above, more than one navigational characteristic of the target vehicle may be analyzed when making an aggressive / defensive determination. Further analysis may determine that, when target vehicle 1703 is following target vehicle 1705 at a non-aggressive distance, vehicle 1703 has a relative speed and / or relative acceleration relative to vehicle 1705 that exceeds one or more thresholds associated with aggressive behavior. In effect, host vehicle 1701 may determine that target vehicle 1703 is accelerating relative to vehicle 1705 and approaching a gap that exists between vehicles 1703 and 1705. Based on further analysis of relative speed, acceleration, and distance (and even the rate at which the gap between vehicles 1703 and 1705 is closing), host vehicle 1701 may determine that target vehicle 1703 is behaving aggressively. Thus, while there may be a sufficient gap that the host vehicle can safely navigate into, host vehicle 1701 may anticipate that a merge in front of target vehicle 1703 will result in an aggressively navigating vehicle immediately behind the host vehicle. Furthermore, based on behavior observed through image analysis or other sensor outputs, target vehicle 1703 may be expected to continue accelerating toward host vehicle 1701 or continue traveling toward host vehicle 1701 at a non-zero relative speed if host vehicle 1701 were to merge in front of vehicle 1703. This situation may be undesirable from a safety perspective and may also cause discomfort to the occupants of the host vehicle. For this reason, Figure 17B As shown, host vehicle 1701 may choose to yield to vehicle 1703 and merge onto the roundabout behind vehicle 1703 and in front of vehicle 1710, which is deemed defensive based on an analysis of one or more of its navigational characteristics.
[0398] Back to Figure 18At step 1809, at least one processing device of the navigation system of the host vehicle may determine a navigation action for the host vehicle (e.g., merging in front of vehicle 1710 and behind vehicle 1703) based on at least one identified navigation characteristic of the target vehicle relative to the identified obstacle. To implement the navigation action (at step 1811), at least one processing device may cause at least one adjustment of a navigation actuator of the host vehicle in response to the determined navigation action. For example, the brakes may be applied to yield to the target vehicle. Figure 17A and the accelerator may be applied along with steering of the host vehicle's wheels to cause the host vehicle to enter the roundabout behind vehicle 1703, as Figure 17B shown.
[0399] As described in the above examples, the navigation of the host vehicle may be based on the navigation characteristics of the target vehicle relative to another vehicle or object. Alternatively, the navigation of the host vehicle may be based solely on the navigation characteristics of the target vehicle without specific reference to another vehicle or object. For example, in Figure 18 At step 1807, analysis of the plurality of images captured from the host vehicle's environment may enable determination of at least one navigational characteristic of the identified target vehicle, the navigational characteristic being indicative of a level of aggression associated with the target vehicle. Navigational characteristics may include speed, acceleration, and the like, which do not require reference to another object or target vehicle in order to make an aggressive / defensive determination. For example, an observed acceleration and / or speed associated with the target vehicle that exceeds a predetermined threshold or falls within or exceeds a numerical range may indicate aggressive behavior. Conversely, an observed acceleration and / or speed associated with the target vehicle that falls below a predetermined threshold or falls within or exceeds a numerical range may indicate defensive behavior.
[0400] Of course, in some cases, observed navigational characteristics (e.g., position, distance, acceleration, etc.) may be referenced relative to the host vehicle for purposes of making an aggressive / defensive determination. For example, observed navigational characteristics of a target vehicle that indicate a level of aggressiveness associated with the target vehicle may include an increase in relative acceleration between the target vehicle and the host vehicle, a following distance of the target vehicle behind the host vehicle, a relative speed between the target vehicle and the host vehicle, etc.
[0401] Navigation based on accident liability constraints
[0402] As described in the sections above, planned navigation actions can be tested against predetermined constraints to ensure that certain rules are met. In some embodiments, this concept can be extended to considerations of potential accident liability. As described below, a primary goal of autonomous navigation is safety. Since absolute safety may not be possible (e.g., at least because a particular host vehicle under autonomous control cannot control other vehicles around it - it can only control its own actions), using potential accident liability as a consideration in autonomous navigation, and indeed as a constraint on planned actions, can help ensure that a particular autonomous vehicle does not take any actions that are considered unsafe - e.g., those actions for which potential accident liability could be attached to the host vehicle. A desired level of accident avoidance (e.g., fewer than 10 accidents per hour driven) can be achieved if the host vehicle only takes actions that are safe and that are determined not to result in an accident for which the host vehicle is itself at fault (fault) or responsible. -9 ).
[0403] Challenges posed by most current approaches to autonomous driving include a lack of safety guarantees (or at least an inability to provide the desired level of safety), as well as a lack of scalability. Consider the problem of ensuring safe multi-agent driving. Since society is unlikely to tolerate machine-caused road fatalities, an acceptable level of safety is crucial for the acceptance of autonomous vehicles. While the goal may be to provide zero accidents, this may not be possible because multiple agents are often involved in an accident, and it is conceivable that an accident occurs entirely due to the fault of other agents. For example, Figure 19 As shown, host vehicle 1901 is driving on a multi-lane highway, and while host vehicle 1901 can control its own movements relative to target vehicles 1903, 1905, 1907, and 1909, it cannot control the movements of the target vehicles around it. As a result, if vehicle 1905, for example, suddenly cuts into the host vehicle's lane on a collision course with the host vehicle, host vehicle 1901 may be unable to avoid an accident with at least one of the target vehicles. To address this challenge, autonomous vehicle practitioners typically respond by adopting a statistically driven approach, in which safety validation becomes more rigorous as more miles of data are collected.
[0404] However, to understand the nature of the problem with data-driven approaches to safety, first consider that the probability of death due to an accident per hour of (human) driving is known to be 10 -6 It is reasonable to assume that in order for society to accept machines replacing humans in driving tasks, the death rate should be reduced by three orders of magnitude, that is, to 10 per hour. -9 This estimate is similar to the assumed airbag mortality rate and is derived from aviation standards. For example, 10 -9is the probability that the wing will spontaneously separate from the aircraft in mid-air. However, it is not practical to try to guarantee safety using data-driven statistical methods that provide additional confidence by accumulating driving miles. -9 The amount of data required to calculate the probability of death is its reciprocal (i.e. 10 9 hours of data), which is on the order of 30 billion miles. Furthermore, multi-agent systems interact with their environment and may not be validated offline (unless a realistic simulator is available that simulates real human driving in all its richness and complexity (such as reckless driving) - but the problem of validating such a simulator is even more difficult than creating a safe autonomous vehicle agent). Any change to the planning and control software will require the same amount of new data collection, which is clearly impractical and unrealistic. Furthermore, developing systems from data always suffers from a lack of interpretability and explainability of the actions taken - if an autonomous vehicle (AV) gets into an accident resulting in a fatality, we need to know why. Therefore, a model-based approach to safety is needed, but existing "functional safety" and ASIL requirements in the automotive industry are not designed to cope with multi-agent environments.
[0405] The second major challenge in developing safe driving models for autonomous vehicles is the need for scalability. The premise of AVs isn't just about "building a better world," but rather that mobility without drivers can be maintained at a lower cost than mobility with drivers. This premise is always coupled with the concept of scalability—that is, supporting the mass production of AVs (in the millions) and, more importantly, supporting negligible incremental costs to enable driving in new cities. So, while the cost of computing and sensing is indeed important, the cost of validation and the ability to drive "everywhere" rather than in a select few cities are also necessary requirements to sustain business if AVs are to be manufactured at scale.
[0406] The problem with most current approaches lies in a "brute force" mentality along three axes: (i) the required "computational density," (ii) the way HD maps are defined and created, and (iii) the required specifications for sensors. Brute force approaches run counter to scalability and shift the focus to a future of unconstrained, ubiquitous in-vehicle computing, where the cost of building and maintaining HD maps becomes negligible and scalable, and exotic, ultra-advanced sensors are developed, produced to automotive grade, and at negligible cost. A future in which any of these scenarios is achieved is indeed possible, but a future in which all of them are achieved is likely a low-probability event. Therefore, there is a need for a formal model that combines safety and scalability into an AV program that is socially acceptable and scalable in the sense of supporting millions of cars driving anywhere in the developed world.
[0407] The disclosed embodiments represent a solution that can provide the target safety level (or even exceed the safety goal) and can also be scaled to systems including millions (or more) of autonomous vehicles. In terms of safety, a model called "Responsibility Sensitive Safety (RSS)" is introduced, which formalizes the concept of "accident attribution", is interpretable and explainable, and incorporates the concept of "responsibility" into the actions of robotic agents. The definition of RSS is agnostic to the way it is implemented - a key feature that promotes the goal of creating a convincing global safety model. RSS is supported by the following observations (e.g. Figure 19 (as shown) incentives: agents play asymmetric roles in an accident, where typically only one of the agents is responsible for the accident and is therefore liable for it. The RSS model also includes a formal treatment of "cautious driving" under limited sensing conditions, where not all agents are always visible (e.g., due to occlusions). A major goal of the RSS model is to guarantee that an agent never has an accident for which it is "attributed" or for which it is responsible. The model can only be useful if it comes with an effective policy that complies with RSS (e.g., a function that maps "sensory states" to actions). For example, actions that seem innocent at the moment may lead to catastrophic events in the distant future (a "butterfly effect"). RSS can help to construct a set of local constraints in the short future that can guarantee (or at least virtually guarantee) that no accidents will occur in the future due to the actions of the host vehicle.
[0408] Another contribution revolves around the introduction of a “semantic” language that includes units, measurements, and action spaces, along with specifications for how to incorporate them into the planning, sensing, and actuation of an AV. To get a sense of semantics, in this context, consider how humans taking driving lessons are instructed to think about “driving strategies.” These instructions are not geometric—they don’t take the form of “drive 13.7 meters at your current speed, then drive at 0.8 m / s.” 2 Instead, the instructions are semantic in nature—“follow the car in front of you” or “pass that car on your left.” The typical language of human driving policy is in terms of longitudinal and lateral targets, not geometric units of acceleration vectors. Formal semantic languages can be useful in several ways, including ensuring that the computational complexity of planning does not scale exponentially with time and the number of agents, how safety and comfort interact, how sensing computations are defined, and the specification of sensor modalities and how they interact in fusion methods. Fusion methods (based on semantic languages) can ensure that RSS models achieve the desired 10 per hour of driving. -9 The probability of death is 10 5 Offline validation was performed on a dataset of 100 hours of driving data.
[0409] For example, in a reinforcement learning setting, one can define a Q function on a semantic space (e.g., for the state When the action is executed Given such a Q function, the natural choice of action is to pick the action with the highest quality , in this semantic space, the number of trajectories to be examined at any given time is 10 4is bounded, regardless of the time horizon used for planning. The signal-to-noise ratio in this space can be high, allowing efficient machine learning methods to successfully model the Q-function. In the case of computing on sensing, semantics can allow distinguishing between errors that affect safety and errors that affect driving comfort. We define a PAC model for sensing (Probably Approximately Correct (PAC), borrowing Valiant's terminology for PAC learning), fit this model to the Q-function, and show how measured error can be incorporated into planning in a way that conforms to RSS while allowing driving comfort to be optimized. The semantic language may be important to the success of certain aspects of the model, as other standard measures of error (such as error relative to a global coordinate system) may not conform to the PAC sensing model. Furthermore, the semantic language may be an important enabler for defining HD maps, which can be built using low-bandwidth sensory data, thereby being built through crowdsourcing and supporting scalability.
[0410] In summary, the disclosed embodiments may include a formal model covering the following important components of an AV: sensing, planning, and action. This model may help ensure that, from a planning perspective, no accidents occur for which the AV is responsible. And through the PAC sensing model, even in the presence of sensing errors, the described fusion approach may only require a very reasonable amount of offline data collection to comply with the described safety model. Furthermore, the model may tie together safety and scalability through a semantic language, thereby providing a complete approach for safe and scalable AVs. Finally, it is worth noting that developing an accepted safety model that will be adopted by industry and regulators may be a necessary condition for the success of AVs.
[0411] The RSS model generally follows the classic sense-plan-act robotic control approach. The sensing system can be responsible for understanding the current state of the host vehicle's environment. The planning part can be responsible for determining what the best next move is based on the available options for achieving the driving goal (for example, how to move from the left lane to the right lane in order to exit the highway). The planning part can be called a "driving policy" and can be implemented by a set of hard-coded instructions, through a trained system (for example, a neural network), or a combination. The action part is responsible for implementing the plan (for example, a system of actuators and one or more controllers for steering, accelerating and / or braking the vehicle in order to implement the selected navigation action). The embodiments described below focus primarily on the sensing and planning parts.
[0412] Accidents can result from sensing errors or planning errors. Planning is a multi-agent effort since there are other road users (humans and machines) that react to the actions of the AV. The described RSS model is designed, among other things, to address the safety of the planning part. This can be referred to as multi-agent safety. In a statistical approach, the estimation of the probability of planning errors can be done “online”. That is, after each update of the software, billions of miles would have to be driven with the new version to provide an estimate of an acceptable level of frequency of planning errors. This is clearly not feasible. As an alternative, the RSS model can provide 100% guarantee (or almost 100% guarantee) that the planning module will not make mistakes that can be attributed to the AV (formally defining the concept of “attribution”). The RSS model can also provide an efficient method for its validation that does not rely on online testing.
[0413] Errors in the sensing system may be easier to verify because sensing can be independent of vehicle motion and thus we can use “offline” data to verify the probability of severe sensing errors. However, even if more than 10 9 Hours of driving offline data is also challenging.As part of the description of the disclosed sensing system, a fusion method is described that can be validated using significantly smaller amounts of data.
[0414] The described RSS system can also be scaled to millions of cars. For example, the described semantic driving policy and the applied safety constraints can be consistent with the sensing and mapping requirements that can be scaled to millions of cars even with today's technology.
[0415] The fundamental building block of such a system is a thorough safety definition—the minimum standards to which an AV system may adhere. In the following technical lemma, statistical methods for verifying AV systems are shown to be infeasible, even for simple statements such as "the system experiences N accidents per hour." This means that a model-based safety definition is the only viable tool for verifying AV systems.
[0416] Lemma 1 Let X be a probability space, and A be Suppose we sample from X (iid) samples, and let .So
[0417] .
[0418] Prove that we use the inequality (proven in Appendix A.1 for completeness), we get
[0419] .
[0420] Corollary 1: Assume that AV1 crashes with a small but insufficient probability p1. Any deterministic verification procedure, given 1 / p1 samples, will be unable to distinguish AV1 from a different AV0 that never crashes, with constant probability.
[0421] To gain perspective on typical values for these probabilities, suppose we expect the probability of an accident to be 10 per hour. -9 , and some AV systems provide only 10 -8 Even if the system obtains 10 8 hours of driving, there is still a constant probability that the validation process will fail to indicate that the system is dangerous.
[0422] Finally, note that the difficulty lies in invalidating a single, specific, dangerous AV system. A complete solution cannot be considered a single system, as new versions, bug fixes, and updates are required. From the verifier's perspective, every change, even to a single line of code, generates a new system. Therefore, a statistically validated solution must be verified online with new samples after each small correction or change to account for the shift in the state distribution observed and derived by the new system. Repeatedly and systematically obtaining such a large number of samples (and even then, with a constant probability of failing to validate the system) is infeasible.
[0423] Furthermore, any statistical claim must be formalized to be measurable. Claiming that a system has a number of accidents is a much weaker statistical property than claiming that it drives in a safe manner. To make this point, one must formally define what safety is.
[0424] Absolute security is impossible
[0425] If no accident occurs at some time in the future following action a, then the action a taken by car c can be considered absolutely safe. Figure 19 The simple driving scenario depicted demonstrates that absolute safety is impossible. From the perspective of vehicle 1901, no action can guarantee that surrounding vehicles will not collide with it. Nor can this problem be addressed by prohibiting autonomous vehicles from engaging in such situations. Since every highway with more than two lanes leads to such a scenario at some point, prohibiting it is tantamount to requiring drivers to stay in their garages. At first glance, this suggestion may seem disappointing. Nothing is absolutely safe. However, as mentioned above, such a demand for absolute safety may be excessive, as evidenced by the fact that human drivers do not adhere to it. Instead, they act according to a concept of safety that relies on responsibility.
[0426] Responsibility-Sensitive Safety (RSS)
[0427] An important aspect missing from the concept of absolute safety is the asymmetric nature of most accidents – usually one of the drivers is responsible for the crash and is therefore held liable. Figure 19 In the example, if, for example, left-hand car 1909 suddenly drives into center car 1901, center car 1901 should not be held responsible. To formalize this lack of responsibility, AV 1901's behavior of staying in its lane can be considered safe. To this end, a formal concept of "accident blame" or accident responsibility is described, which can serve as a prerequisite for a safe driving method.
[0428] As an example, consider two cars c f 、c r A simple case of two cars traveling at the same speed, one behind the other, along a straight road. Assume that the car in front, c f Because of an obstacle on the road, the driver brakes suddenly and tries to avoid it. Unfortunately, the r No with c f Keep enough distance to not be able to respond in time and hit c f Obviously, the blame is on c r ; It is the responsibility of the following vehicle to maintain a safe distance from the vehicle in front and to be prepared for unexpected but reasonable braking.
[0429] Next, consider a broader set of scenarios: driving on a multi-lane road, where cars can freely change lanes, cut into the paths of other cars, travel at different speeds, and so on. To simplify the following discussion, assume a straight road on a plane, where the lateral and longitudinal axes are the x-axis and y-axis, respectively. Under mild conditions, this can be achieved by defining a homomorphism between actual curves and straights. In addition, consider a discrete time space. The definition may help to distinguish two intuitively different groups of cases: simple cases where no significant lateral maneuvers are performed, and more complex cases involving lateral motion.
[0430] Definition 1 (Corridor): The corridor of car c is the range ,in It is the position of the leftmost and rightmost corners of c.
[0431] Definition 2 (cut-in): If car c1 (e.g., Figure 20A and 20B car 2003) at time t-1 is not in contact with car c0 (e.g., Figure 20A and20B The corridor of car 2001) in time t intersects with the corridor of car c0, and at time t it intersects with the corridor of car c0, then car c1 cuts into the corridor of car c0 at time t.
[0432] A further distinction can be made between the front and rear parts of the corridor. The term "entry direction" can describe movement in the direction of the relevant corridor boundary. These definitions can be used to define situations where lateral movement occurs. For simple cases where this does not occur, such as when a car is following another car, a safe longitudinal distance is defined:
[0433] Definition 3 (Safe longitudinal distance) If for c f Any braking command a executed by (car 2105), , if c r (Car 2103) applies its maximum brake from time p until it comes to a complete stop, then it does not f Collision, then the car c r and located at c r Another car in the corridor ahead of f The vertical distance between 2101 ( Figure 21 ) is safe with respect to response time p.
[0434] Lemma 2 below computes d as c r 、c f speed, response time ρ and maximum acceleration function. and Both are constants and should be fixed by regulation to some reasonable values.
[0435] Lemma 2 Let c r The vertical axis is c f The vehicle behind. 、 are the maximum braking and maximum acceleration commands, and c r The response time of is the longitudinal velocity of the car, and let is their length. , and define as well as .make Then, c r The minimum safe longitudinal distance is:
[0436] .
[0437] Prove that d t is the distance at time t. To prevent accidents, for each t, we must let d t>L. To construct d min , we need to find the tightest required lower bound for d0. Clearly, d0 must be at least L. As long as the two cars have not stopped after seconds, the speed of the previous car will be , while the upper bound of the speed of c r will be . Therefore, the lower bound of the distance between the cars after T seconds will be:
[0438] .
[0439] Note that T r is the time when c r reaches a complete stop (speed is 0), and T f is the time when the other vehicle reaches a complete stop. Note that , so if , it is sufficient to require d0 > L. If , then
[0440] .
[0441] Requiring and rearranging the terms leads to the conclusion of the proof.
[0442] Finally, a comparison operator is defined that allows for comparison with the concept of some "margin": when comparing lengths, speeds, etc., it is necessary to accept very similar quantities as "equal".
[0443] Definition 4 (μ-comparison) The μ-comparison of two numbers a, b is: if then ; if a < b – μ then ; and if then .
[0444] The following comparisons (argmin, argmax, etc.) are μ-comparisons for some appropriate μ. Suppose an accident occurs between cars c1 and c2. To consider who is responsible for the accident, relevant moments to be checked are defined. This is a certain point in time before the accident and intuitively is a "point of no return"; after that, nothing can prevent the accident from occurring.
[0445] Definition 5 (Blame Time) The blame time of an accident is the earliest time before the accident:
[0446] · where there is an intersection between one car and the corridor of the other car, and
[0447] · the longitudinal distance is unsafe.
[0448] Clearly, such a time exists, because at the moment of the accident, both conditions hold. The time of attribution can be divided into two different categories:
[0449] • One type is where cut-ins also occur, i.e., they are the first moment that the corridors of one car and another car intersect, and are within an unsafe distance.
[0450] One is where no cut-in occurs, i.e., the corridor has been intersected at a safe longitudinal distance, and this distance has become unsafe at the time of attribution...
Claims
1. An autonomous driving system for a host vehicle, the system comprising: an interface for obtaining image data of an environment proximate to the host vehicle, the image data being captured from at least one image capture device of the host vehicle; as well as At least one processing device, the at least one processing device being configured to: determining a planned navigation maneuver for achieving a navigation goal of the host vehicle; identifying a target vehicle in the host vehicle's environment from the image data; Predicting the distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; identifying a braking rate of the host vehicle, a maximum acceleration capability of the host vehicle, a current longitudinal speed of the host vehicle, and a response time of the host vehicle to apply the braking rate; determining a host vehicle stopping distance to stop the host vehicle based on an evaluation of: (i) an identified braking rate of the host vehicle, (ii) a maximum acceleration capability of the host vehicle, (iii) a current longitudinal speed of the host vehicle, and (iv) a response time of the host vehicle to apply the identified braking rate, wherein the identified braking rate of the host vehicle is a next-largest braking rate that is less than the maximum braking capability of the host vehicle; identifying a target vehicle's braking rate and a target vehicle's current longitudinal speed; determining a target vehicle stopping distance to stop the target vehicle based on an evaluation of: (i) an identified braking rate of the target vehicle, and (ii) a current longitudinal speed of the target vehicle; as well as Allowing the host vehicle to continue planning the navigation maneuver while the predicted distance of the planned navigation maneuver is greater than a minimum safe longitudinal distance calculated based on: (i) the host vehicle stopping distance, (ii) the target vehicle stopping distance, and (iii) an acceleration distance corresponding to a distance the host vehicle can travel within the response time period starting from the host vehicle's current longitudinal speed at the host vehicle's maximum acceleration capability.
2. The system according to claim 1, wherein: The planned navigation maneuver includes the target vehicle traveling toward the host vehicle in a lane occupied by the host vehicle.
3. The system according to claim 1, wherein: The identified braking rate of the host vehicle is limited based on at least one identified characteristic of the host vehicle or at least one weather and road condition of the environment.
4. The system according to claim 1, wherein: The identified braking rate of the target vehicle is limited based on at least one identified characteristic of the target vehicle or at least one weather and road condition of the environment.
5. The system according to claim 1, wherein: The current longitudinal velocity of the target vehicle is determined based on the image data.
6. The system according to claim 1, wherein: The current longitudinal velocity of the target vehicle is determined based on output from at least one of a LIDAR system or a RADAR system of the host vehicle.
7. The system according to claim 1, wherein: The planned navigation maneuver causes at least one of steering, braking, or accelerating in the host vehicle.
8. At least one non-transitory machine-readable storage medium comprising instructions stored thereon, the instructions, when executed by a processor of a navigation system of a host vehicle, causing the processor to perform operations comprising: obtaining image data of an environment proximate to the host vehicle, the image data being captured from at least one image capture device of the host vehicle; determining a planned navigation maneuver for achieving a navigation goal of the host vehicle; identifying a target vehicle in the host vehicle's environment from the image data; Predicting the distance between the host vehicle and the target vehicle that would result if the planned navigation action were taken; identifying a braking rate of the host vehicle, a maximum acceleration capability of the host vehicle, a current longitudinal speed of the host vehicle, and a response time of the host vehicle to apply the braking rate; determining a host vehicle stopping distance to stop the host vehicle based on an evaluation of: (i) an identified braking rate of the host vehicle, (ii) a maximum acceleration capability of the host vehicle, (iii) a current longitudinal speed of the host vehicle, and (iv) a response time of the host vehicle to apply the identified braking rate, wherein the identified braking rate of the host vehicle is a next-largest braking rate that is less than the maximum braking capability of the host vehicle; identifying a target vehicle's braking rate and a target vehicle's current longitudinal speed; determining a target vehicle stopping distance to stop the target vehicle based on an evaluation of: (i) an identified braking rate of the target vehicle, and (ii) a current longitudinal speed of the target vehicle; as well as Allowing the host vehicle to continue planning the navigation maneuver while the predicted distance of the planned navigation maneuver is greater than a minimum safe longitudinal distance calculated based on: (i) the host vehicle stopping distance, (ii) the target vehicle stopping distance, and (iii) an acceleration distance corresponding to a distance the host vehicle can travel within the response time period starting from the host vehicle's current longitudinal speed at the host vehicle's maximum acceleration capability.
9. The machine-readable storage medium according to claim 8, wherein: The planned navigation maneuver includes the target vehicle traveling toward the host vehicle in a lane occupied by the host vehicle.
10. The machine-readable storage medium according to claim 8, wherein: The identified braking rate of the host vehicle is limited based on at least one identified condition of the host vehicle or at least one weather and road condition of the environment.
11. The machine-readable storage medium according to claim 8, wherein: The identified braking rate of the target vehicle is limited based on at least one identified characteristic of the target vehicle or at least one weather and road condition of the environment.
12. The machine-readable storage medium according to claim 8, wherein: The current longitudinal velocity of the target vehicle is determined based on the image data.
13. The machine-readable storage medium according to claim 8, wherein: The current longitudinal velocity of the target vehicle is determined based on output from at least one of a LIDAR system or a RADAR system of the host vehicle.
14. The machine-readable storage medium according to claim 8, wherein: The planned navigation maneuver causes at least one of steering, braking, or accelerating in the host vehicle.
Citation Information
Patent Citations
Collision Warning System
US9168868B2
Machine learning navigational engine with imposed constraints
US20180032082A1
Collision mitigation and avoidance
US20180204460A1