Systems and methods for navigating at a safe distance

The system uses camera-based image analysis and processing for autonomous vehicles to ensure safe navigation by determining braking, acceleration, and turning decisions, addressing safety and scalability challenges in complex scenarios.

JP2026090262APending Publication Date: 2026-06-02MOBILEYE VISION TECH LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MOBILEYE VISION TECH LTD
Filing Date
2026-01-16
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Autonomous vehicles face challenges in navigating safely and efficiently while adhering to liability constraints, requiring interpretable mathematical models for safety assurances that can scale to millions of vehicles without increasing costs, and handling complex scenarios like intersections and pedestrian crossings.

Method used

The system utilizes multiple cameras and processing devices to analyze images, determine navigation actions, and consider GPS and sensor data to ensure safe vehicle operation, including braking, acceleration, and turning decisions based on real-time environmental analysis.

Benefits of technology

Enables safe and efficient autonomous vehicle navigation by accurately assessing distances and obstacles, adhering to traffic rules, and responding to dynamic scenarios, thereby enhancing safety and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026090262000001_ABST
    Figure 2026090262000001_ABST
Patent Text Reader

Abstract

A system and method for navigating a host vehicle are disclosed. [Solution] In one embodiment, at least one processing device may be programmed to receive an image representing the environment of the host vehicle, determine a planned navigation operation for the host vehicle, analyze the image to identify a target vehicle having a direction of travel toward the host vehicle, and determine the distance to the next state between the host vehicle and the target vehicle that would occur if the planned navigation operation were performed. The at least one processing device may further determine the stopping distance of the host vehicle based on the braking rate, maximum acceleration capability, and current speed of the host vehicle, and determine the stopping distance of the target vehicle based on the braking rate, maximum acceleration capability, and current speed of the target vehicle, and if the determined distance to the next state is greater than the sum of the stopping distances of the host vehicle and the target vehicle, the planned navigation operation may be performed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Cross-reference of related applications] This application claims priority to U.S. Provisional Patent Application No. 62 / 718,554 filed on 14 August 2018, U.S. Provisional Patent Application No. 62 / 724,355 filed on 29 August 2018, U.S. Provisional Patent Application No. 62 / 772,366 filed on 28 November 2018, and U.S. Provisional Patent Application No. 62 / 777,914 filed on 11 December 2018. All of the above applications are incorporated herein by reference in their entirety.

[0002] This disclosure generally relates to autonomous vehicle navigation. In addition, this disclosure relates to systems and methods for navigating in accordance with potential liability constraints. [Background technology]

[0003] As technology continues to advance, the goal of fully autonomous vehicles capable of navigating on the road is becoming a reality. Autonomous vehicles may need to consider a variety of factors and make appropriate decisions based on those factors to safely and accurately reach their intended destination. For example, autonomous vehicles may need to process and interpret visual information (e.g., information captured from cameras), information from radar or LiDAR, and may also use information obtained from other sources (e.g., GPS devices, speed sensors, accelerometers, suspension sensors, etc.). At the same time, in order to navigate to a destination, autonomous vehicles may also need to identify their own position on a specific road (e.g., a specific lane on a multi-lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, observe traffic signals and signs, proceed from one road to another at an appropriate intersection or interchange, and respond to any other situations that occur or develop during the vehicle's operation. Furthermore, navigation systems may have to adhere to certain imposed constraints. In some cases, these constraints may relate to the interaction between the host vehicle and one or more other objects, such as other vehicles or pedestrians. In other cases, these constraints may relate to the rules of responsibility that must be followed when performing one or more navigation operations for the host vehicle.

[0004] In the field of autonomous driving, there are two key considerations for viable autonomous vehicle systems. The first is the standardization of safety assurances, including requirements that all autonomous vehicles must meet to guarantee safety, and how those requirements can be verified. The second consideration is scalability, because engineering solutions that increase costs may not scale to millions of vehicles, potentially hindering widespread adoption of autonomous vehicles, or even less widespread adoption. Therefore, there is a need for interpretable mathematical models for safety assurances and system designs that can scale to millions of vehicles while adhering to safety assurance requirements. [Overview of the Initiative]

[0005] Embodiments provided in this disclosure provide systems and methods for autonomous vehicle navigation. The disclosed embodiments may use cameras to provide autonomous vehicle navigation features. For example, according to embodiments of this disclosure, the disclosed system may include one, two, or three or more cameras that monitor the vehicle environment. The disclosed system may provide navigation responses based, for example, on the analysis of images captured by one or more of the cameras. The navigation responses may also take into account other data, including, for example, Global Positioning System (GPS) data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.) and / or other map data.

[0006] In one embodiment, a system for navigating a host vehicle is disclosed. The system may include at least one processing device programmed to: receive at least one image representing the host vehicle's environment from an image capture device; determine a planned navigation action to achieve the host vehicle's navigation objective based on at least one driving policy; analyze at least one image to identify a target vehicle in the host vehicle's environment, wherein the target vehicle's direction of travel is toward the host vehicle; determine the distance to the next state between the host vehicle and the target vehicle that would occur if the planned navigation action were performed; determine the host vehicle's braking rate, maximum acceleration capability, and current speed; determine the host vehicle's current stopping distance based on the host vehicle's braking rate, maximum acceleration capability, and current speed; determine the target vehicle's current speed, maximum acceleration capability, and braking rate; determine the target vehicle's stopping distance based on the target vehicle's braking rate, maximum acceleration capability, and current speed; and if the determined distance to the next state is greater than the sum of the host vehicle's stopping distance and the target vehicle's stopping distance, perform the planned navigation action.

[0007] In one embodiment, a method for navigating a host vehicle is disclosed. The method may include: receiving at least one image representing the host vehicle's environment from an image capture device; determining a planned navigation operation to achieve a navigation objective for the host vehicle based on at least one driving policy; analyzing at least one image to identify a target vehicle in the host vehicle's environment, wherein the target vehicle's direction of travel is toward the host vehicle; determining the distance to the next state between the host vehicle and the target vehicle that would occur if the planned navigation operation were performed; determining the host vehicle's braking rate, maximum acceleration capability, and current speed; determining the host vehicle's current stopping distance based on the host vehicle's braking rate, maximum acceleration capability, and current speed; determining the target vehicle's current speed, maximum acceleration capability, and braking rate; determining the target vehicle's stopping distance based on the target vehicle's braking rate, maximum acceleration capability, and current speed; and, if the determined distance to the next state is greater than the sum of the host vehicle's stopping distance and the target vehicle's stopping distance, performing the planned navigation operation.

[0008] In one embodiment, a system for navigating a host vehicle is disclosed. The system receives at least one image representing the host vehicle's environment from an image capture device; determines a planned navigation action to achieve the host vehicle's navigation objective based on at least one driving policy; analyzes at least one image to identify a target vehicle in the host vehicle's environment; determines the next state lateral distance between the host vehicle and the target vehicle that would occur if the planned navigation action were performed; determines the host vehicle's maximum yaw rate capability, the maximum change in the host vehicle's turning radius capability and the host vehicle's current lateral speed; and determines the host vehicle's maximum yaw rate capability and the host vehicle The system may include at least one processing device programmed to: determine the lateral braking distance of the host vehicle based on the maximum change in the host vehicle's turning radius capability and the host vehicle's current lateral speed; determine the target vehicle's current lateral speed, the target vehicle's maximum yaw rate capability and the maximum change in the target vehicle's turning radius capability; determine the target vehicle's lateral braking distance based on the target vehicle's current lateral speed, the target vehicle's maximum yaw rate capability and the maximum change in the target vehicle's turning radius capability; and perform a planned navigation maneuver if the lateral distance of the next determined state is greater than the sum of the host vehicle's lateral braking distance and the target vehicle's lateral braking distance.

[0009] In one embodiment, a method for navigating a host vehicle is disclosed. The method includes the steps of: receiving at least one image representing the host vehicle's environment from an image capture device; determining a planned navigation action to achieve the host vehicle's navigation objective based on at least one driving policy; analyzing at least one image to identify a target vehicle in the host vehicle's environment; determining the lateral distance between the host vehicle and the target vehicle for the next state that would occur if the planned navigation action were performed; determining the host vehicle's maximum yaw rate capability, the maximum change in the host vehicle's turning radius capability, and the host vehicle's current lateral speed; and the maximum The process may include the steps of: determining the lateral braking distance of the host vehicle based on the yaw rate capability, the maximum change in the host vehicle's turning radius capability, and the host vehicle's current lateral speed; determining the target vehicle's current lateral speed, the target vehicle's maximum yaw rate capability, and the maximum change in the target vehicle's turning radius capability; determining the target vehicle's lateral braking distance based on the target vehicle's current lateral speed, the target vehicle's maximum yaw rate capability, and the maximum change in the target vehicle's turning radius capability; and, if the lateral distance of the next determined state is greater than the sum of the host vehicle's lateral braking distance and the target vehicle's lateral braking distance, performing the planned navigation maneuver.

[0010] In one embodiment, a system for navigating a host vehicle near a crosswalk is disclosed. The system may include at least one processing device programmed to receive at least one image representing the host vehicle's environment from an image capture device; detect a representation of a crosswalk in at least one image based on analysis of at least one image; determine whether a representation of a pedestrian appears in at least one image based on analysis of at least one image; detect the presence of a traffic light in the host vehicle's environment; determine whether the detected traffic light is associated with the host vehicle and the crosswalk; determine the state of the detected traffic light; determine the pedestrian's proximity to the detected crosswalk when a representation of a pedestrian appears in at least one image; and determine a planned navigation action for navigating the host vehicle to the detected crosswalk based on at least one driving policy, the determination of the planned navigation action being further based on a determined state of the detected traffic light and the determined proximity of the pedestrian to the detected crosswalk; and cause one or more actuator systems of the host vehicle to perform the planned navigation action.

[0011] In one embodiment, a method for navigating a host vehicle near a crosswalk is disclosed. The method may include the steps of: receiving at least one image representing the host vehicle's environment from an image capture device; detecting a representation of a crosswalk in at least one image based on analysis of at least one image; determining whether a representation of a pedestrian appears in at least one image based on analysis of at least one image; detecting the presence of a traffic light in the host vehicle's environment; determining whether the detected traffic light is associated with the host vehicle and the crosswalk; determining the state of the detected traffic light; determining the proximity of the pedestrian to the detected crosswalk when a representation of a pedestrian appears in at least one image; determining a planned navigation action for navigating the host vehicle to the detected crosswalk based on at least one driving policy, wherein the determination of the planned navigation action is further based on the determined state of the detected traffic light and the determined proximity of the pedestrian to the detected crosswalk; and causing one or more actuator systems of the host vehicle to perform the planned navigation action.

[0012] In one embodiment, a method for navigating a host vehicle near a crosswalk is disclosed. The method may include the steps of: receiving at least one image representing the host vehicle's environment from an image capture device; detecting the start and end positions of a crosswalk; determining whether pedestrians are present near the crosswalk based on analysis of at least one image; detecting the presence of a traffic light in the host vehicle's environment; determining whether the traffic light is associated with the host vehicle and the crosswalk; determining the state of the traffic light; determining the navigation action of the host vehicle near the detected crosswalk based on the association of the traffic light, the determined state of the traffic light, the presence of pedestrians near the crosswalk, the shortest distance selected between the pedestrian and either the start or end position of the crosswalk, and the pedestrian's motion vector; and causing one or more actuator systems of the host vehicle to perform the navigation action.

[0013] According to other embodiments disclosed, a non-temporary computer-readable storage medium is executable by at least one processing device and can store program instructions for performing any of the steps and / or methods described herein.

[0014] The above overview and the detailed explanation below are illustrative and descriptive, and do not constitute a limitation of the claims. [Brief explanation of the drawing]

[0015] The accompanying drawings incorporated herein and forming part of this specification illustrate various embodiments disclosed.

[0016] [Figure 1] This is a diagrammatic representation of an exemplary system according to the disclosed embodiments.

[0017] [Figure 2A] This is an exemplary side view representation of a vehicle including the system according to the disclosed embodiment.

[0018] [Figure 2B] Figure 2A is a top view representation of the vehicle and system shown in the disclosed embodiment.

[0019] [Figure 2C] This is a top view representation of another embodiment of a vehicle including the system according to the disclosed embodiment.

[0020] [Figure 2D] This is a top view representation of yet another embodiment of a vehicle including the system according to the disclosed embodiment.

[0021] [Figure 2E] This is a top view representation of yet another embodiment of a vehicle including the system according to the disclosed embodiment.

[0022] [Figure 2F]It is a diagrammatic representation of an exemplary vehicle control system according to the disclosed embodiment.

[0023] [Figure 3A] It is a diagrammatic representation of the interior of a vehicle including a rearview mirror and a user interface of a vehicle imaging system according to the disclosed embodiment.

[0024] [Figure 3B] It is a diagram of an example of a camera mount configured to be positioned opposite the vehicle front windshield behind a rearview mirror according to the disclosed embodiment.

[0025] [Figure 3C] It is a diagram of the camera mount shown in FIG. 3B from different viewpoints according to the disclosed embodiment.

[0026] [Figure 3D] It is a diagram of an example of a camera mount configured to be positioned opposite the vehicle front windshield behind a rearview mirror according to the disclosed embodiment.

[0027] [Figure 4] It is an exemplary block diagram of a memory configured to store instructions for performing one or more operations according to the disclosed embodiment.

[0028] [Figure 5A] It is a flowchart showing an exemplary process for generating one or more navigation responses based on monocular image analysis according to the disclosed embodiment.

[0029] [Figure 5B] It is a flowchart showing an exemplary process for detecting one or more vehicles and / or pedestrians within a set of images according to the disclosed embodiment.

[0030] [Figure 5C]A flowchart showing an exemplary process for detecting road markings and / or lane geometry information within a set of images according to the disclosed embodiments.

[0031] [Figure 5D] A flowchart showing an exemplary process for detecting traffic signals within a set of images according to the disclosed embodiments.

[0032] [Figure 5E] A flowchart of an exemplary process for generating one or more navigation responses based on a vehicle route according to the disclosed embodiments.

[0033] [Figure 5F] A flowchart showing an exemplary process for determining whether a leading vehicle is changing lanes according to the disclosed embodiments.

[0034] [Figure 6] A flowchart showing an exemplary process for generating one or more navigation responses based on stereo image analysis according to the disclosed embodiments.

[0035] [Figure 7] A flowchart showing an exemplary process for generating one or more navigation responses based on the analysis of three sets of images according to the disclosed embodiments.

[0036] [Figure 8] A block diagram representation of modules that can be implemented by one or more specifically programmed processing devices of a navigation system for an autonomous vehicle according to the disclosed embodiments.

[0037] [Figure 9] A graph of navigation options according to the disclosed embodiments.

[0038] [Figure 10] This is a graph of navigation options according to the disclosed embodiments.

[0039] [Figure 11A] A schematic diagram of the navigation options for a host vehicle in a merging area, according to the disclosed embodiments, is shown. [Figure 11B] A schematic diagram of the navigation options for a host vehicle in a merging area, according to the disclosed embodiments, is shown. [Figure 11C] A schematic diagram of the navigation options for a host vehicle in a merging area, according to the disclosed embodiments, is shown.

[0040] [Figure 11D] A diagrammatic representation of a dual merging scenario according to the disclosed embodiments is shown.

[0041] [Figure 11E] A graph of potentially useful options in a dual-merging scenario according to the disclosed embodiments is shown.

[0042] [Figure 12] A representative image capturing the host vehicle environment, along with potential navigation constraints, according to the disclosed embodiments is shown.

[0043] [Figure 13] A flowchart of an algorithm for navigating a vehicle according to the disclosed embodiments is shown.

[0044] [Figure 14] A flowchart of an algorithm for navigating a vehicle according to the disclosed embodiments is shown.

[0045] [Figure 15] A flowchart of an algorithm for navigating a vehicle according to the disclosed embodiments is shown.

[0046] [Figure 16] A flowchart of an algorithm for navigating a vehicle according to the disclosed embodiments is shown.

[0047] [Figure 17A] A diagram of a host vehicle navigating within a ring road according to the disclosed embodiment is shown. [Figure 17B] A diagram of a host vehicle navigating within a ring road according to the disclosed embodiment is shown.

[0048] [Figure 18] A flowchart of an algorithm for navigating a vehicle according to the disclosed embodiments is shown.

[0049] [Figure 19] An example of a host vehicle traveling on a multi-lane highway according to the disclosed embodiment is shown.

[0050] [Figure 20A] An example of a vehicle cutting in front of another vehicle according to the disclosed embodiment is shown. [Figure 20B] An example of a vehicle cutting in front of another vehicle according to the disclosed embodiment is shown.

[0051] [Figure 21] An example of a vehicle following another vehicle according to the disclosed embodiment is shown.

[0052] [Figure 22] An example of a vehicle exiting a parking lot and potentially merging onto a busy road is shown according to the disclosed embodiment.

[0053] [Figure 23] The disclosed embodiments show a vehicle traveling on a road.

[0054] [Figure 24A] An example of a scenario according to the disclosed embodiments is shown. [Figure 24B] An example of a scenario according to the disclosed embodiments is shown. [Figure 24C] An example of a scenario according to the disclosed embodiments is shown. [Figure 24D] An example of a scenario according to the disclosed embodiments is shown.

[0055] [Figure 25] An example of a scenario according to the disclosed embodiments is shown.

[0056] [Figure 26] An example of a scenario according to the disclosed embodiments is shown.

[0057] [Figure 27] An example of a scenario according to the disclosed embodiments is shown.

[0058] [Figure 28A] An example of a scenario in which one vehicle follows another vehicle, according to the disclosed embodiments, is shown. [Figure 28B] An example of a scenario in which one vehicle follows another vehicle, according to the disclosed embodiments, is shown.

[0059] [Figure 29A] An example of negligence in an interruption scenario according to the disclosed embodiments is shown. [Figure 29B] An example of negligence in an interruption scenario according to the disclosed embodiments is shown.

[0060] [Figure 30A] An example of negligence in an interruption scenario according to the disclosed embodiments is shown. [Figure 30B] An example of negligence in an interruption scenario according to the disclosed embodiments is shown.

[0061] [Figure 31A] An example of negligence in a drift scenario according to the disclosed embodiments is shown. [Figure 31B]An example of negligence in a drift scenario according to the disclosed embodiments is shown. [Figure 31C] An example of negligence in a drift scenario according to the disclosed embodiments is shown. [Figure 31D] An example of negligence in a drift scenario according to the disclosed embodiments is shown.

[0062] [Figure 32A] An example of negligence in a two-way traffic scenario according to the disclosed embodiments is shown. [Figure 32B] An example of negligence in a two-way traffic scenario according to the disclosed embodiments is shown.

[0063] [Figure 33A] An example of negligence in a two-way traffic scenario according to the disclosed embodiments is shown. [Figure 33B] An example of negligence in a two-way traffic scenario according to the disclosed embodiments is shown.

[0064] [Figure 34A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 34B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0065] [Figure 35A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 35B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0066] [Figure 36A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 36B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0067] [Figure 37A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 37B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0068] [Figure 38A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 38B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0069] [Figure 39A] An example of negligence in a route priority scenario according to the disclosed embodiments is shown. [Figure 39B] An example of negligence in a route priority scenario according to the disclosed embodiments is shown.

[0070] [Figure 40A] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown. [Figure 40B] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown.

[0071] [Figure 41A] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown. [Figure 41B] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown.

[0072] [Figure 42A] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown. [Figure 42B] An example of negligence in a traffic signal scenario according to the disclosed embodiments is shown.

[0073] [Figure 43A]An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 43B] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 43C] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown.

[0074] [Figure 44A] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 44B] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 44C] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown.

[0075] [Figure 45A] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 45B] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 45C] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown.

[0076] [Figure 46A] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 46B] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 46C] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown. [Figure 46D] An example of a vulnerable road user (VRU) scenario according to the disclosed embodiments is shown.

[0077] [Figure 47A] This is a diagram of the point of negligence and the appropriate response according to the disclosed embodiment.

[0078] [Figure 47B] This is a diagram of route priority for routes of different geometric shapes according to the disclosed embodiments.

[0079] [Figure 47C] This is a diagram of the longitudinal ordering of different geometric shapes along a root according to the disclosed embodiment.

[0080] [Figure 47D] This is a diagram showing the safe longitudinal distance between vehicles according to the disclosed embodiment.

[0081] [Figure 47E] This diagram illustrates a situation in the disclosed embodiment where one vehicle cannot predict the path of another vehicle.

[0082] [Figure 47F] This is a diagram of route priority in a traffic signal according to the disclosed embodiment.

[0083] [Figure 47G] This is a diagram of an exemplary unstructured route according to the disclosed embodiments.

[0084] [Figure 47H] This is a diagram illustrating exemplary lateral behavior in an unstructured state according to the disclosed embodiments.

[0085] [Figure 47I] This diagram shows the exposure time and negligence time in the concealed area according to the disclosed embodiment.

[0086] [Figure 48A] An example scenario of two vehicles traveling in opposite directions according to the disclosed embodiments is shown.

[0087] [Figure 48B] An example of a target vehicle traveling toward a host vehicle according to the disclosed embodiment is shown.

[0088] [Figure 49] An example of a host vehicle maintaining a safe longitudinal distance according to the disclosed embodiment is shown.

[0089] [Figure 50A] A flowchart is provided illustrating an exemplary process for maintaining a safe longitudinal distance according to the disclosed embodiments. [Figure 50B] A flowchart is provided illustrating an exemplary process for maintaining a safe longitudinal distance according to the disclosed embodiments.

[0090] [Figure 51A] An example of a scenario in the disclosed embodiment, in which two vehicles are positioned laterally apart from each other, is shown.

[0091] [Figure 51B] An example of a host vehicle maintaining a safe lateral distance according to the disclosed embodiment is shown.

[0092] [Figure 52A] An example of a host vehicle performing a planned navigation operation according to the disclosed embodiments is shown.

[0093] [Figure 52B] This shows an example of a host vehicle that decides whether or not to perform navigation operations.

[0094] [Figure 53A] A flowchart is provided illustrating an exemplary process for maintaining a safe lateral distance according to the disclosed embodiments. [Figure 53B] A flowchart is provided illustrating an exemplary process for maintaining a safe lateral distance according to the disclosed embodiments.

[0095] [Figure 54] This is a schematic diagram of a road including a pedestrian crossing according to the disclosed embodiment.

[0096] [Figure 55A] This is a schematic diagram of possible navigation operations performed by a vehicle traveling along a road, according to the disclosed embodiments. [Figure 55B] This is a schematic diagram of possible navigation operations performed by a vehicle traveling along a road, according to the disclosed embodiments.

[0097] [Figure 55C] An example of estimating the distance from a pedestrian to a crosswalk according to the disclosed embodiment is shown. [Figure 55D] An example of estimating the distance from a pedestrian to a crosswalk according to the disclosed embodiment is shown.

[0098] [Figure 56] This is a flowchart illustrating the process of navigating a vehicle near a pedestrian crossing according to the disclosed embodiment. [Modes for carrying out the invention]

[0099] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the components shown in the drawings, and the exemplary methods described herein may be modified by substitution, reordering, deletion, or addition of steps in the disclosed method. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Rather, the appropriate scope is defined by the appended claims.

[0100] Overview of autonomous vehicles

[0101] When used throughout this disclosure, the term “autonomous vehicle” means a vehicle capable of performing at least one navigation change without driver input. “Navigation change” means one or more changes to the vehicle’s steering, braking, or acceleration / deceleration. To be autonomous, a vehicle does not need to be fully automatic (e.g., fully operational without a driver or driver input). Rather, an autonomous vehicle includes a vehicle that can operate under driver control during certain periods of time and without driver control during other periods of time. An autonomous vehicle may also include a vehicle that controls only certain aspects of vehicle navigation, such as steering (e.g., to maintain a vehicle course between vehicle lane constraints), or a vehicle that controls certain steering actions under certain conditions (but not all conditions), but leaves other aspects (e.g., braking or braking under certain conditions) to the driver. In some cases, an autonomous vehicle may handle some or all aspects of the vehicle’s braking, speed control, and / or steering.

[0102] Since human drivers typically rely on visual cues and observations to control their vehicles, traffic infrastructure is built accordingly, with lane markings, traffic signs, and traffic lights designed to provide drivers with visual information. In light of these design features of traffic infrastructure, autonomous vehicles may include cameras and processing units that analyze visual information captured from the vehicle's environment. Visual information may include, for example, images representing components of the traffic infrastructure observable by the driver (e.g., lane markings, traffic signs, traffic lights, etc.) and other obstacles (e.g., other vehicles, pedestrians, debris, etc.). Furthermore, autonomous vehicles may also use stored information, such as information that provides a model of the vehicle's environment, when navigating. For example, a vehicle may use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.) and / or other map data to provide information related to the vehicle's environment while the vehicle is in motion, and the vehicle (and other vehicles) may use the information to determine its own position in the model. Some vehicles may also be capable of inter-vehicle communication, information sharing, and modification of peer vehicles in response to hazards or changes around the vehicle.

[0103] System Overview

[0104] Figure 1 is a block diagram representation of System 100 according to an exemplary embodiment disclosed. System 100 may include various components depending on the specific implementation requirements. In some embodiments, System 100 may include a processing unit 110, an image acquisition unit 120, a position sensor 130, one or more memory units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. The processing unit 110 may include one or more processing devices. In some embodiments, the processing unit 110 may include an application processor 180, an image processor 190, or any other suitable processing device. Similarly, the image acquisition unit 120 may include any number of image acquisition devices and components depending on the requirements of the specific application. In some embodiments, the image acquisition unit 120 may include one or more image acquisition devices (e.g., a camera, a CCD, or any other type of image sensor), such as image acquisition device 122, image acquisition device 124, image acquisition device 126, etc. The system 100 may also include a data interface 128 that enables communication between the processing unit 110 and the image acquisition unit 120. For example, the data interface 128 may include one or more arbitrary wired and / or wireless links for transmitting image data acquired by the image acquisition unit 120 to the processing unit 110.

[0105] The wireless transceiver 172 may include one or more devices configured to exchange transmissions with one or more networks (e.g., cellular or the Internet) via a wireless interface using radio frequencies, infrared frequencies, magnetic fields, or electric fields. The wireless transceiver 172 may transmit and / or receive data using any known standard (e.g., Wi-Fi®, Bluetooth®, Bluetooth Smart, 802.15.4, ZigBee®, etc.). Such transmissions may include communication from a host vehicle to one or more remotely located servers. Such transmissions may also include (unidirectional or bidirectional) communication between the host vehicle and one or more target vehicles in the host vehicle's environment (e.g., to facilitate the adjustment of the host vehicle's navigation in consideration of or with such target vehicles), as well as broadcast transmissions to unspecified receivers near the transmitting vehicle.

[0106] Both the application processor 180 and the image processor 190 may include various types of hardware-based processing devices. For example, either or both of the application processor 180 and the image processor 190 may include a microprocessor, preprocessor (such as an image preprocessor), graphics processor, central processing unit (CPU), support circuitry, digital signal processor, integrated circuit, memory, or any other type of device suitable for running applications and processing and analyzing images. In some embodiments, the application processor 180 and / or the image processor 190 may include any type of single-core or multi-core processor, mobile device microcontroller, central processing unit, etc. Various processing devices are available, including processors available from manufacturers such as Intel® and AMD®, and may include various architectures (e.g., x86 processor, ARM®, etc.).

[0107] In some embodiments, the application processor 180 and / or image processor 190 may include any EyeQ series processor available from Mobileye®. These processor designs include multiple processing units, each having local memory and instruction sets. Such processors may include video inputs for receiving image data from multiple image sensors, and may also include video output capabilities. In one example, EyeQ2® uses 90nm-micron technology operating at 332MHz. The EyeQ2® architecture consists of two floating-point hyperthreaded 32-bit RISC CPUs (MIPS32® 34K® cores), five vision computing engines (VCEs), three vector microcode processors (VMP®), a Denali 64-bit mobile DDR controller, a 128-bit internal acoustic interconnect, dual 16-bit video input and 18-bit video output controllers, a 16-channel DMA, and several peripherals. The MIPS34K CPU manages five VCEs, three VMPs® and DMAs, a second MIPS34K CPU and multi-channel DMA, and other peripherals. The five VCEs, three VMPs® and MIPS34K CPUs can perform the intensive vision computations required by multi-function bundled applications. In another example, the disclosed embodiment may use a third-generation processor, EyeQ3®, which is six times more powerful than EyeQ2®. In yet another example, EyeQ4® and / or EyeQ5® may be used in the disclosed embodiment. Naturally, newer or future EyeQ processing devices may be used with the disclosed embodiment.

[0108] Any of the processing devices disclosed herein can be configured to perform a specific function. Configuring a processing device, such as the EyeQ processor or any other controller or microprocessor described herein, to perform a specific function may include programming computer executable instructions and providing those instructions to the processing device for execution during the operation of the processing device. In some embodiments, configuring a processing device may include directly programming architectural instructions into the processing device. In other embodiments, configuring a processing device may include storing executable instructions in memory accessible to the processing device during operation. For example, the processing device may access memory during operation to retrieve and execute stored instructions. In any case, processing devices configured to perform sensing, image analysis, and / or navigation functions disclosed herein represent a dedicated hardware-based system that controls multiple hardware-based components of a host vehicle.

[0109] Figure 1 shows two separate processing devices included in processing unit 110, but more or fewer processing devices may be used. For example, in some embodiments, a single processing device may be used to accomplish the tasks of application processor 180 and image processor 190. In other embodiments, these tasks may be performed by three or more processing devices. Furthermore, in some embodiments, system 100 may include one or more processing units 110 and not include other components such as image acquisition unit 120.

[0110] The processing unit 110 may include various types of devices. For example, the processing unit 110 may include various devices such as a controller, an image preprocessor, a central processing unit (CPU), support circuits, a digital signal processor, an integrated circuit, memory, or any other type of device that processes and analyzes images. The image preprocessor may include a video processor that captures, digitizes, and processes images from an image sensor. The CPU may include any number of microcontrollers or microprocessors. The support circuits may include any number of circuits commonly known in the art, including caches, power supplies, clocks, and input / output circuits. The memory may store software that, when executed by the processor, controls the operation of the system. The memory may include databases and image processing software. The memory may include any number of random access memories, read-only memories, flash memories, disk drives, optical memory devices, tape memory devices, removable memory devices, and other types of memory devices. In one example, the memory may be separate from the processing unit 110. In another example, the memory may be integrated into the processing unit 110.

[0111] Each memory unit 140, 150 may contain software instructions that, when executed by a processor (e.g., an application processor 180 and / or an image processor 190), can control various aspects of the operation of the system 100. These memory units may contain various database and image processing software, as well as trained systems such as neural networks or deep neural networks. The memory units may include random access memory, read-only memory, flash memory, disk drives, optical memory, tape memory, removable memory, and / or any other type of memory. In some embodiments, the memory units 140, 150 may be separate from the application processor 180 and / or the image processor 190. In other embodiments, these memory units may be integrated into the application processor 180 and / or the image processor 190.

[0112] The position sensor 130 may include any type of device suitable for determining the position associated with at least one component of the system 100. In some embodiments, the position sensor 130 may include a GPS receiver. Such a receiver can determine the user's position and speed by processing signals broadcast by Global Positioning System satellites. Position information from the position sensor 130 may be provided to the application processor 180 and / or the image processor 190.

[0113] In some embodiments, the system 100 may include components such as a speed sensor (e.g., a speedometer) for measuring the speed of the vehicle 200. The system 100 may also include one or more (single-axis or multi-axis) accelerometers for measuring the acceleration of the vehicle 200 along one or more axes.

[0114] Memory units 140 and 150 may contain data organized in a database or any other format indicating the locations of known landmarks. Environmental sensory information (images, radar signals, depth information obtained by lidar or stereoscopic processing of two or more images, etc.) can be processed together with positional information such as GPS coordinates and the vehicle's own motion to determine the vehicle's current position relative to known landmarks and to refine the vehicle's position. Certain aspects of this technology are included in the positioning technology known as REM (trademark), which is sold by the assignee of this application.

[0115] The user interface 170 may include any device suitable for providing information or receiving input from one or more users of the system 100. In some embodiments, the user interface 170 may include user input devices, such as a touchscreen, microphone, keyboard, pointer device, track wheel, camera, knob, button, etc. Using such input devices, a user may provide information input or commands to the system 100 by typing instructions or information, providing voice commands, using buttons, pointers or eye-tracking functions, or by selecting menu options on the screen through any other suitable technique for communicating information to the system 100.

[0116] The user interface 170 may include one or more processing devices configured to provide information to or receive information from the user and process that information for use, for example, by the application processor 180. In some embodiments, such processing devices may execute commands to recognize and track eye movements, commands to receive and interpret voice commands, commands to recognize and interpret touches and / or gestures made on a touchscreen, commands to respond to keyboard input or menu selections, and so on. In some embodiments, the user interface 170 may include a display, a speaker, a haptic device and / or any other device that provides output information to the user.

[0117] The map database 160 may include any type of database that stores map data useful to system 100. In some embodiments, the map database 160 may include data relating to the location in a reference coordinate system of various items, including roads, water features, geographical features, businesses, points of interest, restaurants, gas stations, etc. The map database 160 may store not only the locations of such items but also descriptors associated with those items, including, for example, names associated with any of the stored features. In some embodiments, the map database 160 may be physically located together with other components of system 100. Alternatively or additionally, the map database 160 or a part thereof may be located remotely with respect to other components of system 100 (e.g., processing unit 110). In such embodiments, information from the map database 160 may be downloaded to a network via a wired or wireless data connection (e.g., via a cellular network and / or the Internet, etc.). In some cases, the map database 160 may store a sparse data model that includes a polynomial representation of specific road features (e.g., lane markings) or the target trajectory of a host vehicle. The map database 160 may also include stored representations of various recognized landmarks that can be used to determine or update the known position of a host vehicle relative to a target track. Landmark representations may include data fields such as the type of landmark and the location of the landmark, among other potential identifiers.

[0118] The image capture devices 122, 124, and 126 may each include any type of device suitable for capturing at least one image from the environment. Furthermore, any number of image capture devices may be used to acquire images to input to the image processor. Some embodiments may include only a single image capture device, while others may include two, three, or even four or more image capture devices. The image capture devices 122, 124, and 126 are further described below with reference to Figures 2B to 2E.

[0119] One or more cameras (e.g., image acquisition devices 122, 124, and 126) may be part of a sensing block included on the vehicle. Various other sensors may be included in the sensing block, and any or all of the sensors may be used to develop the detected navigation state of the vehicle. In addition to cameras (forward, side, rear, etc.), other sensors such as radar, lidar, and acoustic sensors may be included in the sensing block. In addition, the sensing block may include one or more components configured to transmit and receive information relating to the vehicle's environment. For example, such components may include a radio transceiver (RF, etc.) capable of receiving sensor-based information or any other type of information relating to the host vehicle's environment from a remotely located source relative to the host vehicle. Such information may include sensor output information or related information received from vehicle systems other than the host vehicle. In some embodiments, such information may include information received from remote computing devices or centralized servers, etc. Furthermore, the cameras can take many different configurations, such as a single camera unit, multiple cameras, camera clusters, long FOV, short FOV, wide-angle, fisheye, etc.

[0120] System 100 or various components of System 100 can be incorporated into various different platforms. In some embodiments, System 100 may be included in a vehicle 200, as shown in Figure 2A. For example, the vehicle 200 may include the processing unit 110 and any other components of System 100, as described above with respect to Figure 1. In some embodiments, the vehicle 200 may have only a single image acquisition device (e.g., a camera), while in other embodiments, such as those considered in relation to Figures 2B to 2E, multiple image acquisition devices may be available. For example, as shown in Figure 2A, either of the image acquisition devices 122 and 124 of the vehicle 200 may be part of an ADAS (Advanced Driver-Assistance System) imaging set.

[0121] The image acquisition device included in the vehicle 200 as part of the image acquisition unit 120 can be located in any suitable position. In some embodiments, as shown in Figures 2A-2E and 3A-3C, the image acquisition device 122 may be located near the rearview mirror. This position can provide a similar line of sight to the driver of the vehicle 200 and can help the driver determine what is visible and what is not. While the image acquisition device 122 can be located in any position near the rearview mirror, positioning the image acquisition device 122 on the driver side of the mirror can further assist in acquiring images representing the driver's field of view and / or line of sight.

[0122] Other positions can also be used for the image acquisition device of the image acquisition unit 120. For example, the image acquisition device 124 may be placed on or inside the bumper of the vehicle 200. Such a position may be particularly suitable for an image acquisition device having a wide field of view. The line of sight of an image acquisition device placed on the bumper may differ from the line of sight of the driver, and therefore the bumper image acquisition device and the driver are not always looking at the same object. The image acquisition devices (e.g., image acquisition devices 122, 124 and 126) may also be placed in other positions. For example, the image acquisition devices may be placed on one or both of the side mirrors of the vehicle 200, on the roof of the vehicle 200, on the hood of the vehicle 200, on the trunk of the vehicle 200, on the side of the vehicle 200, mounted on any window of the vehicle 200, positioned behind or in front, mounted on or near the front and / or rear lights of the vehicle 200, etc.

[0123] In addition to the image acquisition device, the vehicle 200 may include various other components of the system 100. For example, the processing unit 110 may be integrated into the vehicle's engine control unit (ECU) or included in the vehicle 200 separately from the ECU. The vehicle 200 may also be equipped with position sensors 130 such as a GPS receiver, and may also include a map database 160 and memory units 140 and 150.

[0124] As described above, the wireless transceiver 172 may receive and / or upload data via one or more networks (e.g., a cellular network, the Internet, etc.). For example, the wireless transceiver 172 may upload data collected by the system 100 to one or more servers and download data from one or more servers. Through the wireless transceiver 172, the system 100 may receive updates to data stored in the map database 160, memory 140, and / or memory 150, for example, periodically or on demand. Similarly, the wireless transceiver 172 may upload any data from the system 100 (e.g., images captured by the image acquisition unit 120, data received by the position sensor 130, other sensors, or the vehicle control system, etc.) and / or any data processed by the processing unit 110 to one or more servers.

[0125] System 100 may upload data to a server (e.g., the cloud) based on privacy level settings. For example, System 100 may implement privacy level settings that regulate or restrict data (including metadata) that can uniquely identify a vehicle and / or the vehicle's driver / owner, which is transmitted to the server. Such settings may be configured, for example, by a user via the wireless transceiver 172, initialized by factory default settings, or configured by data received by the wireless transceiver 172.

[0126] In some embodiments, system 100 may upload data according to a “high” privacy level, and under certain settings, system 100 may transmit data that does not contain any details about a specific vehicle and / or driver / owner (e.g., location information related to a route, captured images, etc.). For example, when uploading data according to a “high” privacy level, system 100 may transmit data that does not include the vehicle identification number (VIN) or the name of the vehicle's driver or owner, but instead includes captured images and / or limited location information related to a route.

[0127] Other privacy levels are also intended. For example, system 100 may transmit data to the server according to a “medium” privacy level, which may include additional information not included under a “high” privacy level, such as the manufacturer and / or model of the vehicle and / or vehicle type (e.g., passenger car, sports utility vehicle, truck, etc.). In some embodiments, system 100 may upload data according to a “low” privacy level. Under a “low” privacy level setting, system 100 may upload and include data sufficient to uniquely identify a particular vehicle, its owner / driver and / or part or all of the route the vehicle has traveled. Such “low” privacy level data may include one or more, for example, the VIN, driver / owner name, the vehicle's starting point before departure, the vehicle's intended destination, the vehicle's manufacturer and / or model, and the vehicle type.

[0128] Figure 2A is a side view representation of an exemplary vehicle imaging system according to the disclosed embodiment. Figure 2B is a top view representation of the embodiment shown in Figure 2A. As shown in Figure 2B, the disclosed embodiment may represent a vehicle 200 that includes a system 100 within its body, having a first image acquisition device 122 positioned near the rearview mirror and / or near the driver of the vehicle 200, a second image acquisition device 124 positioned on or within the bumper area (e.g., one of the bumper areas 210) of the vehicle 200, and a processing unit 110.

[0129] As shown in Figure 2C, both image capture devices 122 and 124 can be positioned near the rearview mirror and / or near the driver of the vehicle 200. Furthermore, although two image capture devices 122 and 124 are shown in Figures 2B and 2C, it should be understood that other embodiments may include three or more image capture devices. For example, in the embodiments shown in Figures 2D and 2E, a first image capture device 122, a second image capture device 124, and a third image capture device 126 are included in the system 100 of the vehicle 200.

[0130] As shown in Figure 2D, the image capture device 122 may be positioned near the rearview mirror and / or near the driver of the vehicle 200, and the image capture devices 124 and 126 may be positioned on the bumper area of ​​the vehicle 200 (e.g., one of the bumper areas 210). Also, as shown in Figure 2E, the image capture devices 122, 124 and 126 may be positioned near the rearview mirror and / or near the driver's seat of the vehicle 200. The disclosed embodiments are not limited to any particular number and configuration of image capture devices, and the image capture devices may be positioned in and / or on the vehicle 200 at any suitable location.

[0131] It should be understood that the disclosed embodiments are not limited to vehicles and may be applicable in other situations. It should also be understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and may be applicable to all types of vehicles, including automobiles, trucks, trailers and other types of vehicles.

[0132] The first image capture device 122 may include any suitable type of image capture device. The image capture device 122 may include an optical axis. In one example, the image capture device 122 may include an Aptina M9V024 WVGA sensor with a global shutter. In other embodiments, the image capture device 122 may provide a resolution of 1280 × 960 pixels and may include a rolling shutter. The image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included to provide, for example, a desired focal length and field of view of the image capture device. In some embodiments, a 6 mm lens or a 12 mm lens may be associated with the image capture device 122. In some embodiments, the image capture device 122 may be configured to capture an image having a desired field of view (FOV) 202, as shown in Figure 2D. For example, the image capture device 122 may be configured to have a normal FOV, such as in the range of 40 to 56 degrees, including 46-degree FOV, 50-degree FOV, 52-degree FOV, or degrees exceeding 52 degrees FOV. Alternatively, the image capture device 122 may be configured to have a narrow FOV in the range of 23 to 40 degrees, such as 28-degree FOV or 36-degree FOV. In addition, the image capture device 122 may be configured to have a wide FOV in the range of 100 to 180 degrees. In some embodiments, the image capture device 122 may include a wide-angle bumper camera or a bumper camera having an FOV of up to 180 degrees. In some embodiments, the image capture device 122 may be a 7.2M pixel image capture device with an aspect ratio of about 2:1 (e.g., H×V=3800×1900 pixels) and a horizontal FOV of about 100 degrees. Such an image capture device may be used as an alternative to a three-dimensional image capture device configuration. Due to significant lens distortion, the vertical field of view (FOV) of such an image capture device can be much lower than 50 degrees in implementations where the image capture device uses a radially symmetric lens. For example, such a lens may not be radially symmetric, thereby allowing a vertical FOV greater than 50 degrees with a horizontal FOV of 100 degrees.

[0133] The first image acquisition device 122 can acquire multiple first images of a scene associated with the vehicle 200. Each of the multiple first images may be acquired as a series of image scan lines, which may be captured using a rolling shutter. Each scan line may contain multiple pixels.

[0134] The first image acquisition device 122 may have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate may refer to the rate at which the image sensor can acquire image data associated with each pixel contained in a particular scan line.

[0135] The image acquisition devices 122, 124, and 126 may include any suitable type and number of image sensors, including, for example, CCD sensors or CMOS sensors. In one embodiment, a CMOS image sensor may be used in conjunction with a rolling shutter, so that each pixel in a row is read one at a time, and the row scanning proceeds row by row until the entire image frame is captured. In some embodiments, rows may be captured sequentially from top to bottom relative to the frame.

[0136] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124, and 126) may constitute a high-resolution imager and may have a resolution of more than 5 megapixels, more than 7 megapixels, more than 10 megapixels, or even higher.

[0137] The use of a rolling shutter can result in pixels within different rows being exposed and captured at different times, potentially leading to skew and other image artifacts in the captured image frame. On the other hand, if the image capture device 122 is configured to operate with a global or synchronous shutter, all pixels may be exposed during a common exposure period over the same amount of time. As a result, image data within a frame collected from a system utilizing a global shutter represents a snapshot of the entire FOV (FOV 202, etc.) at a particular time. Conversely, when a rolling shutter is applied, each row within the frame is exposed, and the data is captured at different times. Therefore, moving objects may appear distorted in image capture devices with a rolling shutter. This phenomenon is described in more detail below.

[0138] The second image capture device 124 and the third image capture device 126 can be any type of image capture device. Like the first image capture device 122, each of the image capture devices 124 and 126 may include an optical axis. In one embodiment, each of the image capture devices 124 and 126 may include an Aptina M9V024 WVGA sensor with a global shutter. Alternatively, each of the image capture devices 124 and 126 may include a rolling shutter. Like the image capture device 122, the image capture devices 124 and 126 may be configured to include various lenses and optical elements. In some embodiments, the lenses associated with the image capture devices 124 and 126 may be the same as the FOV associated with the image capture device 122 (FOV 202, etc.) or provide a narrower FOV (FOV 204 and 206, etc.). For example, the image acquisition devices 124 and 126 may have an FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less than 20 degrees.

[0139] Image acquisition devices 124 and 126 may acquire a plurality of second and third images for a scene associated with the vehicle 200. Each of the plurality of second and third images may be acquired as a second and third series of image scan lines, which may be captured using a rolling shutter. Each scan line or each line may have a plurality of pixels. Image acquisition devices 124 and 126 may have second and third scan rates associated with the acquisition of each image scan line contained within the second and third series.

[0140] Each image capture device 122, 124, and 126 can be positioned at any suitable location and in any suitable orientation relative to the vehicle 200. The relative positions of the image capture devices 122, 124, and 126 can be selected to facilitate the fusion of information acquired from the image capture devices. For example, in some embodiments, the FOV associated with image capture device 124 (FOV 204) may partially or completely overlap with the FOV associated with image capture device 122 (FOV 202, etc.) and the FOV associated with image capture device 126 (FOV 206, etc.).

[0141] The image acquisition devices 122, 124, and 126 can be positioned on the vehicle 200 at any suitable relative height. In one example, there may be height differences between the image acquisition devices 122, 124, and 126, and these height differences may provide sufficient parallax information to enable stereoscopic analysis. For example, as shown in Figure 2A, the two image acquisition devices 122 and 124 are at different heights. There may also be lateral displacement differences between the image acquisition devices 122, 124, and 126, which provide additional parallax information for stereoscopic analysis by, for example, the processing unit 110. Lateral displacement differences are as shown in Figures 2C and 2D, d x This can be shown as follows. In some embodiments, a frontal or rearward displacement (e.g., range displacement) may exist between the image acquisition devices 122, 124, and 126. For example, image acquisition device 122 may be positioned 0.5 to 2 meters or more behind image acquisition devices 124 and / or image acquisition devices 126. With this type of displacement, one of the image acquisition devices may be able to cover a potential blind spot of the other image acquisition devices.

[0142] The image capture device 122 may have any suitable resolution capability (e.g., the number of pixels associated with the image sensor), and the resolution of the image sensor associated with the image capture device 122 may be higher, lower, or the same as the resolution of the image sensors associated with the image capture devices 124 and 126. In some embodiments, the image sensors associated with the image capture device 122 and / or the image capture devices 124 and 126 may have a resolution of 640×480, 1024×768, 1280×960, or any other suitable resolution.

[0143] The frame rate (e.g., the rate at which an image capture device acquires a set of pixel data for one image frame before moving on to acquiring the pixel data associated with the next image frame) may be controllable. The frame rate associated with image capture device 122 may be higher, lower, or the same as the frame rates associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 may depend on various factors that may affect the timing of the frame rate. For example, one or more of the image capture devices 122, 124, and 126 may include a selectable pixel delay period that is imposed before or after the acquisition of image data associated with one or more pixels of the image sensors within image capture devices 122, 124, and / or 126. Generally, the image data corresponding to each pixel may be acquired according to the device's clock rate (e.g., one pixel per clock cycle). Furthermore, in embodiments including a rolling shutter, one or more of the image capture devices 122, 124, and 126 may include a selectable horizontal blanking period that is imposed before or after acquisition of image data associated with a pixel row of the image sensor in the image capture devices 122, 124, and / or 126. Furthermore, one or more of the image capture devices 122, 124, and / or 126 may include a selectable vertical blanking period that is imposed before or after acquisition of image data associated with an image frame of the image capture devices 122, 124, and 126.

[0144] These timing controls make it possible to synchronize the frame rates associated with image acquisition devices 122, 124, and 126, even if the line scan rates of each image acquisition device are different. Furthermore, as will be discussed in more detail below, these selectable timing controls, in particular among factors (e.g., image sensor resolution, maximum line scan rate, etc.), make it possible to synchronize image acquisition from areas where the FOV of image acquisition device 122 overlaps with one or more FOVs of image acquisition devices 124 and 126, even if the field of view of image acquisition device 122 is different from the FOV of image acquisition devices 124 and 126.

[0145] The frame rate timing in image acquisition devices 122, 124, and 126 may depend on the resolution of the associated image sensor. For example, assuming that both devices have similar line scan rates, and one device includes an image sensor with a resolution of 640 × 480 and the other device includes an image sensor with a resolution of 1280 × 960, acquiring frames of image data from the sensor with a higher resolution will require a longer time.

[0146] Another factor that may affect the timing of image data acquisition in image capture devices 122, 124, and 126 is the maximum line scan rate. For example, acquiring a line of image data from the image sensors included in image capture devices 122, 124, and 126 requires some minimum amount of time. Assuming no pixel delay period is added, this minimum amount of time to acquire a line of image data will be related to the maximum line scan rate of a particular device. Devices that offer a higher maximum line scan rate have the potential to offer a higher frame rate than devices with a lower maximum line scan rate. In some embodiments, one or both of image capture devices 124 and 126 may have a higher maximum line scan rate than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture device 124 and / or 126 may be 1.25 times, 1.5 times, 1.75 times, or 2 times or more the maximum line scan rate of image capture device 122.

[0147] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may operate at a scan rate less than or equal to its maximum scan rate. The system may be configured so that one or both of image capture devices 124 and 126 operate at a line scan rate equal to that of image capture device 122. In other examples, the system may be configured so that the line scan rates of image capture devices 124 and / or 126 are 1.25 times, 1.5 times, 1.75 times, or 2 times or more the line scan rate of image capture device 122.

[0148] In some embodiments, the image acquisition devices 122, 124, and 126 may be asymmetrical. That is, these image acquisition devices may include cameras having different fields of view (FOV) and focal lengths. The fields of view of the image acquisition devices 122, 124, and 126 may include, for example, any desired area of ​​the environment of the vehicle 200. In some embodiments, one or more of the image acquisition devices 122, 124, and 126 may be configured to acquire image data from the environment in front of the vehicle 200, the environment behind the vehicle 200, the environments on both sides of the vehicle 200, or a combination thereof.

[0149] Furthermore, the focal lengths associated with each image acquisition device 122, 124, and / or 126 may be selectable so that each device acquires images of objects within a desired distance range from the vehicle 200 (e.g., by incorporating an appropriate lens). For example, in some embodiments, the image acquisition devices 122, 124, and 126 may acquire images of nearby objects within a few meters of the vehicle. The image acquisition devices 122, 124, and 126 may also be configured to acquire images of objects at a greater distance from the vehicle (e.g., 25m, 50m, 100m, 150m, or beyond). Furthermore, the focal lengths of the image acquisition devices 122, 124, and 126 can be selected so that one image acquisition device (e.g., image acquisition device 122) can acquire images of objects relatively close to the vehicle (e.g., within 10m or 20m), while the other image acquisition devices (e.g., image acquisition devices 124 and 126) can acquire images of objects further away from the vehicle 200 (e.g., beyond 20m, beyond 50m, beyond 100m, beyond 150m, etc.).

[0150] According to several embodiments, the field of view (FOV) of one or more image acquisition devices 122, 124, and 126 may be wide-angle. For example, it may be advantageous for image acquisition devices 122, 124, and 126, which can be used to acquire images of the immediate vicinity of the vehicle 200, to have a 140-degree FOV. For example, image acquisition device 122 may be used to acquire images of the right or left area of ​​the vehicle 200, and in such embodiments, it may be desirable for image acquisition device 122 to have a wide FOV (e.g., at least 140 degrees).

[0151] The fields of view associated with each of the image acquisition devices 122, 124, and 126 may depend on their respective focal lengths. For example, as the focal length increases, the corresponding field of view decreases.

[0152] Image capture devices 122, 124, and 126 can be configured to have any suitable field of view. In one particular example, image capture device 122 may have a horizontal FOV of 46 degrees, image capture device 124 may have a horizontal FOV of 23 degrees, and image capture device 126 may have a horizontal FOV of 23 to 46 degrees. In another example, image capture device 122 may have a horizontal FOV of 52 degrees, image capture device 124 may have a horizontal FOV of 26 degrees, and image capture device 126 may have a horizontal FOV of 26 to 52 degrees. In some embodiments, the ratio of the FOV of image capture device 122 to the FOV of image capture device 124 and / or image capture device 126 may vary from 1.5 to 2.0. In other embodiments, this ratio may vary from 1.25 to 2.25.

[0153] System 100 may be configured such that the field of view of image capture device 122 at least partially or completely overlaps with the fields of view of image capture device 124 and / or image capture device 126. In some embodiments, System 100 may be configured such that the fields of view of image capture devices 124 and 126 fall within the field of view of image capture device 122 (for example, being smaller than the field of view of image capture device 122) and share a common center with the field of view of image capture device 122. In other embodiments, image capture devices 122, 124 and 126 may capture adjacent FOVs or have partially overlapping FOVs. In some embodiments, the fields of view of image capture devices 122, 124 and 126 may be aligned such that the centers of image capture device 124 and / or 126 with narrower FOVs are located in the lower half of the field of view of device 122 with wider FOVs.

[0154] Figure 2F is a graphical representation of an exemplary vehicle control system according to the disclosed embodiment. As shown in Figure 2F, the vehicle 200 may include a throttle system 220, a brake system 230, and a steering system 240. System 100 may provide input (e.g., control signals) to one or more of the throttle system 220, brake system 230, and steering system 240 via one or more data links (e.g., any wired and / or wireless links or data transmission links). For example, based on the analysis of images acquired by image acquisition devices 122, 124, and / or 126, System 100 may provide control signals to one or more of the throttle system 220, brake system 230, and steering system 240 to navigate the vehicle 200 (e.g., by causing acceleration, turning, lane shifting, etc.). Furthermore, system 100 may receive inputs indicating the operating status of the vehicle 200 (e.g., speed, whether the vehicle 200 is braking and / or turning, etc.) from one or more of the throttle system 220, brake system 230, and steering system 24. Further details are provided below in reference to Figures 4 to 7.

[0155] As shown in Figure 3A, the vehicle 200 may also include a user interface 170 for interacting with the driver or occupants of the vehicle 200. For example, the user interface 170 in the vehicle application may include a touchscreen 320, a knob 330, buttons 340, and a microphone 350. The driver or occupants of the vehicle 200 may also interact with the system 100 using a steering wheel (e.g., including a turn signal handle, located on or near the steering column of the vehicle 200) and buttons (e.g., located on the steering wheel of the vehicle 200), etc. In some embodiments, the microphone 350 may be positioned adjacent to the rearview mirror 310. Similarly, in some embodiments, an image capture device 122 may be located near the rearview mirror 310. In some embodiments, the user interface 170 may also include one or more speakers 360 (e.g., speakers of the vehicle audio system). For example, the system 100 may provide various notifications (e.g., alerts) via the speakers 360.

[0156] Figures 3B to 3D illustrate an exemplary camera mount 370 according to a disclosed embodiment, configured to be positioned behind a rearview mirror (e.g., rearview mirror 310) and facing the vehicle's windshield. As shown in Figure 3B, the camera mount 370 may include image capture devices 122, 124, and 126. The image capture devices 124 and 126 may be positioned behind a glare shield 380, which may be in direct contact with the windshield and may include a composition of film and / or anti-reflective material. For example, the glare shield 380 may be positioned to align with the windshield having a matching incline. In some embodiments, each of the image capture devices 122, 124, and 126 may be positioned behind the glare shield 380, as shown, for example, in Figure 3D. The disclosed embodiments are not limited to any particular configuration of the image capture devices 122, 124, and 126, the camera mount 370, and the glare shield 380. Figure 3C is a view of the camera mount 370 shown in Figure 3B, seen from the front.

[0157] As will be understood by those skilled in the art who benefit from this disclosure, many variations and / or modifications can be made to the embodiments disclosed above. For example, not all components are essential for the operation of system 100. Furthermore, any component may be placed in any suitable part of system 100, and the components may be rearranged into various configurations while providing the functionality of the disclosed embodiments. Thus, the configurations discussed above are examples, and regardless of the configurations described above, system 100 can provide a wide range of functions for analyzing the surroundings of vehicle 200 and navigating vehicle 200 in response to the analysis.

[0158] As will be discussed in more detail below, in various disclosed embodiments, System 100 can provide various features related to autonomous driving and / or driver assistance technologies. For example, System 100 can analyze image data, location data (e.g., GPS location information), map data, speed data and / or data from sensors included in the vehicle 200. System 100 can collect data for analysis from, for example, an image acquisition unit 120, a location sensor 130 and other sensors. Furthermore, System 100 can analyze the collected data to determine whether the vehicle 200 should take a particular action, and then automatically take the determined action without human intervention. For example, if the vehicle 200 is navigating without human intervention, System 100 can automatically control the brakes, acceleration and / or steering of the vehicle 200 (e.g., by transmitting control signals to one or more of the throttle system 220, brake system 230 and steering system 240). Furthermore, System 100 can analyze the collected data and issue warnings and / or alerts to the vehicle's occupants based on the analysis of the collected data. Further details regarding the various embodiments provided by System 100 are provided below.

[0159] Forward-facing multi-imaging system

[0160] As discussed above, system 100 may provide driver assistance functions using a multi-camera system. The multi-camera system may use one or more cameras facing forward of the vehicle. In other embodiments, the multi-camera system may include one or more cameras facing side or rear of the vehicle. In one embodiment, for example, system 100 may use a two-camera imaging system, in which case the first and second cameras (e.g., image acquisition devices 122 and 124) may be positioned at the front and / or side of the vehicle (e.g., vehicle 200). Other camera configurations are also disclosed in embodiments that disclose them, and the configurations disclosed herein are examples. For example, system 100 may include configurations of any number of cameras (e.g., one, two, three, four, five, six, seven, eight, etc.). Furthermore, system 100 may include camera "clusters". For example, a cluster of cameras (including any appropriate number, e.g., one, four, eight, etc.) can be facing forward relative to the vehicle or facing any other direction (e.g., backward, sideways, oblique, etc.). Thus, system 100 may include multiple clusters of cameras, each cluster oriented in a specific direction to capture images from a specific area of ​​the vehicle's environment.

[0161] The first camera may have a field of view that is larger, smaller, or partially overlaps with that of the second camera. Furthermore, the first camera may be connected to a first image processor to perform monocular image analysis of the images provided by the first camera, and the second camera may be connected to a second image processor to perform monocular image analysis of the images provided by the second camera. The outputs of the first and second image processors (e.g., processed information) may be combined. In some embodiments, the second image processor may receive images from both the first and second cameras and perform stereoscopic analysis. In another embodiment, system 100 may use a three-camera imaging system, in which case each camera has a different field of view. Thus, such a system may make decisions based on information derived from objects at various distances both in front of and to the sides of the vehicle. The reference to monocular image analysis may refer to cases where the image analysis is performed based on images captured from a single viewpoint (e.g., a single camera). Stereoscopic image analysis may refer to cases where the image analysis is performed based on two or more images captured with one or more image capture parameters changed. For example, capture images suitable for performing stereoscopic image analysis may include images captured from two or more different locations, images captured from different fields of view, images captured using different focal lengths, images captured with parallax information, and so on.

[0162] For example, in one embodiment, the system 100 may implement a three-camera configuration using image capture devices 122-126. In such a configuration, image capture device 122 may provide a narrow field of view (e.g., 34 degrees or other values ​​selected from the range of about 20-45 degrees), image capture device 124 may provide a wide field of view (e.g., 150 degrees or other values ​​selected from the range of about 100-180 degrees), and image capture device 126 may provide a medium field of view (e.g., 46 degrees or other values ​​selected from the range of about 35-60 degrees). In some embodiments, image capture device 126 may operate as the primary or first camera. The image capture devices 122-126 may be positioned substantially side by side (e.g., 6 cm apart) behind the rearview mirror 310. Furthermore, in some embodiments, as discussed above, one or more of the image capture devices 122-126 may be mounted behind the glare shield 380, which is coplanar with the windshield of the vehicle 200. Such shields may be designed to minimize the impact on any reflective image-capturing devices 122-126 from inside the vehicle.

[0163] In another embodiment, as discussed above in relation to Figures 3B and 3C, a wide-field camera (e.g., image acquisition device 124 in the above example) may be mounted lower than a narrow primary-field camera (e.g., image acquisition devices 122 and 126 in the above example). This configuration may provide a free line of sight from the wide-field camera. To reduce reflections, the camera may be mounted near the windshield of the vehicle 200 and may include a polarizer to reduce reflected light.

[0164] A three-camera system can offer specific performance characteristics. For example, some embodiments may include a function to verify object detection by one camera based on detection results from another camera. In the three-camera configuration discussed above, the processing unit 110 may include, for example, three processing devices (e.g., three EyeQ series processor chips as discussed above), each processing device directed to process images captured by one or more image capture devices 122-126.

[0165] In a three-camera system, the first processing device can receive images from both the main camera and the narrow-field-of-view camera, and perform vision processing on the narrow-field-of-view camera to detect, for example, other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the first processing device can calculate the pixel mismatch between the image from the main camera and the image from the narrow-field-of-view camera and create a 3D reconstruction of the vehicle 200's environment. Next, the first processing device can combine the 3D reconstruction with 3D map data or 3D information calculated based on information from another camera.

[0166] A second processing device may receive images from the main camera and perform vision processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the second processing device may calculate camera displacement and, based on the displacement, calculate pixel mismatches between consecutive images to create a 3D reconstruction of the scene (e.g., structure from motion). The second processing device may transmit the structure from motion based on the 3D reconstruction to the first processing device and combine the structure from motion with a stereoscopic 3D image.

[0167] A third processing device may receive images from a wide-field-of-view camera and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing device may further execute additional processing commands to analyze the images and identify moving objects in the images, such as vehicles changing lanes or pedestrians.

[0168] In some embodiments, independently capturing and processing image-based information streams may provide an opportunity to offer redundancy in the system. Such redundancy may, for example, verify and / or supplement information obtained by capturing and processing image information from at least a second image capturing device using a first image capturing device and the images processed from that device.

[0169] In some embodiments, system 100 may use two image capture devices (e.g., image capture devices 122 and 124) to provide navigation assistance to vehicle 200, and a third image capture device (e.g., image capture device 126) may be used to provide redundancy and verify the analysis of data received from the other two image capture devices. For example, in such a configuration, image capture devices 122 and 124 may provide stereoscopic analysis images by system 100 for navigating vehicle 200, while image capture device 126 may provide images for monocular analysis by system 100 to provide redundancy and verification of information obtained based on images captured from image capture devices 123 and / or image capture devices 124. That is, image capture device 126 (and corresponding processing device) may be considered to provide a redundant subsystem that provides checks on the analysis derived from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Furthermore, in some embodiments, the redundancy and verification of received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors, information received from one or more transceivers outside the vehicle).

[0170] Those skilled in the art will recognize that the above camera configuration, camera arrangement, number of cameras, camera positions, etc., are merely illustrative. These components described in relation to the overall system can be assembled and used in various different configurations without departing from the scope of the disclosed embodiments. Further details regarding the use of the multi-camera system to provide driver assistance and / or autonomous vehicle functions follow below.

[0171] Figure 4 is an exemplary functional block diagram of memories 140 and / or 150 that can store / program instructions for performing one or more operations according to the disclosed embodiments. Hereafter, we will refer to memory 140, but those skilled in the art will recognize that instructions can be stored in memories 140 and / or 150.

[0172] As shown in Figure 4, memory 140 may store a monocular image analysis module 402, a stereoscopic image analysis module 404, a velocity and acceleration module 406, and a navigation response module 408. The disclosed embodiments are not limited to any particular configuration of memory 140. Furthermore, the application processor 180 and / or the image processor 190 may execute instructions stored in any of the modules 402-408 contained in memory 140. Those skilled in the art will understand that the reference to processing unit 110 in the following discussion may refer to the application processor 180 and the image processor 190 individually or collectively. Accordingly, any step of the following process may be performed by one or more processing devices.

[0173] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) to perform monocular image analysis of a set of images acquired by one of the image acquisition devices 122, 124, and 126, when executed by the processing unit 110. In some embodiments, the processing unit 110 may perform monocular image analysis by combining information from the set of images with additional sensory information (e.g., information from radar). As described below in relation to Figures 5A to 5D, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous materials, and any other features related to the vehicle's environment. Based on the analysis, the system 100 may cause one or more navigation responses, such as turns, lane shifts, and acceleration changes, in the vehicle 200 (e.g., via the processing unit 110), as discussed below in relation to the navigation response module 408.

[0174] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) to perform monocular image analysis of a set of images acquired by one of the image acquisition devices 122, 124, and 126, when executed by the processing unit 110. In some embodiments, the processing unit 110 may perform monocular image analysis by combining information from the set of images with additional sensory information (e.g., information from radar or LiDAR). As described below in relation to Figures 5A to 5D, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous materials, and any other features related to the vehicle's environment. Based on the analysis, the system 100 may cause one or more navigation responses, such as turns, lane shifts, and acceleration changes, in the vehicle 200, as discussed below in relation to determining navigation responses (e.g., by the processing unit 110).

[0175] In one embodiment, the stereoscopic image analysis module 404 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform stereoscopic image analysis of first and second sets of images acquired by a combination of image acquisition devices selected from image acquisition devices 122, 124, and 126. In some embodiments, the processing unit 110 may perform stereoscopic image analysis by combining information from the first and second sets of images with additional sensory information (e.g., information from radar). For example, the stereoscopic image analysis module 404 may include instructions to perform stereoscopic image analysis based on a first set of images acquired by image acquisition device 124 and a second set of images acquired by image acquisition device 126. As will be described below with reference to Figure 6, the stereoscopic image analysis module 404 may include instructions to detect sets of features in the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and hazardous materials. Based on the analysis, the processing unit 110 may cause one or more navigation responses, such as turns, lane shifts, and acceleration changes, in the vehicle 200, as described later in relation to the navigation response module 408. Furthermore, in some embodiments, the stereoscopic image analysis module 404 may implement techniques related to trained systems (such as neural networks or deep neural networks) or untrained systems.

[0176] In one embodiment, the velocity and acceleration module 406 may store software configured to analyze data received from one or more computing and electromechanical devices within the vehicle 200 that are configured to change the velocity and / or acceleration of the vehicle 200. For example, the processing unit 110 may execute commands associated with the velocity and acceleration module 406 to calculate the target velocity of the vehicle 200 based on data derived from the execution of the monocular image analysis module 402 and / or stereoscopic image analysis module 404. Such data may include, for example, target position, velocity and / or acceleration, the position and / or velocity of the vehicle 200 relative to nearby vehicles, pedestrians or road objects, and position information of the vehicle 200 relative to road lane markings. In addition, the processing unit 110 may calculate the target velocity of the vehicle 200 based on sensory input (e.g., information from radar) and input from other systems of the vehicle 200, such as the vehicle's throttle system 220, brake system 230 and / or steering system 240. Based on the calculated target speed, the processing unit 110 may transmit electronic signals to the vehicle 200's throttle system 220, brake system 230 and / or steering system 240 to trigger changes in speed and / or acceleration, for example, by physically reducing the brakes or reducing the accelerator of the vehicle 200.

[0177] In one embodiment, the navigation response module 408 is executable by the processing unit 110 and may store software that determines a desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereoscopic image analysis module 404. Such data may include position and speed information associated with nearby vehicles, pedestrians and road objects, as well as target position information for the vehicle 200. Furthermore, in some embodiments, the navigation response may be based (partially or entirely) on map data, a predetermined position of the vehicle 200, and / or relative velocity or relative acceleration between the vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereoscopic image analysis module 404. The navigation response module 408 may also determine a desired navigation response based on sensory input (e.g., information from radar) and input from other systems of the vehicle 200, such as the vehicle's throttle system 220, brake system 230, and steering system 240. Based on the desired navigation response, the processing unit 110 may trigger the desired navigation response by transmitting electronic signals to the vehicle 200's throttle system 220, brake system 230, and steering system 240, for example, by turning the steering wheel of the vehicle 200 to achieve a rotation of a predetermined angle. In some embodiments, the processing unit 110 may use the output of the navigation response module 408 (e.g., the desired navigation response) as input to the execution of the speed and acceleration module 406 for calculating changes in the vehicle 200's speed.

[0178] Furthermore, any of the modules disclosed herein (e.g., modules 402, 404, and 406) can implement techniques related to trained systems (such as neural networks or deep neural networks) or untrained systems.

[0179] Figure 5A is a flowchart illustrating an exemplary process 500A that generates one or more navigation responses based on monocular image analysis according to a disclosed embodiment. In step 510, the processing unit 110 may receive multiple images via a data interface 128 between the processing unit 110 and the image acquisition unit 120. For example, a camera included in the image acquisition unit 120 (such as an image capture device 122 having a field of view 202) may capture multiple images of an area in front of the vehicle 200 (or, for example, the side or rear of the vehicle) and transmit them to the processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). The processing unit 110 may perform monocular image analysis using a monocular image analysis module 402 to analyze the multiple images in step 520, as will be described in more detail below in relation to Figures 5B to 5D. By performing the analysis, the processing unit 110 may detect sets of features within the image set, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, and traffic lights.

[0180] In step 520, the processing unit 110 can also run the monocular image analysis module 402 to detect various road hazards, such as truck tire parts, fallen road signs, loose cargo, and small animals. The structure, shape, size, and color of road hazards can vary, making their detection more difficult. In some embodiments, the processing unit 110 can run the monocular image analysis module 402 to perform multi-frame analysis on multiple images to detect road hazards. For example, the processing unit 110 can estimate camera movement between consecutive image frames and calculate pixel mismatches between frames to construct a 3D map of the road. The processing unit 110 can then use the 3D map to detect the road surface and any hazards present on the road surface.

[0181] In step 530, the processing unit 110 may execute the navigation response module 408 to cause one or more navigation responses to the vehicle 200 based on the analysis performed in step 520 and the techniques described above in relation to Figure 4. Navigation responses may include, for example, turns, lane shifts, and acceleration changes. In some embodiments, the processing unit 110 may use data derived from the execution of the speed and acceleration module 406 to cause one or more navigation responses. Furthermore, the multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof. For example, the processing unit 110 may cause the vehicle 200 to cross one lane and then accelerate, for example, by sequentially sending control signals to the steering system 240 and throttle system 220 of the vehicle 200. Alternatively, the processing unit 110 may cause the vehicle 200 to brake and simultaneously shift lanes by simultaneously sending control signals to the brake system 230 and steering system 240 of the vehicle 200.

[0182] Figure 5B is a flowchart illustrating an exemplary process 500B for detecting one or more vehicles and / or pedestrians in a set of images according to a disclosed embodiment. Processing unit 110 may perform process 500B by running monocular image analysis module 402. In step 540, processing unit 110 may identify a set of candidate objects that may represent vehicles and / or pedestrians. For example, processing unit 110 may scan one or more images, compare the images to one or more predetermined patterns, and identify locations within each image that may contain a target object (e.g., a vehicle, a pedestrian, or a part thereof). The predetermined patterns may be specified to achieve a low rate of "false hits" and a low rate of "misses". For example, processing unit 110 may use a low similarity threshold to a predetermined pattern to identify a candidate object as a possible vehicle or pedestrian. In doing so, processing unit 110 may reduce the probability of missing (e.g., not identifying) a candidate object that represents a vehicle or pedestrian.

[0183] In step 542, the processing unit 110 may filter the set of candidate objects to exclude certain candidates (e.g., irrelevant or unrelated objects) based on classification criteria. Such criteria may be derived from various characteristics associated with object types stored in a database (e.g., a database stored in memory 140). Characteristics may include the shape, dimensions, texture, and location of the object (e.g., relative to the vehicle 200), etc. Thus, the processing unit 110 may use one or more sets of criteria to reject false candidates from the set of candidate objects.

[0184] In step 544, the processing unit 110 may analyze multiple image frames to determine whether an object in a set of candidate images represents a vehicle and / or a pedestrian. For example, the processing unit 110 may track the detected candidate object across consecutive frames and accumulate frame-by-frame data associated with the detected object (e.g., size, position relative to the vehicle 200, etc.). Furthermore, the processing unit 110 may estimate the parameters of the detected object and compare the frame-by-frame position data of the object with a predicted position.

[0185] In step 546, the processing unit 110 may construct a set of measurements of the detected object. Such measurements may include, for example, the position, velocity, and acceleration values ​​(relative to the vehicle 200) associated with the detected object. In some embodiments, the processing unit 110 may construct measurements based on estimation techniques that use a series of time-based observations, such as a Kalman filter or linear quadratic estimation (LQE), and / or modeling data available for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). The Kalman filter is obtained based on a measurement of the scale of the object, where the scale measurement is proportional to the time to collision (e.g., the amount of time until the vehicle 200 reaches the object). Thus, by performing steps 540-546, the processing unit 110 may identify vehicles and pedestrians appearing in the set of captured images and derive information (e.g., position, velocity, size) associated with the vehicles and pedestrians. Based on the identified and derived information, the processing unit 110 may produce one or more navigation responses in the vehicle 200, as described above in relation to Figure 5A.

[0186] In step 548, the processing unit 110 may perform optical flow analysis on one or more images to reduce the probability of detecting a "false hit" and missing a candidate object representing a vehicle or pedestrian. Optical flow analysis may refer to, for example, analyzing a movement pattern separate from the road surface movement for vehicle 200 in one or more images associated with other vehicles and pedestrians. The processing unit 110 may calculate the movement of a candidate object by observing the different positions of the object across multiple image frames captured at different times. The processing unit 110 may calculate the movement of a candidate object by using position and time values ​​as input to a mathematical model. Thus, optical flow analysis may provide an alternative method for detecting vehicles and pedestrians near vehicle 200. The processing unit 110 may perform optical flow analysis in combination with steps 540-546 to provide redundancy in detecting vehicles and pedestrians and increase the reliability of system 100.

[0187] Figure 5C is a flowchart illustrating an exemplary process 500C for detecting road mark and / or lane geometry information within a set of images according to a disclosed embodiment. Processing unit 110 may perform process 500C by running monocular image analysis module 402. In step 550, processing unit 110 may detect sets of objects by scanning one or more images. To detect lane mark segments, lane geometry information and other related road marks, processing unit 110 may filter the sets of objects to exclude those deemed irrelevant (e.g., small holes, small rocks, etc.). In step 552, processing unit 110 may group together segments detected in step 550 that belong to the same road mark or lane mark. Based on the grouping, processing unit 110 may develop a model, such as a mathematical model, to represent the detected segments.

[0188] In step 554, the processing unit 110 may construct a set of measurements associated with the detected segment. In some embodiments, the processing unit 110 may create a projection of the detected segment from the image plane onto the real-world plane. The projection may be characterized using a cubic polynomial with coefficients corresponding to physical properties such as the detected road position, slope, curvature, and curvature derivative. In generating the projection, the processing unit 110 may take into account changes in the road surface as well as the pitch and roll rates associated with the vehicle 200. In addition, the processing unit 110 may model the road height by analyzing the position and motion cues present on the road surface. Furthermore, the processing unit 110 may estimate the pitch and roll rates associated with the vehicle 200 by tracking sets of feature points in one or more images.

[0189] In step 556, the processing unit 110 may perform multi-frame analysis, for example, by tracking detection segments across consecutive image frames and accumulating frame-by-frame data associated with the detection segments. When the processing unit 110 performs multi-frame analysis, the set of measurements constructed in step 554 may become more reliable and can be associated with increasingly higher confidence. Thus, by performing steps 550 to 556, the processing unit 110 may identify road marks appearing in the set of captured images and derive lane geometry information. Based on the identified and derived information, the processing unit 110 may generate one or more navigation responses in the vehicle 200, as described above in relation to Figure 5A.

[0190] In step 558, the processing unit 110 may further develop a safety model of the vehicle 200 in the surrounding environment by considering additional information sources. Using the safety model, the processing unit 110 may define the conditions under which the system 100 can safely perform autonomous control of the vehicle 200. To develop the safety model, in some embodiments, the processing unit 110 may consider the positions and movements of other vehicles, detected road edges and barriers, and / or general road shape descriptions extracted from map data (such as data from the map database 160). By considering additional information sources, the processing unit 110 may provide redundancy for detecting road marks and lane geometry, thereby increasing the reliability of the system 100.

[0191] Figure 5D is a flowchart illustrating an exemplary process 500D for detecting a traffic light in a set of images according to a disclosed embodiment. Processing unit 110 may perform process 500D by running monocular image analysis module 402. In step 560, processing unit 110 may scan the set of images and identify objects appearing at locations in the images that are likely to contain traffic lights. For example, processing unit 110 may filter the identified objects to construct a set of candidate objects that exclude objects that are less likely to correspond to traffic lights. Filtering may be based on various characteristics associated with traffic lights, such as shape, dimensions, texture, and location (e.g., relative to vehicle 200). Such characteristics may be obtained based on many examples of traffic lights and traffic control signals and stored in a database. In some embodiments, processing unit 110 may perform multi-frame analysis on the set of candidate objects that reflect possible traffic lights. For example, processing unit 110 may track candidate objects across consecutive image frames, estimate the real-world location of the candidate objects, and filter out objects that are moving (less likely to be traffic lights). In some embodiments, the processing unit 110 may perform color analysis on candidate objects to identify the relative positions of detected colors that may be represented inside a potential traffic light.

[0192] In step 562, the processing unit 110 may analyze the geometry of the intersection. The analysis may be based on any combination of (i) the number of lanes detected on both sides of the vehicle 200, (ii) marks detected on the road (such as arrow marks), and (iii) descriptions of the intersection extracted from map data (such as data from the map database 160). The processing unit 110 may perform the analysis using information derived from the execution of the monocular analysis module 402. In addition, the processing unit 110 may identify the correspondence between the traffic lights detected in step 560 and the lanes that appear near the vehicle 200.

[0193] As the vehicle 200 approaches the intersection, in step 564, the processing unit 110 may update the confidence level associated with the analyzed intersection geometry and detected traffic lights. For example, the number of traffic lights estimated to appear at the intersection compared to the number actually appearing at the intersection may affect the confidence level. Based on the confidence level, the processing unit 110 may delegate control to the driver of the vehicle 200 to improve the safety situation. By performing steps 560 to 564, the processing unit 110 may identify the traffic lights appearing in the set of captured images and analyze the intersection geometry information. Based on the identification and analysis, the processing unit 110 may produce one or more navigation responses in the vehicle 200 as described above in relation to Figure 5A.

[0194] Figure 5E is a flowchart of an exemplary process 500E according to a disclosed embodiment that generates one or more navigation responses in a vehicle 200 based on a vehicle path. In step 570, the processing unit 110 may construct an initial vehicle path associated with the vehicle 200. The vehicle path may be represented using a set of points represented by coordinates (x,y) and the distance d between any two points in the set of points. i This can be within a range of 1 to 5 meters. In one embodiment, the processing unit 110 may construct an initial vehicle path using two polynomials, such as left and right road polynomials. The processing unit 110 calculates the geometric midpoint between the two polynomials and, if there is a predetermined offset (offset 0 may correspond to driving in the center of the lane), may offset each point included in the resulting vehicle path by a predetermined offset (e.g., smart lane offset). The offset may be perpendicular to the division between any two points in the vehicle path. In another embodiment, the processing unit 110 may use one polynomial and an estimated lane width to offset each point in the vehicle path by half the estimated lane width plus a predetermined offset (e.g., smart lane offset).

[0195] In step 572, the processing unit 110 may update the vehicle route constructed in step 570. The processing unit 110 calculates the distance d between two points within the set of points representing the vehicle route k to be shorter than the above-described distance d i and may reconstruct the vehicle route constructed in 570 using a higher resolution. For example, the distance d k may be in the range of 0.1 to 0.3 meters. The processing unit 110 may reconstruct the vehicle route using a parabolic spline algorithm, which may result in a cumulative distance vector S corresponding to the total length of the vehicle route (i.e., based on the set of points representing the vehicle route).

[0196] In step 574, the processing unit 110 may identify a look-ahead point (represented in coordinates as (x l , z l )) based on the updated vehicle route performed in step 572. The processing unit 110 may extract the look-ahead point from the cumulative distance vector S and may associate a look-ahead distance and a look-ahead time with the look-ahead point. The look-ahead distance may have a lower limit range of 10 to 20 meters and may be calculated as the product of the speed of the vehicle 200 and the look-ahead time. For example, as the speed of the vehicle 200 decreases, the look-ahead distance may also decrease (e.g., until it reaches the lower limit). The look-ahead time, which may be in the range of 0.5 to 1.5 seconds, may be inversely proportional to the gain of one or more control loops associated with causing a navigation response in the vehicle 200, such as a forward error tracking control loop. For example, the gain of the forward error tracking control loop may depend on the bandwidths of a yaw rate loop, a steering actuator loop, and vehicle lateral dynamics. Therefore, the higher the gain of the forward error tracking control loop, the shorter the look-ahead time.

[0197] In step 576, the processing unit 110 may determine a forward error and a yaw rate command based on the look-ahead point identified in step 574. The processing unit 110 calculates the arctangent of the look-ahead point, e.g., arctan(x l / z lThe progress error can be identified by calculating the following. The processing unit 110 can determine the yaw rate command as the product of the progress error and the high-level control gain. The high-level control gain may be equal to (2 / look-ahead time) if the look-ahead distance is not at the lower limit. If the look-ahead distance is at the lower limit, the high-level control gain may be equal to (2*vehicle speed / look-ahead distance).

[0198] Figure 5F is a flowchart illustrating an exemplary process 500F according to a disclosed embodiment for determining whether a preceding vehicle is changing lanes. In step 580, processing unit 110 may identify navigation information associated with the preceding vehicle (e.g., a vehicle moving ahead of vehicle 200). For example, processing unit 110 may identify the position, speed (e.g., direction and velocity) and / or acceleration of the preceding vehicle using the techniques described above in relation to Figures 5A and 5B. Processing unit 110 may also identify one or more road polynomials, lookahead points (associated with vehicle 200) and / or snail trails (e.g., sets of points describing the path taken by the preceding vehicle) using the techniques described above in relation to Figure 5E.

[0199] In step 582, the processing unit 110 may analyze the navigation information identified in step 580. In one embodiment, the processing unit 110 may calculate the distance (e.g., along the trail) between the snail trail and the road polynomial. If the difference in this distance along the trail exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters for straight roads, 0.3 to 0.4 meters for gently curving roads, and 0.5 to 0.6 meters for sharply curving roads), the processing unit 110 may determine that the preceding vehicle is likely changing lanes. If it is detected that multiple vehicles are traveling ahead of vehicle 200, the processing unit 110 may compare the snail trails associated with each vehicle. Based on the comparison, the processing unit 110 may determine that a vehicle whose snail trail does not match the snail trail of another vehicle is likely changing lanes. The processing unit 110 may further compare the curvature of the snail trail (associated with the preceding vehicle) with the expected curvature of the road section the preceding vehicle is traveling on. The expected curvature can be extracted from map data (e.g., data from map database 160), road polynomials, snail trails of other vehicles, and prior knowledge about the road. If the difference between the curvature of the snail trail and the expected curvature of the road section exceeds a predetermined threshold, the processing unit 110 may determine that there is a high probability that the preceding vehicle is changing lanes.

[0200] In another embodiment, the processing unit 110 may compare the instantaneous position of a preceding vehicle with a look-ahead point (associated with the vehicle 200) over a specific time period (e.g., 0.5 to 1.5 seconds). If the difference and cumulative sum of the distance differences between the instantaneous position of the preceding vehicle and the look-ahead point over the specific time period exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a gently curving road, and 1.3 to 1.7 meters on a sharply curving road), the processing unit 110 may determine that the preceding vehicle is likely changing lanes. In another embodiment, the processing unit 110 may analyze the geometric shape of the snail trail by comparing the lateral distance traveled along the trail with the expected curvature of the snail trail. The expected radius of curvature is calculated as: (δz 2 +δ x 2 ) / 2 / (δ x ) can be determined according to the formula, where δ x δ represents the lateral movement distance, z represents the vertical travel distance. If the difference between the horizontal travel distance and the expected curvature exceeds a predetermined threshold (e.g., 500-700 meters), the processing unit 110 may determine that the preceding vehicle is likely changing lanes. In another embodiment, the processing unit 110 may analyze the position of the preceding vehicle. If the position of the preceding vehicle obscures the road polynomial (e.g., the preceding vehicle overlaps the road polynomial), the processing unit 110 may determine that the preceding vehicle is likely changing lanes. If the position of the preceding vehicle is such that another vehicle is detected ahead of the preceding vehicle and the snail trails of the two vehicles are not parallel, the processing unit 110 may determine that the (closer) preceding vehicle is likely changing lanes.

[0201] In step 584, the processing unit 110 may determine whether the preceding vehicle 200 is changing lanes based on the analysis performed in step 582. For example, the processing unit 110 may make this determination based on a weighted average of the individual analyses performed in step 582. Under such a scheme, for example, the processing unit 110 may assign a value of "1" to its determination that the preceding vehicle is likely to be changing lanes based on a particular type of analysis (where "0" represents a determination that the preceding vehicle is unlikely to be changing lanes). Different weights may be assigned to different analyses performed in step 582, and the disclosed embodiments are not limited to any particular combination of analyses and weights. Furthermore, in some embodiments, the analysis may utilize a trained system (e.g., a machine learning or deep learning system) that can estimate the future path of a vehicle beyond its current position based on images captured at its current position.

[0202] Figure 6 is a flowchart illustrating an exemplary process 600 that generates one or more navigation responses based on stereoscopic image analysis according to a disclosed embodiment. In step 610, the processing unit 110 may receive a plurality of first and second images via the data interface 128. For example, a camera included in the image acquisition unit 120 (e.g., image capture devices 122 and 124 having fields of view 202 and 204) may capture a plurality of first and second images of an area in front of the vehicle 200 and transmit them to the processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, the processing unit 110 may receive a plurality of first and second images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.

[0203] In step 620, the processing unit 110 may execute the stereoscopic image analysis module 404 to perform stereoscopic image analysis on the first and second sets of images to create a 3D map of the road in front of the vehicle and to detect features in the images such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and road hazards. The stereoscopic image analysis may be performed in the same manner as the steps described above in relation to Figures 5A to 5D. For example, the processing unit 110 may execute the stereoscopic image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road markings, traffic lights, road hazards, etc.) in the first and second sets of images, filter and exclude subsets of candidate objects based on various criteria, perform multi-frame analysis, construct measurements, and determine the confidence level of the remaining candidate objects. In performing the above steps, the processing unit 110 may consider information from both the first and second sets of images, rather than information from only one set of images. For example, the processing unit 110 may analyze the difference in pixel-level data (or other data subsets from the two streams of captured images) of candidate objects that appear in both the first and second sets of images. As another example, the processing unit 110 may estimate the position and / or velocity (e.g., relative to the vehicle 200) of a candidate object by observing that an object appears in one of the sets of images but not in the others, or by other differences that may exist for objects that appear in the two image streams. For example, the position, velocity, and / or acceleration relative to the vehicle 200 may be determined based on the trajectory, position, motion characteristics, etc., of features associated with the object that appear in one or both of the image streams.

[0204] In step 630, the processing unit 110 may execute the navigation response module 408 to generate one or more navigation responses in the vehicle 200 based on the analysis performed in step 620 and the techniques described above in relation to Figure 4. Navigation responses may include, for example, turns, lane shifts, acceleration changes, speed changes, and braking. In some embodiments, the processing unit 110 may use data derived from the execution of the speed and acceleration module 406 to generate one or more navigation responses. Furthermore, the multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.

[0205] Figure 7 is a flowchart illustrating an exemplary process 700 that generates one or more navigation responses based on the analysis of three sets of images according to the disclosed embodiments. In step 710, the processing unit 110 may receive first, second, and third sets of images via the data interface 128. For example, cameras included in the image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) may capture first, second, and third sets of images of the front and / or side areas of the vehicle 200 and transmit them to the processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, the processing unit 110 may receive first, second, and third sets of images via three or more data interfaces. For example, each of the image capture devices 122, 124, and 126 may have an associated data interface for communicating data to the processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.

[0206] In step 720, the processing unit 110 may analyze the first, second, and third sets of images to detect features within the images such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and road hazards. The analysis may be performed in the same manner as the steps described above in relation to Figures 5A to 5D and Figure 6. For example, the processing unit 110 may perform monocular image analysis on each of the first, second, and third sets of images (e.g., based on the execution of the monocular image analysis module 402 and the steps described above in relation to Figures 5A to 5D). Alternatively, the processing unit 110 may perform stereoscopic image analysis on the first and second sets of images, the second and third sets of images, and / or the first and third sets of images (e.g., via the execution of the stereoscopic image analysis module 404 and based on the steps described above in relation to Figure 6). The processed information corresponding to the analysis of the first, second, and / or third sets of images may be combined. In some embodiments, the processing unit 110 may perform a combination of monocular image analysis and stereoscopic image analysis. For example, the processing unit 110 may perform monocular image analysis on a first set of images (e.g., via the execution of monocular image analysis module 402) and stereoscopic image analysis on second and third sets of images (e.g., via the execution of stereoscopic image analysis module 404). The configuration of the image capture devices 122, 124, and 126—including their respective positions and fields of view 202, 204, and 206—may affect the type of analysis performed on the first, second, and third sets of images. The disclosed embodiments are not limited to the specific configuration of the image capture devices 122, 124, and 126 or the type of analysis performed on the first, second, and third sets of images.

[0207] In some embodiments, the processing unit 110 may perform tests on the system 100 based on the images acquired and analyzed in steps 710 and 720. Such tests may provide an indicator of the overall performance of the system 100 in a particular configuration of the image acquisition devices 122, 124, and 126. For example, the processing unit 110 may identify the percentage of "false hits" (e.g., when the system 100 incorrectly determines the presence of a vehicle or pedestrian) and "misses."

[0208] In step 730, the processing unit 110 may generate one or more navigation responses in the vehicle 200 based on information derived from two of the first, second, and third images. The selection of two of the first, second, and third images may depend on various factors, such as the number, type, and size of objects detected in each of the images. The processing unit 110 may make the selection based on the image quality and resolution, the effective field of view reflected in the image, the number of captured frames, and the extent to which the target object actually appears in the frame (e.g., the percentage of frames in which the object appears, the proportion of each such frame in which the object appears, etc.).

[0209] In some embodiments, the processing unit 110 may select information derived from two of a plurality of first, second, and third images by determining the degree to which information derived from one image source is consistent with information derived from other image sources. For example, the processing unit 110 may combine the processed information derived from each of the image acquisition devices 122, 124, and 126 (whether monocular analysis, stereoscopic analysis, or any combination of the two) to identify visual indicators (e.g., lane markings, detected vehicles and / or their positions and / or routes, detected traffic lights, etc.) that are consistent across images captured by each of the image acquisition devices 122, 124, and 126. The processing unit 110 may also exclude information that is inconsistent across the captured images (e.g., vehicles changing lanes, lane models indicating vehicles too close to vehicle 200, etc.). Thus, the processing unit 110 may select information derived from two of a plurality of first, second, and third images based on the identification of consistent and inconsistent information.

[0210] Navigation responses may include, for example, turns, lane shifts, and acceleration changes. Processing unit 110 may generate one or more navigation responses based on the analysis performed in step 720 and the techniques described above in relation to Figure 4. Processing unit 110 may also generate one or more navigation responses using data derived from the execution of the velocity and acceleration module 406. In some embodiments, processing unit 110 may generate one or more navigation responses based on the relative position, relative velocity, and / or relative acceleration between the vehicle 200 and an object detected in any of the first, second, and third images. Multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.

[0211] Reinforcement learning and trained navigation system

[0212] The following sections discuss autonomous driving, along with systems and methods for achieving autonomous vehicle control, whether the vehicle's autonomous control is fully autonomous (autonomous vehicle) or partially autonomous (e.g., one or more drivers assist the system or functions). As shown in Figure 8, the autonomous driving task can be divided into three main modules, including a sensing module 801, a driving policy module 803, and a control module 805. In some embodiments, modules 801, 803, and 805 can be stored in memory units 140 and / or 150 of system 100, and / or modules 801, 803, and 805 (or parts thereof) can be stored away from system 100 (e.g., stored in a server accessible to system 100, for example, by a wireless transceiver 172). Furthermore, any of the modules disclosed herein (e.g., modules 801, 803, and 805) can implement techniques related to trained systems (such as neural networks or deep neural networks) or untrained systems.

[0213] The detection module 801, which can be implemented using the processing unit 110, can handle a variety of tasks related to detecting the navigation state in the host vehicle's environment. Such tasks may depend on inputs from various sensors and detection systems associated with the host vehicle. These inputs may include images or image streams from one or more onboard cameras, GPS location information, accelerometer output, user feedback, user input to one or more user interface devices, radar, lidar, etc. Detections, which may include data from cameras and / or any other available sensors along with map information, can be collected, analyzed, and systematically represented into a "detection state" that describes the information extracted from the scene in the host vehicle's environment. The detection state may particularly include detection information related to target vehicles, lane markings, pedestrians, traffic lights, road geometry, lane shape, obstacles, distance to other objects / vehicles, relative velocity, and relative acceleration, among any potential detection information. Supervised machine learning can be performed to produce detection state outputs based on the detection data provided to the detection module 801. The output of the detection module can represent the detection navigation "state" of the host vehicle, which can be sent to the driving policy module 803.

[0214] Detection states can be developed based on image data received from one or more cameras or image sensors associated with the host vehicle, but detection states used for navigation can be developed using any suitable sensor or combination of sensors. In some embodiments, detection states can be developed without using captured image data. In fact, any of the navigation principles described herein may be applicable to detection states developed based on captured image data as well as detection states developed using other non-image-based sensors. Detection states can also be determined by sources outside the host vehicle. For example, detection states can be developed entirely or partially based on information received from sources distant from the host vehicle (e.g., based on sensor information or processed state information shared from other vehicles, a central server, or any other source of information related to the navigation state of the host vehicle).

[0215] The driving policy module 803, which will be described in more detail below and can be implemented using the processing unit 110, can implement a desired driving policy for determining one or more navigation actions to be performed by the host vehicle in response to a detected navigation state. If there are no other agents (e.g., target vehicles or pedestrians) in the host vehicle's environment, the detected state input to the driving policy module 803 can be processed in a relatively simple manner. If the detected state requires negotiation with one or more other agents, this task becomes more complex. The techniques used to generate the output of the driving policy module 803 may include reinforcement learning (described in more detail below). The output of the driving policy module 803 may include at least one navigation action for the host vehicle, and among potential desired navigation actions, may particularly include a desired acceleration (which may lead to an updated speed for the host vehicle), a desired yaw rate for the host vehicle, and a desired trajectory.

[0216] Based on the output from the driving policy module 803, a control module 805, which can also be implemented using the processing unit 110, can develop control commands for one or more actuators or controlled devices related to the host vehicle. Such actuators and devices may include accelerators, one or more steering controls, brakes, signal transmitters, displays, or any other actuators or devices that can be controlled as part of navigation operations related to the host vehicle. Aspects of control theory can be used to generate the output of the control module 805. To implement a desired navigation goal or requirement of the driving policy module 803, the control module 805 may be responsible for developing and outputting commands to controllable components of the host vehicle.

[0217] Returning to the driving policy module 803, in some embodiments, the driving policy module 803 can be implemented using a trained system trained by reinforcement learning. In other embodiments, the driving policy module 803 can be implemented without machine learning methods by "manually" addressing various scenarios that may occur during autonomous navigation using a specified algorithm. However, such methods, while feasible, may result in driving policies that are too simplistic and lack the flexibility of a machine learning-based trained system. A trained system may be better equipped to handle complex navigation states, better able to determine whether a taxi is parked or stopped to pick up or drop off passengers, better able to determine whether a pedestrian is about to cross the road ahead of the host vehicle, better able to balance self-preservation with unexpected behavior of other drivers, better able to navigate congested roads including target vehicles and / or pedestrians, better able to determine when to interrupt or supplement certain navigation rules, and better able to anticipate undetected but expected conditions (e.g., whether a pedestrian appears from behind a vehicle or obstacle), etc. Reinforcement learning-based trained systems may be better equipped to handle continuous and high-dimensional state spaces as well as continuous action spaces.

[0218] Training a system using reinforcement learning may involve learning a driving policy to map detected states to navigation actions. The driving policy is a function π:S→A, where S is a set of states.

number

[0219] The system can be trained by exposing it to various navigation states, allowing it to apply policies, and rewarding it (based on a reward function designed to reward desired navigation behaviors). Based on the reward feedback, the system can "learn" policies and become trained to produce desired navigation behaviors. For example, a learning system can learn from its current state s t Observe ∈S and policy

number

[0220] The goal of reinforcement learning (RL) is to find policy π. At time t, state s t Located in, operation a t The reward function r measures the immediate quality of the action. t It is usually assumed that there is an action a at time t. tPerforming an action affects the environment and therefore the future state value. Consequently, when deciding which action to take, not only the current reward but also the future reward should be considered. In some cases, if the system determines that a lower reward option will lead to greater future rewards, then the system should perform that action, even if it is associated with a lower reward than other available options. To formalize this, given policy π and initial state s,

number

[0221]

number

[0222] Instead of limiting the target period to T, we can discount future rewards and define the following equation for some fixed γ∈(0,1).

[0223]

number

[0224] In any case, the optimal policy is,

[0225]

number

[0226] This is the solution, and the expected value is over the initial state s.

[0227] There are several possible methodologies for training a driving policy system. For example, the system can use imitation techniques that learn from state / action pairs (e.g., behavior cloning), where the actions are selected by a good agent (e.g., a human) depending on a particular observed state. Assume a human driver is observed. This observation provides the basis for training the driving policy system, (s t ,a t ) in the form (s t is a state, and a t Many examples of the behavior of a human driver can be obtained, observed, and used. For example, π(s t )≒a t The policy π can be learned using supervised learning such that ||π(s) holds. This method has many potential advantages. Firstly, there is no need to define a reward function. Secondly, the learning is supervised and performed offline (there is no need to apply the agent during the learning process). The disadvantage of this method is that various human drivers, and even the same human driver, are not deterministic in terms of their own policy selection. Therefore, ||π(s) t )-a t It is often impossible to learn functions with very small || values. Furthermore, even small errors can gradually accumulate and lead to large errors.

[0228] Another technique that can be used is policy-based learning. Here, the policy is expressed in parametric form and can be directly optimized using an appropriate optimization technique (e.g., stochastic gradient descent). The method here is

number

[0229] The system can also be trained by value-based learning (learning the Q function or V function). Optimal value function V * We assume that we can learn a good approximation of the value function. We can then construct an optimal policy (for example, by using the Bellman equation). Some versions of value-based learning can be performed offline (called "off-policy" training). Some disadvantages of value-based methods may stem from their heavy reliance on the Markov assumption and the required approximation of a complex function (approximating the value function can be more difficult than directly approximating the policy).

[0230] Another technique may include model-based learning and planning (learning the probabilities of state transitions and solving the optimization problem to find the optimal V). A combination of these techniques can also be used to train the learning system. In this method, the process dynamics, i.e., (s t ,a t ) and then the next state s t+1 It is possible to learn a function that yields a distribution over a certain range. Once this function is learned, it is possible to solve the optimization problem and find the policy π for which its value is optimal. This is called the "plan". One advantage of this method is that the learning part is supervised, and it uses triplets (s t ,a t ,s t+1This could potentially be applied offline by observing the following. Similar to "imitation" methods, one drawback of this method is that small errors in the learning process can accumulate, potentially leading to a policy with insufficient functionality.

[0231] Another approach to training the driving policy module 803 may involve decomposing the driving policy function into semantically important components. Doing so would allow for the manual implementation of parts of the policy that can guarantee the safety of the policy, as well as the implementation of other parts of the policy using reinforcement learning methods that can enable adaptability to many scenarios, a human-like balance between defensive and aggressive behavior, and human-like negotiation with other drivers. From a technical standpoint, reinforcement learning methods can combine several methodologies to provide an easy-to-handle training procedure in which most of the training can be performed using recorded data or a self-built simulator.

[0232] In some embodiments, training of the driving policy module 803 can utilize a “choice” mechanism. To illustrate this, consider a simple scenario of driving policy for a two-lane highway. In the direct RL method, the state

number

[0233] Automatic driving control (ACC) policy, o ACC :S→A: This policy always outputs a yaw ratio of 0 and only changes the speed to ensure smooth and accident-free driving.

[0234] ACC+Left policy, o L:S→A: The longitudinal command for this policy is the same as the ACC command. Yaw rate is a simple implementation of centering the vehicle toward the center of the left lane while ensuring safe lateral movement (for example, not moving to the left if there is a car on the left).

[0235] ACC+Right policy, o R :S→A:o L This is the same, but the vehicle can be centered toward the center of the right lane.

[0236] These policies can be called "options." The policy π depends on these "options" and selects an option. o :S→O can be learned, where O is the set of available choices. In some cases, O={o ACC ,o L ,o R For all s for which} holds,

number

[0237] In practice, the policy function can be decomposed into a graph of choices 901, as shown in Figure 9. Another example of a graph of choices 1000 is shown in Figure 10. The graph of choices may represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There is a special node called the root node 903 of the graph. This node has no input nodes. The decision process starts from the root node and traverses this graph until it reaches a “leaf” node that has no output decision line. As shown in Figure 9, the leaf nodes may include, for example, nodes 905, 907, and 909. When a leaf node is encountered, the driving policy module 803 may output acceleration and steering commands related to the desired navigation action associated with the leaf node.

[0238] For example, internal nodes such as nodes 911, 913, and 915 may result in the implementation of a policy that selects children from among its available options. The available sets of children for an internal node include all nodes associated with a particular internal node by a decision line. For example, internal node 913, shown as "Converge" in Figure 9, includes three child nodes 909, 915, and 917 ("Stay," "Pass Right," and "Pass Left," respectively) that are connected to node 913 by decision lines.

[0239] The flexibility of the decision-making system can be achieved by allowing nodes to adjust their own position within the hierarchy of the choice graph. For example, any node may be allowed to declare itself "critical." Each node can implement a function "is critical" that outputs "true" if the node is in the critical section of its policy implementation. For example, a node responsible for takeovers can declare itself critical during an operation. This may impose constraints on the available sets of children of node u, such children may include all nodes v that are children of node u and for which there is a path that passes through all nodes designated as critical from v to leaf nodes. Such a technique can, on the one hand, allow for the declaration of desired paths on the graph at each time step, while on the other hand, it can maintain the stability of the policy, particularly while critical parts of the policy are being implemented.

[0240] By defining a graph of choices, the problem of learning the driving policy π:S→A can be broken down into the problem of defining the policy for each node in the graph, where the policy for an internal node should be selected from the available child nodes. For some nodes, individual policies can be implemented manually (e.g., by an if-then algorithm specifying a set of actions depending on the observed state), while for others, policies can be implemented using a trained system built by reinforcement learning. The choice between manual or trained / learned methods may depend on the safety aspects related to the task and their relative simplicity. The graph of choices can be constructed in such a way that some nodes are easily implemented while others depend on a trained model. Such methods can ensure the safe operation of the system.

[0241] The following explanation provides further details regarding the role of the choice graph in Figure 9 within the driving policy module 803. As discussed above, the input to the driving policy module is, for example, a “sensing state” that outlines the environmental map obtained from available sensors. The output of the driving policy module 803 is a set of desires (optionally accompanied by a set of strict constraints) that define the trajectory as a solution to the optimization problem.

[0242] As described above, the graph of choices represents a hierarchical set of decisions organized as a DAG. There is a special node in the graph called the "root." The root node is the only node that has no input edges (e.g., decision lines). The decision process starts from the root node and traverses the graph until it reaches a "leaf" node, i.e., a node that has no output edges. Each internal node should implement a policy of choosing one child from its available children. Every leaf node should implement a policy of defining a set of desires (e.g., a set of navigation goals for the host vehicle) based on the entire path from the root to the leaf. Along with a set of strict constraints that are directly determined based on the detected state, the set of desires establishes an optimization problem whose solution is the vehicle's trajectory. Strict constraints can be used to further enhance the safety of the system, and desires can be used to bring about driving comfort and human-like driving behavior of the system. The trajectory provided as a solution to the optimization problem then determines the commands that should be given to the steering, brakes, and / or engine actuators to realize the trajectory.

[0243] Returning to Figure 9, the choice graph 901 represents the choices for a two-lane highway including a merging lane (meaning that at some point a third lane merges into the right or left lane of the highway). The root node 903 first determines whether the host vehicle is in a simple road scenario or approaching a merging scenario. This is an example of a decision that can be made based on the detected state. The simple road node 911 includes three child nodes: a stay node 909, a left overtake node 917, and a right overtake node 915. Stay refers to a situation where the host vehicle wants to continue driving in the same lane. The stay node is a leaf node (it has no output edges / lines). Thus, the stay node defines a set of desires. The first desire defined by this node may include, for example, a desired lateral position as close as possible to the center of the current lane of travel. There may also be a desire to navigate smoothly (e.g., within the range of a default or acceptable maximum acceleration). The stay node can also define how the host vehicle reacts to other vehicles. For example, a stationary node can investigate the detected target vehicle and assign a semantic meaning to each that can be translated into components of the trajectory.

[0244] Various semantic meanings can be assigned to target vehicles within the host vehicle's environment. For example, in some embodiments, the semantic meaning may include any of the following indications: 1) Irrelevant: Indicates that the detected vehicle in the scene is not currently relevant; 2) Adjacent Lane: Indicates that the detected vehicle is in an adjacent lane and should maintain an appropriate offset relative to that vehicle (the exact offset can be calculated in an optimization problem that constructs a trajectory given desires and strict constraints, which may depend on the vehicle, and the leaves remaining on the graph of choices set the type of semantics for the target vehicle that defines the desires toward the target vehicle); 3) Yield: The host vehicle attempts to yield to the detected target vehicle, for example by slowing down, (especially if the host vehicle determines that there is a high probability that the target vehicle will cut into the host vehicle's lane); 4) Proceed: The host vehicle attempts to accept and respond to the right of way by, for example, accelerating; 5) Follow: The host vehicle wishes to follow this target vehicle to maintain smooth driving; 6) Left / Right Overtake: This means that the host vehicle wishes to initiate a lane change to the left or right lane. Left overtaking node 917 and right overtaking node 915 are internal nodes that have not yet defined their desires.

[0245] The next node in the choice graph 901 is the gap selection node 919. This node may be responsible for selecting the gap between two target vehicles in a specific target lane into which the host vehicle wishes to enter. By selecting a node of the form IDj, for some value of j, the host vehicle reaches a leaf specifying a wish related to the trajectory optimization problem, for example, the operation the host vehicle wishes to perform to reach the selected gap. Such an operation may include first accelerating / braking in the current lane and then moving into the target lane at a suitable time to enter the selected gap. If the gap selection node 919 cannot find a suitable gap, the host vehicle returns to the center of the current lane and proceeds to the abort node 921, which defines a wish to cancel the overtake.

[0246] Returning to the merging node 913, as the host vehicle approaches the merge, the host vehicle has several options that may depend on the specific circumstances. For example, as shown in Figure 11A, the host vehicle 1105 is moving along a two-lane road without detecting any other target vehicles in the main lane or merging lane 1111 of the two-lane road. In this situation, the driving policy module 803 may choose to stay at node 909 when it reaches the merging node 913. That is, if no target vehicles are detected as merging onto the road, it may be desirable to stay in its current lane.

[0247] In Figure 11B, this situation is slightly different. Here, the host vehicle 1105 detects one or more target vehicles 1107 entering the main road 1112 from the merging lane 1111. In this situation, when the driving policy module 803 approaches the merging node 913, the driving policy module 803 may decide to initiate a left-side overtaking maneuver to avoid the merging situation.

[0248] In Figure 11C, the host vehicle 1105 encounters one or more target vehicles 1107 entering the main road 1112 from a merging lane 1111. The host vehicle 1105 also detects a target vehicle 1109 moving in a lane adjacent to the host vehicle's lane. The host vehicle also detects one or more target vehicles 1110 moving in the same lane as the host vehicle 1105. In this situation, the driving policy module 803 can decide to adjust the speed of the host vehicle 1105 to yield to target vehicle 1107 and proceed ahead of target vehicle 1115. This can be achieved, for example, by proceeding to a gap selection node 919 that selects a gap between ID0 (vehicle 1107) and ID1 (vehicle 1115) as an appropriate merging gap. In that case, an appropriate gap in the merging situation defines the objective of the trajectory planner optimization problem.

[0249] As discussed above, nodes in the choice graph can declare themselves "critical," and such a declaration can ensure that the selected choice passes through critical nodes. Formally, each node can implement the function IsCritical. After performing a forward path from root to leaf on the choice graph and solving the trajectory planner optimization problem, a reverse path can be performed from leaf to root. Along this reverse path, the IsCritical function of all nodes in the path can be called, and a list of all critical nodes can be saved. The driving policy module 803 may be required to select a path from the root node to the leaf node that passes through all critical nodes within the forward path corresponding to the next time frame.

[0250] Figures 11A–11C can be used to illustrate the potential benefits of this technique. For example, in a situation where an overtaking maneuver has begun and the driving policy module 803 has reached the leaf corresponding to IDk, it would be undesirable for the host vehicle to select, for example, the stay node 909, if it is in the middle of the overtaking operation. To avoid this jump, the IDj node can designate itself as critical. During the operation, the success of the trajectory planner can be monitored, and the function IsCritical returns a value of "true" if the overtaking operation is proceeding as intended. This technique can ensure that the overtaking operation continues within the next timeframe (rather than jumping to another potentially inconsistent operation before completing the initially selected operation). On the other hand, if monitoring of the operation indicates that the selected operation is not proceeding as intended, or if the operation has become unnecessary or impossible, the function IsCritical can return a value of "false". This may allow the gap selection node to select a different gap within the next timeframe, or to completely abort the overtaking operation. This approach, on the one hand, may allow for the declaration of a desired path on the graph of choices at each time step, while on the other hand, it may help promote policy stability during critical parts of execution.

[0251] Strict constraints, which will be discussed in more detail below, can be distinguished from navigation desires. For example, strict constraints can ensure safe driving by applying an additional filtering layer to planned navigation actions. Involved strict constraints, which can be manually programmed and defined rather than by using a trained system built on reinforcement learning, can be determined from the detected state. However, in some embodiments, the trained system can learn applicable strict constraints that are applied and followed. Such a technique can facilitate the driving policy module 803 arriving at selected actions that already comply with applicable strict constraints, thereby reducing or eliminating selected actions that might later require modification to comply with applicable strict constraints. Nevertheless, as a redundant safety measure, strict constraints can still be applied to the output of the driving policy module 803 even if the driving policy module 803 is trained to take a given strict constraint into consideration.

[0252] There are many examples of potential strict constraints. For example, a strict constraint can be set in relation to a guardrail at the edge of a road. Under no circumstances is a host vehicle permitted to cross the guardrail. Such a rule creates a lateral strict constraint on the trajectory of the host vehicle. Another example of a strict constraint can include road bumps (e.g., speed control bumps), which can create a strict constraint on the driving speed before or while crossing the bump. Strict constraints can be considered to prioritize safety and therefore can be set manually rather than relying solely on a trained system that learns constraints during training.

[0253] In contrast to strict constraints, the goal of a desire may be to enable or achieve comfortable driving. As discussed above, one example of a desire may include the goal of positioning the host vehicle laterally within the lane, corresponding to the center of the lane. Another desire may include the ID of a gap to enter. It should be noted that the desire for the host vehicle to be as close to the center of the lane as possible, rather than strictly in the center of the lane, can ensure that the host vehicle can easily move back to the center of the lane if it deviates from the center. A desire does not have to prioritize safety above all else. In some embodiments, a desire may require negotiation with other drivers and pedestrians. One technique for constructing a desire may utilize a graph of choices, and policies implemented within at least some nodes of the graph may be based on reinforcement learning.

[0254] For a graph of 901 or 1000 nodes of choices implemented as nodes trained based on learning, the training process may include breaking down the problem into supervised learning phases and reinforcement learning phases. In the supervised learning phase,

number

number

number

number

[0255] A key element that may be provided in some scenarios is a differentiable path back to the decision about the action from the future loss / reward. In the structure of the choice graph, the implementation of the choice, including safety constraints, is usually not differentiable. To overcome this problem, the selection of children at a node of the learned policy can be probabilistic. That is, a node can output a probability vector p, which assigns the probability used when selecting each of the children of a particular node. Assume that a node has k children, a (1) ,...,a (k) This represents the movement of the pathway from each child to the leaf. Therefore, the resulting predicted movement is:

number

number

[0256] s t a t Given,

number

[0257] Furthermore, in some embodiments, the system can implement a multi-agent approach. For example, the system can consider data from various sources and / or images captured from multiple angles. In addition, it can consider predicting events that do not directly involve the host vehicle but may affect it, and it may also consider predicting events that may lead to unpredictable situations involving other vehicles (for example, radar can "foresee" inevitable events that will affect preceding vehicles and the host vehicle, and the high probability of such events), so some embodiments disclosed may result in energy savings.

[0258] A trained system with imposed navigation constraints.

[0259] In relation to autonomous driving, a critical concern is how to ensure that the learned policies of a trained navigation network are safe. In some embodiments, constraints can be used to train the driving policy system so that the actions selected by the trained system may already take applicable safety constraints into account. In addition, in some embodiments, an additional layer of safety can be provided by passing the selected actions of the trained system through one or more strict constraints involved by specific sensing scenes in the host vehicle's environment. Such a technique can ensure that the actions taken by the host vehicle are limited to those that are confirmed to satisfy applicable safety constraints.

[0260] At its core, the navigation system may include a learning algorithm based on a policy function that maps observed states to one or more desired actions. In some implementations, the learning algorithm is a deep learning algorithm. The desired actions may include at least one action that is expected to maximize the expected reward for the vehicle. In some cases, the actual action performed by the vehicle may correspond to one of the desired actions, while in other cases, the actual action performed may be determined based on the observed states, one or more desired actions, and non-learned strict constraints (e.g., safety constraints) imposed on the learning navigation engine. These constraints may include non-driving zones surrounding various types of detected objects (e.g., target vehicles, pedestrians, static objects on the shoulder or in the road, moving objects on the shoulder or in the road, guardrails, etc.). In some cases, the size of the zones may vary based on the detected movement of the detected objects (e.g., speed and / or direction). Other constraints may include the maximum speed when passing through areas affected by pedestrians, the maximum deceleration (to accommodate the distance between the host vehicle and the target vehicle), and mandatory stops at detected crosswalks or railway crossings.

[0261] Strict constraints used with systems trained by machine learning can provide a degree of safety in autonomous driving that may exceed the level of safety obtainable based solely on the output of the trained system. For example, a machine learning system can be trained using a desired set of constraints as training guidelines, and thus the trained system can select actions that adhere to the limits of applicable navigation constraints, depending on the detected navigation state. However, even then, the trained system has some flexibility in selecting navigation actions, and there may be at least some situations in which the action selected by the trained system does not strictly adhere to the relevant navigation constraints. Therefore, in order to ensure that the selected action strictly adheres to the relevant navigation constraints, non-machine learning components that guarantee the strict application of the relevant navigation constraints outside the learning / training framework can be used to combine, compare, filter, adjust, modify, etc., the output of the trained system.

[0262] The following explanation provides further details on the potential benefits (particularly in terms of safety) that can be obtained from combining trained systems and trained systems with algorithmic components outside the training / learning framework. As discussed earlier, the policy-based reinforcement learning objective can be optimized by ascending the stochastic gradient. The objective (e.g., expected reward) is:

number

[0263] In machine learning scenarios, objectives that include expectations can be used. However, such objectives may not return behavior that is strictly constrained by navigation constraints, without being constrained by those constraints. For example, for a trajectory representing a rare "turning point" event that should be avoided (e.g., an accident),

number

number

number

number

[0264]

number

number

number

number

[0265] Lemma: π o Let be the policy, and let p and r be scalars, and thereafter, for probability p,

number

number

[0266]

number

[0267] The following holds true, and the final approximation corresponds to the case where r ≥ 1 / p.

[0268] This explanation is about the format

number

number

number

number

number

number

number

number

[0269] The double-merging navigation situation shown in Figure 11D provides an example that further illustrates these concepts. In a double-merging situation, vehicles arrive at the merging area 1130 from both the left and right sides. From each side, it can be determined whether a vehicle such as vehicle 1133 or vehicle 1135 merges into the lane on the opposite side of the merging area 1130. In congested traffic, successfully executing a double-merging can require significant negotiation skills and experience, and may be difficult to do using heuristic or brute-force methods that enumerate all possible trajectories that all agents in the scene can take. In this example of a double-merging situation, a set of desires D suitable for the double-merging operation can be defined. D is the Cartesian product of the following sets D=[0,v max ]xLx{g,t,o} n This can be done, however, [0,v max ] is the desired target speed of the host vehicle, L={1,1.5,2,2.5,3,3.5,4} is the desired lateral position in each lane, where integers indicate the center of the lane and fractions indicate the lane boundaries, and {g,t,o} are classification labels to be assigned to each of the other n vehicles. If the host vehicle should yield to other vehicles, it can be assigned "g", if the host vehicle should gain the right of way to other vehicles, it can be assigned "t", or if the host vehicle should maintain an offset distance from other vehicles, it can be assigned "o".

[0270] The following is a set of desires (v,l,c1,...,c n This explains how )∈D can be transformed into a cost function over the operating track. The operating track is (x1,y1),...,(x k ,yk ) can be represented by, (x i , y i ) is the position (in the lateral and longitudinal directions) of the host vehicle at time τ·i (in ego - centric units). In some experiments, τ = 0.1 second and k = 10. Of course, other values can also be selected. The cost assigned to a trajectory can include a weighted sum of the individual costs assigned to the desired speed, lateral position, and the labels assigned to each of the other n vehicles.

[0271] Given the desired speed v ∈ [0, v max , the cost of the trajectory related to the speed

Number

[0272] Given the desired lateral position, the cost related to the desired lateral position

Number

[0273] is, where dist(x, y, l) is the distance from the point (x, y) to the position l of the lane. Regarding the cost caused by other vehicles, for any other vehicle, (x ' 1, y ' 1),..., (x ' k , y ' k ) can represent other vehicles in the ego - centric units of the host vehicle, and i is the (x i , y i ) and (x ' j , y ' jIt can be the earliest point where j exists such that the distance to + is small. If there is no such point, i can be set to i = ∞. If another vehicle is classified as "yield", it may be desirable that τi > τj + 0.5, which means that the host vehicle reaches the same point at least 0.5 seconds after the other vehicle reaches the intersection of the trajectories. A possible formula for converting the above constraints into cost is [τ(j - i)+0.5]

[0274] Similarly, if another vehicle is classified as "take the road", it may be desirable that τj > τi + 0.5, which can be converted into cost [τ(i - j)+0.5] + If another vehicle is classified as "offset", it may be desirable that i = ∞, which means that the trajectory of the host vehicle and the trajectory of the offset vehicle do not intersect. This condition can be converted into cost by imposing a penalty on the distance between the trajectories.

[0275] Assigning weights to each of these costs can give a single objective function π (T) for the trajectory planner. Costs can be added for the purpose of promoting smooth driving. Strict constraints can be added for the purpose of ensuring the functional safety of the trajectory. For example, (x i、 y i ) can be prohibited from deviating from the road, and (x i、 y i ) can be prohibited from approaching any trajectory point (x ' j、 y ' j ) of any other vehicle when |i - j| is small at (x ' j、 y ' j ).

[0276] In summary, policy π θThis can be broken down into a mapping from an agnostic state to a set of desires and a mapping from desires to actual trajectories. The latter mapping can be performed by solving an optimization problem that is not based on learning, whose cost depends on the desires, and whose strict constraints guarantee the functional safety of the policy.

[0277] The following explanation describes the mapping from agnostic states to sets of desires. As explained above, in order to adhere to functional safety, a system that relies solely on reinforcement learning is reward

number

[0278] For various reasons, decision-making can be further broken down into semantically significant components. For example, D may be large and continuous. In the double merging scenario described above with respect to Figure 11D, D=[0,v max ]xLx{g,t,o} n ) holds. In addition, the gradient estimator is,

number

[0279] Returning to the concept of the choice graph, Figure 11E shows a choice graph that can represent the double-convergence scenario shown in Figure 11D. As discussed earlier, the choice graph can represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There may be a special node in the graph called the "root" node 1140, which is the only node that has no input edges (e.g., decision lines). The decision process can traverse the graph starting from the root node and reaching a "leaf" node, i.e., a node that has no output edges. Each internal node can implement a policy function that selects one child from its available children. There may be a default mapping from a set of choices traversed on the choice graph to a set of desires D. In other words, a traverse on the choice graph can be automatically converted to desires in D. Given a node v in the graph, the parameter vector θ v This can define a policy for selecting children of v. v If it is a connection of θ, then at each node v v Using the policy defined by, by traversing the graph from the root to the leaves while selecting child nodes,

number

[0280] In the double merge options graph 1139 in Figure 11E, the root node 1140 can first determine whether the host vehicle is within the merging area (e.g., area 1130 in Figure 11D), or whether the host vehicle is approaching the merging area and needs to prepare for a possible merge. In either case, the host vehicle may need to decide whether to change lanes (e.g., to the left or to the right) or remain in its current lane. If the host vehicle decides to change lanes, it may need to determine whether the conditions are suitable for continuing and performing the lane change operation (e.g., at the “Go” node 1142). If changing lanes is not possible, the host vehicle may attempt to “push forward” toward the desired lane (e.g., at node 1144 as part of negotiations with vehicles in the desired lane) by aiming to be on the lane mark. Alternatively, the host vehicle may choose to “remain” in the same lane (e.g., at node 1146). This process can determine the host vehicle’s lateral position in a natural way. For example.

[0281] This could allow the desired lateral position to be determined in a natural way. For example, if the host vehicle changes lanes from lane 2 to lane 3, the "go" node can set the desired lateral position to 3, the "stay" node can set the desired lateral position to 2, and the "push forward" node can set the desired lateral position to 2.5. The host vehicle can then decide whether to maintain the "same" speed (node ​​1148), "accelerate" (node ​​1150), or "decelerate" (node ​​1152). The host vehicle can then enter a "chain" structure 1154 that examines other vehicles and sets their semantic meanings to values ​​in the pair {g,t,o}. This process can set desires toward other vehicles. The parameters of all nodes in this chain can be shared (similar to a regression neural network).

[0282] One potential benefit of the choices is the interpretability of the results. Another potential benefit is that the decomposable structure of set D can be utilized, and thus the policy at each node can be chosen from a small number of possibilities. In addition, the structure may allow for a reduction in the variance of the policy gradient estimator.

[0283] As discussed above, the length of an episode in a double-merging scenario can be approximately T = 250 steps. This value (or any other appropriate value depending on the specific navigation scenario) provides sufficient time to acknowledge the consequences of the host vehicle's actions (for example, if the host vehicle decides to change lanes in preparation for merging, it only acknowledges the benefit after the merge is completed without incident). On the other hand, due to the dynamics of driving, the host vehicle must make decisions at a sufficiently fast frequency (e.g., 10 Hz in the above example).

[0284] The graph of choices may allow for a reduction in the effective value of T in at least two ways. First, given high-level decisions, rewards for low-level decisions can be determined while taking shorter episodes into account. For example, if the host vehicle has already selected the “change lane” and “go” nodes, a policy for assigning semantic meaning to the vehicle can be learned by looking at a 2-3 second episode (meaning T would be 20-30 instead of 250). Second, for high-level decisions (such as changing lanes or staying in the same lane), the host vehicle may not need to make a decision every 0.1 seconds. Instead, the host vehicle may make decisions less frequently (e.g., every second), or a “choice end” function may be implemented, where the gradient is calculated only after each choice has ended. In either case, the effective value of T may be an order of magnitude smaller than its original value. Overall, the estimators at all nodes can rely on a value of T that is an order of magnitude smaller than the original 250 steps, which can immediately translate to a smaller variance.

[0285] As discussed above, strict constraints can promote safer driving, and there can be several different types of constraints. For example, static strict constraints can be determined directly from the detected state. These may include deceleration bumps, speed limits, road curvature, intersections, etc., in the host vehicle's environment, which may involve one or more constraints on the vehicle's speed, direction of travel, acceleration, braking (deceleration), etc. Static strict constraints can also include semantic free space, where the host vehicle is prohibited, for example, from going outside of free space and navigating too close to physical barriers. Static strict constraints can also restrict (e.g., prohibit) operations that do not conform to various aspects of the vehicle's kinematic motion. For example, static strict constraints can be used to prohibit operations that could lead to the host vehicle overturning, skidding, or otherwise losing control.

[0286] Strict constraints can also relate to vehicles. For example, a constraint can be used requiring a vehicle to maintain a longitudinal distance of at least 1 meter from other vehicles and a lateral distance of at least 0.5 meters from other vehicles. Constraints can also be applied to prevent the host vehicle from maintaining a collision course with one or more other vehicles. For example, time τ can be a measure of time based on a particular scene. The predicted trajectories of the host vehicle and one or more other vehicles from the current time to time τ can be considered. If the two trajectories intersect,

number

number

number

number

[0287] The time τ for tracking the trajectories of the host vehicle and one or more other vehicles may vary. However, in intersection scenarios where speeds may be slow, τ can be longer, and τ can be set so that the host vehicle enters and exits the intersection in less than τ seconds.

[0288] Naturally, applying strict constraints to vehicle trajectories requires that those trajectories be predictable. For a host vehicle, predicting the trajectory can be relatively straightforward, as the host vehicle generally already understands and actually plans its intended trajectory at any given time. For other vehicles, predicting their trajectories may be less straightforward. For other vehicles, the baseline calculation for determining the predicted trajectory may depend, for example, on the current speed and direction of travel of the other vehicle, which are determined by analyzing image streams captured by one or more cameras and / or other sensors (radar, lidar, acoustics, etc.) mounted on the host vehicle.

[0289] However, there may be some exceptions that simplify the problem or at least provide greater reliability to the predicted trajectory for another vehicle. For example, with respect to structured roads where lane markings exist and yielding rules may be in place, the trajectory of another vehicle can be based at least partially on the other vehicle's position relative to the lane and on applicable yielding rules. Thus, in some situations, where an observed lane structure exists, it can be assumed that a vehicle in an adjacent lane will adhere to the lane boundary. That is, a host vehicle can assume that a vehicle in an adjacent lane will remain in its own lane unless there is an observed reason (e.g., a signal light, strong lateral movement, movement crossing the lane boundary) indicating that the vehicle in the adjacent lane is cutting into the host vehicle's lane.

[0290] Other circumstances can also provide clues about the expected trajectory of other vehicles. For example, at stop signs, traffic lights, roundabouts, etc., where the host vehicle may have priority, it can be assumed that other vehicles will respect that priority. Therefore, unless there is evidence that the rules have been broken, it can be assumed that other vehicles will proceed along a trajectory that respects the priority of the host vehicle.

[0291] Strict constraints can also be applied to pedestrians within the host vehicle's environment. For example, a buffer distance for pedestrians can be established to prohibit the host vehicle from navigating even slightly closer than a specified buffer distance to any observed pedestrian. The buffer distance for pedestrians can be any appropriate distance. In some embodiments, the buffer distance may be at least 1 meter from the observed pedestrian.

[0292] Similar to the vehicle's situation, strict constraints can also be applied to the relative motion between a pedestrian and a host vehicle. For example, a pedestrian's trajectory can be monitored (based on direction and speed) relative to the host vehicle's predicted trajectory. Given a particular pedestrian's trajectory, for all points p on the trajectory, t(p) may represent the time it takes for the pedestrian to reach point p. To maintain the required buffer distance of at least 1 meter from the pedestrian, t(p) must be greater than the time it takes for the host vehicle to reach point p (with a sufficient time difference so that the host vehicle passes in front of the pedestrian by a distance of at least 1 meter), or (for example, if the host vehicle brakes and yields to the pedestrian) t(p) must be less than the time it takes for the host vehicle to reach point p. Furthermore, in the latter example, the strict constraint may require that the host vehicle arrives at point p at a time sufficiently later than the pedestrian so that the host vehicle passes behind the pedestrian and maintains the required buffer distance of at least 1 meter. Naturally, there may be exceptions to the pedestrian's strict constraints. For example, if the host vehicle has priority or is traveling at a very slow speed and there is no observed grounds for the pedestrian to refuse to yield to the host vehicle or to navigate toward the host vehicle, the strict restriction on pedestrians may be relaxed (for example, to a narrower buffer of at least 0.75 meters or 0.50 meters).

[0293] In some cases, constraints can be relaxed if it is determined that all conditions cannot be met. For example, in situations where the road is too narrow to maintain a desired distance (e.g., 0.5 meters) from both curves or from both a curve and a parked vehicle, one or more constraints may be relaxed if there are mitigating circumstances. For example, if there are no pedestrians (or other objects) on the sidewalk, a vehicle may proceed slowly at 0.1 meters from a curve. In some embodiments, constraints can be relaxed if doing so improves the user experience. For example, to avoid a pothole, a constraint may be relaxed to allow a vehicle to navigate closer to the edge of the lane, a curve, or a pedestrian than would normally be permitted. Furthermore, when deciding which constraints to relax, in some embodiments, the one or more constraints decided to relax are those considered to have the least adverse impact on safety. For example, a constraint on how close a vehicle can move to a curve or concrete barrier may be relaxed before a constraint dealing with proximity to other vehicles is relaxed. In some embodiments, pedestrian constraints may be the last to be relaxed, or may never be relaxed depending on the circumstances.

[0294] Figure 12 shows an example of a scene that may be captured and analyzed during the navigation of a host vehicle. For example, the host vehicle may include the above-described navigation system (e.g., system 100) which can receive multiple images representing the host vehicle's environment from cameras associated with the host vehicle (e.g., at least one of image capture devices 122, 124, and 126). The scene shown in Figure 12 is an example of one image that may be captured at time t from the environment of a host vehicle moving in lane 1210 along a predicted trajectory 1212. The navigation system may include at least one processing device (e.g., including either the EyeQ processor or other device) that is specifically programmed to receive multiple images and analyze those images to determine actions depending on the scene. In particular, the at least one processing device may implement the detection module 801, driving policy module 803, and control module 805 shown in Figure 8. The detection module 801 may be responsible for collecting and outputting image information gathered from the camera and providing that information to the driving policy module 803 in the form of identified navigation states, and the driving policy module 803 may constitute a trained navigation system trained by machine learning methods such as supervised learning or reinforcement learning. Based on the navigation state information provided to the driving policy module 803 by the detection module 801, the driving policy module 803 may generate desired navigation actions to be performed by the host vehicle according to the identified navigation state (for example, by implementing the graph method of the above choices).

[0295] In some embodiments, at least one processing device, for example, using a control module 805, can directly translate a desired navigation operation into a navigation command. However, in other embodiments, strict constraints can be applied to test the desired navigation operation provided by the driving policy module 803 against various predetermined navigation constraints that may be involved by the scene and the desired navigation operation. For example, if the driving policy module 803 outputs a desired navigation operation that causes the host vehicle to follow a track 1212, that navigation operation can be tested against one or more strict constraints related to various aspects of the host vehicle's environment. For example, the captured image 1201 may determine a curve 1213, a pedestrian 1215, a target vehicle 1217, and a static object (e.g., an overturned box) present in the scene. Each of these may be associated with one or more strict constraints. For example, the curve 1213 may be associated with a static constraint that prevents the host vehicle from navigating on the sidewalk 1214 in or beyond the curve. Curve 1213 may also relate to a road barrier envelope that defines a predetermined distance (e.g., a buffer zone) that extends along and away from the curve (e.g., 0.1 m, 0.25 m, 0.5 m, 1 m, etc.) that defines a non-navigable area for the host vehicle. Naturally, static constraints may also relate to other types of roadside boundaries (e.g., guardrails, concrete posts, cones, pylons, or any other type of roadside barrier).

[0296] It should be noted that distance and range measurement can be determined by any suitable method. For example, in some embodiments, distance information may be provided by an onboard radar and / or LiDAR system. Alternatively, distance information may be derived by analyzing one or more images captured from the host vehicle's environment. For example, the number of pixels of a recognized object represented in an image can be determined and compared with the geometry of a known field of view and focal length of the image acquisition device to determine its scale and distance. For example, velocity and acceleration can be determined by observing the change in scale between objects from image to image over a known time interval. This analysis may indicate the direction of movement toward or away from the host vehicle, along with how quickly the object is moving away from or approaching the host vehicle. Crossing velocity can be determined by analyzing the change in the X-coordinate position of an object from one image to another over a known time period.

[0297] Pedestrian 1215 may be associated with a pedestrian envelope that defines a buffer zone 1216. In some cases, imposed strict constraints may prohibit host vehicles from navigating within 1 meter of pedestrian 1215 (in any direction relative to the pedestrian). Pedestrian 1215 may also define the location of a pedestrian influence zone 1220. This influence zone may be associated with constraints that limit the speed of host vehicles within the influence zone. The influence zone may extend 5 meters, 10 meters, 20 meters, etc., from pedestrian 1215. Different speed limits may be associated with each grade of the influence zone. For example, within the area from 1 meter to 5 meters from pedestrian 1215, host vehicles may be limited to a first speed (e.g., 10 mph or 20 mph, etc.) which may be below the speed limit in the pedestrian influence zone extending from 5 meters to 10 meters. Any grades can be used for the various stages of the influence zone. In some embodiments, the first stage can be narrower than 1 to 5 meters, and may extend only to 1 to 2 meters. In other embodiments, the first stage of the affected area may extend from 1 meter (the boundary of the non-navigable area around the pedestrian) to at least 10 meters. The second stage may then extend from 10 meters to at least about 20 meters. The second stage may be related to the maximum travel speed of the host vehicle, which exceeds the maximum travel speed associated with the first stage of the pedestrian affected area.

[0298] Depending on the detected scene in the host vehicle's environment, one or more static object constraints may also be involved. For example, in image 1201, at least one processing device may detect a static object such as a box 1219 located on the road. The detected static object may include a variety of objects such as trees, poles, road signs, and at least one of these objects on the road. One or more default navigation constraints may be associated with the detected static object. For example, such constraints may include a static object envelope, which defines a buffer zone around the object within which the host vehicle's navigation may be prohibited. At least a portion of the buffer zone may extend a predetermined distance from the edge of the detected static object. For example, in the scene represented by image 1201, a buffer zone of at least 0.1 meters, 0.25 meters, 0.5 meters or more may be associated with the box 1219, so that the host vehicle passes to the right or left of the box by at least some distance (e.g., the distance of the buffer zone) to avoid a collision with the detected static object.

[0299] The default strict constraints may also include one or more target vehicle constraints. For example, a target vehicle 1217 may be detected in image 1201. One or more strict constraints can be used to ensure that the host vehicle does not collide with the target vehicle 1217. In some cases, the target vehicle envelope may be associated with the distance of a single buffer zone. For example, the buffer zone may be defined by a distance of 1 meter surrounding the target vehicle in all directions. The buffer zone may define an area extending at least 1 meter from the target vehicle into which the host vehicle is prohibited from navigating.

[0300] However, the envelope surrounding the target vehicle 1217 does not need to be defined by a fixed buffer distance. In some cases, the predetermined strict constraints related to the target vehicle (or any other movable object detected within the host vehicle's environment) may depend on the orientation of the host vehicle relative to the detected target vehicle. For example, in some cases, the distance of the longitudinal buffer zone (extending from the target vehicle to the front or rear of the host vehicle, for example, when the host vehicle is traveling toward the target vehicle) may be at least 1 meter. The distance of the lateral buffer zone (extending from the target vehicle to either side of the host vehicle, for example, when the host vehicle is moving in the same or opposite direction as the target vehicle, and therefore the side of the host vehicle passes directly alongside the side of the target vehicle) may be at least 0.5 meters.

[0301] As explained above, other constraints can also be introduced by detecting target vehicles or pedestrians in the host vehicle's environment. For example, the predicted trajectories of the host vehicle and target vehicle 1217 can be considered, and when the two trajectories intersect (for example, at intersection 1230), the strict constraint is:

number

[0302] Other strict constraints can also be used. For example, in at least some cases, the maximum deceleration rate of the host vehicle can be used. This maximum deceleration rate can be determined based on the detected distance to the target vehicle following the host vehicle (for example, using images collected from a rear-facing camera). Strict constraints may include mandatory stops at detected pedestrian crossings or railway crossings or other applicable constraints.

[0303] If the scene analysis in the host vehicle's environment indicates that one or more default navigation constraints may be involved, those constraints can be imposed on one or more planned navigation actions of the host vehicle. For example, if the scene analysis results in the driving policy module 803 returning a desired navigation action, that desired navigation action can be tested against one or more involved constraints. If it is determined that the desired navigation action violates any aspect of the involved constraints (for example, if a default strict constraint requires the host vehicle to remain at least 1.0 meter from pedestrian 1215, and the desired navigation action moves the host vehicle within 0.7 meters of pedestrian 1215), then at least one modification can be made to the desired navigation action based on one or more default navigation constraints. Adjusting the desired navigation action in this manner can result in the host vehicle's actual navigation action conforming to the constraints involved by a particular scene detected in the host vehicle's environment.

[0304] After determining the actual navigation operation of the host vehicle, the navigation operation can be implemented by causing adjustment of at least one of the host vehicle's navigation actuators in accordance with the determined actual navigation operation of the host vehicle. This navigation actuator may include at least one of the host vehicle's steering mechanism, brake, or accelerator.

[0305] Prioritized constraints

[0306] As described above, various strict constraints can be used with the navigation system to ensure the safe operation of the host vehicle. These constraints may include, among others, the minimum safe driving distance to pedestrians, target vehicles, road barriers, or detected objects, the maximum speed when passing through an area affected by a detected pedestrian, or the maximum deceleration rate of the host vehicle. These constraints can be imposed by a trained system trained on machine learning (supervised, reinforcement, or a combination thereof), but they can also be useful by an untrained system (for example, using an algorithm that directly addresses expected situations occurring within the host vehicle's environment scene).

[0307] In any case, a hierarchy of constraints is possible. In other words, some navigation constraints may take precedence over others. Therefore, if a situation arises where a navigation action that satisfies all the constraints involved is unavailable, the navigation system can determine the available navigation action that first fulfills the highest priority constraint. For example, the system may cause a vehicle to avoid a pedestrian first, even if pedestrian avoidance navigation would cause a collision with another vehicle or object detected on the road. In another example, the system may cause a vehicle to drive onto a curve to avoid a pedestrian.

[0308] Figure 13 shows a flowchart illustrating an algorithm for implementing a hierarchy of constraints to be involved, determined based on an analysis of the scene in the host vehicle's environment. For example, in step 1301, at least one processor associated with the navigation system (e.g., an EyeQ processor) can receive multiple images representing the host vehicle's environment from the host vehicle's onboard cameras. In step 1303, the navigation state associated with the host vehicle can be determined by analyzing one or more images representing the scene in the host vehicle's environment. For example, the navigation state may indicate, among various characteristics of the scene, that the host vehicle is moving along a two-lane road 1210, that a target vehicle 1217 is proceeding to an intersection ahead of the host vehicle, that a pedestrian 1215 is waiting to cross the road the host vehicle is traveling on, and that there is an object 1219 ahead of the host vehicle's lane, as shown in Figure 12.

[0309] In step 1305, one or more navigation constraints involved by the navigation state of the host vehicle can be determined. For example, at least one processing device can analyze a scene in the host vehicle's environment represented by one or more captured images, and then determine one or more navigation constraints involved by objects, vehicles, pedestrians, etc., recognized by the image analysis of the captured images. In some embodiments, at least one processing device can determine at least a first default navigation constraint and a second default navigation constraint involved by the navigation state, the first default navigation constraint may differ from the second default navigation constraint. For example, the first navigation constraint may relate to one or more target vehicles detected in the host vehicle's environment, and the second navigation constraint may relate to pedestrians detected in the host vehicle's environment.

[0310] In step 1307, at least one processing device can determine the priority of the constraints identified in step 1305. In the example described, a second default navigation constraint relating to pedestrians may have a higher priority than a first default navigation constraint relating to a target vehicle. The priority of navigation constraints may be determined or assigned based on a variety of factors, but in some embodiments, the priority of navigation constraints may relate to their relative importance from a safety perspective. For example, it may be important to follow or satisfy all implemented navigation constraints in as many situations as possible, but some constraints may relate to a greater safety risk than others and therefore may be assigned a higher priority. For example, a navigation constraint requiring a host vehicle to maintain a distance of at least 1 meter from a pedestrian may have a higher priority than a constraint requiring a host vehicle to maintain a distance of at least 1 meter from a target vehicle. This may be because a collision with a pedestrian may have more serious consequences than a collision with another vehicle. Similarly, maintaining a safe distance between the host vehicle and the target vehicle may take precedence over constraints requiring the host vehicle to avoid boxes in the road, travel below a certain speed over deceleration bumps, or expose the occupants of the host vehicle to levels below the maximum acceleration level.

[0311] The driving policy module 803 is designed to maximize safety by satisfying navigation constraints involved by a particular scene or navigation state, but it may be physically impossible to satisfy all constraints involved by a given situation. As shown in step 1309, in such a situation, the priority of each constraint involved can be used to determine which of the constraints should be satisfied first. Continuing the above example, in a situation where both the pedestrian clearance constraint and the target vehicle clearance constraint cannot be satisfied, and only one of the constraints can be satisfied, the higher priority of the pedestrian clearance constraint may result in that the constraint being satisfied before an attempt is made to maintain clearance to the target vehicle. Thus, in normal circumstances, as shown in step 1311, if both the first default navigation constraint and the second default navigation constraint can be satisfied, at least one processing device can determine a first navigation action for the host vehicle that satisfies both the first default navigation constraint and the second default navigation constraint, based on the identified navigation state of the host vehicle. However, in other situations where not all constraints involved can be met, as shown in step 1313, if both the first and second default navigation constraints cannot be met, at least one processing device may determine, based on the identified navigation state, a second navigation operation for the host vehicle that satisfies the second default navigation constraint (i.e., the constraint with higher priority) but does not satisfy the first default navigation constraint (which has a lower priority than the second navigation constraint).

[0312] Next, in step 1315, in order to perform a navigation operation for the determined host vehicle, at least one processing device can cause adjustment of at least one navigation actuator of the host vehicle in response to a determined first navigation operation or a determined second navigation operation for the host vehicle. As in the previous example, the navigation actuator may include at least one of the steering mechanism, brake, or accelerator.

[0313] Relaxation of constraints

[0314] As discussed above, navigation constraints can be imposed for safety. These constraints may include, among other things, the minimum safe driving distance to pedestrians, target vehicles, road barriers, or detected objects, the maximum speed when passing through the area of ​​influence of a detected pedestrian, or the maximum deceleration rate of the host vehicle. These constraints can be imposed in a learning navigation system or a non-learning navigation system. In certain situations, these constraints can be relaxed. For example, if a host vehicle slows down or stops near a pedestrian and moves slowly to communicate its intention to pass by the pedestrian, the pedestrian's response can be detected from the acquired image. If the pedestrian's response is to remain still or stop moving (and / or if eye contact with the pedestrian is detected), it can be understood that the pedestrian has recognized the navigation system's intention to pass by the pedestrian. In such situations, the system may relax one or more default constraints and implement less stringent constraints (for example, allowing the vehicle to navigate within a 0.5-meter range of the pedestrian rather than within a stricter 1-meter boundary).

[0315] Figure 14 shows a flowchart for implementing control of a host vehicle based on relaxing one or more navigation constraints. In step 1401, at least one processing device can receive multiple images representing the host vehicle's environment from cameras associated with the host vehicle. Analyzing the images in step 1403 may enable the identification of a navigation state associated with the host vehicle. In step 1405, at least one processor can determine navigation constraints associated with the host vehicle's navigation state. Navigation constraints may include a first default navigation constraint involved by at least one aspect of the navigation state. In step 1407, analyzing the multiple images may determine the presence of at least one navigation constraint relaxation factor.

[0316] Navigation constraint mitigation factors may include any appropriate indicators that cause one or more navigation constraints to be interrupted, modified, or otherwise mitigated in at least one aspect. In some embodiments, at least one navigation constraint mitigation factor may include a determination (based on image analysis) that a pedestrian's eyes are looking in the direction of the host vehicle. In this case, it can be more safely assumed that the pedestrian is aware of the host vehicle. As a result, there can be a greater degree of confidence that the pedestrian will not be involved in any unexpected action that would cause the pedestrian to move into the host vehicle's path. Other constraint mitigation factors may also be used. For example, at least one navigation constraint mitigation factor may include a pedestrian who is determined to be stationary (e.g., one who is estimated to be unlikely to enter the host vehicle's path) or a pedestrian whose movement is determined to be slowing down. Navigation constraint mitigation factors may also include more complex actions, such as a pedestrian who is determined to be stationary when the host vehicle stops and then resumes movement. In such situations, it can be assumed that the pedestrian understands that the host vehicle has priority, and the stopping pedestrian may indicate their intention to yield to the host vehicle. Other circumstances that may ease one or more of the constraints include the type of curb (for example, a low curb or a gently sloping ramp may allow for a relaxation of the distance constraint), the absence of pedestrians or other objects on the sidewalk, a vehicle whose engine is not running may have a relaxed distance, or a situation in which the pedestrian is facing away from the area ahead of the host vehicle and / or a situation in which the pedestrian is at a distance.

[0317] Once the presence of a navigation constraint relaxation factor is identified (for example, in step 1407), a second navigation constraint can be determined or developed in response to the detection of the constraint relaxation factor. This second navigation constraint may differ from the first navigation constraint and may include at least one characteristic that is relaxed compared to the first navigation constraint. The second navigation constraint may include a newly generated constraint based on the first constraint, which includes at least one modification that relaxes the first constraint in at least one respect. Alternatively, the second constraint may constitute a predetermined constraint that is less strict than the first navigation constraint in at least one respect. In some embodiments, this second constraint may be reserved for use only in situations where a constraint relaxation factor is identified within the host vehicle environment. Whether the second constraint is newly generated or selected from a predetermined set of constraints that are fully or partially available, the application of the second navigation constraint as a substitute for a stricter first navigation constraint (which might be applied if no relevant navigation constraint relaxation factor is detected) can be called constraint relaxation and can be achieved in step 1409.

[0318] If at least one constraint relaxation factor is detected in step 1407 and at least one constraint is relaxed in step 1409, a navigation operation for the host vehicle can be determined in step 1411. The navigation operation for the host vehicle can be based on the identified navigation state and can satisfy the second navigation constraint. In step 1413, the determined navigation operation can be implemented by causing at least one adjustment of the host vehicle's navigation actuator in accordance with the determined navigation operation.

[0319] As discussed above, the use of navigation constraints and relaxed navigation constraints can be used with trained navigation systems (e.g., by machine learning) or untrained navigation systems (e.g., systems programmed to respond with predetermined actions depending on a particular navigation state). When using a trained navigation system, the availability of relaxed navigation constraints for a particular navigation state may represent a switch in mode from the trained system's response to the untrained system's response. For example, a trained navigation network may determine the original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle may differ from the navigation action that satisfies the first navigation constraint. Rather, the action taken may satisfy a more relaxed second navigation constraint and may be an action developed by an untrained system (in response to detecting certain conditions in the host vehicle's environment, such as the presence of navigation constraint relaxation factors).

[0320] There are many examples of navigation constraints that can be relaxed in response to the detection of constraint relaxation factors within the host vehicle's environment. For example, if a default navigation constraint includes a buffer zone associated with a detected pedestrian, and at least a portion of the buffer zone extends a predetermined distance from the detected pedestrian, then a relaxed navigation constraint (newly developed, called from memory from a predetermined set, or generated as a relaxed version of an existing constraint) may include a different or modified buffer zone. For example, the different or modified buffer zone may have a shorter distance to the pedestrian than the original or unmodified buffer zone relative to the detected pedestrian. As a result, if appropriate constraint relaxation factors are detected within the host vehicle's environment, the relaxed constraint may be taken into account, and the host vehicle may be permitted to navigate closer to the detected pedestrian.

[0321] The characteristics of the navigation constraints that are relaxed may include the reduced width of the buffer zone associated with at least one pedestrian, as described above. However, the relaxed characteristics may also include the reduced width of the buffer zone associated with the target vehicle, detected objects, roadside barriers, or any other objects detected within the host vehicle's environment.

[0322] At least one relaxed property may also include other types of modifications to the navigation constraint properties. For example, a relaxed property may include a velocity increase associated with at least one default navigation constraint. A relaxed property may also include an increase in the maximum acceptable deceleration / acceleration associated with at least one default navigation constraint.

[0323] As described above, constraints can be relaxed in certain situations, while navigation constraints can be augmented in others. For example, in some situations, the navigation system may determine that conditions justify augmenting the standard set of navigation constraints. Such augmentation may include adding new constraints to the default set or adjusting one or more aspects of the default constraints. This addition or adjustment may result in more careful navigation than the default set of constraints applicable under normal driving conditions. Conditions that may justify augmenting constraints may include sensor failures or unfavorable environmental conditions (rain, snow, fog, or other conditions related to reduced visibility or reduced static friction of the vehicle), etc.

[0324] Figure 15 shows a flowchart for implementing control of a host vehicle based on augmenting one or more navigation constraints. In step 1501, at least one processing device can receive multiple images representing the host vehicle's environment from cameras associated with the host vehicle. Analyzing the images in step 1503 may enable the identification of a navigation state associated with the host vehicle. In step 1505, at least one processor can determine navigation constraints associated with the host vehicle's navigation state. The navigation constraints may include a first default navigation constraint involved by at least one aspect of the navigation state. In step 1507, analyzing the multiple images may determine the presence of at least one navigation constraint augmentation factor.

[0325] The navigation constraints involved may include any of the navigation constraints discussed above (for example, with respect to Figure 12) or any other suitable navigation constraints. Navigation constraint augmentants may include any indicators that can supplement / enhance one or more navigation constraints in at least one aspect. Supplementation or augmentation of navigation constraints may be done on a set basis (for example, by adding a new navigation constraint to a given set of constraints) or on a constraint basis (for example, by modifying a particular constraint so that the modified constraint is more restrictive than the original, or by adding a new constraint corresponding to a given constraint, where the new constraint is more restrictive than the corresponding constraint in at least one aspect). In addition, or alternatively, supplementation or augmentation of navigation constraints may refer to selecting from a given set of constraints based on a hierarchy. For example, a set of augmented constraints may be provided for selection based on whether navigation augmentants are detected in or against the host vehicle's environment. Under normal conditions where no augmentants are detected, the navigation constraints involved may be derived from constraints applicable under normal conditions. On the other hand, if one or more constraint augmentation factors are identified, the constraints involved can be derived from augmented constraints generated for one or more augmentation factors or from predetermined augmented constraints. The augmented constraints may be more restrictive in at least one aspect than the corresponding constraints applicable under normal conditions.

[0326] In some embodiments, at least one navigation constraint augmentation factor may include detecting the presence of ice, snow, or water on the road surface within the host vehicle's environment (e.g., based on image analysis). This determination may be based on detecting, for example, areas with higher reflectivity than expected on a dry road (e.g., indicating ice or water on the road), white areas on the road surface indicating the presence of snow, shadows on the road that coincide with the presence of longitudinal grooves on the road (e.g., tire tracks in the snow), water droplets or small particles of ice / snow on the host vehicle's windshield, or any other suitable indicator of the presence of water or ice / snow on the road surface.

[0327] At least one navigation constraint augmentation factor may include the detection of small particles on the outer surface of the host vehicle's windshield. Such small particles may impair the image quality of one or more image capture devices associated with the host vehicle. While the description has been made in relation to the host vehicle's windshield in relation to a camera mounted on the back of the host vehicle's windshield, it can also be shown that the detection of small particles on other surfaces associated with the host vehicle (e.g., camera lenses or lens covers, headlight lenses, rear windows, taillight lenses, or any other surface of the host vehicle visible to (or detected by sensors of) the image capture device) may also be a navigation constraint augmentation factor.

[0328] Navigation constraint augmentation factors can also be detected as characteristics of one or more image acquisition devices. For example, a detected degradation in the image quality of one or more images captured by an image acquisition device (e.g., a camera) associated with the host vehicle may also constitute a navigation constraint augmentation factor. The degradation in image quality may be related to a hardware failure or partial hardware failure related to the image acquisition device or an assembly associated with the image acquisition device. Such degradation in image quality may also be caused by environmental conditions. For example, the presence of smoke, fog, rain, snow, etc., in the air surrounding the host vehicle may contribute to a degradation in image quality of roads, pedestrians, target vehicles, etc., that may be in the host vehicle's environment.

[0329] Navigation constraint compensators may also relate to other aspects of the host vehicle. For example, in some situations, navigation constraint compensators may include detected failures or partial failures of systems or sensors associated with the host vehicle. Such compensators may include, for example, detecting failures or partial failures of speed sensors, GPS receivers, accelerometers, cameras, radar, lidar, brakes, tires, or any other systems associated with the host vehicle that may affect the host vehicle's ability to navigate in relation to navigation constraints related to the host vehicle's navigation state.

[0330] If the presence of navigation constraint augmentation factors is identified (e.g., in step 1507), a second navigation constraint can be determined or developed in response to the detection of the constraint augmentation factors. This second navigation constraint may differ from the first navigation constraint and may include at least one characteristic that augments the first navigation constraint. The second navigation constraint may be more restrictive than the first navigation constraint because detecting constraint augmentation factors in or related to the host vehicle's environment may suggest that at least one of the host vehicle's navigation capabilities may be reduced compared to normal operating conditions. Such reductions in capability may include reduced static friction of the road (e.g., ice, snow, or water on the road, reduced tire pressure, etc.), impaired visibility (e.g., rain, snow, dust, smoke, fog, etc. that reduce capture image quality), impaired detection capability (e.g., sensor failure or partial failure, reduced sensor performance, etc.), or any other reduction in the host vehicle's ability to navigate depending on the detected navigation state.

[0331] If at least one constraint augmentation factor is detected in step 1507 and at least one constraint is augmented in step 1509, a navigation operation for the host vehicle can be determined in step 1511. The navigation operation for the host vehicle can be based on the identified navigation state and can satisfy the second navigation (i.e., augmented) constraint. In step 1513, the determined navigation operation can be implemented by causing at least one adjustment of the host vehicle's navigation actuator in accordance with the determined navigation operation.

[0332] As discussed earlier, the use of navigation constraints and augmented navigation constraints can be used with trained navigation systems (e.g., by machine learning) or untrained navigation systems (e.g., systems programmed to respond with predetermined actions depending on a particular navigation state). When using a trained navigation system, the availability of augmented navigation constraints for a particular navigation state may represent a switch in mode from the trained system's response to the untrained system's response. For example, a trained navigation network can determine the original navigation action for a host vehicle based on a first navigation constraint. However, the action taken by the vehicle may differ from the navigation action that satisfies the first navigation constraint. Rather, the action taken may satisfy a second augmented navigation constraint and may be an action developed by an untrained system (in response to detecting certain conditions in the host vehicle's environment, such as the presence of navigation constraint augmentation factors).

[0333] There are many examples of navigation constraints that can be generated, supplemented, or augmented in response to the detection of constraint augmentation factors within the host vehicle's environment. For example, if a default navigation constraint includes a buffer zone related to a detected pedestrian, object, vehicle, etc., and at least a portion of the buffer zone extends a predetermined distance from the detected pedestrian / object / vehicle, then an augmented navigation constraint (newly developed, recalled from memory from a predetermined set, or generated as an augmented version of an existing constraint) may include a different or modified buffer zone. For example, a different or modified buffer zone may have a longer distance to the detected pedestrian / object / vehicle than the original or unmodified buffer zone. As a result, if appropriate constraint augmentation factors are detected within or related to the host vehicle's environment, the host vehicle may be compelled to consider the augmented constraint and navigate further away from the detected pedestrian / object / vehicle.

[0334] At least one augmented characteristic may also include other types of modifications to the navigation constraint characteristics. For example, an augmented characteristic may include a velocity reduction associated with at least one default navigation constraint. An augmented characteristic may also include a reduction in the maximum acceptable deceleration / acceleration associated with at least one default navigation constraint.

[0335] Navigation based on long-term plans

[0336] In some embodiments, the disclosed navigation system can not only respond to detected navigation states in the host vehicle's environment but also determine one or more navigation actions based on a long-term plan. For example, the system can consider the potential impact of one or more available navigation actions as options for navigating with respect to a detected navigation state on future navigation states. Considering the effect of available actions on future states may enable the navigation system to determine navigation actions not only based on the currently detected navigation state but also on a long-term plan. Navigation using long-term planning techniques may be particularly applicable when the navigation system uses one or more reward functions as a technique for selecting navigation actions from available options. Potential rewards can be analyzed with respect to available navigation actions that can be performed in accordance with the current detected navigation state of the host vehicle. However, potential rewards can also be analyzed in relation to actions that can be performed in accordance with future navigation states predicted to result from the available actions for the current navigation state. As a result, in some cases, the disclosed navigation system may select a navigation action in accordance with a detected navigation state, even if the selected navigation action may not yield the highest reward among the available actions that can be performed in accordance with the current navigation state. This may be particularly true when the system determines that the selected action could lead to a future navigation state that triggers one or more potential navigation actions that yield a higher reward than the selected action or, in some cases, any of the actions available for the current navigation state. This principle can be more simply expressed as performing a less advantageous action now in order to bring about a higher-reward option in the future. Thus, a disclosed navigation system capable of long-term planning may choose a second-best short-term action if long-term forecasts indicate that a short-term loss of reward may result in a long-term increase in reward.

[0337] Generally, applications of autonomous driving may involve a set of planning problems in which a navigation system can determine immediate actions to optimize a long-term objective. For example, if a vehicle faces a situation where it needs to merge into a ring road, the navigation system can determine an immediate acceleration or braking command to begin navigation into the ring road. The immediate action for the navigation state detected in the ring road may include an acceleration or braking command depending on the detected state, but the long-term objective is to successfully merge, and the long-term effect of the selected command is the success or failure of the merge. The planning problem can be addressed by breaking it down into two phases. First, supervised learning can be applied to predict the near future based on the present (assuming the predictors are differentiable with respect to the current representation). Second, a regression neural network can be used to model the agent's complete trajectory, with unexplained factors modeled as (additional) input nodes. This may allow for finding a solution to the long-term planning problem using supervised learning methods and direct optimization on the regression neural network. Such methods can also enable the learning of robust policies by incorporating adversarial factors against the environment.

[0338] The two most fundamental elements of an autonomous driving system are sensing and planning. Sensing deals with finding a compact representation of the current state of the environment, while planning deals with deciding which actions to take to optimize future objectives. Supervised machine learning methods are useful for solving the sensing problem. For the planning part, machine learning algorithm frameworks, particularly reinforcement learning (RL) frameworks as described above, can also be used.

[0339] RL can be executed in a series of consecutive rounds. In round t, the planner (also known as the agent or driving policy module 803) checks the state s representing the agent and the environment. t∈S can be observed. Then the planner performs action a t The agent should determine ∈A. After performing the action, the agent receives an immediate reward.

number

number

number

[0340] Supervised learning (SL) can be considered a special case of RL, and in supervised learning, from some distribution over S, s t The sample is taken, and the reward function is r t =-l(a t ,y t ) can have the form, where l is the loss function, and the learning side is state s t The optimal (and sometimes noisy) value of y to perform when acknowledging this is the value of the action that is best to perform. tObserve the value of . There may be some differences between the general RL model and specific cases of SL, and these differences can make the general RL problem more difficult.

[0341] In some SL (System Learning) scenarios, the actions (or predictions) performed by the learning side may not have any impact on the environment. In other words, s t+1 and a t and are independent. This can have two important implications. First, in SL, the samples (s1, y1), ..., (s m ,y m ) can be collected in advance, thereby allowing us to begin searching for policies (or predictors) that have excellent accuracy for the sample for the first time. In contrast, with RL, state s t+1 The action performed usually depends on the action that is performed (and the state prior to it), and the action that is performed, in turn, depends on the policy used to generate the action. This links the data generation process to the policy learning process. Secondly, in SL, since the action does not affect the environment, a of the performance of π t The contribution of the selection is local. Specifically, a t This only affects the immediate reward value. In contrast, in RL, actions taken in round t may have long-term effects on reward values ​​in future rounds.

[0342] In SL, the shape of the reward is r t =-l(a t ,y t ) along with the knowledge of the "correct" solution yt, a t It can provide complete knowledge of the rewards for all possible options, and it is a tThis can enable the calculation of the derivative of the reward with respect to [a certain factor]. In contrast, in RL, the "one-shot" value of the reward can be all that can be observed about a particular choice of action to be taken. This can be called "bandit" feedback. In RL-based systems, if only "bandit" feedback is available, the system may not always know whether the action taken was the best action to take, which is one of the most important reasons why "exploration" is necessary as part of a long-term navigation plan.

[0343] Many RL algorithms rely, at least partially, on a mathematically clear model of a Markov decision process (MDP). The Markov assumption is that s t and a t Given s t+1 The distribution of is completely determined. This yields a closed form of the cumulative reward of a given policy with respect to the stationary distribution across the states of the MDP. The stationary distribution of the policy can be expressed as a solution to a linear programming problem. This yields two sets of algorithms: 1) optimization on the primal problem, which can be called policy search, and 2) optimization on the dual problem, where the variable is called the value function Vπ. The value function determines the expected cumulative reward when the MDP starts from an initial state s and actions are selected from there according to π. The quantity in question is the state-action value function Q π (s,a) is a function that determines the cumulative reward, assuming a starting state s, an action a immediately chosen, and actions chosen therefrom according to π. This Q function can lead to the characterization of the optimal policy (using the Bellman equation). Specifically, this Q function can be shown to be a deterministic function of the optimal policy from S to A (in fact, the optimal policy can be characterized as a "greedy" policy with respect to the optimal Q function).

[0344] One potential advantage of the MDP model is that it allows the MDP model to use the Q function to connect the future to the present. For example, given that the host vehicle is currently in state s, Q πThe value of (s,a) can indicate the effect of performing action a on the future. Thus, the Q function provides a local measure of the quality of action a, thereby making the RL problem more similar to the SL scenario.

[0345] Many RL algorithms approximate the V function or Q function in some way. Value iteration algorithms, such as Q-learning algorithms, can utilize the fact that the V and Q functions of the optimal policy may be fixed points of some operator obtained from the Bellman equations. Actor-critic policy iteration algorithms aim to learn the policy in an iterative manner, where in iteration t, "critic" is...

number

[0346] Despite the mathematical simplicity of MDP and the convenience of switching to Q-function representations, this method can have several limitations. For example, in some cases, the concept of approximating Markov behavioral states may be all that can be found. Furthermore, state transitions may depend not only on the agent's behavior but also on the behavior of other players in the environment. For example, in the ACC example above, the dynamics of the autonomous vehicle may be Markov-like, but the next state may depend on the behavior of other car drivers, which is not necessarily Markov-like. One possible solution to this problem is to use a partially observed MDP, in which it is assumed that there are Markov states, but what can be seen are observations distributed according to the hidden states.

[0347] A more direct approach would be to consider a game-theoretic generalization of MDP (e.g., within a probabilistic game framework). In fact, algorithms for MDP can be generalized to multi-agent games (e.g., minimax Q-learning or Nash Q-learning). Other approaches may include explicit modeling of other players and vanishing regret learning algorithms. Learning in a multi-agent setting can be more complex than learning in a single-agent setting.

[0348] A second limitation of Q-function representation can arise from deviating from tabular configurations. Tabular configurations are applicable when the number of states and actions is small, and therefore Q can be represented as a |S| row and |A| column table. However, if the natural representations of S and A involve Euclidean space and the state and action spaces are discretized, the number of states / actions can be exponential in scale. In such cases, adopting a tabular configuration may not be practical. Instead, the Q-function can be approximated by some function from a parametric hypothesis class (e.g., a neural network of a specific architecture). For example, the Deep Q Network (DQN) learning algorithm can be used. In DQNs, the state space can be continuous, but the action space can remain a small discrete set. Techniques for handling continuous action spaces are possible, but they may rely on approximating the Q-function. In any case, the Q-function can be complex and noise-sensitive, and therefore may be difficult to learn.

[0349] A different approach might be to address the RL problem using a recurrent neural network (RNN). In some cases, RNNs can be combined with the concept of multi-agent games and robustness to adversarial environments from game theory. Furthermore, this approach may not be strictly dependent on the Markov assumption.

[0350] The following describes in more detail a method for navigation using prediction-based planning. In this method, the state space S is:

number

number

[0351] As a first step in this approach, one might observe an interesting problem where the bandit nature of the reward is not the issue. For example, the reward value for applications of ACC (discussed in more detail below) may be differentiable with respect to the current state and action. In fact, even if the reward is given in a "bandit" form,

number

number

number

number

number

[0352] Similar concepts can be used to address the connection between the past and the future. For example,

number

number

number

number

number

number

number

[0353] As discussed above, the learning system can benefit from robustness to adversarial environments, such as the host vehicle environment which may include multiple other drivers that may behave in unexpected ways. t In models that do not impose probabilistic assumptions on v, t We can consider environments in which μ is selected in an adversarial manner. In some cases, μ t Restrictions can be placed on this, otherwise the enemy may make the planning problem difficult or even impossible. One natural constraint is ||μ t || may require that the constraint be bounded by another constraint.

[0354] Robustness to adversarial environments can be useful in autonomous driving applications. tChoosing this option can even accelerate the learning process because it allows the learning system to focus on a robust, optimal policy. This concept can be illustrated using a simple game. The states are:

number

number

[0355] Such techniques can be applied to virtually all possible navigation situations. Below, we describe one example, a technique applied to adaptive cruise control (ACC). In the ACC problem, the host vehicle may attempt to maintain a sufficient distance from the target vehicle ahead (e.g., 1.5 seconds to the target vehicle). Another objective may be to drive as smoothly as possible while maintaining a desired gap. A model representing this situation can be defined as follows: The state space is:

number

number

number

number

[0356] The complete dynamics of the system can be described by the following equation:

number

[0357] This can be described as the sum of two vectors.

number

[0358] The first vector represents the predictable portion, and the second vector represents the unpredictable portion. The reward for round t is determined as follows:

number

[0359] The first term may result in a penalty for non-zero acceleration, thus encouraging smoother driving. The second term is the target car x t Distance to and desired distance

number

[0360] By implementing the method outlined above, the host vehicle's navigation system can select an action in response to an observed state (for example, through the operation of the driving policy module 803 in the navigation system's processing unit 110). The selected action may be based not only on an analysis of the rewards associated with the available response actions for the detected navigation state, but also on an analysis of future states, potential actions in response to those future states, and the rewards associated with those potential actions.

[0361] Figure 16 illustrates an algorithmic approach to navigation based on detection and long-term planning. For example, in step 1601, at least one processing device 110 of the navigation system for the host vehicle may receive multiple images. These images can capture scenes representing the environment of the host vehicle and may be supplied by any of the image capture devices (e.g., cameras, sensors, etc.) described above. Analyzing one or more of these images in step 1603 may enable at least one processing device 110 to identify the current navigation state related to the host vehicle (as described above).

[0362] In steps 1605, 1607, and 1609, various potential navigation actions can be determined depending on the detected navigation state. (For example, to complete a merge, to smoothly follow a preceding vehicle, to overtake a target vehicle, to avoid an object in the road, to slow down for a detected stop sign, to avoid an intruding target vehicle, or to complete any other navigation action that may aid the system's navigation goal.) These potential navigation actions (e.g., from the first navigation action to the Nth available navigation action) can be determined based on the detected state and the navigation system's long-term goal.

[0363] For each of the potential navigation actions determined, the system can determine the expected reward. The expected reward can be determined according to any of the techniques described above and may include an analysis of a particular potential action against one or more reward functions. For each of the potential navigation actions (e.g., the first, second, and Nth) determined in steps 1605, 1607, and 1609, the expected rewards 1606, 1608, and 1610 can be determined.

[0364] In some cases, the host vehicle's navigation system can select from available potential actions based on values ​​related to expected rewards 1606, 1608, and 1610 (or any other type of indicator of expected reward). For example, in some situations, the action that yields the highest expected reward may be selected.

[0365] In particular, in other instances where a navigation system is involved in long-term planning to determine navigation actions for a host vehicle, the system may not select the potential action that yields the highest expected reward. Rather, the system can look to the future and analyze whether there may be an opportunity to realize a higher reward later if it selects a lower-reward action in accordance with the current navigation state. For example, future states can be determined for any or all of the potential actions determined in steps 1605, 1607, and 1609. Each future state determined in steps 1613, 1615, and 1617 may represent a future navigation state that is expected to occur based on the current navigation state modified by each potential action (e.g., the potential actions determined in steps 1605, 1607, and 1609).

[0366] For each of the future states predicted in steps 1613, 1615, and 1617, one or more future actions can be determined and evaluated (as navigation options available depending on the future state determined). In steps 1619, 1621, and 1623, for example, expected reward values ​​or any other type of indicator associated with one or more future actions can be developed (for example, based on one or more reward functions). Expected rewards associated with one or more future actions can be evaluated by comparing the values ​​of the reward functions associated with each future action or by comparing any other indicator associated with the expected rewards.

[0367] In step 1625, the navigation system for the host vehicle may select a navigation action for the host vehicle based on comparing expected rewards, not only on potential actions identified for the current navigation state (e.g., in steps 1605, 1607, and 1609), but also on expected rewards determined as a result of available future potential actions depending on the predicted future state (e.g., determined in steps 1613, 1615, and 1617). The selection in step 1625 may be based on the analysis of options and rewards performed in steps 1619, 1621, and 1623.

[0368] The selection of a navigation action in step 1625 may be based solely on comparing the expected rewards associated with the options for future actions. In this case, the navigation system may select an action for the current state based solely on comparing the expected rewards resulting from actions for potential future navigation states. For example, the system may select a potential action identified in steps 1605, 1607, or 1609 that is associated with the highest future reward value determined by the analysis in steps 1619, 1621, and 1623.

[0369] The selection of a navigation action in step 1625 may be based solely on comparing the current action options (as described above). In this situation, the navigation system may select a potential action identified in steps 1605, 1607, or 1609 that is associated with the highest expected reward 1606, 1608, or 1610. This selection may be made with little or no consideration of the future expected rewards for the available navigation actions depending on the future navigation state or the expected future navigation state.

[0370] On the other hand, in some cases, the selection of a navigation action in step 1625 may be based on comparing the expected rewards associated with both future action options and current action options. This can, in fact, be one of the principles of long-term planning-based navigation. To achieve potentially higher rewards depending on subsequent navigation actions that are expected to become available depending on future navigation states, the expected rewards for future actions can be analyzed to determine whether it is justifiable to select an action with a lower reward depending on the current navigation state. For example, the value of expected reward 1606 or other indicator may indicate the highest expected reward among rewards 1606, 1608, and 1610. On the other hand, expected reward 1608 may indicate the lowest expected reward among rewards 1606, 1608, and 1610. Rather than simply selecting the potential action determined in step 1605 (i.e., the action that produces the highest expected reward 1606), the analysis of future states, potential future actions, and future rewards can be used when selecting a navigation action in step 1625. In one example, it may be determined that the reward identified in step 1621 (depending on at least one future action for a future state determined in step 1615 based on a second potential action determined in step 1607) may be higher than the expected reward 1606. Based on this comparison, even though the expected reward 1606 is higher than the expected reward 1608, the second potential action determined in step 1607 may be selected instead of the first potential action determined in step 1605. In one example, the potential navigation action determined in step 1605 may include merging in front of the detected target vehicle, while the potential navigation action determined in step 1607 may include merging behind the target vehicle.The expected reward 1606 for merging in front of the target vehicle may be higher than the expected reward 1608 associated with merging behind the target vehicle, but it may be determined that merging behind the target vehicle could lead to a future state in which there are action options that may give an even higher potential reward than expected rewards 1606, 1608, or other rewards based on the actions available depending on the currently detected navigation state.

[0371] The selection of potential actions in step 1625 may be based on any appropriate comparison of expected rewards (or any other measure or indicator of the benefit associated with a particular potential action that outperforms another potential action). In some cases, as described above, the second potential action may be chosen over the first potential action if it is predicted that the second potential action will result in at least one future action associated with a higher expected reward than the reward associated with the first potential action. In other cases, more complex comparisons may be used. For example, the rewards associated with action choices depending on the predicted future state may be compared with the expected rewards associated with multiple potential actions that are decided upon.

[0372] In some scenarios, if at least one of the future actions is expected to yield a reward higher than any of the expected rewards resulting from the potential actions for the current state (e.g., expected rewards 1606, 1608, 1610, etc.), then actions and expected rewards based on the predicted future state may influence the selection of the potential action for the current state. In some cases, the option of the future action that yields the highest expected reward (e.g., among the expected rewards associated with the potential action for the current state and the expected rewards associated with the options of potential future actions for potential future navigation states) can be used as a guide for selecting the potential action for the current navigation state. That is, after identifying the option of the future action that yields the highest expected reward (or a reward exceeding a predetermined threshold, etc.), the potential action leading to the future state associated with the identified future action that yields the highest expected reward can be selected in step 1625.

[0373] In other cases, the selection of available actions can be based on the determined difference between expected rewards. For example, if the difference between the expected reward associated with the future action determined in step 1621 and expected reward 1606 is greater than the difference between expected reward 1608 and expected reward 1606 (assuming a positive sign difference), then the second potential action determined in step 1607 can be selected. In another example, if the difference between the expected reward associated with the future action determined in step 1621 and the expected reward associated with the future action determined in step 1619 is greater than the difference between expected reward 1608 and expected reward 1606, then the second potential action determined in step 1607 can be selected.

[0374] We have described several examples of selecting from potential actions in relation to the current navigation state. However, any other appropriate comparative techniques or criteria can be used to select available actions by long-term planning based on an analysis of actions and rewards extending to predicted future states. In addition, while Figure 16 shows two layers in long-term planning analysis (for example, a first layer that considers rewards arising from potential actions in relation to the current state and a second layer that considers rewards arising from future action options depending on predicted future states), analyses based on more layers may also be possible. For example, instead of basing long-term planning analysis on one or two layers, three, four, or even more layers of analysis can be used when selecting from potential actions available in relation to the current navigation state.

[0375] After selecting from potential actions in accordance with the detected navigation state, in step 1627, at least one processor may trigger adjustment of at least one navigation actuator of the host vehicle in accordance with the selected potential navigation action. The navigation actuator may include any suitable device for controlling at least one aspect of the host vehicle. For example, the navigation actuator may include at least one of a steering mechanism, a brake, or an accelerator.

[0376] Navigation based on the inferred aggression of others

[0377] To determine indicators of driving aggression, the target vehicle can be monitored by analyzing the acquired image stream. Herein, aggression is described as a qualitative or quantitative parameter, but other characteristics, namely the detected level of attention (potential driver defects, distractions - cell phone use, drowsiness, etc.), may be used. In some cases, the target vehicle can be considered to have a defensive attitude, and in other cases, the target vehicle can be considered to have a more aggressive attitude. Navigation actions can be selected or developed based on indicators of aggression. For example, in some cases, it is possible to determine whether the target vehicle is aggressive or defensive by tracking relative speed, relative acceleration, increase in relative acceleration, pursuit distance, etc., relative to the host vehicle. If it is determined that the target vehicle has an aggression level above a threshold, for example, the host vehicle may be inclined to yield to the target vehicle. The level of aggression of the target vehicle can also be determined based on the determined behavior of the target vehicle towards one or more obstacles in or near the route (e.g., preceding vehicles, road obstacles, traffic lights, etc.).

[0378] As an introduction to this concept, we will describe an example experiment concerning a host vehicle merging into a roundabout, where the navigation objective is to exit the roundabout. This situation can begin with the host vehicle reaching the entrance of the roundabout and end with the host vehicle reaching the exit of the roundabout (e.g., the second exit). Success can be measured based on whether the host vehicle maintains a safe distance from all other vehicles at all times, whether the host vehicle completes the route as quickly as possible, and whether the host vehicle adheres to a smooth acceleration policy. In this explanation, N TSeveral target vehicles may be randomly placed on a roundabout. To model the confusion between adversarial and typical behavior, a target vehicle can be modeled with an "aggressive" driving policy with probability p, so that when the host vehicle attempts to merge in front of the target vehicle, the aggressive target vehicle will accelerate. A target vehicle can be modeled with a "defensive" driving policy with probability 1-p, so that the target vehicle will decelerate and allow the host vehicle to merge. In this experiment, p=0.5, and information about other driver types may not be provided to the host vehicle's navigation system. Other driver types may be randomly selected at the start of the episode.

[0379] Navigation status can be represented as the speed and position of the host vehicle (agent) and the position, speed, and acceleration of the target vehicle. To distinguish between aggressive and defensive drivers based on the current state, it may be important to continuously observe the target's acceleration. All target vehicles can move along a one-dimensional curve that outlines the path of the ring road. The host vehicle can move along its own one-dimensional curve that intersects with the target vehicle's curve at the merging point, which is the origin of both curves. To model reasonable driving, an upper limit can be set by a constant on the absolute value of the acceleration of all vehicles. Since driving in the reverse direction is not permitted, speed can also pass through ReLU. Note that by not allowing driving in the reverse direction, agents cannot regret their past actions, which may necessitate long-term planning.

[0380] As explained above, the following state s t+1 is the predictable part

number

number

number

[0381] Figure 18 shows a flowchart illustrating an example of an algorithm for navigating a host vehicle based on the predicted aggressiveness of other vehicles. In the example in Figure 18, the level of aggressiveness associated with at least one target vehicle can be inferred based on the observed behavior of the target vehicle toward objects in the target vehicle's environment. For example, in step 1801, at least one processing device of the host vehicle's navigation system (e.g., processing device 110) can receive multiple images representing the host vehicle's environment from cameras associated with the host vehicle. In step 1803, analyzing one or more of the received images may enable at least one processor to identify a target vehicle (e.g., vehicle 1703) in the environment of host vehicle 1701. In step 1805, analyzing one or more of the received images may enable at least one processing device to identify at least one obstacle for the target vehicle in the host vehicle's environment. Objects may include debris in the road, stop signals / traffic lights, pedestrians, other vehicles (e.g., vehicles moving ahead of the target vehicle or parked vehicles), boxes in the road, road barriers, curves, or any other type of object that may be encountered within the host vehicle's environment. Step 1807 may allow at least one processing device to determine at least one navigation characteristic of the target vehicle relative to at least one identified obstacle for the target vehicle by analyzing one or more of the received images.

[0382] To develop an appropriate navigation response to a target vehicle, the level of aggression of the detected target vehicle can be inferred using various navigation characteristics. For example, such navigation characteristics may include the relative acceleration between the target vehicle and at least one identified obstacle, the distance of the target vehicle from the obstacle (e.g., the distance a target vehicle is following behind another vehicle), and / or the relative velocity between the target vehicle and the obstacle.

[0383] In some embodiments, the navigation characteristics of a target vehicle can be determined based on outputs from sensors associated with the host vehicle (e.g., radar, speed sensors, GPS, etc.). However, in some cases, the navigation characteristics of a target vehicle can be determined in part or entirely based on the analysis of images of the host vehicle's environment. For example, an image analysis technique described above and incorporated herein by reference can be used to recognize a target vehicle in the host vehicle's environment. Monitoring the position of the target vehicle in captured images over a period of time, and / or monitoring the position of one or more features associated with the target vehicle (e.g., taillights, headlights, bumpers, wheels, etc.) in captured images may enable the determination of relative distances, speeds, and / or accelerations between the target vehicle and the host vehicle, or between the target vehicle and one or more other objects in the host vehicle's environment.

[0384] The level of aggression of an identified target vehicle can be inferred from any appropriate observed navigation characteristics of the target vehicle or any combination of observed navigation characteristics. For example, aggression can be determined from any observed characteristics and one or more predetermined threshold levels or any other appropriate qualitative or quantitative analysis. In some embodiments, a target vehicle can be considered aggressive if it is observed to be following a host vehicle or another vehicle at a distance less than a predetermined aggressive distance threshold. On the other hand, a target vehicle can be considered defensive if it is observed to be following a host vehicle or another vehicle at a distance greater than a predetermined aggressive distance threshold. The predetermined aggressive distance threshold does not have to be the same as the predetermined defensive distance threshold. In addition, either or both of the predetermined aggressive distance threshold and the predetermined defensive distance threshold may include a range of values ​​rather than a clear boundary line. Furthermore, neither the predetermined aggressive distance threshold nor the predetermined defensive distance threshold has to be fixed. Rather, these values ​​or ranges may shift over time, and various threshold / threshold ranges can be applied based on the observed characteristics of the target vehicle. For example, the threshold applied may depend on one or more other characteristics of the target vehicle. Higher observed relative velocity and / or acceleration may justify the application of a larger threshold / threshold range. Conversely, lower relative velocity and / or acceleration, including zero, may justify the application of a smaller distance threshold / threshold range when making aggressive / defensive inferences.

[0385] Aggressive / defensive inferences can also be based on thresholds of relative velocity and / or relative acceleration. A target vehicle can be considered aggressive if its observed relative velocity and / or relative acceleration relative to another vehicle exceeds a predetermined level or range. A target vehicle can be considered defensive if its observed relative velocity and / or relative acceleration relative to another vehicle falls below a predetermined level or range.

[0386] The determination of aggressive / defensive behavior can be based solely on any observed navigational characteristics, but this determination can also depend on any combination of observed characteristics. For example, as mentioned above, in some cases, a target vehicle can be considered aggressive simply because it is observed to be pursuing another vehicle at a distance below a certain threshold or range. However, in other cases, a target vehicle can be considered aggressive if it is pursuing another vehicle at a distance below a certain amount (which may be the same as or different from the threshold applied when the determination is based solely on distance) and has a relative speed and / or relative acceleration above a certain amount or range. Similarly, a target vehicle can be considered defensive simply because it is observed to be pursuing another vehicle at a distance above a certain threshold or range. However, in other cases, a target vehicle can be considered defensive if it is pursuing another vehicle beyond a certain amount (which may be the same as or different from the threshold applied when the determination is based solely on distance) and has a relative speed and / or relative acceleration below a certain amount or range. System 100 may perform offensive / defensive actions, for example, when a vehicle exceeds an acceleration or deceleration of 0.5G (e.g., a jerk of 5 m / s³), when a vehicle has a lateral acceleration of 0.5G on a lane change or curve, when a vehicle causes another vehicle to do any of the above, when a vehicle changes lanes and causes another vehicle to yield the right of way by exceeding a deceleration of 0.3G or a jerk of 3 m / s³, and / or when a vehicle changes lanes by two without stopping.

[0387] It should be understood that mentioning a quantity above a certain range may indicate that the quantity is above or within the range of all values ​​related to that range. Similarly, mentioning a quantity below a certain range may indicate that the quantity is below or within the range of all values ​​related to that range. In addition, while the examples described for making aggressive / defensive reasoning have been explained using distance, relative acceleration, and relative velocity, any other appropriate quantity can be used. For example, time to collision, which can be used in calculations, or any indirect indicator of the target vehicle's distance, acceleration, and / or velocity can be used. While the above examples focus on a target vehicle relative to other vehicles, it should also be noted that aggressive / defensive reasoning can be performed by observing the navigation characteristics of a target vehicle relative to any other type of obstacle (e.g., pedestrians, road barriers, traffic lights, debris, etc.).

[0388] Returning to the example shown in Figures 17A and 17B, as the host vehicle 1701 approaches the ring road, the navigation system, which includes at least one of its own processing devices, can receive an image stream from a camera associated with the host vehicle. Based on the analysis of one or more of the received images, it can identify any of the target vehicles 1703, 1705, 1706, 1708, and 1710. Furthermore, the navigation system can analyze the navigation characteristics of one or more of the identified target vehicles. The navigation system can recognize that the gap between target vehicles 1703 and 1705 represents a first opportunity for a potential merge into the ring road. The navigation system can analyze target vehicle 1703 to determine an indicator of aggression associated with target vehicle 1703. If target vehicle 1703 is deemed aggressive, the host vehicle's navigation system can decide to yield to vehicle 1703 rather than merge ahead of it. On the other hand, if target vehicle 1703 is deemed to be acting in self-defense, the host vehicle's navigation system may attempt to complete a merging maneuver in front of vehicle 1703.

[0389] When the host vehicle 1701 reaches a roundabout, at least one processing device of the navigation system can analyze the captured image to determine navigation characteristics related to the target vehicle 1703. For example, based on the image, it may be determined that vehicle 1703 is following vehicle 1705 at a distance that gives host vehicle 1701 enough space to enter safely. In fact, it may be determined that vehicle 1703 is following vehicle 1705 at a distance exceeding an aggressive distance threshold, and therefore, based on this information, the host vehicle's navigation system may be inclined to identify target vehicle 1703 as self-defensive. However, in some situations, multiple navigation characteristics of the target vehicle can be analyzed when making an aggressive / self-defensive decision, as discussed above. Expanding the analysis, the host vehicle's navigation system may determine that target vehicle 1703 is following target vehicle 1705 at a non-aggressive distance, but vehicle 1703 has a relative velocity and / or relative acceleration to vehicle 1705 that exceeds one or more thresholds associated with aggressive behavior. In fact, the host vehicle 1701 can determine that the target vehicle 1703 is accelerating toward vehicle 1705, narrowing the gap between vehicles 1703 and 1705. Based on further analysis of relative speed, acceleration, and distance (and the speed at which the gap between vehicles 1703 and 1705 is narrowing), the host vehicle 1701 can determine that the target vehicle 1703 is behaving aggressively. Therefore, although there may be enough space for the host vehicle to navigate safely, the host vehicle 1701 can anticipate that merging in front of the target vehicle 1703 will result in a vehicle aggressively navigating directly behind the host vehicle. Furthermore, if the host vehicle 1701 merges in front of vehicle 1703, based on behavior observed by image analysis or other sensor outputs, it can be anticipated that the target vehicle 1703 will continue to accelerate toward the host vehicle 1701, or proceed toward the host vehicle 1701 at a non-zero relative speed. Such situations can be undesirable from a safety standpoint and may also cause discomfort to the occupants of the host vehicle.For these reasons, as shown in Figure 17B, host vehicle 1701 may decide to yield to vehicle 1703 and merge into the ring road behind vehicle 1703 and in front of vehicle 1710, which is deemed to be in a self-defensive position based on one or more analyses of its navigation characteristics.

[0390] Returning to Figure 18, in step 1809, at least one processing device of the host vehicle's navigation system may determine a navigation action for the host vehicle (e.g., merging in front of vehicle 1710 and behind vehicle 1703) based on at least one identified navigation characteristic of the target vehicle relative to an identified obstacle. To carry out the navigation action (step 1811), at least one processing device may trigger at least one adjustment of the host vehicle's navigation actuator in accordance with the determined navigation action. For example, the brakes may be applied to yield to vehicle 1703 in Figure 17A, and the accelerator may be pressed along with steering of the host vehicle's wheels to bring the host vehicle into the roundabout behind vehicle 1703, as shown in Figure 17B.

[0391] As explained in the example above, the host vehicle's navigation can be based on the target vehicle's navigation characteristics relative to another vehicle or object. In addition, the host vehicle's navigation can be based solely on the target vehicle's navigation characteristics without specifically referencing another vehicle or object. For example, in step 1807 of Figure 18, analyzing multiple images captured from the host vehicle's environment may enable the determination of at least one navigation characteristic of an identified target vehicle that indicates the level of aggression associated with the target vehicle. Navigation characteristics may include speed, acceleration, etc., which do not require reference to another object or target vehicle to make an aggressive / defensive determination. For example, observed acceleration and / or speed associated with a target vehicle that exceeds a predetermined threshold, falls within a certain range, or exceeds a certain range may indicate aggressive behavior. Conversely, observed acceleration and / or speed associated with a target vehicle that falls below a predetermined threshold, falls within a certain range, or exceeds a certain range may indicate defensive behavior.

[0392] Naturally, in some cases, observed navigation characteristics (e.g., position, distance, acceleration, etc.) can be referenced for the host vehicle to make an aggressive / defensive decision. For example, observed navigation characteristics of a target vehicle indicating the level of aggression associated with the target vehicle may include the increase in relative acceleration between the target vehicle and the host vehicle, the distance the target vehicle is following behind the host vehicle, and the relative speed between the target vehicle and the host vehicle.

[0393] Navigation based on liability constraints

[0394] As described in the section above, planned navigation actions can be tested against predetermined constraints to ensure compliance with specific rules. In some embodiments, this concept can be extended to the consideration of potential accident liability. As will be discussed below, the primary goal of autonomous navigation is safety. Since absolute safety may be impossible (for example, because a particular host vehicle under autonomous control cannot control other vehicles around it and can only control its own actions), using potential accident liability as a consideration in autonomous navigation and as a constraint on planned actions can facilitate ensuring that a particular autonomous vehicle does not perform any actions that are considered dangerous (for example, actions for which potential accident liability could be attributed to the host vehicle). If the host vehicle performs only actions that are determined to be safe and not cause an accident for which the host vehicle is at fault or responsible, then the desired level of accident avoidance (e.g., 10 per hour of driving time) can be achieved. -9 It is possible to achieve less than ( ).

[0395] The challenges posed by the latest approaches to autonomous driving are the lack of safety assurance (or at least the inability to provide the desired level of safety) and the lack of scalability. We will examine the problem of ensuring safe multi-agent driving. Since society is unlikely to accept fatal traffic accidents caused by machines, an acceptable level of safety is paramount to the acceptance of autonomous vehicles. The goal may be to have no accidents at all, but this may be impossible, as we can imagine situations where multiple agents are generally involved in an accident, and the accident occurs solely due to the negligence of other agents. For example, as shown in Figure 19, host vehicle 1901 is traveling on a multi-lane highway, and host vehicle 1901 can control its own actions relative to target vehicles 1903, 1905, 1907, and 1909, but cannot control the actions of the target vehicles surrounding it. As a result, for example, if vehicle 1905 suddenly cuts into the host vehicle's lane on a collision course with the host vehicle, host vehicle 1901 may not be able to avoid a collision with at least one of the target vehicles. To address this challenge, a typical response from autonomous vehicle experts is to employ a data-driven approach where safety verification becomes more rigorous as data is collected over longer distances.

[0396] However, to understand the problematic nature of data-driven approaches to safety, the fatality rate caused by accidents per hour of (human) driving is 10 -6 First, consider that it is known that for society to accept machines replacing humans in driving tasks, the fatality rate must be in the triple digits, i.e., 10 per hour. -9 It is reasonable to assume that the probability should decrease to 10. This estimate is similar to the assumed fatality rates for airbags and aviation standards. For example, 10 -9 This is the probability that a wing will spontaneously detach from an aircraft in mid-air. However, attempts to guarantee safety using data-driven methods that become more reliable in proportion to the cumulative distance traveled are not practical. 10 per hour of operation -9 The amount of data needed to guarantee the mortality rate is the reverse (i.e., 109 The data is proportional to the time, which is approximately 30 billion miles. Furthermore, multi-agent systems interact with their environment and are unlikely to be verifiable offline unless a realistic simulator is available that emulates real human driving with all its richness and complexities such as reckless driving (however, the problem of verifying a simulator is far more difficult than creating a safe autonomous vehicle agent). Any changes to the planning and control software would require the collection of new data on the same scale, which is clearly unmanageable and impractical. Moreover, developing systems by data always faces a lack of interpretability and explainability of the actions performed, and if an autonomous vehicle (AV) causes an accident resulting in a fatality, the reason must be known. As a result, model-based methods for safety are needed, but existing "functional safety" and ASIL requirements in the automotive industry are not designed to deal with multi-agent environments.

[0397] The second major challenge in developing safe driving models for autonomous vehicles is the need for scalability. The premise underlying AVs goes beyond "creating a better world" and is rather based on the premise that driverless mobility can be maintained at a lower cost than driver-driven mobility. This premise is always linked to the concept of scalability in the sense of supporting the mass production of AVs (in the millions) and, more importantly, supporting the negligible incremental costs of enabling them to operate in new cities. Therefore, computation and sensing costs are important, and if AVs are mass-produced, the costs of validation and the ability to operate "anywhere," rather than just in a select few cities, are also necessary requirements to sustain the business.

[0398] The problem with recent methodologies lies in their "brute-force" approach along three axes: (i) the required "computational density," (ii) the method by which high-resolution maps are defined and created, and (iii) the required specifications of the sensors. Brute-force methods are contrary to scalability, leading to an ubiquitous and infinite onboard computation, making the cost of building and maintaining HD maps trivial and scalable, and shifting the burden to a future where novel and ultra-advanced sensors are developed and commercialized at trivial costs to match the grade of the vehicle. While any of the above scenarios are indeed reasonable, achieving all of these effects is likely to be a low-probability event. Therefore, there is a need to provide a formal model that is scalable in the sense that it can be accepted by society and support millions of cars operating in any part of an advanced country, integrating safety and scalability within AV programs.

[0399] The disclosed embodiments represent a solution that can provide a target level of safety (and even exceed it) and can scale to systems involving millions (or more) autonomous vehicles. In terms of safety, a model called “Responsibility-Sensitive Safety” (RSS) is introduced, which formalizes the concept of “accidental negligence” as interpretable and explainable, and incorporates a sense of “responsibility” into the actions of robotic agents. The definition of RSS is agnostic by the way it is implemented, which is a key feature for facilitating the goal of creating a human-convincing global safety model. RSS is motivated by the idea that agents play an asymmetric role in an accident (as seen in Figure 19), in which typically only one agent is responsible for the accident and therefore responsible. The RSS model also includes a formal treatment of “cautious driving” under limited detection conditions where not all agents are always visible (e.g., due to occlusion). One of the main goals of the RSS model is to ensure that agents never cause an accident in which they are “negligent” or responsible. The model may only be useful if it is accompanied by an efficient policy that conforms to RSS (e.g., a function that maps "detection states" to actions). For example, an action that seems harmless at the present time may lead to a catastrophe in the distant future ("butterfly effect"). RSS can help construct a set of local constraints on the short-term future that can guarantee (or at least effectively guarantee) that future accidents will not occur as a result of the host vehicle's actions.

[0400] Another contribution revolves around the introduction of a “semantic” language consisting of units, measurements, and operating spaces, and specifications regarding how they are incorporated into AV planning, detection, and operation. To understand what semantics is, we examine how a person receiving driving instruction in this regard is instructed to think about “driving policy.” These instructions are not geometric (e.g., “travel 13.7 meters at the current speed, then 0.8 m / s 2(It does not take the form of "accelerate at a certain rate"). Rather, these instructions are of a semantic nature ("follow the car in front" or "overtake the car on the left"). Typical language of human driving policy concerns longitudinal and lateral targets rather than geometric units of acceleration vectors. Formal semantic language can be useful in several ways: it is linked to the complexity of planning calculations that do not increase exponentially with time and the number of agents; it is linked to how safety and comfort interact; it is linked to how sensing calculations are defined; and it is linked to the specifications of sensor modalities and how they interact within a fusion methodology. A fusion methodology (based on semantic language) is 10 5 While only offline validation is performed on datasets of driving data lasting a few hours, the RSS model requires 10 hours of driving time per hour. -9 This can guarantee that the mortality rate will be achieved.

[0401] For example, in a reinforcement learning setup, regardless of the time period used for planning, the number of trajectories to be examined at any given time is 10. 4 A Q-function (for example, a function that evaluates the long-term quality of an action a∈A performed by an agent in a state s∈S) on a semantic space constrained by . Given such a Q-function, the natural choice of action is the one of the highest quality, π(s)=argmax aA Q(s,a) can be chosen to define a Q function. The signal-to-noise ratio in this space can be high, allowing effective machine learning methods to successfully model the Q function. In the case of detection computation, semantics can allow for the distinction between errors that affect safety and errors that affect driving comfort. A PAC model (Probabilistically Approximately Correct (PAC)), borrowing Valiant's PAC learning terminology, is defined for detection, which is tied to the Q function and shows how to incorporate measurement errors into the plan in a way that conforms to RSS but allows for the optimization of driving comfort. Since other standard error measures, such as errors to the global coordinate system, may not conform to the PAC detection model, a semantic language may be important for the success of aspects of this model's characteristics. In addition, the semantic language can be constructed using low-bandwidth detection data and therefore can be constructed by crowdsourcing, and can be an important enabler for defining HD maps that support scalability.

[0402] In summary, the disclosed embodiments may include formal models encompassing key elements of AV, namely detection, planning, and operation. These models can facilitate ensuring the absence of accidents for which the AV is responsible, from a planning perspective. Furthermore, the PAC detection model may allow the described fused methodology to require only a reasonably sized set of offline data to comply with the described safety model, even with detection errors. Moreover, this model can link safety and scalability through semantic language, thereby providing a complete methodology for safe and scalable AV. Finally, it should be noted that developing approved safety models adopted by industry and regulatory bodies may be a necessary condition for the success of AV.

[0403] The RSS model can generally follow the classical sense-plan-act robot control method. The sensing system may be responsible for understanding the current state of the host vehicle's environment. The planning part, which can be called the "driving policy" and can be implemented as a set of instructions hardcoded by a trained system (e.g., a neural network) or a combination, may be responsible for determining what the best next move is in terms of the available options to achieve the driving objective (e.g., how to move from the left lane to the right lane to exit a highway). The acting part is responsible for implementing the plan (e.g., a system of actuators and one or more controllers for steering, accelerating, and / or braking the vehicle to perform selected navigation actions). The embodiments described below will primarily focus on the sensing and planning parts.

[0404] Accidents can stem from detection errors or planning errors. Planning is a multi-agent endeavor because there are other road users (humans and machines) that react to the operation of the AV. The RSS model described is designed to address safety, particularly with respect to the planning portion. This can be called multi-agent safety. Statistical methods can estimate the probability of planning errors "online," meaning that billions of miles must be driven with each new version of the software after each update to obtain an acceptable level of estimation for the frequency of planning errors. This is clearly impractical. As an alternative, the RSS model can provide a 100% guarantee (or virtually 100% guarantee) that the planning module will not make mistakes for which the AV is at fault (the concept of "fault" is formally defined). The RSS model can also provide an efficient means of verification that does not rely on online testing.

[0405] Detection can be independent of vehicle operation, and therefore the probability of serious detection errors can be verified using "offline" data, so errors in the detection system may be easier to verify. However, 10 9Even collecting offline data on driving over extended periods is difficult. As part of the description of the detection system we are disclosing, we have described a fusion technique that can be validated using significantly smaller amounts of data.

[0406] The RSS system described can also be scaled to accommodate millions of vehicles. For example, the semantic driving policies and applicable safety constraints described can match detection and mapping requirements that can scale to millions of vehicles, even with today's technology.

[0407] The fundamental component of such a system is the safety definition, which is the minimum standard that the AV system may need to adhere to. The following technical lemma demonstrates that statistical methods for verifying AV systems are impractical, even for verifying simple claims such as "the system experiences N accidents per hour." This implies that only model-based safety definitions are feasible tools for verifying AV systems. Lemma 1 Let X be a probability space, and A be an event for which Pr(A)=p1<0.1 holds.

number

number

Claims

1. A system for navigating a host vehicle, wherein the system The process involves receiving at least one image representing the environment of the host vehicle from an image capture device, The planned navigation operation for achieving the navigation objective of the host vehicle is determined based on at least one driving policy, To identify a target vehicle in the environment of the host vehicle, the at least one image is analyzed, and the direction of travel of the target vehicle is toward the host vehicle. To determine the distance between the host vehicle and the target vehicle in the following states, which will occur when the planned navigation operation is performed: To determine the host vehicle braking ratio, the host vehicle's maximum acceleration capability, and the host vehicle's current speed, The stopping distance of the host vehicle is determined based on the host vehicle braking ratio, the host vehicle's maximum acceleration capability, and the host vehicle's current speed. The current speed of the target vehicle, the target vehicle's maximum acceleration capability, and the target vehicle's braking ratio are determined. The stopping distance of the target vehicle is determined based on the target vehicle braking ratio, the target vehicle's maximum acceleration capability, and the target vehicle's current speed. If the distance to the next determined state is greater than the sum of the stopping distance of the host vehicle and the stopping distance of the target vehicle, the planned navigation operation shall be performed. A system comprising at least one processing device programmed to perform the following.

2. The system according to claim 1, wherein the target vehicle and the host vehicle are driving within the parking lot.

3. The system according to claim 1, wherein the target vehicle travels toward the host vehicle within a lane occupied by the host vehicle.

4. The system according to claim 1 or 2, wherein the host vehicle braking ratio corresponds to the maximum braking capacity of the host vehicle.

5. The system according to any one of claims 1 to 3, wherein the host vehicle braking ratio corresponds to a predetermined quasi-maximum braking capacity of the host vehicle.

6. The system according to any one of claims 1 to 5, wherein the target vehicle braking rate corresponds to the maximum braking capacity of the target vehicle determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

7. The system according to any one of claims 1 to 5, wherein the target vehicle braking rate corresponds to a predetermined quasi-maximum braking capacity of the target vehicle.

8. The system according to any one of claims 1 to 7, wherein the host vehicle braking rate is smaller than the target vehicle braking rate.

9. The system according to any one of claims 1 to 8, wherein the current speed of the target vehicle is determined based on a review of two or more received images representing the environment of the host vehicle.

10. The system according to any one of claims 1 to 8, wherein the current speed of the target vehicle is determined based on the output of a radar unit or ridiculous unit associated with the host vehicle.

11. The system according to any one of claims 1 to 10, wherein the stopping distance of the host vehicle includes a host vehicle acceleration distance, which corresponds to the distance the host vehicle can travel during the host vehicle response time at the host vehicle's maximum acceleration capability, starting from the current speed of the host vehicle as determined.

12. The system according to any one of claims 1 to 11, wherein the stopping distance of the target vehicle includes a target vehicle acceleration distance, which corresponds to the distance the target vehicle can travel during the target vehicle response time at the target vehicle's maximum acceleration capability, starting from the current speed of the target vehicle as determined.

13. The system according to any one of claims 1 to 12, wherein at least one of the host vehicle braking rate or the target vehicle braking rate is determined based on detected road surface conditions.

14. The system according to any one of claims 1 to 13, wherein at least one of the host vehicle braking rate or the target vehicle braking rate is determined based on detected weather conditions.

15. The system according to any one of claims 1 to 14, wherein the maximum acceleration capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

16. The system according to claim 15, wherein the at least one characteristic of the target vehicle includes one or more of the vehicle type, vehicle size, or vehicle model.

17. The system according to any one of claims 1 to 16, wherein the planned navigation operation includes at least one of lane change operation, merging operation, passing operation, maintaining forward operation, parking navigation operation, or maintaining throttle operation.

18. The system according to any one of claims 1 to 17, wherein the at least one processing device is configured to perform the planned navigation operation when the distance to the next determined state is greater by at least a predetermined minimum distance than the sum of the stopping distance of the host vehicle and the stopping distance of the target vehicle.

19. The system according to claim 18, wherein the predetermined minimum distance corresponds to a predetermined separation distance maintained between the host vehicle and other vehicles.

20. The system according to claim 18, wherein the predetermined separation distance is at least 1 meter.

21. A method for navigating a host vehicle, wherein the method is The steps include receiving at least one image representing the environment of the host vehicle from an image capture device, The steps include determining a planned navigation operation to achieve the navigation objective of the host vehicle based on at least one driving policy, A step of analyzing at least one image in order to identify a target vehicle in the environment of the host vehicle, wherein the direction of travel of the target vehicle is toward the host vehicle, The steps include determining the distance between the host vehicle and the target vehicle in the following states, which will occur when the planned navigation operation is performed, The steps include determining the host vehicle braking rate, the host vehicle's maximum acceleration capability, and the current speed of the host vehicle, A step of determining the stopping distance of the host vehicle based on the host vehicle braking ratio, the host vehicle's maximum acceleration capability, and the host vehicle's current speed. The steps include determining the current speed of the target vehicle, the target vehicle's maximum acceleration capability, and the target vehicle's braking ratio, A step of determining the stopping distance of the target vehicle based on the target vehicle braking rate, the target vehicle's maximum acceleration capability, and the target vehicle's current speed. If the distance to the next determined state is greater than the sum of the stopping distance of the host vehicle and the stopping distance of the target vehicle, the step of performing the planned navigation operation is performed. Methods that include...

22. The method according to claim 21, wherein the host vehicle braking ratio corresponds to the maximum braking capacity of the host vehicle.

23. The method according to claim 21, wherein the host vehicle braking ratio corresponds to a predetermined quasi-maximum braking capacity of the host vehicle.

24. The method according to any one of claims 21 to 23, wherein the target vehicle braking rate corresponds to the maximum braking capacity of the target vehicle determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

25. The method according to any one of claims 21 to 23, wherein the target vehicle braking rate corresponds to a predetermined quasi-maximum braking capacity of the target vehicle.

26. The method according to any one of claims 21 to 25, wherein the host vehicle braking rate is smaller than the target vehicle braking rate.

27. The method according to any one of claims 21 to 26, wherein the current speed of the target vehicle is determined based on analysis of two or more received images representing the environment of the host vehicle.

28. The method according to any one of claims 21 to 26, wherein the current speed of the target vehicle is determined based on the output of a radar unit or ridiculous unit associated with the host vehicle.

29. The method according to any one of claims 21 to 28, wherein the maximum acceleration capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

30. The method according to claim 29, wherein the at least one characteristic of the target vehicle includes one or more of the vehicle type, vehicle size, or vehicle model.

31. The method according to any one of claims 21 to 30, wherein the planned navigation operation includes at least one of lane change operation, merging operation, passing operation, maintaining forward operation, parking navigation operation, or maintaining throttle operation.

32. A system for navigating a host vehicle, wherein the system The process involves receiving at least one image representing the environment of the host vehicle from an image capture device, The planned navigation operation for achieving the navigation objective of the host vehicle is determined based on at least one driving policy, To identify the target vehicle in the environment of the host vehicle, the at least one image is analyzed, To determine the lateral distance between the host vehicle and the target vehicle in the following states, which will occur when the planned navigation operation is performed: The maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, and the current lateral speed of the host vehicle are determined. The lateral braking distance of the host vehicle is determined based on the host vehicle's maximum yaw rate capability, the maximum change in the host vehicle's turning radius capability, and the host vehicle's current lateral speed. The current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in the turning radius capability of the target vehicle are determined. The lateral braking distance of the target vehicle is determined based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in the turning radius capability of the target vehicle. If the lateral distance of the next determined state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle, the planned navigation operation shall be performed. A system including at least one processing device programmed to perform the following.

33. The system according to claim 32, wherein the lateral braking distance of the host vehicle corresponds to the distance required for the lateral speed of the host vehicle to reach zero, starting from the current lateral speed of the host vehicle and following the lateral acceleration of the host vehicle as a result of the maximum change in the maximum yaw rate capability and turning radius capability of the host vehicle during the response time associated with the host vehicle.

34. The system according to claim 32 or 33, wherein the lateral braking distance of the target vehicle corresponds to the distance required for the target vehicle's lateral speed to reach zero, starting from the current lateral velocity of the target vehicle and occurring at the maximum change in the target vehicle's lateral velocity and turning radius capability, after the lateral acceleration of the target vehicle as a result of a rotational operation by the target vehicle toward the host vehicle, performed at the maximum change in the target vehicle's lateral velocity and turning radius capability, during the response time associated with the target vehicle.

35. The system according to any one of claims 32 to 34, wherein the at least one processing device is programmed to perform the planned navigation operation regardless of the lateral distance of the next state if the planned navigation operation results in a safe longitudinal distance between the host vehicle and the target vehicle.

36. The aforementioned safe longitudinal distance is greater than or equal to the sum of the longitudinal stopping distance of the host vehicle and the longitudinal stopping distance of the target vehicle. The stopping distance of the host vehicle includes, based on the maximum longitudinal braking capacity of the host vehicle and the current longitudinal speed of the host vehicle, a host vehicle acceleration distance corresponding to the distance the host vehicle can travel at its maximum acceleration capacity during a response time associated with the host vehicle, starting from the current longitudinal speed of the host vehicle. The system according to claim 35, wherein the stopping distance of the target vehicle includes, based on the maximum longitudinal braking capacity of the target vehicle and the current longitudinal speed of the target vehicle, a target vehicle acceleration distance corresponding to the distance the target vehicle can travel at its maximum acceleration capacity during a response time associated with the target vehicle, starting from the current longitudinal speed of the target vehicle.

37. The system according to any one of claims 32 to 36, wherein at least one of the maximum yaw rate capability of the target vehicle and the maximum change in the turning radius capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

38. The system according to claim 37, wherein the at least one characteristic of the target vehicle includes one or more of the vehicle type, vehicle size, or vehicle model.

39. The system according to any one of claims 32 to 38, wherein the maximum change in the host vehicle's maximum yaw rate capability and the host vehicle's turning radius capability is based on predetermined constraints of the host vehicle.

40. The system according to any one of claims 32 to 39, wherein the current lateral speed of the target vehicle is determined based on the output of a radar unit or ridiculous unit associated with the host vehicle.

41. The system according to any one of claims 32 to 40, wherein at least one of the maximum yaw rate capability or the maximum change in the turning radius capability of the host vehicle is determined based on detected road surface conditions.

42. The system according to any one of claims 32 to 41, wherein at least one of the maximum yaw rate capability of the target vehicle or the maximum change in the turning radius capability of the target vehicle is determined based on detected road surface conditions.

43. The system according to any one of claims 32 to 42, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in the turning radius capability of the target vehicle is determined based on the weather conditions on which the detection occurred.

44. The system according to any one of claims 32 to 43, wherein the planned navigation operation includes at least one of lane change operation, merging operation, passing operation, parking navigation operation, or turning operation.

45. The system according to any one of claims 32 to 44, wherein the at least one processing device is configured to perform the planned navigation operation when the lateral distance of the next determined state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle by at least a predetermined minimum distance.

46. The system according to claim 45, wherein the predetermined minimum distance corresponds to a predetermined lateral separation distance maintained between the host vehicle and other vehicles.

47. The system according to claim 46, wherein the predetermined lateral separation distance is at least 1 meter.

48. A method for navigating a host vehicle, wherein the method is The steps include receiving at least one image representing the environment of the host vehicle from an image capture device, The steps include determining a planned navigation operation to achieve the navigation objective of the host vehicle based on at least one driving policy, The steps include analyzing the at least one image in order to identify a target vehicle in the environment of the host vehicle, The steps include determining the lateral distance between the host vehicle and the target vehicle in the following states, which will occur when the planned navigation operation is performed, The steps include determining the maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, and the current lateral speed of the host vehicle. A step of determining the lateral braking distance of the host vehicle based on the maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, and the current lateral speed of the host vehicle. The steps include determining the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in the turning radius capability of the target vehicle, A step of determining the lateral braking distance of the target vehicle based on the current lateral speed of the target vehicle, the maximum yaw rate capability of the target vehicle, and the maximum change in the turning radius capability of the target vehicle. If the lateral distance of the next determined state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle, the stage of performing the planned navigation operation is performed. Methods that include...

49. The method of claim 48, wherein the lateral braking distance of the host vehicle corresponds to the distance required for the lateral speed of the host vehicle to reach zero, starting from the current lateral speed of the host vehicle and following the lateral acceleration of the host vehicle as a result of the maximum change in the maximum yaw rate capability and turning radius capability of the host vehicle during the response time associated with the host vehicle.

50. The method according to claim 48 or 49, wherein the lateral braking distance of the target vehicle corresponds to the distance required for the target vehicle's lateral speed to reach zero, starting from the current lateral velocity of the target vehicle and occurring at the maximum change in the target vehicle's lateral acceleration toward the host vehicle as a result of a rotational operation by the target vehicle toward the host vehicle, performed at the maximum change in the target vehicle's lateral velocity and turning radius capability during a response time related to the target vehicle.

51. The method according to any one of claims 48 to 50, wherein the at least one processing device is programmed to perform the planned navigation operation regardless of the lateral distance of the next state if the planned navigation operation results in a safe longitudinal distance between the host vehicle and the target vehicle.

52. The aforementioned safe longitudinal distance is greater than or equal to the sum of the longitudinal stopping distance of the host vehicle and the longitudinal stopping distance of the target vehicle. The stopping distance of the host vehicle includes, based on the maximum longitudinal braking capacity of the host vehicle and the current longitudinal speed of the host vehicle, a host vehicle acceleration distance corresponding to the distance the host vehicle can travel at its maximum acceleration capacity during a response time associated with the host vehicle, starting from the current longitudinal speed of the host vehicle. The method according to claim 51, wherein the stopping distance of the target vehicle includes, based on the maximum longitudinal braking capacity of the target vehicle and the current longitudinal speed of the target vehicle, a target vehicle acceleration distance corresponding to the distance the target vehicle can travel at its maximum acceleration capacity during a response time associated with the target vehicle, starting from the current longitudinal speed of the target vehicle.

53. The method according to any one of claims 48 to 52, wherein at least one of the maximum yaw rate capability and the maximum change in the turning radius capability of the target vehicle is determined based on at least one characteristic of the target vehicle identified through the analysis of the at least one image.

54. The method according to claim 53, wherein the at least one characteristic of the target vehicle includes one or more of the vehicle type, vehicle size, or vehicle model.

55. The method according to any one of claims 48 to 54, wherein the maximum change in the host vehicle's maximum yaw rate capability and the host vehicle's turning radius capability is based on predetermined constraints of the host vehicle.

56. The method according to any one of claims 48 to 55, wherein the current lateral speed of the target vehicle is determined based on the output of a radar unit or ridiculous unit associated with the host vehicle.

57. The method according to any one of claims 48 to 56, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in the turning radius capability of the target vehicle is determined based on the detected road surface conditions.

58. The method according to any one of claims 48 to 57, wherein at least one of the maximum yaw rate capability of the host vehicle, the maximum change in the turning radius capability of the host vehicle, the maximum yaw rate capability of the target vehicle, or the maximum change in the turning radius capability of the target vehicle is determined based on the weather conditions under which it is detected.

59. The method according to any one of claims 48 to 58, wherein the planned navigation operation includes at least one of a lane change operation, a merging operation, a passing operation, a parking navigation operation, or a turning operation.

60. The method according to any one of claims 48 to 59, wherein the at least one processing device is configured to perform the planned navigation operation when the lateral distance of the next determined state is greater than the sum of the lateral braking distance of the host vehicle and the lateral braking distance of the target vehicle by at least a predetermined minimum distance.

61. The method according to claim 60, wherein the predetermined minimum distance corresponds to a predetermined lateral separation distance maintained between the host vehicle and other vehicles.

62. A system for navigating a host vehicle near a pedestrian crossing, wherein the system The process involves receiving at least one image representing the environment of the host vehicle from an image capture device, Based on the analysis of the at least one image, the representation of a pedestrian crossing in the at least one image is detected. Based on the analysis of the at least one image, it is determined whether a representation of a pedestrian appears in the at least one image. The host vehicle detects the presence of traffic lights in the environment, To determine whether the detected traffic signal is related to the host vehicle and the pedestrian crossing, To determine the state of the detected traffic signal, When a representation of a pedestrian appears in at least one of the images, the proximity of the pedestrian to the detected crosswalk is determined, Determining a planned navigation action to navigate the host vehicle to the detected crosswalk based on at least one driving policy, wherein the determination of the planned navigation action is further based on the determined state of the detected traffic signal and the determined proximity of the pedestrian to the detected crosswalk. To cause one or more actuator systems of the host vehicle to perform the planned navigation operation. A system comprising at least one processing device programmed to perform the following.

63. The system according to claim 62, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk without adjustments guided by the pedestrian if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk indicates that there is no possibility of the pedestrian entering the crosswalk before the host vehicle and the crosswalk intersect.

64. The system according to claim 62 or 63, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk without any adjustments induced by the pedestrian if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk, in combination with the detected direction and speed associated with the pedestrian, indicates that the pedestrian cannot enter the crosswalk before the host vehicle and the crosswalk intersect.

65. The system according to any one of claims 62 to 64, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk with adjustments guided by the pedestrian if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk indicates that the pedestrian can enter the crosswalk before the host vehicle and the crosswalk intersect.

66. The system according to claim 65, wherein the adjustments induced by the pedestrian are determined based on the detected direction and speed of the pedestrian.

67. The system according to claim 65 or 66, wherein the adjustment induced by the pedestrian includes a lateral movement away from the pedestrian.

68. The system according to any one of claims 65 to 67, wherein the adjustment induced by the pedestrian includes slowing down the host vehicle before the host vehicle crosses the pedestrian crossing.

69. The system according to any one of claims 62 to 68, wherein the planned navigation operation of the host vehicle includes stopping the host vehicle in front of the crosswalk when the traffic light is green and a pedestrian is detected at the crosswalk.

70. The system according to any one of claims 62 to 69, wherein the determination of whether the detected traffic light is associated with the host vehicle and the pedestrian crossing is based on traffic light-related information contained in one or more maps accessible by the at least one processing device.

71. The system according to any one of claims 62 to 70, wherein the detection of the presence of the traffic light in the environment of the host vehicle is based on the analysis of the at least one image.

72. The system according to any one of claims 62 to 70, wherein the detection of the presence of the traffic light in the environment of the host vehicle is based on one or more signals transmitted by the detected traffic light.

73. The system according to any one of claims 62 to 72, wherein the determination of the state of the detected traffic signal is based on the analysis of at least one image.

74. The system according to any one of claims 62 to 72, wherein the determination of the state of the detected signal is based on one or more signals transmitted by the detected signal.

75. A method for navigating a host vehicle near a pedestrian crossing, wherein the method is The steps include receiving at least one image representing the environment of the host vehicle from an image capture device, The steps include detecting the representation of a pedestrian crossing in the at least one image based on the analysis of the at least one image, A step of determining whether a representation of a pedestrian appears in the at least one image based on the analysis of the at least one image, The steps include detecting the presence of traffic lights in the environment of the host vehicle, The steps include determining whether the detected traffic signal is related to the host vehicle and the pedestrian crossing, The step of determining the state of the detected traffic signal, When a representation of a pedestrian appears in at least one of the images, the steps include determining the proximity of the pedestrian to the detected crosswalk, A step of determining a planned navigation action to navigate the host vehicle to the detected crosswalk based on at least one driving policy, wherein the determination of the planned navigation action is further based on the determined state of the detected traffic signal and the determined proximity of the pedestrian to the detected crosswalk. The steps include causing one or more actuator systems of the host vehicle to perform the planned navigation operation. A method that includes [a certain feature].

76. The method according to claim 75, wherein the detection of the presence of the traffic light in the environment of the host vehicle is based on one or more signals transmitted by the detected traffic light.

77. The method according to claim 75 or 76, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk without adjustments guided by the pedestrian if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk indicates that there is no possibility of the pedestrian entering the crosswalk before the host vehicle and the crosswalk intersect.

78. The method according to any one of claims 75 to 77, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk without any adjustments induced by the pedestrian if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk, in combination with the detected direction and speed associated with the pedestrian, indicates that the pedestrian cannot enter the crosswalk before the host vehicle intersects with the crosswalk.

79. The method according to any one of claims 75 to 78, wherein the planned navigation operation of the host vehicle includes crossing the crosswalk with adjustments guided by the pedestrian, if the detected traffic light is green and the determined proximity of the pedestrian to the detected crosswalk indicates that the pedestrian can enter the crosswalk before the host vehicle and the crosswalk intersect.

80. A method for navigating a host vehicle near a pedestrian crossing, wherein the method is The steps include receiving at least one image representing the environment of the host vehicle from an image acquisition device, Steps include detecting the start and end positions of the pedestrian crossing, The steps include determining whether a pedestrian is present near the crosswalk based on the analysis of at least one of the aforementioned images, The steps include detecting the presence of traffic lights in the environment of the host vehicle, The steps include determining whether the traffic signal is related to the host vehicle and the pedestrian crossing, The step of determining the state of the aforementioned traffic signal, The relationship between the aforementioned traffic signals, The determined state of the aforementioned signal, The presence of the aforementioned pedestrian near the aforementioned crosswalk, The shortest distance selected between the pedestrian and either the starting position or the ending position of the crosswalk, The motion vector of the aforementioned pedestrian, Based on this, the steps include determining the navigation behavior of the host vehicle near the detected pedestrian crossing, The step of causing one or more actuator systems of the host vehicle to perform the navigation operation. Methods that include...

81. The method according to claim 80, wherein the detection of the presence of the traffic light in the environment of the host vehicle is based on one or more signals transmitted by the detected traffic light.