PREDICTION OF THE DRIVER'S PATH BASED ON THE VIEW DIRECTION AND ITS APPLICATIONS

The system uses cameras and sensor data to predict and control vehicle paths based on driver gaze, enhancing safety and scalability in autonomous navigation.

DE102025150375A1Pending Publication Date: 2026-06-03MOBILEYE VISION TECH LTD

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
MOBILEYE VISION TECH LTD
Filing Date
2025-12-03
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Autonomous vehicles face challenges in navigating safely and efficiently while adhering to liability rules and ensuring scalability, requiring an interpretable mathematical model for safety assurance.

Method used

A system utilizing cameras to monitor the vehicle's surroundings, process GPS and sensor data, and predict the driver's path based on gaze direction to control the vehicle's navigation.

Benefits of technology

Enhances safety and scalability by accurately predicting and controlling vehicle paths based on driver intent, addressing navigation complexities and liability considerations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This document provides system, equipment, device, process, and / or computer program product implementations and / or combinations and sub-combinations thereof for predicting the driver path of a vehicle. In this process, an image of the vehicle's driver is received. Based on this image, the driver's gaze direction is determined. The vehicle's path is then predicted, at least partially, based on this gaze direction. The vehicle is then controlled based on the predicted path.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND Technical area

[0001] This disclosure relates generally to autonomous vehicle navigation. Furthermore, this disclosure relates to navigation systems and methods, taking into account potential limitations on liability in accidents. Background information

[0002] With the steady advancement of technology, the goal of a fully autonomous vehicle capable of navigating roads is drawing closer. Autonomous vehicles may need to consider a multitude of factors and make appropriate decisions based on these factors to reach an intended destination safely and accurately. For example, an autonomous vehicle may need to process and interpret visual information (e.g., information captured by a camera) and information from radar or lidar, and may also use information from other sources (e.g., from a GPS device, a speed sensor, an accelerometer, a suspension sensor, etc.). At the same time, in order to navigate to a destination, an autonomous vehicle may also need to determine its position within a specific lane (e.g.,The navigation system can identify a specific lane within a multi-lane road, navigate alongside other vehicles, avoid obstacles and pedestrians, observe traffic signals and signs, change lanes at appropriate intersections or junctions, and respond to any other situation that occurs or arises during the vehicle's operation. Furthermore, the navigation system may be required to adhere to certain imposed limitations. In some cases, these limitations may relate to the interaction between a host vehicle and one or more other objects, such as other vehicles, pedestrians, etc. In other cases, the limitations may relate to liability rules that must be followed when performing one or more navigation actions for a host vehicle.

[0003] In the field of autonomous driving, two key considerations for viable autonomous vehicle systems are paramount. The first is the standardization of safety assurance, including the requirements that every self-driving car must meet to guarantee safety and how these requirements can be verified. The second is scalability, as engineering solutions that lead to exorbitant costs are not scalable to millions of cars and could hinder widespread, or even less widespread, acceptance of autonomous vehicles. Therefore, there is a need for an interpretable mathematical model for safety assurance and the design of a system that meets safety assurance requirements while remaining scalable to millions of cars. SUMMARY

[0004] Embodiments in accordance with the present disclosure provide systems and methods for autonomous vehicle navigation. The disclosed embodiments can use cameras to provide autonomous vehicle navigation functions. For example, in accordance with the disclosed embodiments, the disclosed systems can include one, two, or more cameras that monitor the vehicle's surroundings. The disclosed systems can, for example, provide a navigation response based on an analysis of images captured by one or more cameras. The navigation response can also take into account other data, including, for example, global positioning system (GPS) data, sensor data (e.g., from an accelerometer, a velocity sensor, a suspension sensor, etc.), and / or other map data.

[0005] In some embodiments, a non-transitory computer-readable device has instructions stored on it which, when executed by at least one computing device, cause at least one computing device to perform operations such as receiving an image of the driver of the vehicle; determining, based on the image, a direction of the driver's gaze; predicting, at least partially based on the direction of gaze, a path of the vehicle; and controlling the vehicle based on the predicted path.

[0006] In some embodiments, a system for predicting a vehicle's driver path includes one or more outward-facing sensors that monitor the vehicle's external environment; one or more inward-facing cameras that monitor the vehicle's driver; a memory; and at least one processor connected to the memory and configured to receive an image of the vehicle's driver; determine the driver's viewing direction based on the image; predict the vehicle's path, at least partially, based on the viewing direction; and control the vehicle based on the predicted path.

[0007] In some embodiments, a system for predicting a vehicle's driver path includes a vehicle control unit; a memory; and at least one processor connected to the memory and configured to receive an image of the vehicle's driver; determine the driver's viewing direction based on the image; predict the vehicle's path, at least partially, based on the viewing direction; and control the vehicle based on the predicted path.

[0008] In some embodiments, a computer-implemented method for predicting a vehicle's driver path includes receiving an image of the vehicle's driver; determining the driver's viewing direction based on the image; predicting the vehicle's path, at least partially, based on the viewing direction; and controlling the vehicle based on the predicted path.

[0009] In accordance with other disclosed embodiments, non-transitory computer-readable storage media can store program instructions that can be executed by at least one processing device and perform any of the steps and / or procedures described herein.

[0010] The foregoing general description and the following detailed description serve solely for illustration and explanation and do not limit the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are included in and form part of this disclosure, illustrate various disclosed embodiments. They show: Fig. Figure 1 shows a schematic representation of an exemplary system, in accordance with the disclosed embodiments. Fig. Figure 2A shows a schematic side view of an exemplary vehicle including a system, in accordance with the disclosed embodiments. Fig. 2B shows a schematic top view of the in Fig. 2A vehicle and system shown, in accordance with the disclosed embodiments. Fig. Figure 2C shows a schematic top view of another embodiment of a vehicle including a system, in accordance with the disclosed embodiments. Fig. Figure 2D shows a schematic top view of yet another embodiment of a vehicle including a system, in accordance with the disclosed embodiments. Fig. 2E shows a schematic top view of yet another embodiment of a vehicle including a system, in accordance with the disclosed embodiments. Fig. Figure 2F shows a schematic representation of exemplary vehicle control systems, in accordance with the disclosed embodiments. Fig. Figure 3A shows a schematic representation of the interior of a vehicle including a rearview mirror and a user interface for a vehicle imaging system, in accordance with the disclosed embodiments. Fig. Figure 3B shows an illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, in accordance with the disclosed embodiments. Fig. 3C shows an illustration of the in Fig. 3B camera mount shown from a different perspective, in accordance with the disclosed embodiments. Fig. Figure 3D shows an illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, in accordance with the disclosed embodiments. Fig. Figure 4 shows an exemplary block diagram of a memory configured to store instructions for performing one or more operations, in accordance with the disclosed embodiments. Fig. Figure 5A shows a flowchart illustrating an exemplary process for initiating one or more navigation responses based on monocular image analysis, in accordance with the disclosed embodiments. Fig. Figure 5B shows a flowchart illustrating an exemplary process for detecting one or more vehicles and / or pedestrians in a set of images, in accordance with the disclosed embodiments. Fig. Figure 5C shows a flowchart illustrating an exemplary process for recognizing road markings and / or lane geometry information in a set of images, in accordance with the disclosed embodiments. Fig. Figure 5D shows a flowchart illustrating an exemplary process for detecting traffic lights in a set of images, in accordance with the disclosed embodiments. Fig. Figure 5E shows a flowchart illustrating an exemplary process for initiating one or more navigation responses based on a vehicle path, in accordance with the disclosed embodiments. Fig. Figure 5F shows a flowchart illustrating an exemplary process for determining whether a vehicle ahead is changing lanes, in accordance with the disclosed embodiments. Fig. Figure 6 shows a flowchart illustrating an exemplary process for initiating one or more navigation responses based on a stereo image analysis, in accordance with the disclosed embodiments. Fig. Figure 7 shows a flowchart illustrating an exemplary process for initiating one or more navigation responses based on an analysis of three sets of images, in accordance with the disclosed embodiments. Fig. Figure 8 shows a block diagram representation of modules that can be implemented by one or more specially programmed processing devices of a navigation system for an autonomous vehicle, in accordance with the disclosed embodiments. Fig. Figure 9 shows a navigation options graph, in accordance with the disclosed embodiments. Fig. Figure 10 shows a navigation options graph, in accordance with the disclosed embodiments. Fig. 11A, Fig. 11B and Fig. Figure 11C shows a schematic representation of the navigation options of a host vehicle in a merging zone, in accordance with the disclosed embodiments. Fig. Figure 11D shows a schematic representation of a double threading scenario, in accordance with the disclosed embodiments. Fig. Figure 11E shows an options graph that is potentially useful in a double-threading scenario, in accordance with the disclosed embodiments. Fig. Figure 12 shows a diagram of a representative image captured from the environment of a host vehicle, together with potential navigation restrictions, in accordance with the disclosed embodiments. Fig. Figure 13 shows an algorithmic flowchart for navigating a vehicle, in accordance with the disclosed embodiments. Fig. Figure 14 shows an algorithmic flowchart for navigating a vehicle, in accordance with the disclosed embodiments. Fig. Figure 15 shows an algorithmic flowchart for navigating a vehicle, in accordance with the disclosed embodiments. Fig. Figure 16 shows an algorithmic flowchart for navigating a vehicle, in accordance with the disclosed embodiments. Fig. 17A and Fig. Figure 17B shows a schematic illustration of a host vehicle navigating into a roundabout, in accordance with the disclosed embodiments. Fig. Figure 18 shows an algorithmic flowchart for navigating a vehicle, in accordance with the disclosed embodiments. Fig. Figure 19 shows an example of a host vehicle traveling on a multi-lane highway in accordance with the disclosed embodiments. Fig. 20A and Fig. Figure 20B shows examples of a vehicle merging in front of another vehicle in accordance with the disclosed embodiments. Fig. Figure 21 shows an example of a vehicle following another vehicle in accordance with the disclosed embodiments. Fig. Figure 22 shows an example of a vehicle leaving a parking lot and merging onto a potentially busy road, in accordance with the disclosed embodiments. Fig. Figure 23 shows a vehicle driving on a road in accordance with the disclosed embodiments. Fig. Figures 24A-24D show four example scenarios, in accordance with the disclosed embodiments. Fig. Figure 25 shows an example scenario, in accordance with the disclosed embodiments. Fig. Figure 26 shows an example scenario, in accordance with the disclosed embodiments. Fig. Figure 27 shows an example scenario, in accordance with the disclosed embodiments. Fig. 28A and Fig. Figure 28B shows an example of a scenario in which one vehicle follows another vehicle, in accordance with the disclosed embodiments. Fig. 29A and Fig. 29B shows an example of the question of fault in lane-change scenarios, in accordance with the disclosed embodiments. Fig. 30A and Fig. Figure 30B shows an example of the question of fault in lane-change scenarios, in accordance with the disclosed embodiments. Fig. Figures 31A-31D show an example of the question of fault in drift scenarios, in accordance with the disclosed embodiments. Fig. 32A and Fig. Figure 32B shows an example of the question of fault in oncoming traffic scenarios, in accordance with the disclosed embodiments. Fig. 33A and Fig. Figure 33B shows an example of the question of fault in oncoming traffic scenarios, in accordance with the disclosed embodiments. Fig. 34A and Fig. Figure 34B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 35A and Fig. Figure 35B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 36A and Fig. Figure 36B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 37A and Fig. Figure 37B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 38A and Fig. Figure 38B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 39A and Fig. Figure 39B shows an example of the question of fault in route priority scenarios, in accordance with the disclosed embodiments. Fig. 40A and Fig. Figure 40B shows an example of the question of fault in traffic light scenarios, in accordance with the disclosed embodiments. Fig. 41A and Fig. Figure 41B shows an example of the question of fault in traffic light scenarios, in accordance with the disclosed embodiments. Fig. 42A and Fig. Figure 42B shows an example of the question of fault in traffic light scenarios, in accordance with the disclosed embodiments. Fig. Figures 43A-43C show example scenarios involving vulnerable road users (VRUs), in accordance with the disclosed embodiments. Fig. Figures 44A-44C show example scenarios involving vulnerable road users (VRUs) in accordance with the disclosed embodiments. Fig. 45A-45C show example scenarios involving vulnerable road users (VRUs), in accordance with the disclosed embodiments. Fig. Figures 46A-46D show example scenarios involving vulnerable road users (VRUs) in accordance with the disclosed embodiments. Fig. 47A and Fig. 47B shows example scenarios in which one vehicle follows another vehicle, in accordance with the disclosed embodiments. Fig. Figure 48 shows a flowchart illustrating an exemplary process for navigating a host vehicle in accordance with the disclosed embodiments. Fig. Figures 49A-49D show example scenarios in which one vehicle follows another vehicle, in accordance with the disclosed embodiments. Fig. Figure 50 shows a flowchart illustrating an exemplary process for decelerating a host vehicle in accordance with the disclosed embodiments. Fig. Figure 51 shows a flowchart illustrating an exemplary process for navigating a host vehicle in accordance with the disclosed embodiments. Fig. Figures 52A-52D show exemplary approximation buffers for a host vehicle, in accordance with the disclosed embodiments. Fig. 53A and Fig. Figure 53B shows example scenarios including an approximation buffer, in accordance with the disclosed embodiments. Fig. 54A and Fig. Figure 54B shows example scenarios including an approximation buffer, in accordance with the disclosed embodiments. Fig. Figure 55 shows a flowchart for selectively displacing the control of a host vehicle by a human driver, in accordance with the disclosed embodiments. Fig. Figure 56 shows a flowchart illustrating an exemplary process for navigating a host vehicle in accordance with the disclosed embodiments. Fig. Figures 57A-57C show example scenarios in accordance with the disclosed embodiments. Fig. Figure 58 shows a flowchart illustrating an exemplary process for navigating a host vehicle in accordance with the disclosed embodiments. Fig. Figure 59 shows an environment illustrating how the viewing direction can affect path prediction, according to some embodiments. Fig. Figure 60 shows a flowchart illustrating a procedure for predicting a path and steering a vehicle based on the direction of view. Fig. Figure 61 illustrates a scene showing how the view of a driver is captured by an inward-facing camera, according to one embodiment. Fig. Figure 62 shows a model of a Purkinje image according to some embodiments. Fig. Figures 63A-C show the detection of the head and eye position for classifying the driver's gaze according to one embodiment. Fig. Figure 64 shows object detection in an outward-facing camera, according to one embodiment. Fig. Figure 65 shows a scene captured by outward-facing sensors and a heat map showing the driver's direction of view as captured by inward-facing sensors, according to some embodiments, according to one embodiment. Fig. Figure 66 shows a flowchart illustrating a procedure for predicting a path and steering a vehicle based on the direction of gaze. Fig. 67 Flowchart illustrating a procedure for predicting a path and steering a vehicle based on the direction of gaze. Fig. Figure 68 is a block diagram of an ADAS according to some embodiments. Fig. Figure 69 shows a computer system according to exemplary embodiments of the present disclosure. DETAILED DESCRIPTION

[0012] The following detailed description refers to the accompanying drawings. Where possible, the same reference numbers are used in the drawings and the following description to refer to identical or similar parts. While several illustrative embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the components illustrated in the drawings, and the illustrative methods described herein may be modified by replacing, rearranging, removing, or adding steps to the disclosed methods. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the correct scope of protection is defined by the accompanying claims. Overview of autonomous vehicles

[0013] As used throughout this disclosure, the term “autonomous vehicle” refers to a vehicle capable of implementing at least one navigation change without driver input. A “navigation change” refers to a change in one or more of the vehicle’s steering, braking, or acceleration / deceleration. To be autonomous, a vehicle need not be fully automatic (e.g., operating completely without a driver or without driver input). Rather, autonomous vehicles include those capable of operating under driver control during certain periods and without driver control during other periods. Autonomous vehicles may also include vehicles that control only some aspects of vehicle navigation, such as steering (e.g.,Autonomous vehicles can, for example, maintain a vehicle course between lane restrictions or perform some steering operations under certain circumstances (but not all circumstances), but may leave other aspects to the driver (e.g., braking or stopping under certain conditions). In some cases, autonomous vehicles can handle some or all aspects of braking, speed control, and / or steering.

[0014] Since human drivers typically rely on visual cues and observations to control a vehicle, traffic infrastructures are designed accordingly, with lane markings, traffic signs, and traffic lights all intended to provide visual information to the driver. Given these design characteristics of traffic infrastructures, an autonomous vehicle can include a camera and a processing unit that analyzes visual information captured from the vehicle's surroundings. This visual information can include, for example, components of the traffic infrastructure (e.g., lane markings, traffic signs, traffic lights, etc.) that can be observed by drivers, as well as other obstacles (e.g., other vehicles, pedestrians, debris, etc.).Additionally, an autonomous vehicle can also use stored information, such as information provided by a model of the vehicle's surroundings when navigating. For example, the vehicle can use GPS data, sensor data (e.g., from an accelerometer, speed sensor, suspension sensor, etc.), and / or other map data to provide information about its environment while driving, and the vehicle (as well as other vehicles) can use this information to locate itself on the model. Some vehicles are also capable of communicating with each other, exchanging information, warning oncoming vehicles of hazards or changes in the vehicles' surroundings, and so on. System overview

[0015] Fig. Figure 1 is a block diagram representation of a system 100, in accordance with the exemplary disclosed embodiments. The system 100 may include various components, depending on the requirements of a particular implementation. In some embodiments, the system 100 may include a processing unit 110, an image acquisition unit 120, a position sensor 130, one or more storage units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. The processing unit 110 may include one or more processing devices. In some embodiments, the processing unit 110 may include an application processor 180, an image processor 190, or another suitable processing device.Similarly, the image acquisition unit 120 can include any number of image acquisition devices and components, depending on the requirements of a particular application. In some embodiments, the image acquisition unit 120 can include one or more image acquisition devices (e.g., cameras, CCDs, or any other type of image sensor), such as the image acquisition device 122, the image acquisition device 124, and the image acquisition device 126. The system 100 can also include a data interface 128 that provides communication between the processing device 110 and the image acquisition unit 120. For example, the data interface 128 can include one or more wired and / or wireless connections for transmitting the image data acquired by the image acquisition device 120 to the processing unit 110.

[0016] The Wireless Transceiver 172 can include one or more devices configured to exchange transmissions over an air interface with one or more networks (e.g., cellular, internet, etc.) using a radio frequency, an infrared frequency, a magnetic field, or an electric field. The Wireless Transceiver 172 can use any known standard for transmitting and / or receiving data (e.g., Wi-Fi, Bluetooth®, Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions can include communications from the host vehicle to one or more remote servers. Such transmissions can also include communications (one-way or two-way) between the host vehicle and one or more target vehicles in the vicinity of the host vehicle (e.g.,, to facilitate the coordination of the host vehicle's navigation with respect to or together with target vehicles in the vicinity of the host vehicle) or even include a broadcast transmission to unspecified receivers in an environment of the transmitting vehicle.

[0017] Both the Application Processor 180 and the Image Processor 190 can include various types of processing devices. For example, one or both of the Application Processor 180 and the Image Processor 190 can include a microprocessor, preprocessors (such as an image preprocessor), graphics processors, a central processing unit (CPU), support circuitry, digital signal processors, integrated circuits, memory, or any other type of device suitable for running applications and for image processing and analysis. In some embodiments, the Application Processor 180 and / or the Image Processor 190 can include any type of single- or multi-core processor, mobile device microcontroller, central processing unit, etc. Various processing devices can be used, for example, processors from manufacturers such as Intel®, AMD®, etc., which can have different architectures (e.g. x86 processor, ARM®, etc.).

[0018] In some embodiments, the Application Processor 180 and / or the Image Processor 190 can include any of the EyeQ series of processor chips available from Mobileye®. These processor configurations each include multiple processing units with local memory and instruction sets. Such processors can include video inputs for receiving image data from multiple image sensors and can also include video output capabilities. In one example, the EyeQ2® uses 90 nm micrometer technology operating at 332 MHz. The EyeQ2® architecture consists of two floating-point hyper-threaded 32-bit RISC CPUs (MIPS32® 34K® cores), five Vision Computing Engines (VCE), three Vector Microcode Processors (VMP®), Denali 64-bit Mobile DDR control, 128-bit internal Sonics Interconnect, dual 16-bit video input and 18-bit video output controls, 16-channel DMA, and multiple peripherals.The MIPS34K CPU manages the five VCEs, three VMP™ and the DMA, the second MIPS34K CPU and the multi-channel DMA, as well as the other peripherals. The five VCEs, three VMP®, and the MIPS34K CPU can perform intensive vision calculations required by multi-function bundle applications. In another example, the EyeQ3®, a third-generation processor six times more powerful than the EyeQ2®, can be used in the disclosed embodiments. In other examples, the EyeQ4® and / or EyeQ5® can be used in the disclosed embodiments. Of course, any newer or future EyeQ processing devices can also be used in conjunction with the disclosed embodiments.

[0019] Any of the processing devices disclosed herein can be configured to perform specific functions. Configuring a processing device, such as any of the described EyeQ processors or other controller or microprocessor, to perform specific functions can include programming computer-executable instructions and making these instructions available to the processing device for execution during its operation. In some embodiments, configuring a processing device can include programming the processing device directly with architectural instructions. In other embodiments, configuring a processing device can include storing executable instructions in memory accessible to the processing device during operation.For example, the processing device can access memory to retrieve and execute stored instructions during operation. In any case, the processing device, configured to perform the sensing, image analysis, and / or navigation functions disclosed herein, constitutes a specialized hardware-based system that controls multiple hardware-based components of a host vehicle.

[0020] Although Fig. While Figure 1 represents two separate processing devices enclosed within the processing unit 110, more or fewer processing devices can also be used. For example, in some embodiments, a single processing device can be used to perform the tasks of the application processor 180 and the image processor 190. In other embodiments, these tasks can be performed by more than two processing devices. Furthermore, in some embodiments, the system 100 can include one or more processing units 110 without including other components, such as the image acquisition unit 120.

[0021] The processing unit 110 can comprise various types of devices. For example, the processing unit 110 can include various devices such as a controller, an image preprocessor, a central processing unit (CPU), support circuits, digital signal processors, integrated circuits, memory, or any other type of device for image processing and analysis. The image preprocessor can include a video processor for capturing, digitizing, and processing the image data from the image sensors. The CPU can comprise any number of microcontrollers or microprocessors. The support circuits can be any number of circuits that are generally well-known in the prior art, including cache, power supply, clock, and input / output circuits.Memory can store software that, when executed by the processor, controls the operation of the system. Memory can include databases and image processing software. Memory can comprise any number of random-access storage devices, read-only storage devices, flash memory, disk drives, optical storage devices, tape storage devices, removable storage devices, and other types of storage. In one case, memory can be separate from the Processing Unit 110. In another case, memory can be integrated into the Processing Unit 110.

[0022] Each memory unit 140, 150 can contain software instructions which, when executed by a processor (e.g., application processor 180 and / or image processor 190), can control the operation of various aspects of the system 100. These memory units can include various databases and image processing software, as well as a trained system, such as a neural network or a deep neural network. The memory units can include random-access memory, read-only memory, flash memory, disk drives, optical storage, tape storage, removable storage, and / or any other type of storage. In some embodiments, the memory units 140, 150 can be separate from the application processor 180 and / or image processor 190. In other embodiments, these memory units can be integrated into the application processor 180 and / or the image processor 190.

[0023] The position sensor 130 can include any type of device suitable for determining a location associated with at least one component of the system 100. In some embodiments, the position sensor 130 can include a GPS receiver. Such receivers can determine a user's position and speed by processing signals broadcast by satellites of the global positioning system. Position information from the position sensor 130 can be made available to the application processor 180 and / or image processor 190.

[0024] In some embodiments, the system 100 may include components such as a speed sensor (e.g., a speedometer) for measuring the speed of the vehicle 200. The system 100 may also include one or more accelerometers (either single-axis or multi-axis) for measuring the accelerations of the vehicle 200 along one or more axes.

[0025] The storage units 140 and 150 can include a database or otherwise organized data indicating the location of known landmarks. Sensory information (such as images, radar signals, depth information from lidar, or stereo processing of two or more images) of the environment can be processed together with positional information, such as a GPS coordinate, the vehicle's own movement, etc., to determine the vehicle's current location relative to the known landmarks and to refine the vehicle's location. Certain aspects of this technology are incorporated into a localization technology known as REM™, which is marketed by the applicant of the present application.

[0026] The user interface 170 can include any device suitable for providing information to or receiving input from one or more users of the system 100. In some embodiments, the user interface 170 can include user input devices such as a touchscreen, microphone, keyboard, pointing devices, trackwheels, cameras, knobs, buttons, etc. With such input devices, a user can provide information inputs or commands to the system 100 by typing instructions or information, providing voice commands, selecting menu options on a screen using buttons, pointers, or eye-tracking capabilities, or by any other suitable techniques for communicating information to the system 100.

[0027] The user interface 170 can be equipped with one or more processing devices configured to provide and receive information to and from a user and to process this information for use by, for example, the application processor 180. In some embodiments, such processing devices can execute instructions for detecting and tracking eye movements, receiving and interpreting speech commands, detecting and interpreting touches and / or gestures on a touchscreen, responding to keyboard input or menu selections, etc. In some embodiments, the user interface 170 can include a display, a speaker, a tactile device, and / or any other devices for providing output information to a user.

[0028] The map database 160 can include any type of database for storing map data useful to the system 100. In some embodiments, the map database 160 can include data relating to the position of various elements in a reference coordinate system, including roads, water features, geographic features, businesses, landmarks, restaurants, gas stations, etc. The map database 160 can store not only the locations of such elements but also descriptors relating to those elements, including, for example, names associated with the stored features. In some embodiments, the map database 160 can be physically located at the same site as other components of the system 100. Alternatively or additionally, the map database 160, or a part thereof, can be located in relation to other components of the system 100 (e.g.,The processing unit 110) may be located remotely. In such embodiments, information from the map database 160 can be downloaded to a network (e.g., via a cellular network and / or the internet, etc.) via a wired or wireless data connection. In some cases, the map database 160 may store a sparse data model that includes polynomial representations of certain road features (e.g., lane markings) or target trajectories for the host vehicle. The map database 160 may also include stored representations of various detected landmarks that can be used to determine or update a known position of the host vehicle with respect to a target trajectory. The landmark representations may include data fields such as the landmark type and the landmark location, among other potential identifiers.

[0029] The image acquisition devices 122, 124, and 126 can each include any type of device suitable for capturing at least one image from an environment. Furthermore, any number of image acquisition devices can be used to capture images for input into the image processor. Some embodiments may include only a single image acquisition device, while other embodiments may include two, three, or even four or more image acquisition devices. The image acquisition devices 122, 124, and 126 are described below with reference to Fig. 2B-E described.

[0030] One or more cameras (e.g., image acquisition devices 122, 124, and 126) can be part of a sensor acquisition block on a vehicle. Various other sensors can be included in the sensor acquisition block, and any or all of these sensors can be used to develop a sensor-aware navigation state of the vehicle. In addition to cameras (forward, sideways, rearward, etc.), other sensors, such as radar, lidar, and acoustic sensors, can also be included in the sensor acquisition block. Additionally, the sensor acquisition block can include one or more components configured to transmit / receive information relating to the vehicle's environment. For example, such components can be wireless transceivers (RF, etc.).This includes cameras that can receive sensor-based information or any other type of information relating to the host vehicle's environment from a source located remotely from the host vehicle. Such information may include sensor output information or related information received from vehicle systems other than the host vehicle. In some embodiments, such information may include information received from a remote computing device, a central server, etc. Furthermore, the cameras can adopt many different configurations: single camera units, multiple cameras, camera clusters, long FOV, short FOV, wide-angle, fisheye, etc.

[0031] System 100, or various components thereof, can be integrated into different platforms. In some embodiments, System 100 can be enclosed within a Vehicle 200, as in Fig. 2A shown. For example, the vehicle 200 can be equipped with a processing unit 110 and any of the other components of the system 100, as above in relation to Fig. 1 described. While in some embodiments the vehicle 200 may only be equipped with a single image recording device (e.g. a camera), in other embodiments, such as in conjunction with Fig. Sections 2B-2E discuss the use of multiple image acquisition devices. For example, each of the image acquisition devices 122 and 124 of vehicle 200, as described in Fig. 2A shown, is part of an ADAS (Advanced Driver Assistance Systems) imaging set.

[0032] The image acquisition devices enclosed in the vehicle 200 as part of the image acquisition unit 120 can be positioned at any suitable location. In some embodiments, as in Fig. As shown in Figures 2A-2E and 3A-3C, the image capture device 122 can be located near the rearview mirror. This position can provide a line of sight similar to that of the driver of vehicle 200, which can help determine what is visible and not visible to the driver. The image capture device 122 can be positioned anywhere near the rearview mirror, but placing the image capture device 122 on the driver's side of the mirror can further help in obtaining images that are representative of the driver's field of vision and / or line of sight.

[0033] Other locations for the image capture devices of the image acquisition unit 120 can also be used. For example, the image capture device 124 can be located on or in a bumper of the vehicle 200. Such a location can be particularly suitable for image capture devices with a wide field of view. The line of sight of the image capture devices located on the bumper may differ from that of the driver, and therefore the bumper image capture device and the driver may not always see the same objects. The image capture devices (e.g., image capture devices 122, 124, and 126) can also be located in other positions.For example, the image recording devices may be located on or in one or both of the side mirrors of the vehicle 200, on the roof of the vehicle 200, on the hood of the vehicle 200, on the trunk of the vehicle 200 and on the sides of the vehicle 200, mounted on, positioned behind or in front of one of the windows of the vehicle 200 and mounted in or near lighting fixtures on the front and / or rear of the vehicle 200, etc.

[0034] In addition to the image acquisition devices, the vehicle 200 can include various other components of the system 100. For example, the processing unit 110 can be included in the vehicle 200, either integrated into or separate from the vehicle's engine control unit (ECU). The vehicle 200 can also be equipped with a position sensor 130, such as a GPS receiver, and can also include a map database 160 and storage units 140 and 150.

[0035] As previously discussed, the wireless transceiver 172 can transmit and / or receive data over one or more networks (e.g., mobile networks, the internet, etc.). For example, the wireless transceiver 172 can upload data collected by the system 100 to one or more servers and download data from those servers. The wireless transceiver 172 allows the system 100 to receive, for example, periodically or on-demand updated data stored in the map database 160, memory 140, and / or memory 150. Similarly, the wireless transceiver 172 can upload any data (e.g., images captured by the image acquisition unit 120, data received from the position sensor 130 or other sensors, vehicle control systems, etc.) from the system 100 and / or any data processed by the processing unit 110 to the one or more servers.

[0036] The System 100 can upload data to a server (e.g., the cloud) based on a data protection level setting. For example, the System 100 can implement data protection level settings to regulate or limit the types of data (including metadata) sent to the server that could uniquely identify a vehicle and / or a driver / owner of a vehicle. Such settings can be configured by the user, for example, via the Wireless Receiver 172, initialized through factory default settings, or through data received by the Wireless Receiver 172.

[0037] In some embodiments, the system can upload 100 data points according to a "high" privacy level, and under a specific setting, the system can transmit 100 data points (e.g., location information related to a route, captured images, etc.) without any details about the specific vehicle and / or driver / owner. For example, when data is uploaded according to a "high" privacy setting, the system cannot include the vehicle identification number (VIN) or the name of a driver or owner of the vehicle and can instead transmit data such as captured images and / or limited location information related to a route.

[0038] Other data protection levels are also considered. For example, the system can transmit 100 data points to a server according to a "medium" data protection level and include additional information not included in a "high" data protection level, such as a vehicle's make and / or model and / or vehicle type (e.g., a passenger car, an SUV, a truck, etc.). In some embodiments, the system can upload 100 data points according to a "low" data protection level. Under a "low" data protection setting, the system can upload 100 data points and include information sufficient to uniquely identify a specific vehicle, a specific owner / driver, and / or a section or the entirety of a route traveled by the vehicle.Such “low” level data protection data may include one or more of, for example, a VIN, a driver / owner's name, a vehicle's point of origin before departure, a vehicle's intended destination, a vehicle's make and / or model, a vehicle type, etc.

[0039] Fig. Figure 2A is a schematic side view representation of an exemplary vehicle imaging system, in accordance with the disclosed embodiments. Fig. 2B is a schematic top-view illustration of the in Fig. 2A embodiment shown. As in Fig. As illustrated in Figure 2B, the disclosed embodiments may include a vehicle 200 which incorporates in its body a system 100 comprising a first image acquisition device 122, which is positioned near the rearview mirror and / or near the driver of the vehicle 200, a second image acquisition device 124, which is positioned on or in a bumper area (e.g. one of the bumper areas 210) of the vehicle 200, and a processing unit 110.

[0040] As in Fig. As illustrated in Figure 2C, the image capture devices 122 and 124 can both be positioned near the rearview mirror and / or near the driver of the vehicle 200. Even if in Fig. 2B and Fig. Figure 2C shows two image acquisition devices 122 and 124; it is understood that other embodiments may include more than two image acquisition devices. In the Fig. 2D and Fig. In the embodiments shown in Figure 2E, for example, a first, a second and a third image acquisition device (122, 124 and 126) are included in the system 100 of the vehicle 200.

[0041] As in Fig. As illustrated in 2D, the image capture device 122 can be positioned near the rearview mirror and / or near the driver of the vehicle 200, and the image capture devices 124 and 126 can be positioned on or in a bumper area (e.g., one of the bumper areas 210) of the vehicle 200. And as shown in Fig. As shown in Figure 2E, the image capture devices 122, 124, and 126 can be positioned near the rearview mirror and / or near the driver's seat of the vehicle 200. The disclosed embodiments are not limited to a specific number and configuration of the image capture devices, and the image capture devices can be positioned at any suitable location inside and / or on the vehicle 200.

[0042] It is understood that the disclosed embodiments are not limited to vehicles and could be applied in other contexts. It is also understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and may be applicable to all types of vehicles, including automobiles, trucks, trailers, and other vehicle types.

[0043] The first image acquisition device 122 can include any suitable type of image acquisition device. The image acquisition device 122 can include an optical axis. In one instance, the image acquisition device 122 can include a WVGA sensor of the Aptina M9V024 type with a global shutter. In other embodiments, the image acquisition device 122 can provide a resolution of 1280×960 pixels and include a rolling shutter. The image acquisition device 122 can include various optical elements. In some embodiments, one or more lenses can be included to provide, for example, a desired focal length and field of view for the image acquisition device. In some embodiments, the image acquisition device 122 can be associated with a 6 mm lens or a 12 mm lens.In some embodiments, the image acquisition device 122 can be configured to capture images with a desired field of view (FOV) 202, as shown in . Fig. 2D illustration. For example, the image capture device 122 can be configured to have a regular FOV, such as within a range of 40 degrees to 56 degrees, including a 46-degree FOV, 50-degree FOV, 52-degree FOV, or larger. Alternatively, the image capture device 122 can be configured to have a narrow FOV in the range of 23 to 40 degrees, such as a 28-degree FOV or 36-degree FOV. Additionally, the image capture device 122 can be configured to have a wide FOV in the range of 100 to 180 degrees. In some embodiments, the image capture device 122 can include a wide-angle bumper camera or a camera with an FOV of up to 180 degrees. In some embodiments, the image acquisition device 122 can be an image acquisition device with 7.2 M pixels with an aspect ratio of about 2:1 (e.g. HxV = 3800x1900 pixels) with a horizontal FOV of about 100 degrees.Such an image capture device can be used instead of a configuration with three image capture devices. Due to significant lens distortion, the vertical field of view (FOV) of such an image capture device can be significantly less than 50 degrees in implementations where the image capture device uses a radially symmetrical lens. For example, such a lens may not be radially symmetrical, which would allow a vertical FOV greater than 50 degrees with a horizontal FOV of 100 degrees.

[0044] The first image acquisition device 122 can capture a plurality of first images relative to a scene associated with the vehicle 200. Each plurality of first images can be captured as a series of image scan lines, which can be recorded using a rolling shutter. Each scan line can encompass a plurality of pixels.

[0045] The first image acquisition device 122 can have a sampling rate associated with the acquisition of each of the first row of image scan lines. The sampling rate can refer to the rate at which an image sensor can acquire image data associated with each pixel enclosed in a given scan line.

[0046] The image acquisition devices 122, 124, and 126 can include any suitable type and number of image sensors, including, for example, CCD or CMOS sensors. In one embodiment, a CMOS image sensor can be used together with a rolling shutter, such that each pixel in a row is read sequentially, and the scanning of the rows is performed on a row-by-row basis until an entire single image has been acquired. In some embodiments, the rows can be acquired sequentially from top to bottom with respect to the single image.

[0047] In some embodiments, one or more of the image acquisition devices disclosed herein (e.g. image acquisition devices 122, 124 and 126) can form a high-resolution image transmitter and have a resolution of more than 5 M pixels, 7 M pixels, 10 M pixels or more.

[0048] The use of a rolling shutter can cause pixels in different rows to be exposed and captured at different times, which can lead to distortion and other image artifacts in the captured frame. Conversely, if the image capture device 122 is configured to operate with a global or synchronous shutter, all pixels can be exposed for the same amount of time and during a common exposure period. As a result, the image data in a frame collected by a system using a global shutter represents a snapshot of the entire field of view (such as FOV 202) at a specific point in time. In contrast, with a rolling shutter application, each row in a frame is exposed, and the data is captured at different times. Therefore, moving objects may appear distorted in an image capture device using a rolling shutter.This phenomenon is described in more detail below.

[0049] The second image acquisition device 124 and the third image acquisition device 126 can be any type of image acquisition device. Like the first image acquisition device 122, each of the image acquisition devices 124 and 126 can include an optical axis. In one embodiment, each of the image acquisition devices 124 and 126 can include an Aptina M9V024 WVGA sensor with a global shutter. Alternatively, each of the image acquisition devices 124 and 126 can include a rolling shutter. Like the image acquisition device 122, the image acquisition devices 124 and 126 can be configured to include various lenses and optical elements. In some embodiments, the lenses associated with the image acquisition devices 124 and 126 can provide FOVs (such as FOVs 204 and 206) that are equal to or narrower than an FOV (such as FOV 202) associated with the image acquisition device 122.For example, the image capture devices 124 and 126 can have FOVs of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees or less.

[0050] The image acquisition devices 124 and 126 can capture a plurality of second and third images relative to a scene associated with vehicle 200. Each plurality of second and third images can be captured as a second and third row of image scan lines, which can be captured using a rolling shutter. Each scan line or row can have a plurality of pixels. The image acquisition devices 124 and 126 can have second and third sampling rates associated with capturing each of the image scan lines enclosed in the second and third rows.

[0051] Each image acquisition device 122, 124, and 126 can be positioned at any suitable location and in any suitable orientation relative to the vehicle 200. The relative positioning of the image acquisition devices 122, 124, and 126 can be selected to aid in combining the information captured by the image acquisition devices. For example, in some embodiments, an FOV (such as FOV 204) associated with image acquisition device 124 can partially or completely overlap with an FOV (such as FOV 202) associated with image acquisition device 122 and an FOV (such as FOV 206) associated with image acquisition device 126.

[0052] The image acquisition devices 122, 124, and 126 can be located on the vehicle 200 at any suitable relative heights. In one case, there can be a height difference between the image acquisition devices 122, 124, and 126 that can provide sufficient parallax information to enable stereo analysis. For example, as in Fig. As shown in Figure 2A, the two image acquisition devices 122 and 124 are located at different heights. There can also be a lateral displacement difference between the image acquisition devices 122, 124, and 126, which, for example, provides additional parallax information for stereo analysis by the processing unit 110. The difference in lateral displacement can be represented by d x are referred to as in Fig. 2C and Fig. 2D shown. In some embodiments, a forward or backward displacement (e.g., distance displacement) may exist between the image acquisition devices 122, 124, and 126. For example, the image acquisition device 122 may be located 0.5 to 2 meters or more behind the image acquisition device 124 and / or the image acquisition device 126. This type of displacement may allow one of the image acquisition devices to cover potential blind spots of the other image acquisition device(s).

[0053] The image acquisition devices 122 can have any suitable resolution (e.g., number of pixels associated with the image sensor), and the resolution of the image sensor(s) associated with the image acquisition device 122 can be higher, lower, or equal to the resolution of the image sensor(s) associated with the image acquisition devices 124 and 126. In some embodiments, the image sensor(s) associated with the image acquisition device 122 and / or the image acquisition devices 124 and 126 can have a resolution of 640 x 480, 1024 x 768, 1280 x 960, or any other suitable resolution.

[0054] The frame rate (e.g., the rate at which an image capture device acquires a set of pixel data from a single frame before moving to acquire pixel data associated with the next frame) can be controllable. The frame rate associated with image capture device 122 can be higher, lower, or equal to the frame rate associated with image capture devices 124 and 126. The frame rate associated with image capture devices 122, 124, and 126 can depend on a variety of factors that can affect the timing of the frame rate. For example, one or more of the image capture devices 122, 124, and 126 can include a selectable pixel delay period imposed before or after the acquisition of image data associated with one or more pixels of an image sensor in the image capture device 122, 124, and / or 126.In general, image data corresponding to each pixel can be acquired according to a clock rate for the device (e.g., one pixel per clock cycle). Additionally, in embodiments that include a rolling shutter, one or more of the image acquisition devices 122, 124, and 126 can include a selectable horizontal blanking period imposed before or after the acquisition of image data associated with a set of pixels of an image sensor in the image acquisition device 122, 124, and / or 126. Furthermore, one or more of the image acquisition devices 122, 124, and / or 126 can include a selectable vertical blanking period imposed before or after the acquisition of image data associated with a single frame of the image acquisition device 122, 124, and 126.

[0055] These timing controls can enable the synchronization of frame rates associated with the image acquisition devices 122, 124, and 126, even if the line sampling rates differ. Additionally, as discussed in more detail below, these selectable timing controls, along with other factors (e.g., image sensor resolution, maximum line sampling rates, etc.), can enable the synchronization of image acquisition from an area where the field of view (FOV) of the image acquisition device 122 overlaps with one or more FOVs of the image acquisition devices 124 and 126, even if the field of view of the image acquisition device 122 differs from the FOVs of the image acquisition devices 124 and 126.

[0056] The frame rate timing in the image acquisition device 122, 124, and 126 can depend on the resolution of the associated image sensors. For example, assuming similar line sampling rates for both devices, if one device includes an image sensor with a resolution of 640 x 480 and another device includes an image sensor with a resolution of 1280 x 960, then more time is required to acquire a single frame of image data from the sensor with the higher resolution.

[0057] Another factor that can influence the timing of image data acquisition in the image acquisition devices 122, 124, and 126 is the maximum line sampling rate. For example, acquiring a series of image data from an image sensor enclosed in the image acquisition devices 122, 124, and 126 requires a minimum time. Assuming no pixel delay periods are added, this minimum time for acquiring a series of image data refers to the maximum line sampling rate for a given device. Devices offering higher maximum line sampling rates have the potential to provide higher frame rates than devices with lower maximum line sampling rates. In some embodiments, one or more of the image acquisition devices 124 and 126 may have a maximum line sampling rate that is higher than the maximum line sampling rate associated with the image acquisition device 122.In some embodiments, the maximum line sampling rate of the image acquisition device 124 and / or 126 may be 1.25, 1.5, 1.75 or 2 times or more of a maximum line sampling rate of the image acquisition device 122.

[0058] In another embodiment, the image acquisition devices 122, 124, and 126 can have the same maximum line sampling rate, but the image acquisition device 122 can be operated at a sampling rate that is less than or equal to its maximum sampling rate. The system can be configured such that one or more of the image acquisition devices 124 and 126 operate at a line sampling rate that is equal to the line sampling rate of the image acquisition device 122. In other cases, the system can be configured such that the line sampling rate of the image acquisition device 124 and / or the image acquisition device 126 can be 1.25, 1.5, 1.75, or 2 times or more of the line sampling rate of the image acquisition device 122.

[0059] In some embodiments, the image acquisition devices 122, 124, and 126 can be asymmetrical. That is, they can include cameras with different fields of view (FOV) and focal lengths. The fields of view of the image acquisition devices 122, 124, and 126 can, for example, encompass any desired area in relation to the environment of the vehicle 200. In some embodiments, one or more of the image acquisition devices 122, 124, and 126 can be configured to acquire image data from an environment in front of the vehicle 200, behind the vehicle 200, at the sides of the vehicle 200, or combinations thereof.

[0060] Furthermore, the focal length associated with each image-capturing device 122, 124, and / or 126 can be selectable (e.g., by including suitable lenses, etc.) so that each device captures images of objects within a desired range of distances relative to the vehicle 200. For example, in some embodiments, the image-capturing devices 122, 124, and 126 can capture images of close-up objects within a few meters of the vehicle. The image-capturing devices 122, 124, and 126 can also be configured to capture images of objects at distances farther from the vehicle (e.g., 25 m, 50 m, 100 m, 150 m, or more). Furthermore, the focal lengths of the image-capturing devices 122, 124, and 126 can be selected such that one image-capturing device (e.g., image-capturing device 122) captures images of objects relatively close to the vehicle (e.g.,within 10 m or within 20 m) can take pictures, while the other image-taking devices (e.g. image-taking devices 124 and 126) can take pictures of objects further away (e.g. greater than 20 m, 50 m, 100 m, 150 m etc.) from the vehicle 200.

[0061] According to some embodiments, the field of view (FOV) of one or more image capture devices 122, 124, and 126 can have a wide angle. For example, it may be advantageous to have an FOV of 140 degrees, particularly for the image capture devices 122, 124, and 126, which can be used to capture images of the area near the vehicle 200. For example, the image capture device 122 can be used to capture images of the area to the right or left of the vehicle 200, and in such embodiments, it may be desirable for the image capture device 122 to have a wide FOV (e.g., at least 140 degrees).

[0062] The field of view associated with each of the image-taking devices 122, 124 and 126 can depend on the respective focal lengths. For example, as the focal length increases, the corresponding field of view decreases.

[0063] The image acquisition devices 122, 124, and 126 can be configured to have any suitable fields of view. In one particular example, the image acquisition device 122 can have a horizontal FOV of 46 degrees, the image acquisition device 124 can have a horizontal FOV of 23 degrees, and the image acquisition device 126 can have a horizontal FOV between 23 and 46 degrees. In another case, the image acquisition device 122 can have a horizontal FOV of 52 degrees, the image acquisition device 124 can have a horizontal FOV of 26 degrees, and the image acquisition device 126 can have a horizontal FOV between 26 and 52 degrees. In some embodiments, the ratio of the FOV of the image acquisition device 122 to the FOVs of the image acquisition devices 124 and / or 126 can vary from 1.5 to 2.0. In other embodiments, this ratio can vary between 1.25 and 2.25.

[0064] System 100 can be configured such that a field of view of the image acquisition device 122 overlaps at least partially or completely with a field of view of the image acquisition device 124 and / or the image acquisition device 126. In some embodiments, system 100 can be configured such that the fields of view of the image acquisition devices 124 and 126, for example, fall within the field of view of the image acquisition device 122 (e.g., are narrower than it) and share a common center with it. In other embodiments, the image acquisition devices 122, 124, and 126 can capture adjacent fields of view or have a partial overlap in their fields of view. In some embodiments, the fields of view of the image acquisition devices 122, 124 and 126 can be aligned such that a center of the narrower FOV image acquisition devices 124 and / or 126 can be located in a lower half of the field of view of the wider FOV device 122.

[0065] Fig. Figure 2F is a schematic representation of exemplary vehicle control systems, in accordance with the disclosed embodiments. As in Fig. As specified in 2F, the vehicle 200 can include a throttle system 220, a brake system 230, and a steering system 240. The system 100 can provide inputs (e.g., control signals) to one or more of a throttle system 220, a brake system 230, and a steering system 240 via one or more data connections (e.g., any wired and / or wireless connection or data transmission connections). For example, based on an analysis of images captured by the image acquisition devices 122, 124, and / or 126, the system 100 can provide control signals to one or more of a throttle system 220, a brake system 230, and a steering system 240 to navigate the vehicle 200 (e.g., by initiating acceleration, turning, lane changes, etc.).Furthermore, the system can receive 100 inputs from one or more of a throttle system 220, a brake system 230, and a steering system 24, indicating the operating conditions of the vehicle 200 (e.g., speed, whether the vehicle 200 is braking and / or turning, etc.). Further details will be provided in conjunction with [reference to be added]. Fig. 4-7 provided below.

[0066] As in Fig. As shown in Figure 3A, the vehicle 200 can also include a user interface 170 for interacting with a driver or passenger of the vehicle 200. For example, in a vehicle application, the user interface 170 can include a touchscreen 320, knobs 330, buttons 340, and a microphone 350. A driver or passenger of the vehicle 200 can also use handles (located, for example, on or near the steering column of the vehicle 200, including, for example, turn signal handles), buttons (located, for example, on the steering wheel of the vehicle 200), and the like to interact with the system 100. In some embodiments, the microphone 350 can be positioned next to a rearview mirror 310. Similarly, in some embodiments, the image capture device 122 can be located near the rearview mirror 310. In some embodiments, the user interface 170 can also include one or more loudspeakers 360 (e.g.,This includes the speakers of a vehicle audio system. For example, the system can provide 100 different notifications (e.g., warnings) via the 360-degree speakers.

[0067] Fig. Figures 3B-3D are illustrations of an exemplary camera mount 370 configured to be positioned behind a rearview mirror (e.g., rearview mirror 310) and against a vehicle windshield, in accordance with the disclosed embodiments. As shown in Fig. As shown in Figure 3B, the camera mount 370 can include image capture devices 122, 124, and 126. The image capture devices 124 and 126 can be positioned behind a glare shield 380, which can be flush with the vehicle windshield and can include a composition of film and / or antireflective materials. For example, the glare shield 380 can be positioned so that it aligns with a vehicle windshield that has a matching slope. In some embodiments, each of the image capture devices 122, 124, and 126 can be positioned behind the glare shield 380, as shown, for example, in Fig. 3D representation. The disclosed embodiments are not limited to a specific configuration of the image acquisition devices 122, 124 and 126, the camera mount 370 and the glare shield 380. Fig. 3C is an illustration of the 370 camera mount, which is in Fig. 3B is shown from a front perspective.

[0068] As a person skilled in the art, benefiting from this disclosure, will recognize, numerous variations and / or modifications can be made to the embodiments disclosed above. For example, not all components are essential for the operation of the system 100. Furthermore, each component can be located at any suitable position within the system 100, and the components can be rearranged in a multitude of configurations while still providing the functionality of the disclosed embodiments. Therefore, the configurations given above are examples, and regardless of the configurations discussed above, the system 100 can provide a wide range of functionality to analyze the environment of the vehicle 200 and to navigate the vehicle 200 in response to the analysis.

[0069] As discussed in more detail below and in accordance with various disclosed embodiments, the system 100 can provide a variety of features relating to autonomous driving and / or driver assistance technology. For example, the system 100 can analyze image data, position data (e.g., GPS location information), map data, speed data, and / or data from sensors included in the vehicle 200. The system 100 can collect the data for analysis from, for example, the image acquisition unit 120, the position sensor 130, and other sensors. Furthermore, the system 100 can analyze the collected data to determine whether or not the vehicle 200 should perform a specific action and then automatically execute the specific action without human intervention.For example, if the vehicle 200 is navigating without human intervention, the system 100 can automatically control the braking, acceleration, and / or steering of the vehicle 200 (e.g., by sending control signals to one or more of a throttle system 220, a braking system 230, and a steering system 240). Furthermore, the system 100 can analyze the collected data and, based on this analysis, issue warnings and / or alerts to the vehicle occupants. Additional details regarding the various embodiments provided by the system 100 are given below. Forward-facing multi-imaging system

[0070] As discussed above, the System 100 can provide driver assistance functionality that uses a multi-camera system. The multi-camera system can use one or more cameras facing forward in the direction of a vehicle. In other embodiments, the multi-camera system can include one or more cameras facing to the side or rear of a vehicle. In one embodiment, for example, the System 100 can use a two-camera imaging system in which a first camera and a second camera (e.g., image capture devices 122 and 124) can be positioned at the front and / or sides of a vehicle (e.g., vehicle 200). Other camera configurations are consistent with the disclosed embodiments, and the configurations described herein are examples. For example, the System 100 can include a configuration with any number of cameras (e.g.,(one, two, three, four, five, six, seven, eight, etc.). Furthermore, the system can include 100 "clusters" of cameras. For example, a cluster of cameras (including a suitable number of cameras, e.g., one, four, eight, etc.) can be oriented forward or in any other direction relative to a vehicle (e.g., towards the vehicle, to the side, at an angle, etc.). Accordingly, the system can include 100 multiple clusters of cameras, with each cluster oriented in a specific direction to capture images from a particular area of ​​the vehicle's surroundings.

[0071] The first camera can have a field of view that is larger than, smaller than, or partially overlapping with the field of view of the second camera. Additionally, the first camera can be connected to a first image processor to perform monocular image analysis of images provided by the first camera, and the second camera can be connected to a second image processor to perform monocular image analysis of images provided by the second camera. The outputs (e.g., processed information) of the first and second image processors can be combined. In some embodiments, the second image processor can receive images from both the first and second cameras to perform stereo analysis. In another embodiment, the system 100 can use a three-camera imaging system in which each camera has a different field of view.Such a system can therefore make decisions based on information from objects located at varying distances both in front of and to the sides of the vehicle. References to monocular image analysis may refer to cases where image analysis is performed based on images captured from a single viewpoint (e.g., from a single camera). Stereo image analysis may refer to cases where image analysis is performed based on two or more images captured with one or more variations of an image capture parameter. For example, captured images suitable for performing stereo image analysis may include images captured from two or more different positions, from different fields of view, using different focal lengths, along with parallax information, and so on.

[0072] For example, in one embodiment, the system 100 can implement a three-camera configuration using the image capture devices 122-126. In such a configuration, the image capture device 122 can provide a narrow field of view (e.g., 34 degrees or other values ​​selected from a range of about 20 to 45 degrees, etc.), the image capture device 124 can provide a wide field of view (e.g., 150 degrees or other values ​​selected from a range of about 100 to about 180 degrees), and the image capture device 126 can provide a medium field of view (e.g., 46 degrees or other values ​​selected from a range of about 35 to about 60 degrees). In some embodiments, the image capture device 126 can function as the main or primary camera. The image capture devices 122-126 can be positioned behind the rearview mirror 310 and essentially side by side (e.g.,(6 cm apart). Furthermore, in some embodiments, as discussed above, one or more of the image-capturing devices 122-126 can be mounted behind the glare shield 380, which is flush with the windshield of the vehicle 200. Such shielding can serve to minimize the influence of reflections from inside the car on the image-capturing devices 122-126.

[0073] In another embodiment, as above in conjunction with Fig. 3B and Fig. As discussed in Section 3C, the wide-field camera (e.g., image acquisition device 124 in the preceding example) can be mounted lower than the narrow- and wide-field cameras (e.g., image acquisition devices 122 and 126 in the preceding example). This configuration can provide an unobstructed line of sight from the wide-field camera. To reduce reflections, the cameras can be mounted near the windshield of the vehicle 200 and include polarizers on the cameras to attenuate reflected light.

[0074] A three-camera system can provide certain performance characteristics. For example, some embodiments may include the ability to validate object detection by one camera based on the detection results of another camera. In the three-camera configuration discussed above, the processing unit 110 may, for example, include three processing devices (e.g., three EyeQ series processor chips, as discussed above), each processing device dedicated to processing images acquired by one or more of the image acquisition devices 122-126.

[0075] In a three-camera system, a first processing unit can receive images from both the main camera and the narrow-field-of-view camera and perform image processing from the narrow-field-of-view camera to detect, for example, other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the first processing unit can calculate pixel disparities between the images from the main camera and the narrow-field camera and create a 3D reconstruction of the vehicle's surroundings. The first processing unit can then combine the 3D reconstruction with 3D-

[0076] Combine data or 3D information calculated based on information from another camera.

[0077] The second processing unit can receive images from the main camera and perform image processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the second processing unit can calculate camera displacement and, based on this displacement, determine pixel disparity between successive images and create a 3D reconstruction of the scene (e.g., a structure from motion). The second processing unit can then send the structure from this motion-based 3D reconstruction to the first processing unit for merging with the stereo 3D images.

[0078] The third processing unit can receive images from the wide-field-of-view (FOV) camera and process them to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing unit can also execute additional processing instructions to analyze images and identify objects moving within the frame, such as vehicles changing lanes, pedestrians, etc.

[0079] In some embodiments, streams of image-based information that are independently acquired and processed can provide a means of providing redundancy in the system. Such redundancy can, for example, include using a first image acquisition device and the images processed by that device to validate and / or supplement information obtained by acquiring and processing image information from at least one second image acquisition device.

[0080] In some embodiments, the system 100 can use two image acquisition devices (e.g., image acquisition devices 122 and 124) to provide navigational support for the vehicle 200 and a third image acquisition device (e.g., image acquisition device 126) to provide redundancy and validate the analysis of the data received by the other two image acquisition devices. For example, in such a configuration, image acquisition devices 122 and 124 can provide images for stereo analysis by the system 100 for navigating the vehicle 200, while image acquisition device 126 can provide images for monocular analysis by the system 100 to provide redundancy and validation of information obtained based on images acquired by image acquisition device 122 and / or image acquisition device 124.This means that the image acquisition device 126 (and a corresponding processing device) can be considered a redundant subsystem that provides verification of the analysis derived from the image acquisition devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Furthermore, in some embodiments, the redundancy and validation of the received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors, information received from one or more transceivers outside a vehicle, etc.).

[0081] A person skilled in the art will recognize that the foregoing camera configurations, camera placements, number of cameras, camera positions, etc., are merely examples. These components and others described in relation to the overall system can be combined and used in a multitude of different configurations without deviating from the scope of protection of the disclosed embodiments. Further details regarding the use of a multi-camera system to provide driver assistance and / or autonomous vehicle functionality follow.

[0082] Fig. Figure 4 is an exemplary functional block diagram of memory 140 and / or 150, which can be stored / programmed with instructions for performing one or more operations, in accordance with the disclosed embodiments. Although the following refers to memory 140, a person skilled in the art will recognize that instructions can be stored in memory 140 and / or 150.

[0083] As in Fig. As shown in Figure 4, the memory 140 can store a monocular image analysis module 402, a stereo image analysis module 404, a speed and acceleration module 406, and a navigation response module 408. The disclosed embodiments are not limited to a specific configuration of the memory 140. Furthermore, the application processor 180 and / or the image processor 190 can execute the instructions stored in any one of the modules 402-408 included in the memory 140. A person skilled in the art will understand that references in the following discussion to the processing unit 110 may refer individually or jointly to the application processor 180 and the image processor 190. Accordingly, steps of any of the following processes can be performed by one or more processing units.

[0084] In one embodiment, the monocular image analysis module 402 can store instructions (such as computer vision software) which, when executed by the processing unit 110, perform monocular image analysis of a set of images captured by one of the image acquisition devices 122, 124, and 126. In some embodiments, the processing unit 110 can combine information from a set of images with additional sensory information (e.g., information from radar) to perform the monocular image analysis. As described in conjunction with Fig. As described below in sections 5A-5D, the monocular image analysis module 402 can include instructions for recognizing a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with a vehicle's environment. Based on the analysis, the system 100 (e.g., via the processing unit 110) can initiate one or more navigation responses in the vehicle 200, such as a turning maneuver, a lane change, a change in acceleration, and the like, as discussed below in conjunction with the navigation response module 408.

[0085] In one embodiment, the monocular image analysis module 402 can store instructions (such as computer vision software) which, when executed by the processing unit 110, perform monocular image analysis of a set of images captured by one of the image acquisition devices 122, 124, and 126. In some embodiments, the processing unit 110 can combine information from a set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform the monocular image analysis. As described in conjunction with Fig. As described below in sections 5A-5D, the monocular image analysis module 402 can include instructions for recognizing a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with a vehicle's environment. Based on the analysis, the system 100 (e.g., via the processing unit 110) can initiate one or more navigation responses in the vehicle 200, such as turning, changing lanes, altering acceleration, and the like, as discussed below in connection with determining a navigation response.

[0086] In one embodiment, the stereo image analysis module 404 can store instructions (such as computer vision software) which, when executed by the processing unit 110, perform a stereo image analysis of a first and second set of images acquired by a combination of image capture devices selected from any one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 can combine information from the first and second sets of images with additional sensory information (e.g., information from radar) to perform the stereo image analysis. For example, the stereo image analysis module 404 can include instructions for performing a stereo image analysis based on a first set of images acquired by the image capture device 124 and a second set of images acquired by the image capture device 126.As in connection with the following . Fig. As described in Section 6, the stereo image analysis module 404 can include instructions for recognizing a set of features within the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and the like. Based on the analysis, the processing unit 110 can initiate one or more navigation responses in the vehicle 200, such as a turning maneuver, a lane change, a change in acceleration, and the like, as discussed below in connection with the navigation response module 408. Furthermore, in some embodiments, the stereo image analysis module 404 can implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.

[0087] In one embodiment, the speed and acceleration module 406 can store software configured to analyze data received from one or more computing devices and electromechanical devices in the vehicle 200, which are configured to cause a change in the speed and / or acceleration of the vehicle 200. For example, the processing unit 110 can execute instructions associated with the speed and acceleration module 406 to calculate a target speed for the vehicle 200 based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404.Such data can include, for example, a target position, speed, and / or acceleration; the position and / or speed of the vehicle 200 relative to a nearby vehicle, pedestrian, or road object; positional information for the vehicle 200 relative to lane markings; and the like. Additionally, the processing unit 110 can calculate a target speed for the vehicle 200 based on sensor inputs (e.g., information from radar) and inputs from other systems of the vehicle 200, such as the throttle system 220, the braking system 230, and / or the steering system 240.Based on the calculated target speed, the processing unit 110 can transmit electronic signals to the throttle system 220, the brake system 230 and / or the steering system 240 of the vehicle 200 to trigger a change in speed and / or acceleration, for example by physically pressing down the brake or releasing the accelerator pedal of the vehicle 200.

[0088] In one embodiment, the navigation response module 408 can store software that can be executed by the processing unit 110 to determine a desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data can include position and velocity information associated with nearby vehicles, pedestrians, and road objects, target position information for the vehicle 200, and the like.Additionally, in some embodiments, the navigation response can be based (partially or completely) on map data, a predetermined position of the vehicle 200, and / or a relative speed or acceleration between the vehicle 200 and one or more objects detected by the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 can also determine a desired navigation response based on sensor inputs (e.g., information from radar) and inputs from other systems of the vehicle 200, such as the throttle system 220, the braking system 230, and the steering system 240 of the vehicle 200.Based on the desired navigation response, the processing unit 110 can transmit electronic signals to the throttle system 220, the brake system 230, and the steering system 240 of the vehicle 200 to trigger a desired navigation response, for example, by turning the steering wheel of the vehicle 200 to achieve a rotation of a predetermined angle. In some embodiments, the processing unit 110 can use the output of the navigation response module 408 (e.g., the desired navigation response) as an input to execute the speed and acceleration module 406 to calculate a change in the speed of the vehicle 200.

[0089] Furthermore, each of the modules disclosed herein (e.g., modules 402, 404, and 406) can implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.

[0090] Fig. Figure 5A is a flowchart illustrating an exemplary process 500A for initiating one or more navigation responses based on monocular image analysis, in accordance with the disclosed embodiments. In step 510, the processing unit 110 can receive a plurality of images via the data interface 128 between the processing unit 110 and the image acquisition unit 120. For example, a camera enclosed in the image acquisition unit 120 (such as the image capture device 122 with the field of view 202) can capture a plurality of images of an area in front of the vehicle 200 (or, for example, to the sides or rear of a vehicle) and transmit them to the processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.).The processing unit 110 can execute the monocular image analysis module 402 to analyze the multitude of images at step 520, as shown below in conjunction with . Fig. 5B-5D are described in more detail. By performing the analysis, the processing unit 110 can recognize a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and the like.

[0091] The processing unit 110 can also execute the monocular image analysis module 402 to detect various road hazards at step 520, such as parts of a truck tire, fallen road signs, loose cargo, small animals, and the like. Road hazards can vary in structure, shape, size, and color, which can complicate their detection. In some embodiments, the processing unit 110 can execute the monocular image analysis module 402 to perform multi-frame analysis on the multitude of images to detect road hazards. For example, the processing unit 110 can estimate the camera movement between successive frames and calculate the pixel disparities between the frames to create a 3D map of the road.The processing unit 110 can then use the 3D map to detect the road surface as well as the hazards existing above the road surface.

[0092] At step 530, the processing unit 110 can execute the navigation response module 408 to initiate one or more navigation responses in the vehicle 200 based on the analysis performed at step 520 and the techniques described above in conjunction with Fig. 4. Navigation responses can include, for example, a turning maneuver, a lane change, a change in acceleration, and the like. In some embodiments, the processing unit 110 can use data derived from the execution of the speed and acceleration module 406 to initiate one or more navigation responses. Additionally, multiple navigation responses can occur simultaneously, sequentially, or in any combination thereof. For example, the processing unit 110 can cause the vehicle 200 to change lanes laterally and then accelerate by, for example, sequentially transmitting control signals to the steering system 240 and the throttle system 220 of the vehicle 200.Alternatively, the processing unit 110 can cause the vehicle 200 to brake while simultaneously changing lanes, for example by simultaneously transmitting control signals to the braking system 230 and the steering system 240 of the vehicle 200.

[0093] Fig. Figure 5B is a flowchart illustrating an exemplary process 500B for detecting one or more vehicles and / or pedestrians in a set of images, in accordance with the disclosed embodiments. The processing unit 110 can execute the monocular image analysis module 402 to implement process 500B. At step 540, the processing unit 110 can determine a set of candidate objects representing possible vehicles and / or pedestrians. For example, the processing unit 110 can scan one or more images, compare the images against one or more predetermined patterns, and identify possible locations within each image that may contain objects of interest (e.g., vehicles, pedestrians, or parts thereof). The predetermined patterns can be designed to achieve a high rate of "false hits" and a low rate of "errors."For example, processing unit 110 can use a low threshold of similarity to predetermined patterns to identify candidate objects as possible vehicles or pedestrians. This can allow processing unit 110 to reduce the probability of a candidate object representing a vehicle or pedestrian being overlooked (e.g., not identified).

[0094] In step 542, processing unit 110 can filter the set of candidate objects to exclude certain candidates (e.g., irrelevant or less relevant objects) based on classification criteria. Such criteria can be derived from various properties associated with object types stored in a database (e.g., a database stored in memory 140). Properties can include object shape, dimensions, texture, position (e.g., relative to vehicle 200), and the like. Thus, processing unit 110 can use one or more sets of criteria to exclude incorrect candidates from the set of candidate objects.

[0095] In step 544, the processing unit 110 can analyze multiple frames of images to determine whether objects in the set of candidate objects represent vehicles and / or pedestrians. For example, the processing unit 110 can track a detected candidate object across successive frames and accumulate image-by-image data associated with the detected object (e.g., size, position relative to vehicle 200, etc.). Additionally, the processing unit 110 can estimate parameters for the detected object and compare the object's frame-by-frame position data with a predicted position.

[0096] In step 546, the processing unit 110 can generate a set of measurements for the detected objects. Such measurements can include, for example, position, velocity, and acceleration values ​​(relative to vehicle 200) associated with the detected objects. In some embodiments, the processing unit 110 can generate the measurements based on estimation techniques using a range of time-based observations, such as Kalman filters or linear quadratic estimation (LQE), and / or based on available modeling data for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). The Kalman filters can be based on a measurement of an object's scale, where the scale measurement is proportional to a time to collision (e.g., the time it takes vehicle 200 to reach the object).Thus, by performing steps 540-546, the processing unit 110 can identify vehicles and pedestrians appearing within the set of captured images and derive information (e.g., position, speed, size) associated with the vehicles and pedestrians. Based on the identification and the derived information, the processing unit 110 can initiate one or more navigation responses in the vehicle 200, as described above in conjunction with [the above]. Fig. 5A described.

[0097] In step 548, the processing unit 110 can perform an optical flow analysis of one or more images to reduce the probability of a "false hit" being detected and a candidate object representing a vehicle or pedestrian being missing. The optical flow analysis might involve, for example, analyzing motion patterns relative to vehicle 200 in one or more images that are associated with other vehicles and pedestrians and differ from the movement of the road surface. The processing unit 110 can calculate the motion of candidate objects by observing the different positions of the objects across multiple frames taken at different times. The processing unit 110 can use the position and time values ​​as inputs into mathematical models to calculate the motion of the candidate objects.Thus, optical flow analysis can provide an alternative method for detecting vehicles and pedestrians in the vicinity of vehicle 200. Processing unit 110 can perform optical flow analysis in combination with steps 540-546 to provide redundancy for vehicle and pedestrian detection and increase the reliability of system 100.

[0098] Fig. Figure 5C is a flowchart illustrating an exemplary process 500C for detecting road markings and / or lane geometry information in a set of images, in accordance with the disclosed embodiments. The processing unit 110 can execute the monocular image analysis module 402 to implement process 500C. In step 550, the processing unit 110 can detect a set of objects by scanning one or more images. To detect segments of lane markings, lane geometry information, and other relevant road markings, the processing unit 110 can filter the set of objects to exclude those determined to be irrelevant (e.g., small potholes, small stones, etc.). In step 552, the processing unit 110 can group together the segments detected in step 550 that belong to the same road marking or lane marking.Based on the grouping, the processing unit 110 can develop a model to represent the detected segments, such as a mathematical model.

[0099] In step 554, the processing unit 110 can construct a set of measurements associated with the detected segments. In some embodiments, the processing unit 110 can create a projection of the detected segments from the image plane onto the real plane. The projection can be characterized using a third-degree polynomial that has coefficients corresponding to physical properties such as the position, inclination, curvature, and curvature derivative of the detected road. When generating the projection, the processing unit 110 can take into account changes in the road surface as well as pitch and roll rates associated with the vehicle 200. Furthermore, the processing unit 110 can model the road height by analyzing the position and motion cues present on the road surface.Furthermore, the processing unit 110 can estimate the pitch and roll rates associated with the vehicle 200 by tracking a set of feature points in the one or more images.

[0100] In step 556, the processing unit 110 can perform a multi-frame analysis by, for example, tracking the detected segments across successive frames and accumulating frame data associated with detected segments. Because the processing unit 110 performs a multi-frame analysis, the set of measurements created in step 554 can become more reliable and associated with an increasingly higher level of confidence. Thus, by performing steps 550-556, the processing unit 110 can identify road markings appearing within the set of captured images and derive lane geometry information. Based on the identification and the derived information, the processing unit 110 can initiate one or more navigation responses in the vehicle 200, as described above in conjunction with Fig. 5A described.

[0101] In step 558, the processing unit 110 can consider additional information sources to further develop a safety model for the vehicle 200 within the context of its environment. The processing unit 110 can use the safety model to define a context in which the system 100 can safely perform autonomous control of the vehicle 200. To develop the safety model, the processing unit 110, in some embodiments, can consider the position and movement of other vehicles, detected road edges and obstacles, and / or general road shape descriptions extracted from map data (such as data from the map database 160). By considering additional information sources, the processing unit 110 can provide redundancy for detecting road markings and lane geometry, thereby increasing the reliability of the system 100.

[0102] Fig. Figure 5D is a flowchart illustrating an exemplary process 500D for detecting traffic lights in a set of images, in accordance with the disclosed embodiments. The processing unit 110 can execute the monocular image analysis module 402 to implement process 500D. In step 560, the processing unit 110 can scan the set of images and identify objects that appear in locations within the images likely to contain traffic lights. For example, the processing unit 110 can filter the identified objects to create a set of candidate objects, excluding those that are unlikely to be traffic lights. The filtering can be based on various properties associated with traffic lights, such as shape, dimensions, texture, position (e.g., relative to vehicle 200), and the like.These properties can be based on numerous examples of traffic lights and traffic control signals and stored in a database. In some embodiments, the processing unit 110 can perform a multi-frame analysis on the set of candidate objects that represent possible traffic lights. For example, the processing unit 110 can track the candidate objects across successive frames, estimate the actual position of the candidate objects, and filter out those objects that are moving (which are unlikely to be traffic lights). In some embodiments, the processing unit 110 can perform a color analysis on the candidate objects and identify the relative position of the detected colors that appear within possible traffic lights.

[0103] At step 562, the processing unit 110 can analyze the geometry of an intersection. The analysis can be based on any combination of the following factors: (i) the number of lanes detected on both sides of the vehicle 200, (ii) road markings detected (such as arrow markings), and (iii) descriptions of the intersection extracted from map data (e.g., data from the map database 160). The processing unit 110 can perform the analysis using information derived from the execution of the monocular analysis module 402. Additionally, the processing unit 110 can determine a match between the traffic lights detected at step 560 and the lanes appearing near the vehicle 200.

[0104] As vehicle 200 approaches the intersection, processing unit 110 can update the confidence level associated with the analyzed intersection geometry and the detected traffic lights at step 564. For example, the number of traffic lights expected to be present at the intersection, compared to the number actually present, can influence the confidence level. Thus, based on the confidence level, processing unit 110 can delegate control to the driver of vehicle 200 to improve safety conditions. By performing steps 560-564, processing unit 110 can identify traffic lights appearing within the set of captured images and analyze intersection geometry information. Based on this identification and analysis, processing unit 110 can initiate one or more navigation responses in vehicle 200, as described above in conjunction with Fig. 5A described.

[0105] Fig. Figure 5E is a flowchart illustrating an exemplary process 500E for initiating one or more navigation responses in the vehicle 200 based on a vehicle path, in accordance with the disclosed embodiments. At step 570, the processing unit 110 can construct an initial vehicle path associated with the vehicle 200. The vehicle path can be represented using a set of points expressed in coordinates (x, z), and the distance d iThe distance between any two points in the set of points can fall within the range of 1 to 5 meters. In one embodiment, the processing unit 110 can construct the initial vehicle path using two polynomials, such as left and right road polynomials. The processing unit 110 can compute the geometric midpoint between the two polynomials and offset each point enclosed in the resulting vehicle path by a predetermined offset (e.g., an offset of a smart lane), if any (an offset of zero can correspond to driving in the center of a lane). The offset can be in a direction perpendicular to any segment between any two points on the vehicle path.In another embodiment, the processing unit 110 can use a polynomial and an estimated lane width to offset each point of the vehicle path by half the estimated lane width plus a predetermined offset (e.g., an offset of a smart lane).

[0106] At step 572, processing unit 110 can update the vehicle path constructed at step 570. Processing unit 110 can reconstruct the vehicle path constructed at step 570 using a higher resolution, so that the distance d k The distance between two points in the set of points representing the vehicle path is smaller than the distance d described above. i For example, the distance d kThe distance falls within the range of 0.1 to 0.3 meters. The processing unit 110 can reconstruct the vehicle path using a parabolic spline algorithm, which can provide a cumulative distance vector S corresponding to the total length of the vehicle path (i.e., based on the set of points representing the vehicle path).

[0107] At step 574, the processing unit 110 can provide a preview point (expressed in coordinates as (xl)). , z l)) based on the updated vehicle path constructed in step 572. Processing unit 110 can extract the lookout point from the cumulative distance vector S, and the lookout point can be associated with a lookout distance and a lookout time. The lookout distance, which can have a lower limit in the range of 10 to 20 meters, can be calculated as the product of the vehicle 200's speed and the lookout time. For example, if the vehicle 200's speed decreases, the lookout distance can also decrease (e.g., until it reaches the lower limit). The lookout time, which can be in the range of 0.5 to 1.5 seconds, can be inversely proportional to the gain of one or more control loops associated with initiating a navigation response in vehicle 200, such as the driving direction error tracking loop.For example, the gain of the control loop for tracking driving direction errors can depend on the bandwidth of a yaw rate loop, a steering actuator loop, the vehicle's lateral dynamics, and the like. Therefore, the higher the gain of the control loop for tracking driving direction errors, the shorter the predictive time.

[0108] At step 576, processing unit 110 can determine a direction error and a yaw rate command based on the lookout point determined at step 574. Processing unit 110 can determine the direction error by calculating the arctangent of the lookout point, e.g., arctan(x). l / z lThe processing unit 110 can determine the yaw rate command as the product of the direction error and a high control gain. The high control gain can be equal to: (2 / foresight time) if the foresight distance is not at the lower limit. Otherwise, the high control gain can be equal to: (2 * vehicle speed 200 / foresight distance).

[0109] Fig. Figure 5F is a flowchart showing an exemplary process 500F for determining whether a vehicle ahead is changing lanes, in accordance with the disclosed embodiments. At step 580, the processing unit 110 can determine navigation information associated with a vehicle ahead (e.g., a vehicle traveling in front of vehicle 200). For example, the processing unit 110 can determine the position, speed (e.g., direction and velocity), and / or acceleration of the vehicle ahead using the information described above in conjunction with Fig. 5A and Fig. The processing unit 110 can also determine one or more road polynomials, a lookout point (associated with the vehicle 200) and / or a driven route (e.g., a set of points describing a path taken by the vehicle ahead) using the techniques described above in conjunction with Fig. Determine the techniques described in section 5E.

[0110] In step 582, the processing unit 110 can analyze the navigation information determined in step 580. In one embodiment, the processing unit 110 can calculate the distance between a driven route and a road polynomial (e.g., along the route). If the variance of this distance along the route exceeds a predetermined threshold (for example, 0.1 to 0.2 meters on a straight road, 0.3 to 0.4 meters on a moderately curved road, and 0.5 to 0.6 meters on a road with sharp bends), the processing unit 110 can determine that the vehicle ahead is likely to change lanes. If multiple vehicles are detected driving ahead of vehicle 200, the processing unit 110 can compare the driven routes associated with each vehicle.Based on this comparison, processing unit 110 can determine that a vehicle whose route does not match the routes of other vehicles is likely to change lanes. Processing unit 110 can additionally compare the curvature of the route traveled (associated with the vehicle ahead) with the expected curvature of the road segment in which the vehicle ahead is traveling. The expected curvature can be extracted from map data (e.g., data from map database 160), road polynomials, routes traveled by other vehicles, prior knowledge of the road, and the like. If the difference between the curvature of the route traveled and the expected curvature of the road segment exceeds a predetermined threshold, processing unit 110 can determine that the vehicle ahead is likely to change lanes.

[0111] In another embodiment, the processing unit 110 can compare the instantaneous position of the vehicle ahead with the look-ahead point (associated with vehicle 200) over a specific period (e.g., 0.5 to 1.5 seconds). If the distance between the instantaneous position of the vehicle ahead and the look-ahead point varies during this specific period, and the cumulative sum of the variation exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a moderately curved road, and 1.3 to 1.7 meters on a road with sharp curves), the processing unit 110 can determine that the vehicle ahead is likely to change lanes.In another embodiment, the processing unit 110 can analyze the geometry of the driven route by comparing the lateral distance traveled along the route with the expected curvature of the driven route. The expected radius of curvature can be determined according to the following calculation: (δ. z 2 + δ x 2 ) / 2 / (δ x ), where δ x for the lateral distance traveled and δ zThis represents the distance traveled in the longitudinal direction. If the difference between the lateral distance traveled and the expected curvature exceeds a predetermined threshold (e.g., 500 to 700 meters), the processing unit 110 can determine that the vehicle ahead is likely to change lanes. In another embodiment, the processing unit 110 can analyze the position of the vehicle ahead. If the position of the vehicle ahead obscures a road polynomial (e.g., the vehicle ahead is positioned over the road polynomial), then the processing unit 110 can determine that the vehicle ahead is likely to change lanes.In the event that the position of the vehicle ahead is such that another vehicle is detected in front of the vehicle ahead and the routes traveled by the two vehicles are not parallel, the processing unit 110 can determine that the (closer) vehicle ahead is likely to change lanes.

[0112] In step 584, the processing unit 110 can determine, based on the analysis performed in step 582, whether the vehicle 200 ahead is changing lanes or not. For example, the processing unit 110 can make the determination based on a weighted average of the individual analyses performed in step 582. Under such a scheme, for example, a decision by the processing unit 110 that the vehicle ahead is likely to change lanes, based on a certain type of analysis, can be assigned a value of "1" (and "0" to represent a determination that the vehicle ahead is unlikely to change lanes). Different analyses performed in step 582 can be assigned different weights, and the disclosed embodiments are not limited to a specific combination of analyses and weights.Furthermore, in some embodiments, the analysis can use a trained system (e.g., a machine learning or deep learning system) that can, for example, estimate a future path ahead of a vehicle's current location based on an image captured at the current location.

[0113] Fig. Figure 6 is a flowchart illustrating an exemplary process 600 for initiating one or more navigation responses based on stereo image analysis in accordance with the disclosed embodiments. In step 610, the processing unit 110 can receive a first and second set of images via the data interface 128. For example, cameras enclosed in the image acquisition unit 120 (such as the image capture devices 122 and 124 with fields of view 202 and 204) can capture a first and second set of images of an area in front of the vehicle 200 and transmit them to the processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, the processing unit 110 can receive the first and second sets of images via two or more data interfaces.The disclosed embodiments are not limited to specific data interface configurations or protocols.

[0114] In step 620, the processing unit 110 can execute the stereo image analysis module 404 to perform a stereo image analysis of the first and second sets of images in order to create a 3D map of the road in front of the vehicle and to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, and the like. The stereo image analysis can be performed in a similar manner to the steps described above in conjunction with Fig. 5A-5D and 6 are described. For example, the Processing Unit 110 can execute the Stereo Image Analysis Module 404 to identify candidate objects (e.g., vehicles, pedestrians, road markings, traffic lights, road hazards, etc.) within the first and second sets of images, filter out a subset of the candidate objects based on various criteria, perform a multi-frame analysis, construct measurements, and determine a confidence level for the remaining candidate objects. In performing the above steps, the Processing Unit 110 can consider information from both the first and second sets of images, as well as information from a single set of images.For example, processing unit 110 can analyze the differences in pixel-level data (or other data subsets from the two streams of captured images) for a candidate object that appears in both the first and second sets of images. As another example, processing unit 110 can estimate the position and / or velocity of a candidate object (e.g., relative to vehicle 200) by observing that the object appears in one set of images but not the other, or relative to other differences that may exist with respect to objects appearing in the two image streams. For example, position, velocity, and / or acceleration with respect to vehicle 200 can be determined based on trajectories, positions, motion characteristics, etc., of features associated with an object that appears in one or both of the image streams.

[0115] At step 630, the processing unit 110 can execute the navigation response module 408 to perform the analysis carried out in step 620 and the above in conjunction with Fig. The techniques described in section 4 are used to initiate one or more navigation responses in the vehicle 200. Navigation responses can include, for example, turning, changing lanes, changing acceleration, changing speed, braking, and the like. In some embodiments, the processing unit 110 can use data derived from the execution of the speed and acceleration module 406 to initiate the one or more navigation responses. Additionally, multiple navigation responses can occur simultaneously, sequentially, or in any combination thereof.

[0116] Fig. Figure 7 is a flowchart illustrating an exemplary process 700 for initiating one or more navigation responses based on an analysis of three sets of images in accordance with disclosed embodiments. In step 710, the processing unit 110 can receive a first, second, and third plurality of images via the data interface 128. For example, cameras enclosed in the image acquisition unit 120 (such as the image capture devices 122, 124, and 126 with fields of view 202, 204, and 206) can capture a first, second, and third plurality of images of an area in front of and / or to the side of the vehicle 200 and transmit them to the processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, the processing unit 110 can receive the first, second, and third plurality of images via three or more data interfaces.For example, each of the image acquisition devices 122, 124, 126 can have an associated data interface for communicating data to the processing unit 110. The disclosed embodiments are not limited to specific data interface configurations or protocols.

[0117] In step 720, the processing unit 110 can analyze the first, second, and third sets of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, and the like. The analysis can be performed in a similar manner to the steps described above in conjunction with Fig. 5A-5D and 6 are described. For example, the processing unit 110 can perform monocular image analysis (e.g., via the execution of the monocular image analysis module 402 and based on the above in conjunction with Fig. (Steps 5A-5D described) for each of the first, second, and third sets of images. Alternatively, the processing unit 110 can perform a stereo image analysis (e.g., by executing the stereo image analysis module 404 and based on the steps described above in conjunction with Fig. (6 described steps) on the first and second sets of images, the second and third sets of images, and / or the first and third sets of images. The processed information corresponding to the analysis of the first, second, and / or third sets of images can be combined. In some embodiments, the processing unit 110 can perform a combination of monocular image analysis and stereo image analysis. For example, the processing unit 110 can perform a monocular image analysis (e.g., by executing the monocular image analysis module 402) on the first set of images and a stereo image analysis (e.g., by executing the stereo image analysis module 404) on the second and third sets of images.The configuration of the image acquisition devices 122, 124, and 126—including their respective positions and fields of view 202, 204, and 206—can influence the types of analyses performed on the first, second, and third sets of images. The disclosed embodiments are not limited to a specific configuration of the image acquisition devices 122, 124, and 126 or to the types of analyses performed on the first, second, and third sets of images.

[0118] In some embodiments, the processing unit 110 can perform tests on the system 100 based on the images acquired and analyzed in steps 710 and 720. Such tests can provide an indicator of the overall performance of the system 100 for specific configurations of the image acquisition devices 122, 124, and 126. For example, the processing unit 110 can determine the proportion of "false hits" (e.g., cases in which the system 100 incorrectly determined the presence of a vehicle or pedestrian) and the proportion of "errors."

[0119] At step 730, the processing unit 110 can initiate one or more navigation responses in the vehicle 200 based on information derived from two of the first, second, and third sets of images. The selection of two of the first, second, and third sets of images can depend on various factors, such as the number, types, and sizes of the objects detected in each set of images. The processing unit 110 can also make the selection based on image quality and resolution, the effective field of view represented in the images, the number of frames captured, the extent to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which an object appears, the proportion of the object appearing in each of these frames, etc.), and the like.

[0120] In some embodiments, the processing unit 110 can select information derived from any two of the first, second, and third sets of images by determining the degree to which information derived from one image source corresponds to information derived from other image sources. For example, the processing unit 110 can combine the processed information derived from each of the image acquisition devices 122, 124, and 126 (whether by monocular analysis, stereo analysis, or any combination of the two) and determine visual indicators (e.g., lane markings, a detected vehicle and its location and / or path, a detected traffic light, etc.) that are consistent across the images captured by each of the image acquisition devices 122, 124, and 126. The processing unit 110 can also exclude information that is inconsistent across the captured images (e.g.,For example, a vehicle changing lanes, a lane model indicating a vehicle too close to vehicle 200, etc.). Thus, the processing unit can select 110 pieces of information derived from two of the first, second, and third sets of images, based on the determination of consistent and inconsistent information.

[0121] Navigation responses can include, for example, a turning maneuver, a lane change, a change in acceleration, and the like. The processing unit 110 can perform one or more navigation responses based on the analysis carried out in step 720 and the above in conjunction with Fig. The processing unit 110 can initiate one or more navigation responses using the techniques described in section 4. It can also use data derived from the execution of the velocity and acceleration module 406. In some embodiments, the processing unit 110 can initiate one or more navigation responses based on a relative position, relative velocity, and / or relative acceleration between the vehicle 200 and an object detected in any of the first, second, and third sets of images. Multiple navigation responses can occur simultaneously, sequentially, or in any combination thereof. Reinforcement learning and trained navigation systems

[0122] The following sections discuss autonomous driving as well as systems and methods for implementing autonomous vehicle control, regardless of whether it is fully autonomous control (a self-driving vehicle) or semi-autonomous control (e.g., with one or more driver assistance systems or functions). As in Fig. As shown in Figure 8, the autonomous driving task can be divided into three main modules, including a sensor acquisition module 801, a driving guidance module 803, and a control module 805. In some embodiments, modules 801, 803, and 805 can be stored in the memory unit 140 and / or memory unit 150 of the system 100, or modules 801, 803, and 805 (or parts thereof) can be stored remotely from the system 100 (e.g., stored on a server that the system 100 can access, for example, via the wireless transceiver 172). Furthermore, each of the modules disclosed herein (e.g., modules 801, 803, and 805) can implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system.

[0123] The Sensor Acquisition Module 801, which can be implemented using the Processing Unit 110, can perform various tasks related to sensor-acquiring a navigation state in the environment of a host vehicle. Such tasks can be based on inputs from various sensors and sensor acquisition systems associated with the host vehicle. These inputs can include images or image streams from one or more onboard cameras, GPS position data, accelerometer outputs, user feedback or input to one or more user interface devices, radar, lidar, etc. Data from cameras and / or other available sensors can be collected, analyzed, and formulated, along with map information, into a "sensor-acquired state" that describes information from a scene in the environment of the host vehicle.The sensor-detected state can include information relating to target vehicles, lane markings, pedestrians, traffic lights, road geometry, lane shape, obstacles, distances to other objects / vehicles, relative velocities, and relative accelerations, among other potential sensor-detected information. Supervised machine learning can be implemented to generate an output of the sensor-detected state based on the sensor-detected data provided to the Sensor Acquisition Module 801. The output of the Sensor Acquisition Module can represent a sensor-detected navigation "state" of the host vehicle, which can be forwarded to the Driving Guidance Module 803.

[0124] While a sensor-detected state can be developed based on image data received from one or more cameras or image sensors associated with a host vehicle, a sensor-detected state for use in navigation can be developed using any suitable sensor or combination of sensors. In some embodiments, the sensor-detected state can be developed without recourse to acquired image data. In fact, each of the navigation principles described herein can be applied to sensor-detected states developed based on acquired image data as well as to sensor-detected states developed using other, non-image-based sensors. The sensor-detected state can also be determined from sources outside the host vehicle.For example, a sensor-detected state can be developed wholly or partially based on information received from sources remote from the host vehicle (e.g., based on sensor information, processed state information, etc., shared by other vehicles, from a central server, or from any other information source relevant to a navigation state of the host vehicle).

[0125] The Driving Policy Module 803, which is discussed in more detail below and can be implemented using the Processing Unit 110, can implement a desired driving policy to decide on one or more navigation actions of the host vehicle in response to the sensor-detected navigation state. If no other actors (e.g., target vehicles or pedestrians) are present in the vicinity of the host vehicle, the input of the sensor-detected state to the Driving Policy Module 803 can be handled in a relatively straightforward manner. The task becomes more complex if the sensor-detected state requires negotiation with one or more other actors. The technology used to generate the output of the Driving Policy Module 803 can include reinforcement learning (discussed in more detail below).The output of the driving guidance module 803 can include at least one navigation action for the host vehicle and can include a desired acceleration (which can translate into an updated speed for the host vehicle), a desired yaw rate for the host vehicle, and a desired trajectory, among other potential desired navigation actions.

[0126] Based on the output of the steering module 803, the control module 805, which can also be implemented using the processing unit 110, can develop control instructions for one or more actuators or controlled devices associated with the host vehicle. Such actuators and devices can include an accelerator, one or more steering controls, a brake, a signal generator, a display, or any other actuator or device that can be controlled as part of a navigation operation associated with a host vehicle. Control theory principles can be used to generate the output of the control module 805. The control module 805 can be responsible for developing and issuing instructions to controllable components of the host vehicle to implement the desired navigation objectives or requirements of the steering module 803.

[0127] Referring again to the Driving Guidance Module 803, in some embodiments a reinforcement learning-trained system can be used to implement the Driving Guidance Module 803. In other embodiments, the Driving Guidance Module 803 can be implemented without a machine learning approach by using specified algorithms to “manually” address the various scenarios that may occur during autonomous navigation. While such an approach is feasible, it may result in an overly simplistic driving guideline and lack the flexibility of a machine learning-trained system.A trained system can, for example, better handle complex navigation states and determine whether a taxi is parking or stopping to pick up or drop off a passenger; determine whether a pedestrian intends to cross the road in front of the host vehicle; compensate for unexpected behavior from other drivers by adopting a defensive posture; negotiate with target vehicles and / or pedestrians in heavy traffic; decide when to suspend certain navigation rules or extend other rules; and anticipate non-sensor-detected but expected conditions (e.g., whether a pedestrian will emerge from behind a car or obstacle), etc. A system trained using reinforcement learning may also be better able to manage a continuous and high-dimensional state space as well as a continuous action space.

[0128] Training the system using reinforcement learning can include learning a driving guideline to map sensor-detected states to navigation actions. A driving guideline is a function π : S→ A, where S is a set of states and A ⊂ ℝ 2 The action space is (e.g., desired velocity, acceleration, yaw commands, etc.). The state space is S = S s × S p , where S s the sensor detection status and S p The additional information is about the state stored by the strategy. When working in discrete time intervals, the current state s at time t can be determined. t ∈ S can be observed, and the policy can be applied to perform a desired action, a t = π(s t ), to obtain.

[0129] The system can be trained by exposing it to different navigation states, applying the policy, and providing a reward (based on a reward function designed to reward the desired navigation behavior). Based on the reward feedback, the system can "learn" the policy and is trained to produce the desired navigation actions. For example, the learning system can use the current state s t ∈ S observe and choose an action a t ∈ A based on a guideline π:S→D(A) decide. Based on the decided action (and the implementation of the action), the environment moves to the next state s. t +1 ∈ S is passed to the learning system for observation. For each action developed in response to the observed state, the feedback to the learning system is a reward signal r1. r2····.

[0130] The goal of reinforcement learning (RL) is to find a guideline π. It is generally assumed that at time t there exists a reward function r. t there are factors that describe the current quality of the condition s t and taking action a t measures. Taking action a tHowever, at a given time t, the environment influences and therefore affects the value of future states. Consequently, when deciding on the action to take, not only the current reward but also future rewards should be considered. In some cases, the system should take a particular action, even if it is associated with a lower reward than another available option, if the system determines that a greater reward can be realized in the future by taking the lower-reward option now. To formalize this, let's state that a policy π and an initial state s represent a distribution over ℝ. T induce, where the probability of a vector (r1, ... , r T ) the probability of observing the rewards r1, ... ,r TThis occurs when the agent starts in state s0 = s and follows the directive π from there. The value of the initial state s can be defined as follows: Vπ(s)=E[∑t=1Trt|s0=s,∀t≥1,at=π(st)].

[0131] Instead of restricting the time horizon to T, the future rewards can be reduced to define the following for a fixed ∈ (0, 1): Vπ(s)=E[∑t=1∞γtrt|s0=s,∀t≥1,at=π(st)].

[0132] In any case, the optimal policy is the solution of argmaxπE[Vπ(s)] where the expectation is higher than the initial state s.

[0133] There are several possible methods for training the driving guidance system. For example, an imitation approach (e.g., behavior cloning) can be used, where the system learns from state / action pairs, the actions being those chosen by a good agent (e.g., a human) in response to a given observed state. Suppose a human driver is observed. Through this observation, many examples of the form (s t , a t ), where s t the condition and a t The action of the human driver is obtained, observed, and used as a basis for training the driving policy system. For example, supervised learning can be used to learn a policy π such that π(s) t ) ≈ a tThere are many potential advantages to this approach. First, it eliminates the need to define a reward function. Second, the learning is supervised and occurs offline (there is no need to deploy the agent in the learning process). One disadvantage of this method is that different human drivers, and even the same human drivers, are not deterministic in their policy decisions. Therefore, it is often not possible to learn a function for which ||π(s t - a t || is very small. And even small errors can add up to large errors over time.

[0134] Another technique that can be used is policy-based learning. Here, the policy can be expressed in parametric form and directly optimized using a suitable optimization technique (e.g., stochastic gradient descent). The approach involves... argmaxπE[Vπ(s)] to solve the given problem directly. There are, of course, many ways to solve the problem. One advantage of this approach is that it addresses the problem directly and therefore often leads to good practical results. A potential disadvantage is that it frequently requires "on-policy" training; that is, learning π is an iterative process where, at iteration j, we encounter a less-than-perfect policy π. j have, and to determine the next guideline π j To create, we need to interact with the environment while working based on π. j act.

[0135] The system can also be trained using value-based learning (learning Q- or V-functions). Assume that a good approximation to the optimal value function V* can be learned. An optimal policy can be constructed (e.g., by invoking the Bellman equation). Some versions of value-based learning can be implemented offline (called "off-policy" training). Some disadvantages of the value-based approach can arise from its strong reliance on Markov assumptions and the need to approximate a complicated function (it can be more difficult to approximate the value function than to directly approximate the policy).

[0136] Another technique can include model-based learning and planning (learning the probability of state transitions and solving the optimization problem to find the optimal V). Combinations of these techniques can also be used to train the learning system. With this approach, the dynamics of the process can be learned, i.e., the function that (s t , a t ) assumes and a distribution over the nearest state s t+1 This results in... Once this function is learned, the optimization problem can be solved to find the guideline π whose value is optimal. This is called "planning". One advantage of this approach could be that the learning process is monitored and can be performed offline by observing triplets (s t , a t , s t+1) can be applied. One disadvantage of this approach, similar to the "imitation" approach, could be that small errors accumulate in the learning process and lead to strategies that do not function well.

[0137] Another approach to training the Driving Policy Module 803 can involve decomposing the driving policy function into semantically meaningful components. This allows for the manual implementation of parts of the policy, which can ensure policy safety, and the implementation of other parts of the policy using reinforcement learning techniques, which can enable adaptability to many scenarios, a human-like balance between defensive / aggressive behavior, and human-like negotiation with other drivers. From a technical perspective, a reinforcement learning approach can combine several approaches and provide a traceable training process, where the majority of the training can be conducted either using recorded data or a custom-built simulator.

[0138] In some embodiments, training of the driving guide module 803 can rely on an "options" mechanism. To illustrate this, consider a simple driving guide scenario for a two-lane highway. In a direct driving guide approach, a guide π represents the state in A ⊂ ℝ. 2 from, where the first component of π(s) is the desired acceleration command and the second component of π(s) is the yaw rate. With a modified approach, the following guidelines can be constructed:

[0139] Guideline for automatic cruise control (ACC), o ACC : S → A: this policy always outputs a yaw rate of 0 and only changes the speed to implement smooth and accident-free driving.

[0140] ACC+Links policy, o LS → A: The longitudinal command of this policy is the same as the ACC command. The yaw rate is a simple implementation of centering the vehicle on the center of the left lane while ensuring safe lateral movement (e.g., not moving to the left if a car is already on the left).

[0141] ACC+ Legal Policy, o R : S → A: How o L , but the vehicle can be centered towards the middle of the right lane.

[0142] These guidelines can be referred to as "options". Based on these "options", a guideline can be learned that includes options π. o : selects S → O, where O is the set of available options. In one case, () = {o ACC , o L , o R}. The option selection policy π o defines an actual guideline π : S - A by specifying for each s π(s) = o πo (s) (s) determines.

[0143] In practice, the policy function can be decomposed into an options graph 901, as shown in Fig. Figure 9 shows another example options graph for 1000. Fig. Figure 10 shows that the options graph can represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There is a special node called the root node 903 of the graph. This node has no incoming nodes. The decision process traverses the graph, starting from the root node, until it reaches a "leaf" node that points to a node with no outgoing decision lines. As shown in Fig. As shown in Figure 9, the leaf nodes can include, for example, nodes 905, 907, and 909. When the driving guidance module 803 encounters a leaf node, it can issue the acceleration and steering commands associated with a desired navigation action associated with that leaf node.

[0144] Internal nodes, such as nodes 911, 913, and 915, can lead to the implementation of a policy that selects a child from among the available options. The set of available children of an internal node includes all nodes associated with that particular internal node via decision lines. For example, internal node 913, which is in Fig. 9, designated as “threading”, three child knots 909, 915 and 917 (“Stay”, “Overtake on the right” and “Overtake on the left”), each connected to knot 913 by a decision line, are included.

[0145] The flexibility of the decision-making system can be increased by allowing nodes to adjust their position in the hierarchy of the options graph. For example, each node may declare itself as "critical." Each node can implement a function "is critical" that returns "true" if the node is in a critical section of its policy implementation. For example, a node responsible for overtaking might declare itself as critical while in the middle of a maneuver. This can impose restrictions on the set of available children of a node u, which can include all nodes v that are children of node u and for which there is a path from v to a leaf node that passes through all nodes marked as critical.Such an approach can, on the one hand, allow the declaration of the desired path in the graph at each time step, while on the other hand, the stability of a policy can be maintained, especially while critical parts of the policy are being implemented.

[0146] Defining an options graph decomposes the problem of learning the driving policy π: S → A into a problem of defining a policy for each node of the graph, where the policy should choose from the available child nodes at internal nodes. For some of the nodes, the respective policy can be implemented manually (e.g., by if-then algorithms that specify a set of actions in response to an observed state), while for others, the policies can be implemented using a trained system built through reinforcement learning. The choice between manual or trained / learned approaches may depend on safety considerations related to the task and their relative simplicity. The options graphs can be constructed such that some of the nodes are easy to implement, while other nodes can rely on trained models.Such an approach can ensure the safe operation of the system.

[0147] The following discussion provides further details regarding the role of the options graph of Fig. 9 within the guidance module 803. As discussed above, the input to the guidance module is a "sensor-detected state" that summarizes the environment map, as obtained, for example, from the available sensors. The output of the guidance module 803 is a set of wishes (optionally along with a set of hard constraints) that define a trajectory as the solution to an optimization problem.

[0148] As described above, the options graph represents a hierarchical set of decisions organized as a DAG. There is a special node called the "root" of the graph. The root node is the only node that has no incoming edges (e.g., decision lines). The decision process traverses the graph, starting from the root node, until it reaches a "leaf" node, that is, a node that has no outgoing edges. Each internal node should implement a policy that selects a child from its available children. Each leaf node should implement a policy that defines a set of wishes (e.g., a set of navigation destinations for the host vehicle) based on the entire path from the root to the leaf.The set of desired outcomes, together with a set of hard constraints defined directly based on the sensor-detected state, constitutes an optimization problem whose solution is the vehicle's trajectory. The hard constraints can be used to further enhance the system's safety, while the desired outcomes can be used to provide ride comfort and human-like driving characteristics. The trajectory provided as the solution to the optimization problem, in turn, defines the commands that should be sent to the steering, braking, and / or engine actuators to achieve the trajectory.

[0149] With renewed reference to Fig. In Figure 9, the options graph 901 represents an options graph for a two-lane highway, including merging lanes (meaning that at some points a third lane merges into the right or left lane of the highway). The root node 903 first decides whether the host vehicle is in a simple road scenario or approaching a merging scenario. This is an example of a decision that can be implemented based on the sensor detection state. The "Simple Road" node 911 includes three child nodes: the "Stay" node 909, the "Pass Left" node 917, and the "Pass Right" node 915. "Stay" refers to a situation where the host vehicle wants to continue driving in the same lane. The "Stay" node is a leaf node (no outgoing edges / lines). Therefore, the "Stay" node defines a set of wishes.The first wish it defines can include the desired lateral position—for example, as close as possible to the center of the current lane. It can also include the wish to navigate smoothly (e.g., within predefined or permissible maximum acceleration values). The Stay node can also define how the host vehicle should react to other vehicles. For example, the Stay node can check the sensor-detected target vehicles and assign each a semantic meaning that can be translated into trajectory components.

[0150] The target vehicles in the vicinity of the host vehicle can be assigned various semantic meanings. For example, in some embodiments, the semantic meaning may include one of the following labels: 1) not relevant: indicates that the detected vehicle is currently not relevant in the scene; 2) nearest lane: indicates that the detected vehicle is in an adjacent lane and a reasonable offset should be maintained from that vehicle (the exact offset can be computed in the optimization problem that constructs the trajectory taking into account the desires and hard constraints, and may possibly be vehicle-dependent—the "Stay" leaf of the options graph determines the semantic type of the target vehicle, which defines the desire relative to the target vehicle); 3) yield: the host vehicle attempts to yield to the detected target vehicle, for example, by...4) Reduces speed (especially if the host vehicle determines that the target vehicle is likely to enter the host vehicle's lane); 5) Take right of way: the host vehicle attempts to take right of way, for example, by increasing its speed; 6) Follow: the host vehicle wants to maintain a smooth journey by following the target vehicle; 7) Pass left / right: this means the host vehicle wants to initiate a lane change to the left or right lane. The pass-left node 917 and the pass-right node 915 are internal nodes that do not yet define any wishes.

[0151] The next node in options graph 901 is the gap-select node 919. This node can be responsible for selecting a gap between two target vehicles in a specific target lane into which the host vehicle wishes to enter. By selecting a node of the form IDj for a certain value of j, the host vehicle arrives at a leaf that specifies a wish for the trajectory optimization problem—for example, the host vehicle wants to perform a maneuver to reach the selected gap. Such a maneuver might involve first accelerating / decelerating in the current lane and then heading towards the target lane at a suitable time to enter the selected gap. If the gap-select node 919 does not find a suitable gap, it switches to the abort node 921, which defines the wish to return to the center of the current lane and abort the overtaking maneuver.

[0152] Referring again to the threading node 913, the host vehicle, when approaching a threading point, has several options that may depend on the specific situation. As in Fig. As shown in Figure 11A, the host vehicle 1105 is driving on a two-lane road where no other target vehicles are detected, neither in the main lanes of the two-lane road nor in the merging lane 1111. In this situation, the driving guidance module 803 can select the stay node 909 upon reaching the merging node 913. This means that remaining within its current lane may be desirable if no target vehicles are detected as merging onto the roadway.

[0153] In Fig. In 11B, the situation is slightly different. Here, the host vehicle 1105 detects one or more target vehicles 1107 merging from the merging lane 1111 onto the main carriageway 1112. In this situation, the driving guidance module 803, as soon as it reaches the merging node 913, can initiate a left overtaking maneuver to avoid the merging situation.

[0154] In Fig. In scenario 11C, the host vehicle 1105 encounters one or more target vehicles 1107 merging from the merging lane 1111 onto the main carriageway 1112. The host vehicle 1105 also detects target vehicles 1109 traveling in a lane adjacent to the host vehicle's lane. The host vehicle also detects one or more target vehicles 1110 traveling in the same lane as the host vehicle 1105. In this situation, the driving guidance module 803 can decide to adjust the speed of the host vehicle 1105 to give priority to target vehicle 1107 and proceed ahead of target vehicle 1115. This can be achieved, for example, by advancing to the gap-select node 919, which in turn selects a gap between ID0 (vehicle 1107) and ID1 (vehicle 1115) as a suitable merging gap. In such a case, the suitable gap in the threading situation defines the goal for an optimization problem of the trajectory planner.

[0155] As discussed above, nodes in the options graph can self-declare themselves as "critical," which can ensure that the selected option passes through the critical nodes. Formally, each node can implement an IsCritical function. After performing a forward traverse of the options graph, from the root to a leaf, and solving the trajectory planner's optimization problem, a reverse traverse can be performed from the leaf back to the root. Along this reverse traverse, the IsCritical function of all nodes in the traverse can be called, and a list of all critical nodes can be stored. On the forward path corresponding to the next timeframe, the trajectory guidance module 803 can be instructed to choose a path from the root node to a leaf that passes through all critical nodes.

[0156] Fig. 11A-11C can be used to demonstrate a potential benefit of this approach. For example, in a situation where an overtaking maneuver is initiated and the driving policy module 803 reaches the sheet corresponding to IDk, it would be undesirable to select, say, the Stay node 909 if the host vehicle is in the middle of the overtaking maneuver. To avoid such abruptness, the IDj node can classify itself as critical. During the maneuver, the success of the trajectory planner can be monitored, and the IsCritical function returns the value "True" if the overtaking maneuver is progressing as intended. This approach can ensure that the overtaking maneuver continues in the next timeframe (rather than jumping to a different, potentially inconsistent maneuver before the initially selected one is completed).Conversely, if the manipulation monitoring indicates that the selected maneuver is not progressing as intended, or if the maneuver has become unnecessary or impossible, the `IsCritical` function can return a "False" value. This can allow the gap-select node to select a different gap in the next timeframe or to abort the overtaking maneuver altogether. This approach allows the desired path to be declared in the options graph at each timestep and can also help to promote policy stability during critical parts of execution.

[0157] Hard constraints, which will be discussed in more detail below, can be distinguished from navigational requests. For example, hard constraints can ensure safe driving by applying an additional layer of filtering to a planned navigation action. The implied hard constraints, which can be manually programmed and defined rather than by using a reinforcement learning-based trained system, can be determined from the sensor-detected state. In some embodiments, however, the trained system can learn the applicable hard constraints to be applied and followed.Such an approach can help the Driving Guide Module 803 arrive at a selected action that is already consistent with the applicable hard constraints, thereby reducing or eliminating selected actions that might later require modification to comply with the applicable hard constraints. Nevertheless, as a redundant safety measure, hard constraints can be applied to the output of the Driving Guide Module 803, even if the Driving Guide Module 803 has been trained to consider predetermined hard constraints.

[0158] There are many examples of potential hard constraints. For instance, a hard constraint can be defined in conjunction with a guardrail at the edge of a road. Under no circumstances may the host vehicle cross the guardrail. Such a rule induces a hard lateral constraint on the host vehicle's trajectory. Another example of a hard constraint could be a speed bump (e.g., a speed limit bump), which can induce a hard restriction on the vehicle's speed before and while crossing the bump. Hard constraints can be considered safety-critical and may therefore be defined manually rather than relying solely on a trained system that learns the constraints during training.

[0159] Unlike hard constraints, the goal of wishes can be to enable or achieve comfortable driving. As discussed above, an example of a wish might involve positioning the host vehicle in a lateral position within a lane that corresponds to the center of the host vehicle's lane. Another wish might involve identifying a gap into which to merge. It is important to note that there is no requirement for the host vehicle to be exactly in the center of the lane, but instead, a wish to be as close to it as possible can ensure that the host vehicle tends to drift toward the center of the lane, even in the event of deviations from it. Wishes may not be safety-critical. In some embodiments, wishes may require negotiation with other drivers and pedestrians.One approach to constructing desires can rely on the options graph, and the policy implemented in at least some nodes of the graph can be based on reinforcement learning.

[0160] For the nodes of option graph 901 or 1000, which are implemented as nodes trained based on learning, the training process can include decomposing the problem into a supervised learning phase and a reinforcement learning phase. In the supervised learning phase, a differentiable mapping of (s t , a t ) on ŝ t+1 to be learned in such a way that s t+1 ≈ s t+1 This can be similar to "model-based" reinforcement learning. In the forward loop of the network, ŝ t+1 however, through the actual value of s t+1 be replaced, thereby eliminating the problem of error accumulation. The role of predicting ŝ t+1It consists of propagating messages from the future back to past actions. In this sense, the algorithm can be a combination of "model-based" reinforcement learning and "policy-based learning".

[0161] An important element that can be provided in some scenarios is a differentiable path from future losses / rewards back to decisions about actions. With the options graph structure, the implementation of options that include safety constraints is usually not differentiable. To overcome this problem, the selection of a child in a learned policy node can be stochastic. That is, a node can output a probability vector p that assigns the probabilities used when selecting each of the children of that node. Suppose a node has k children, and a (1) ... a (k)are actions along the path from each child to a leaf. The resulting predicted action is therefore a=∑i=1kpia(i), which can lead to a differentiable path from the action to p. In practice, an action a can be represented as a (i) for i ~ p, and the difference between a and â can be called additive noise.

[0162] For the training of ŝ t+1 at s t , a tSupervised learning can be used in conjunction with real-world data. The guidelines from node simulators can be used for training. Later, a guideline can be fine-tuned using real-world data. Two concepts can make the simulation more realistic. First, an initial guideline can be constructed using imitation, following the "behavioral cloning" paradigm, and using large datasets from the real world. In some cases, the resulting agents may be suitable. In other cases, the resulting agents will at least provide very good initial guidelines for the other agents on the streets. Second, using self-play, our own guideline can be used to extend the training. For example, in an initial implementation of the other agents (cars / pedestrians) that might be encountered, a guideline can be trained based on a simulator.Some of the other agents can be replaced by the new policy, and the process can be repeated. As a result, the policy can be further improved, as it has to respond to a greater variety of other agents with varying levels of complexity.

[0163] Furthermore, in some embodiments, the system can implement a multi-agent approach. For example, the system can consider data from various sources and / or images captured from multiple angles. Additionally, some disclosed embodiments can provide energy savings by considering the anticipation of an event that does not directly affect the host vehicle but may have an impact on it, or even the anticipation of an event that may lead to unforeseen circumstances involving other vehicles (e.g., the radar can "see through" the vehicle ahead and anticipate an unavoidable or even highly probable event that would affect the host vehicle). Trained system with imposed navigation restrictions

[0164] In the context of autonomous driving, the question arises as to how to ensure the safety of a learned policy taught by a trained navigation network. In some embodiments, the driving policy system can be trained using constraints, allowing the actions selected by the trained system to take existing safety constraints into account. Additionally, in some embodiments, an extra layer of safety can be provided by guiding the selected actions of the trained system through one or more hard constraints implied by a specific sensor-detected scene in the environment of the host vehicle. Such an approach can ensure that the actions taken by the host vehicle are limited to those that have been confirmed to satisfy the applicable safety constraints.

[0165] At its core, the navigation system may include a learning algorithm based on a policy function that maps an observed state to one or more desired actions. In some implementations, the learning algorithm is a deep learning algorithm. The desired actions may include at least one action expected to maximize a vehicle's anticipated reward. While in some cases the actual action taken by the vehicle may correspond to one of the desired actions, in others the actual action may be determined based on the detected state, one or more desired actions, and unlearned, hard constraints (e.g., safety restrictions) imposed on the learning navigation system. These constraints may include no-entry zones around various types of detected objects (e.g.,Target vehicles, pedestrians, stationary objects at the edge of a road or on a roadway, moving objects at the edge of a road or on a roadway, guardrails, etc.). In some cases, the size of the zone may vary depending on the detected movement (e.g., speed and / or direction) of a detected object. Further restrictions may include a maximum speed when passing through a pedestrian's zone of influence, a maximum deceleration (to account for a target vehicle's distance behind the host vehicle), a mandatory stop at a sensor-detected crosswalk or railroad crossing, etc.

[0166] Strict constraints, used in conjunction with a machine learning-trained system, can provide a level of safety in autonomous driving that exceeds that available based solely on the output of the trained system. For example, the machine learning system can be trained using a desired set of constraints as training guidelines, and therefore, in response to a sensor-detected navigation state, the trained system can select an action that considers and adheres to the applicable navigation constraints. However, the trained system exhibits some flexibility in selecting navigation actions, and therefore at least some situations may arise where an action selected by the trained system does not strictly adhere to the relevant navigation constraints.To ensure that a selected action strictly adheres to the relevant navigation constraints, the output of the trained system can therefore be combined, compared, filtered, adapted, modified, etc. using a non-machine learning component outside the learning / training framework, which guarantees the strict application of the relevant navigation constraints.

[0167] The following discussion provides additional details regarding the trained system and the potential benefits (particularly from a safety perspective) that can be gained by combining a trained system with an algorithmic component outside the training / learning framework. As discussed, the goal of reinforcement learning through guidelines can be optimized via stochastic gradient descent. The goal (e.g., the expected reward) can be defined as Es˜∼P0R(s¯) be defined.

[0168] Goals that include an expectation can be used in machine learning scenarios. However, such a goal, which is not bound by navigation constraints, cannot return actions that are strictly bound by those constraints. For example, consider a reward function where R(s) - -r holds for trajectories representing a rare, avoidable "corner" event (e.g., a crash), and R(s) ∈ |1, 1| for the rest of the trajectories. A goal for the learning system might be to learn how to perform an overtaking maneuver. Normally, for a crash-free trajectory, R(s) would reward successful, smooth overtaking maneuvers and punish remaining in a lane without completing the overtaking maneuver—hence the range [-1, 1]. If a sequence s represents a crash, the reward r should provide a sufficiently high penalty to deter such an incident.The question is, what should the value of r be to ensure accident-free driving?

[0169] Note that the impact of an accident on E[R(s¯)] The additive term pr is where p is the probability mass of trajectories with a collision. If this term is negligible, i.e., p << 1 / r, then the learning system may favor a policy that causes a collision (or more generally, adopts a reckless driving policy) to successfully complete the overtaking maneuver more often than a more defensive policy that accepts some takeovers not being completed successfully. In other words, if the collision probability is to be at most p, then r must be set such that r >> 1 / p. It may be desirable to make p extremely small (e.g., on the order of p = 10). -9Therefore, r should be large. In a guideline gradient, the gradient of E[R(s¯)] can be estimated. The following lemma shows that the variance of the random variable R(s) with pr 2 The gradient increases, which is larger than r for r >> 1 / p. Therefore, estimating the target can be difficult, and estimating its gradient can be even more difficult.

[0170] Lemma: Let π o Let a guideline be given, and let p and r be scalars such that R(s) - -r is obtained with probability p, and R(s) [-1, 1] is obtained with probability 1 - p. Then the following holds: Var[R(s¯)]≥pr2−(pr+(1−p))2=(p−p2)r2−2p(1−p)r−(1−p)2≈pr2 where the last approximation is valid for the case r ≥ 1 / p.

[0171] This discussion shows that one goal of the form E[R(s¯)] Functional safety cannot be guaranteed without introducing a variance problem. Baseline subtraction to reduce the variance may not adequately address this issue, as it would shift the problem from a high variance in R(s) to an equally high variance in the baseline constant, the estimation of which would also be susceptible to numerical instabilities. Furthermore, if the probability of an accident is p, then on average at least 1 / p sequences should be sampled before an accident event is encountered. This implies a lower bound of 1 / p samples of sequences for a learning algorithm that aims to E[R(s¯)] to minimize. The solution to this problem is more likely to be found in the architectural design described herein than in numerical conditioning techniques. The approach here is based on the idea that hard constraints should be introduced outside the learning framework. In other words, the policy function can be decomposed into a learnable part and a non-learnable part. Formally, the policy function can be described as πθ=π(T)∘πθ(D) be structured, whereby πθ(D) maps the (agnostic) state space to a series of desires (e.g., desired navigation destinations, etc.), while π (T) The function maps the wishes onto a trajectory (which can determine how the vehicle should move within a short area). πθ(D) It is responsible for driving comfort and making strategic decisions, such as which other vehicles should be overtaken, which should be given right of way, and what the desired position of the host vehicle is within its lane, etc. Mapping the sensor-detected navigation status to the desired position serves as a guideline. πθ(D) , which can be learned by maximizing an expected reward from experience. The through π0(D) The generated desires can be translated into a cost function based on the travel trajectories. The function π (T) , which is not a learned function, can be implemented,

[0172] By finding a trajectory that minimizes costs and is subject to the strict limitations of functional safety, this decomposition can ensure functional safety while also providing a comfortable ride.

[0173] A navigation situation involving double merging, as in Fig. Figure 11D illustrates this concept. In a double merge, vehicles approach merge area 1130 from both the left and right sides. From each side, a vehicle, such as vehicle 1133 or vehicle 1135, can decide whether to merge into the lanes on the far side of merge area 1130. Successfully executing a double merge in heavy traffic can require considerable negotiation skills and experience, and can be difficult to accomplish using a heuristic or brute-force approach by enumerating all possible trajectories that could be taken by all agents in the scene. In this example of a double merge, sets of wishes, D, to define which are suitable for the double threading maneuver. D The Cartesian product of the following sentences can be: D=[0,vmax]×L×{g,to}n, where [0, V max ] is the desired target speed of the host vehicle, L = {1, 1.5, 2, 2.5, 3, 3.5, 4} is the desired lateral position in lane units, where integers denote a lane center and fractional numbers denote lane boundaries, and {g, t, o} are classification labels assigned to each of the n other vehicles. The other vehicles can be assigned "g" if the host vehicle is to yield to the other vehicle, "t" if the host vehicle is to take the other vehicle's right of way, or "o" if the host vehicle is to maintain an offset distance from the other vehicle.

[0174] Below is a description of how a set of wishes, (v,l,c1,…,cn)∈D, This can be translated into a cost function using travel trajectories. A travel trajectory can be defined by (x1,y1),...,(x k ,y k ) are represented, where (x i , y i ) is the (lateral, longitudinal) location of the host vehicle (in egocentric units) at time τ·i. In some experiments, τ = 0.1 s and k = 10. Of course, other values ​​can be chosen. The costs assigned to a trajectory can include a weighted sum of the individual costs assigned to the desired velocity, lateral position, and the marking assigned to each of the other n vehicles.

[0175] For a desired speed υ ∈ [0, υ max ] are the costs associated with the speed of a trajectory ∑i=2k(v‖(xi,yi)−(xi−1,yi−1)‖ / τ)2.

[0176] For the desired lateral position, l∈ L, the costs associated with the desired lateral position are ∑i=1kdist(xi,yi,l) where dist(x, y, l) is the distance between the point (x, y) and the lane position l. Regarding the costs caused by other vehicles, for each other vehicle (x1',y1'),…,(xk',yk') represent the other vehicle in egocentric units of the host vehicle, and i can be the earliest point for which there exists j, such that the distance between (x i , y ı ) and (x' j , y' j) is small. If there is no such point, i can be set to i = ∞. If another vehicle is classified as "giving way", it may be desirable that τ i > rj + 0.5, meaning that the host vehicle arrives at the intersection of the trajectories at least 0.5 seconds after the other vehicle has arrived at the same point. A possible formula to translate the above constraint into cost is [τ (j ··· i) + 0.5] + .

[0177] If another vehicle is classified as "taking right of way", it may be equally desirable that τ j > τ i+ 0.5, which is in the cost [τ (i - j) +0.5] +This can be translated. If another vehicle is classified as an "offset," it may be desirable for i = ∞, meaning that the trajectory of the host vehicle and the trajectory of the offset vehicle do not intersect. This condition can be translated into cost by penalizing the distance between the trajectories.

[0178] Assigning a weight to each of these costs can provide a single objective function for the trajectory planner, π (T) Costs that encourage smooth driving can be added to the destination. And to ensure the functional safety of the trajectory, strict constraints can be added to the destination. For example, it can (x i ,y i ) are prohibited from being off the roadway, and (x i ,y i ) it may be forbidden to choose for each trajectory point (x' j ., y' j ) near (x' j, y' j ) of another vehicle if |i - j| is small.

[0179] In summary, the guideline π9 can be decomposed into a mapping from the agnostic state to a set of desires and a mapping from the desires to an actual trajectory. The latter mapping is not based on learning and can be implemented by solving an optimization problem whose cost depends on the desires and whose hard constraints can guarantee the functional safety of the guideline.

[0180] The following discussion describes the mapping from an agnostic state to a set of desires. As described above, a system relying solely on reinforcement learning may exhibit a high and unwieldy variance in reward R(s) to satisfy functional safety. This outcome can be avoided by decomposing the problem into a mapping from an (agnostic) state space to a set of desires using policy gradient iterations, followed by a mapping to an actual trajectory that does not involve a machine-learning-trained system.

[0181] For various reasons, decision-making can also be broken down into semantically meaningful components. For example, the size of D large and even continuous. In the double-threading scenario described above with reference to Fig. 11 D applies D=[0,vmax]×L×{g,to}n) Additionally, the gradient estimator can calculate the term ∑t=1T∇θπθ(at|st) This expression includes the following: In such a formula, the variance can increase with the time horizon T. In some cases, the value of T can be roughly 250, which may be high enough to generate significant variance. Assuming the sampling rate is in the range of 10 Hz and the merging area 1130 is 100 meters, then preparation for merging can begin approximately 300 meters before the merging area. If the host vehicle is traveling at 16 meters per second (about 60 km per hour), then the value of T for an episode can be roughly 250.

[0182] Referring again to the concept of an options graph, in Fig. 11E shows an options graph that is representative of the one in Fig. The double-threading scenario shown in Figure 11D can be used. As discussed above, an options graph can represent a hierarchical set of decisions organized as a directed acyclic graph (DAG). There can be a special node in the graph called the "root" node, which may be the only node that has no incoming edges (e.g., decision lines). The decision process can traverse the graph, starting from the root node, until it reaches a "leaf" node, i.e., a node that has no outgoing edges. Each internal node can implement a policy function that selects a child from its available children. There can be a predefined mapping from the set of traverses over the options graph to the sets of wishes. D. In other words, a pass through the options graph can automatically result in a wish in D can be translated. At a node v in the graph, a parameter vector θ can be used. υ Specify the policy for selecting a child of v. If θ is the concatenation of all θ v is, can π0(D) defined by traversing from the root of the diagram to a leaf, where at each node v the line is defined by θ v The defined policy is used to select a child node.

[0183] In the double-threading options graph 1139 of Fig. 11E allows the root node 1140 to first decide whether the host vehicle is within the threading area (e.g., area 1130 of Fig. 11D) or whether the host vehicle is instead approaching the merging area and needs to prepare for a possible merge. In either case, the host vehicle may need to decide whether to change lanes (e.g., move to the left or right) or remain in its current lane. If the host vehicle decides to change lanes, it may need to decide whether conditions are suitable to proceed and perform the lane-change maneuver (e.g., at the "Drive" node 1142). If changing lanes is not possible, the host vehicle may attempt to "push" toward the desired lane (e.g., at node 1144 as part of a negotiation with vehicles already in the desired lane) by aiming to be on the lane marking. Alternatively, the host vehicle may choose to "stay" in the same lane (e.g.,at node 1146). Such a process can naturally determine the lateral position of the host vehicle. For example:

[0184] This can enable the determination of the desired lateral position naturally. For example, if the host vehicle changes lanes from lane 2 to lane 3, the "Drive" node can set the desired lateral position to 3, the "Stay" node to 2, and the "Push" node to 2.5. Next, the host vehicle can decide whether to maintain its speed (node ​​1148), accelerate (node ​​1150), or decelerate (node ​​1152). The host vehicle can then input a chain-like structure 1154 that passes over the other vehicles, setting its semantic meaning to a value in the set {g, t, o}. This process can establish the vehicle's preferences relative to the other vehicles. The parameters of all nodes in this chain can be shared (similar to recurrent neural networks).

[0185] One potential advantage of options is the interpretability of the results. Another potential advantage is that they depend on the decomposable structure of the set. D This ensures reliability, allowing the policy to be selected from a small number of possibilities at each node. Additionally, the structure can reduce the variance of the policy gradient estimator.

[0186] As previously discussed, the length of an episode in the double-merging scenario can be roughly T = 250 steps. Such a value (or any other suitable value depending on a specific navigation scenario) can provide sufficient time to recognize the consequences of the host vehicle's actions (e.g., if the host vehicle has decided to change lanes in preparation for merging, it will only recognize the benefit of this after the merge has been successfully completed). On the other hand, due to vehicle dynamics, the host vehicle must make decisions at a sufficiently rapid frequency (e.g., 10 Hz in the case described above).

[0187] The options graph can enable a reduction in the effective value of T in at least two ways. First, for higher-level decisions, a reward for lower-level decisions can be defined, taking shorter episodes into account. For example, if the host vehicle has already chosen a "lane change" and the "drive" node, a policy for assigning semantic meanings to vehicles can be learned by considering episodes of 2-3 seconds (meaning T becomes 20-30 instead of 250). Second, for high-level decisions (such as whether to change lanes or stay in the same lane), the host vehicle may not need to make decisions every 0.1 seconds. Instead, the host vehicle may be able to make decisions at a lower frequency (e.g.,to hit a node every second) or to implement an "option termination" function, in which case the gradient can only be computed after each option termination. In either case, the effective value of T may be an order of magnitude smaller than its original value. All in all, the estimator at each node may depend on a value of T that is an order of magnitude smaller than the original 250 steps, which may immediately translate to a smaller variance.

[0188] As discussed above, hard constraints can promote safer driving, and there can be several different types of constraints. For example, static hard constraints can be defined directly from the sensor detection state. These can include speed bumps, speed limits, road curvatures, intersections, etc., in the environment of the host vehicle, which may impose one or more restrictions on the vehicle's speed, course, acceleration, braking (deceleration), etc. Static hard constraints can also include semantic free spaces, prohibiting the host vehicle from driving outside the free space and, for example, from driving too close to physical obstacles. Static hard constraints can also restrict maneuvers (e.g.,prohibit maneuvers that do not correspond to various aspects of a vehicle's kinematic movement; for example, a static hard restriction can be used to prohibit maneuvers that could cause the host vehicle to roll over, skid, or otherwise lose control.

[0189] Hard constraints can also be associated with vehicles. For example, a constraint can be used that requires a vehicle to maintain a longitudinal distance of at least one meter and a lateral distance of at least 0.5 meters from other vehicles. Constraints can also be applied to prevent the host vehicle from maintaining a collision course with one or more other vehicles. For example, a time τ can be a measure of time based on a specific scene. The predicted trajectories of the host vehicle and one or more other vehicles can be considered from a current time to time τ. Where the two trajectories intersect, (tia,til) The arrival and departure times of vehicle i at the intersection point are represented. This means that each vehicle arrives at the point where the first part of the vehicle passes the intersection, and a certain amount of time is required before the last part of the vehicle passes the intersection. This time interval separates the arrival time from the departure time. Assuming that t1a <t2a (i.e., that the arrival time of vehicle 1 is less than the arrival time of vehicle 2), we will want to ensure that vehicle 1 has left the intersection before vehicle 2 arrives. Otherwise, a collision would result. Therefore, a hard constraint can be implemented such that t1l>t2a. To ensure that vehicle 1 and vehicle 2 do not miss each other by a minimal amount, an additional safety margin can be created by including a buffer time in the constraint (e.g., 0.5 seconds or another suitable value). A hard constraint relating to the predicted intersection trajectories of two vehicles can be described as t1l>t2a+0.5 be expressed.

[0190] The time interval τ over which the trajectories of the host vehicle and one or more other vehicles are tracked can vary. However, in intersection scenarios where speeds may be lower, τ can be longer, and τ can be defined such that a host vehicle enters and exits the intersection in less than τ seconds.

[0191] Applying hard constraints to vehicle trajectories naturally requires that these trajectories can be predicted. For the host vehicle, trajectory prediction can be relatively straightforward, as the host vehicle generally already understands an intended trajectory and actually plans to follow it at any given time. Compared to other vehicles, predicting their trajectories can be more complex. For other vehicles, calculating the baseline to determine the predicted trajectories may rely on the current speed and direction of travel of the other vehicles, as determined, for example, by analyzing an image stream captured by one or more cameras and / or other sensors (radar, lidar, acoustics, etc.) on board the host vehicle.

[0192] However, there may be some exceptions that simplify the problem or at least provide additional confidence in a trajectory predicted for another vehicle. For example, on structured roads where lane markings are indicated and right-of-way rules may apply, the trajectories of other vehicles may be based, at least in part, on the position of those vehicles relative to the lanes and on the applicable right-of-way rules. Therefore, in some situations where lane markings are observed, it can be assumed that vehicles in the adjacent lane will respect the lane boundaries. That is, the host vehicle may assume that a vehicle in the adjacent lane will remain in its lane unless signs are observed (e.g.,(a signal light, a strong lateral movement, a movement over a lane boundary), which indicate that the vehicle in the adjacent lane is about to merge into the lane of the host vehicle.

[0193] Other situations can also provide clues about the expected trajectories of other vehicles. For example, at stop signs, traffic lights, roundabouts, etc., where the host vehicle may have the right of way, it can be assumed that other vehicles will respect this right of way. Therefore, unless there is observed evidence of a rule violation, it can be assumed that other vehicles will proceed along a trajectory that respects the right of way of the host vehicle.

[0194] Strict restrictions can also be applied with respect to pedestrians in the vicinity of the host vehicle. For example, a buffer distance can be specified with respect to pedestrians, prohibiting the host vehicle from navigating closer than the prescribed buffer distance relative to an observed pedestrian. The pedestrian buffer distance can be any suitable distance. In some embodiments, the buffer distance can be at least one meter relative to an observed pedestrian.

[0195] Similar to the situation with vehicles, hard constraints can also be applied to the relative movement between pedestrians and the host vehicle. For example, a pedestrian's trajectory (based on direction and speed) can be monitored relative to the projected trajectory of the host vehicle. For a given pedestrian trajectory, t(p), at each point p on the trajectory, can represent the time the pedestrian takes to reach that point. To maintain the required buffer distance of at least 1 meter from the pedestrian, either t(p) must be greater than the time the host vehicle takes to reach point p (with a sufficient time difference so that the host vehicle passes the pedestrian at least one meter away) or t(p) must be less than the time the host vehicle takes to reach point p (e.g.,(e.g., if the host vehicle brakes to yield to the pedestrian). Nevertheless, the hard restriction in the latter example may require that the host vehicle reach point p a sufficient time later than the pedestrian so that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter. Of course, there may be exceptions to the hard restriction for pedestrians. For example, where the host vehicle has the right of way, or where speeds are very low and there is no observed indication that the pedestrian will refuse to yield to the host vehicle or is otherwise moving toward the host vehicle, the hard restriction for pedestrians may be relaxed (e.g., to a smaller buffer of at least 0.75 meters or 0.50 meters).

[0196] In some examples, restrictions may be relaxed if it is determined that not all of them can be met. For instance, in situations where a road is too narrow to maintain the required clearance (e.g., 0.5 meters) to both curbs, or to a curb and a parked vehicle, one or more of the restrictions may be relaxed if mitigating circumstances exist. For example, if there are no pedestrians (or other objects) on the sidewalk, one may proceed slowly with a clearance of 0.1 meters to the curb. In some embodiments, restrictions may be relaxed if this improves the user experience. For example, to avoid a pothole, restrictions may be relaxed to allow a vehicle to navigate closer to the edges of the lane, a curb, or a pedestrian than would normally be permitted.Furthermore, in some embodiments, when determining which restrictions should be relaxed, the one or more restrictions deemed to have the least negative impact on safety are selected. For example, a restriction relating to how close the vehicle may drive to the curb or a concrete barrier may be relaxed before a restriction relating to proximity to other vehicles. In some embodiments, restrictions for pedestrians may be relaxed last, or they may never be relaxed in some situations.

[0197] Fig. Figure 12 shows an example of a scene that can be captured and analyzed during the navigation of a host vehicle. For example, a host vehicle may include a navigation system (e.g., System 100), as described above, which can receive a multitude of images representative of the host vehicle's environment from a camera associated with the host vehicle (e.g., at least one of image acquisition device 122, image acquisition device 124, and image acquisition device 126). The images shown in Figure 12 are used to analyze the surrounding environment of the host vehicle. Fig. Scene 12 shown is an example of one of the images that can be captured at time t of the environment of a host vehicle traveling in lane 1210 along a predicted trajectory 1212. The navigation system can include at least one processing device (e.g., including any of the EyeQ processors or other devices described above) that is specifically programmed to receive the multitude of images and analyze the images to determine an action in response to the scene. In particular, the at least one processing device can implement the sensor acquisition module 801, the driving guidance module 803, and the control module 805, as shown in Fig. Figure 8 shows that the sensor acquisition module 801 can be responsible for collecting and outputting the image information gathered by the cameras and providing this information to the driving guidance module 803 in the form of an identified navigation state. The driving guidance module 803 can represent a trained navigation system trained using machine learning techniques such as supervised learning, reinforcement learning, etc. Based on the navigation state information provided to the driving guidance module 803 by the sensor acquisition module 801, the driving guidance module 803 can (e.g., by implementing the options graph approach described above) generate a desired navigation action for execution by the host vehicle in response to the identified navigation state.

[0198] In some embodiments, the at least one processing device can directly translate the desired navigation action into navigation commands, for example, using the control module 805. In other embodiments, however, hard constraints can be applied so that the desired navigation action provided by the driving guidance module 803 is checked against various predetermined navigation constraints that may be implied by the scene and the desired navigation action. For example, where the driving guidance module 803 outputs a desired navigation action that would cause the host vehicle to follow trajectory 1212, this navigation action can be checked against one or more hard constraints associated with various aspects of the host vehicle's environment.For example, a captured image 1201 might show a curb 1213, a pedestrian 1215, a target vehicle 1217, and a stationary object present in the scene (e.g., a fallen cardboard box). Each of these elements can be associated with one or more hard constraints. For example, the curb 1213 might be associated with a static constraint that prohibits the host vehicle from navigating into the curb or beyond the curb and onto a sidewalk 1214. The curb 1213 might also be associated with a road boundary envelope that defines a distance (e.g., a buffer zone) extending away from the curb (e.g., by 0.1 meter, 0.25 meters, 0.5 meters, 1 meter, etc.) and along the curb, defining a no-navigation zone for the host vehicle. Of course, static constraints can also be associated with other types of roadside boundaries (e.g.,Guardrails, concrete pillars, traffic cones, masts or any other type of road boundary).

[0199] It should be noted that distances and range measurements can be determined using any suitable method. For example, in some embodiments, distance information can be provided by onboard radar and / or lidar systems. Alternatively or additionally, distance information can be derived from the analysis of one or more images taken from the environment of the host vehicle. For example, the number of pixels of a detected object represented in an image can be determined and compared with known field-of-view and focal length geometries of the image-capturing devices to determine the scale and distances. Velocities and accelerations can be determined, for example, by observing changes in scale between objects from image to image over known time intervals.This analysis can reveal the direction of movement towards or away from the host vehicle, as well as the speed at which the object is moving away from or towards the host vehicle. The crossing speed can be determined by analyzing the change in the X-coordinate position of an object from one image to the next over known time periods.

[0200] Pedestrian 1215 may be associated with a pedestrian envelope curve that defines a buffer zone 1216. In some cases, an imposed hard restriction may prevent the host vehicle from navigating within 1 meter of pedestrian 1215 (in any direction relative to the pedestrian). Pedestrian 1215 may also define the location of a pedestrian influence zone 1220. Such an influence zone may be associated with a restriction that limits the speed of the host vehicle within the influence zone. The influence zone may extend 5 meters, 10 meters, 20 meters, etc., from pedestrian 1215. Each increment of the influence zone may be associated with a different speed limit. For example, the host vehicle may be limited to a first speed (e.g., 10 mph, 20 mph, etc.) within a zone of 1 meter to 5 meters from pedestrian 1215.The distance to the pedestrian influence zone can be limited to a distance less than the speed limit in a pedestrian influence zone, which extends from 5 to 10 meters. Any number of increments can be used for the different stages of the influence zone. In some embodiments, the first stage can be narrower than the 1-to-5-meter range and may extend only from 1 to 2 meters. In other embodiments, the first stage of the influence zone can extend from 1 meter (the boundary of the non-navigation zone around a pedestrian) to a distance of at least 10 meters. A second stage can, in turn, extend from 10 meters to at least approximately 20 meters. The second stage can be associated with a maximum speed for the host vehicle that is greater than the maximum speed associated with the first stage of the pedestrian influence zone.

[0201] One or more stationary object constraints can also be implied by the scene detected in the vicinity of the host vehicle. For example, the at least one processing device in Figure 1201 can detect a stationary object, such as the cardboard box 1219 located on the roadway. Detected stationary objects can include various objects, such as at least one tree, a post, a road sign, or an object on a roadway. One or more predefined navigation constraints can be associated with the detected stationary object. For example, such constraints can include an envelope of the stationary object, the envelope of the stationary object defining a buffer zone around the object within which navigation by the host vehicle may be prohibited. At least part of the buffer zone can extend a predetermined distance from an edge of the detected stationary object.For example, in the scene depicted by Figure 1201, a buffer zone of at least 0.1 m, 0.25 m, 0.5 m or more can be associated with the carton 1219, so that the host vehicle will pass the carton at a distance of at least a certain amount (e.g. the buffer zone distance) to the right or left in order to avoid a collision with the detected stationary object.

[0202] The predefined hard constraints can also include one or more target vehicle constraints. For example, a target vehicle 1217 can be detected in Figure 1201. To ensure that the host vehicle does not collide with target vehicle 1217, one or more hard constraints can be used. In some cases, a target vehicle envelope can be associated with a single buffer zone distance. For example, the buffer zone can be defined as a distance of 1 meter surrounding the target vehicle in all directions. The buffer zone can define an area extending at least one meter from the target vehicle into which the host vehicle must not navigate.

[0203] The envelope curve surrounding target vehicle 1217 need not be defined by a fixed buffer distance. In some cases, the predefined hard constraints associated with target vehicles (or any other moving objects detected in the vicinity of the host vehicle) may depend on the orientation of the host vehicle relative to the detected target vehicle. In some cases, the longitudinal buffer zone distance (e.g., between the target vehicle and the front or rear of the host vehicle, such as when the host vehicle is approaching the target vehicle) may be at least one meter. A lateral buffer zone distance (e.g., one extending from the target vehicle to either side of the host vehicle—such as when the host vehicle is traveling in the same or opposite direction to the target vehicle, so that one side of the host vehicle passes alongside one side of the target vehicle) may be at least 0.5 meters.

[0204] As described above, other constraints may also be implied by the detection of a target vehicle or a pedestrian in the vicinity of the host vehicle. For example, the predicted trajectories of the host vehicle and the target vehicle 1217 may be taken into account, and where the two trajectories intersect (e.g., at intersection 1230), a hard constraint may be imposed. t1l>t2a or t1l>t2a+0.5 This is required, where the host vehicle is Vehicle 1 and the target vehicle 1217 is Vehicle 2. Similarly, the trajectory of pedestrian 1215 (based on a direction of travel and speed) can be monitored relative to the projected trajectory of the host vehicle. For a given trajectory of a pedestrian, t(p) at each point p on the trajectory represents the time the pedestrian takes to reach point p (i.e., point 1231 in ). Fig. 12) To maintain the required buffer distance of at least 1 meter from the pedestrian, either t(p) must be greater than the time the host vehicle takes to reach point p (with a sufficient time difference so that the host vehicle passes at least one meter in front of the pedestrian) or t(p) must be less than the time the host vehicle takes to reach point p (e.g., if the host vehicle brakes to yield to the pedestrian). However, the hard constraint in the latter example requires that the host vehicle reach point p a sufficient time later than the pedestrian so that the host vehicle can pass behind the pedestrian and maintain the required buffer distance of at least one meter.

[0205] Other hard constraints can also be applied. For example, a maximum deceleration rate of the host vehicle can be used in at least some cases. Such a maximum deceleration rate can be determined based on a detected distance to a target vehicle following the host vehicle (e.g., using images collected by a rear-facing camera). The hard constraints may include a mandatory stop at a sensor-detected crosswalk or railroad crossing, or other applicable restrictions.

[0206] If the analysis of a scene in the vicinity of the host vehicle reveals that one or more predefined navigation restrictions may be implied, these restrictions can be imposed relative to one or more planned navigation actions for the host vehicle. For example, if the analysis of a scene causes the Driving Policy Module 803 to return a desired navigation action, this desired navigation action can be checked against one or more implied restrictions. If it is determined that the desired navigation action violates any aspect of the implied restrictions (e.g.,If the desired navigation action would bring the host vehicle to within 0.7 meters of pedestrian 1215, where a predefined hard constraint requires the host vehicle to remain at least 1.0 meter away from pedestrian 1215, then at least one modification can be made to the desired navigation action based on one or more predefined navigation constraints. Adjusting the desired navigation action in this way can provide an actual navigation action for the host vehicle that complies with the constraints implied by a particular scene detected in the host vehicle's environment.

[0207] After determining the actual navigation action for the host vehicle, this navigation action can be implemented by initiating at least one adjustment of a navigation actuator of the host vehicle in response to the determined actual navigation action for the host vehicle. Such a navigation actuator can include at least one from a steering mechanism, a brake, or an accelerator of the host vehicle. Prioritized restrictions

[0208] As described above, various hard constraints can be used in a navigation system to ensure the safe operation of a host vehicle. These constraints can include, among other things, a minimum safety distance to a pedestrian, a target vehicle, a guardrail, or a detected object; a maximum speed when passing through the influence zone of a detected pedestrian; or a maximum deceleration rate for the host vehicle. These constraints can be imposed by a trained system based on machine learning (supervised, reinforcement, or a combination thereof), but they can also be useful for untrained systems (e.g., those that use algorithms to directly address foreseeable situations that arise in scenes from a host vehicle environment).

[0209] In both cases, there can be a hierarchy of constraints. In other words, some navigation constraints may take priority over others. Therefore, if a situation arises where no navigation action is available that would satisfy all implied constraints, the navigation system can determine the available navigation action that will satisfy the highest-priority constraints first. For example, the system might direct the vehicle to avoid a pedestrian first, even if navigating to avoid the pedestrian would result in a collision with another vehicle or an object detected on the road. In another example, the system might direct the vehicle to drive onto a curb to avoid a pedestrian.

[0210] Fig. Figure 13 provides a flowchart illustrating an algorithm for implementing a hierarchy of implied constraints based on the analysis of a scene in a host vehicle's environment. For example, at step 1301, at least one processing device associated with the navigation system (e.g., an EyeQ processor, etc.) can receive a variety of images from a camera mounted on the host vehicle, representing an environment of the host vehicle. By analyzing one or more images representative of the host vehicle's environment scene at step 1303, a navigation state associated with the host vehicle can be identified. For example, a navigation state might indicate that the host vehicle is traveling along a two-lane road 1210, as shown in Figure 1210. Fig. 12, that a target vehicle 1217 is driving through an intersection in front of the host vehicle, that a pedestrian 1215 is waiting to cross the road on which the host vehicle is driving, that an object 1219 is in front of the host vehicle on the lane, as well as various other attributes of the scene.

[0211] In step 1305, one or more navigation constraints implied by the navigation state of the host vehicle can be determined. For example, after analyzing a scene in the environment of the host vehicle, represented by one or more captured images, the at least one processing device can determine one or more navigation constraints implied by objects, vehicles, pedestrians, etc., detected through image analysis of the captured images. In some embodiments, the at least one processing device can determine at least one first predefined navigation constraint and one second predefined navigation constraint implied by the navigation state, and the first predefined navigation constraint can differ from the second predefined navigation constraint.For example, the first navigation restriction may refer to one or more target vehicles detected in the vicinity of the host vehicle, and the second navigation restriction may refer to a pedestrian detected in the vicinity of the host vehicle.

[0212] In step 1307, the at least one processing device can determine a priority associated with the constraints identified in step 1305. In the described example, the second predefined navigation constraint, which relates to pedestrians, can have a higher priority than the first predefined navigation constraint, which relates to target vehicles. While priorities associated with navigation constraints can be determined or assigned based on various factors, in some embodiments, the priority of a navigation constraint can relate to its relative importance from a safety perspective. For example, while it may be important that all implemented navigation constraints are followed or fulfilled in as many situations as possible, some constraints may be associated with greater safety risks than others and can therefore be assigned higher priorities.For example, a navigation restriction requiring the host vehicle to maintain a distance of at least one meter from a pedestrian may have a higher priority than a restriction requiring the host vehicle to maintain a distance of at least one meter from a destination vehicle. This may be because a collision with a pedestrian can have more serious consequences than a collision with another vehicle. Similarly, maintaining a distance between the host vehicle and a destination vehicle may have a higher priority than a restriction requiring the host vehicle to avoid a cardboard box in the road, to drive more slowly over a speed bump at a certain speed, or not to expose the host vehicle's occupants to more than a maximum acceleration level.

[0213] While the Driving Policy Module 803 is designed to maximize safety by fulfilling navigation constraints implied by a particular scene or navigation state, in some situations it may be physically impossible to fulfill every implied constraint. In such situations, the priority of each implied constraint can be used to determine which of the implied constraints should be fulfilled first, as shown in Step 1309. Continuing the preceding example, in a situation where it is not possible to fulfill both the pedestrian distance constraint and the target vehicle distance constraint, but only one of the constraints can be fulfilled, the higher priority of the pedestrian distance constraint may cause that constraint to be fulfilled before any attempt is made to maintain a target vehicle distance.Therefore, in normal situations, the at least one processing device can determine a first navigation action for the host vehicle based on the identified navigation state of the host vehicle. This action satisfies both the first and second predefined navigation constraints, as shown in step 1311. However, in other situations where not all implied constraints can be satisfied, the at least one processing device can determine a second navigation action for the host vehicle based on the identified navigation state. This second navigation action satisfies the second predefined navigation constraint (i.e.,the higher priority constraint), but does not satisfy the first predefined navigation constraint (which has a lower priority than the second navigation constraint), where the first predefined navigation constraint and the second predefined navigation constraint cannot both be satisfied, as shown in step 1313.

[0214] Next, in step 1315, to implement the specified navigation actions for the host vehicle, at least one processing device can cause at least one adjustment of a navigation actuator of the host vehicle in response to the specified first navigation action or the specified second navigation action for the host vehicle. As in the previous example, the navigation actuator can include at least one from a steering mechanism, a brake, or an accelerator. Easing of restrictions

[0215] As discussed above, navigation restrictions can be imposed for safety reasons. These restrictions can include, among other things, a minimum safety distance to a pedestrian, a target vehicle, a guardrail, or a detected object; a maximum speed when passing through the influence zone of a detected pedestrian; or a maximum deceleration rate for the host vehicle. These restrictions can be imposed in a learning or non-learning navigation system. In certain situations, these restrictions can be relaxed. For example, if the host vehicle slows down or stops near a pedestrian and then slowly resumes driving to convey the intention to pass the pedestrian, the captured images can be used to detect a reaction from the pedestrian.If the pedestrian's reaction is to stop or remain motionless (and / or if eye contact with the pedestrian is detected by the sensor), it can be assumed that the pedestrian recognizes the navigation system's intention to drive around them. In such situations, the system can relax one or more predefined restrictions and implement a less stringent limit (e.g., allowing the vehicle to navigate within 0.5 meters of a pedestrian instead of a stricter 1-meter limit).

[0216] Fig. Figure 14 provides a flowchart for implementing the control of the host vehicle based on the relaxation of one or more navigation constraints. In step 1401, the at least one processing device can receive a multitude of images from a camera associated with the host vehicle, which are representative of the host vehicle's environment. The analysis of the images in step 1403 can enable the identification of a navigation state associated with the host vehicle. In step 1405, the at least one processor can determine navigation constraints associated with the navigation state of the host vehicle. The navigation constraints can include an initial predefined navigation constraint implied by at least one aspect of the navigation state. In step 1407, the analysis of the multitude of images can reveal the presence of at least one navigation constraint relaxation factor.

[0217] A navigation restriction relaxation factor may include any suitable indicator that one or more navigation restrictions may be suspended, modified, or otherwise relaxed in at least one aspect. In some embodiments, the at least one navigation restriction relaxation factor may include a determination (based on image analysis) that a pedestrian's eyes are looking in the direction of the host vehicle. In such cases, it can be assumed with greater confidence that the pedestrian perceives the host vehicle. Consequently, a higher level of confidence may be established that the pedestrian will not take any unexpected actions that would cause the pedestrian to move into the path of the host vehicle. Other restriction relaxation factors may also be used.For example, the at least one navigation restriction relaxation factor may include: a pedestrian determined to remain stationary (e.g., a pedestrian deemed less likely to enter the host vehicle's path); or a pedestrian determined to slow down. The navigation restriction relaxation factor may also include more complex actions, such as a pedestrian determined to remain stationary after the host vehicle has come to a stop and then resumed its movement. In such a situation, the pedestrian may be deemed to understand that the host vehicle has the right of way, and the pedestrian's remaining stationary may indicate an intention to yield to the host vehicle.Other situations that may cause one or more restrictions to be relaxed include the type of curb (e.g., a low curb or one with a gentle slope might permit a relaxed distance restriction), the absence of pedestrians or other objects on the sidewalk, a vehicle with its engine off being allowed a relaxed distance, or a situation in which a pedestrian is looking away and / or moving away from the area toward which the host vehicle is heading.

[0218] If the presence of a navigation constraint relaxation factor is identified (e.g., at step 1407), a second navigation constraint can be determined or developed in response to the detection of the restriction relaxation factor. This second navigation constraint can differ from the first navigation constraint and can include at least one property that is relaxed compared to the first navigation constraint. The second navigation constraint can include a newly generated constraint based on the first constraint, where the newly generated constraint includes at least one modification that relaxes the first constraint in at least one respect. Alternatively, the second constraint can be a predetermined constraint that is less strict than the first navigation constraint in at least one respect.In some embodiments, such second constraints may be reserved for use only in situations where a constraint relaxation factor is identified in the host vehicle's environment. Regardless of whether the second constraint is newly generated or selected from a set of fully or partially available predetermined constraints, the application of a second navigation constraint instead of a stricter first navigation constraint (which may be applied in the absence of the detection of relevant navigation constraint relaxation factors) may be termed constraint relaxation and may be achieved in step 1409.

[0219] If at least one restriction relaxation factor is detected in step 1407 and at least one restriction has been relaxed in step 1409, a navigation action for the host vehicle can be determined in step 1411. The navigation action for the host vehicle can be based on the identified navigation state and satisfy the second navigation restriction. The navigation action can be implemented in step 1413 by initiating at least one adjustment of a navigation actuator of the host vehicle in response to the determined navigation action.

[0220] As discussed above, the use of navigation constraints and relaxed navigation constraints can be employed in navigation systems that are either trained (e.g., through machine learning) or untrained (e.g., systems programmed to respond with predetermined actions in response to specific navigation states). When trained navigation systems are used, the availability of relaxed navigation constraints for certain navigation situations can represent a mode switch from a trained system response to an untrained system response. For example, a trained navigation network might determine an initial navigation action for the host vehicle based on the first navigation constraint. However, the action taken by the vehicle might differ from the navigation action that satisfies the first navigation constraint.Instead, the action taken can satisfy the relaxed second navigation constraint and be an action developed by an untrained system (e.g., in response to the detection of a specific condition in the environment of the host vehicle, such as the presence of a navigation constraint relaxation factor).

[0221] There are many examples of navigation constraints that can be relaxed in response to the detection of a constraint relaxation factor in the host vehicle's environment. For instance, if a predefined navigation constraint includes a buffer zone associated with a detected pedestrian, and at least part of the buffer zone extends at a distance from the detected pedestrian, a relaxed navigation constraint (either newly created, retrieved from memory from a predetermined set, or generated as a relaxed version of an existing constraint) can include a different or modified buffer zone. For example, the different or modified buffer zone might have a distance relative to the pedestrian that is smaller than the distance of the original or unmodified buffer zone relative to the detected pedestrian.As a result, the host vehicle may be permitted to navigate closer to a detected pedestrian, given the relaxed restriction, if a suitable restriction relaxation factor is detected in the vicinity of the host vehicle.

[0222] A relaxed property of a navigation constraint can include a reduced width in a buffer zone associated with at least one pedestrian, as noted above. However, the relaxed property can also include a reduced width in a buffer zone associated with a target vehicle, a detected object, a road boundary, or any other object detected in the vicinity of the host vehicle.

[0223] The at least one relaxed property can also include other types of modifications to the navigation restriction properties. For example, the relaxed property can include an increase in the speed associated with at least one predefined navigation restriction. The relaxed property can also include an increase in the maximum permissible deceleration / acceleration associated with at least one predefined navigation restriction.

[0224] While restrictions can be relaxed in certain situations, as described above, navigation restrictions can be extended in other situations. For example, a navigation system may determine in some situations that conditions warrant an extension of a normal set of navigation restrictions. Such an extension may involve adding new restrictions to a predefined set of restrictions or adjusting one or more aspects of a predefined restriction. The addition or adjustment may result in more conservative navigation compared to the predefined restrictions that apply under normal driving conditions. Conditions that may justify an extension of the restrictions include sensor failure, adverse environmental conditions (rain, snow, fog, or other conditions associated with limited visibility or reduced vehicle traction), and so on.include.

[0225] Fig. Figure 15 provides a flowchart for implementing the control of the host vehicle based on the extension of one or more navigation constraints. In step 1501, the at least one processing device can receive a multitude of images from a camera associated with the host vehicle, which are representative of the host vehicle's environment. The analysis of the images in step 1503 can enable the identification of a navigation state associated with the host vehicle. In step 1505, the at least one processor can determine navigation constraints associated with the navigation state of the host vehicle. The navigation constraints can include an initial predefined navigation constraint implied by at least one aspect of the navigation state. In step 1507, the analysis of the multitude of images can reveal the presence of at least one navigation constraint extension factor.

[0226] An implied navigation restriction can be any of the navigation restrictions discussed above (e.g., regarding Fig. 12) or any other suitable navigation constraint. A navigation constraint extension factor may include any indicator that one or more navigation constraints can be augmented / extended in at least one aspect. The augmentation or extension of navigation constraints may be performed on a per-set basis (e.g., by adding new navigation constraints to a predetermined set of constraints) or may be performed on a per-constraint basis (e.g., by modifying a particular constraint so that the modified constraint is more restrictive than the original, or by adding a new constraint corresponding to a predetermined constraint, where the new constraint is more restrictive than the corresponding constraint in at least one aspect).Additionally or alternatively, the addition or extension of navigation constraints can involve selection from a set of predefined constraints based on a hierarchy. For example, a set of extended constraints can be available for selection based on whether a navigation extension factor is detected in the environment of the host vehicle or relative to it. Under normal conditions, where no extension factor is detected, the implied navigation constraints can be derived from constraints that apply under normal conditions. On the other hand, if one or more constraint extension factors are detected, the implied constraints can be derived from the extended constraints that were either generated or predefined with respect to the one or more extension factors.The extended restrictions may be more restrictive in at least one respect than the corresponding restrictions that apply under normal conditions.

[0227] In some embodiments, the at least one navigation restriction enhancement factor may include detection (e.g., based on image analysis) of the presence of ice, snow, or water on a road surface in the vicinity of the host vehicle. Such a determination may, for example, be based on the detection of: areas with a higher reflectance than expected for dry road surfaces (e.g., an indication of ice or water on the road surface); white areas on the road surface indicating the presence of snow; shadows on the road surface corresponding to the presence of longitudinal ruts (e.g., tire tracks in the snow) on the road surface; water droplets or ice / snow particles on a windshield of the host vehicle; or any other suitable indicator of the presence of water or ice / snow on a road surface.

[0228] The at least one navigation restriction enhancement factor may also include the detection of particles on an outer surface of the host vehicle's windshield. Such particles can impair the image quality of one or more image-capturing devices associated with the host vehicle. While the description refers to the host vehicle's windshield, which is relevant for cameras mounted behind the windshield, the detection of particles on other surfaces (e.g., a camera lens or lens cover, a headlight lens, a rear window, a taillight lens, or any other surface of the host vehicle visible to (or detected by) an image-capturing device associated with the host vehicle) may also indicate the presence of a navigation restriction enhancement factor.

[0229] The navigation restriction enhancement factor can also be recognized as an attribute of one or more image acquisition devices. For example, a detected decrease in the image quality of one or more images captured by an image acquisition device (e.g., a camera) associated with the host vehicle can also represent a navigation restriction enhancement factor. Deterioration in image quality can be associated with a hardware failure or partial hardware failure of the image acquisition device or an assembly associated with the image acquisition device. Such deterioration in image quality can also be caused by environmental conditions. For example, the presence of smoke, fog, rain, snow, etc., in the air surrounding the host vehicle can also lead to reduced image quality with respect to the road, pedestrians, target vehicles, etc.contribute, which may be present in the vicinity of the host vehicle.

[0230] The navigation restriction extension factor can also relate to other aspects of the host vehicle. For example, in some situations, the navigation restriction extension factor may include a detected failure or partial failure of a system or sensor associated with the host vehicle. Such an extension factor might include, for example, the detection of a failure or partial failure of a speed sensor, GPS receiver, accelerometer, camera, radar, lidar, brakes, tires, or any other system associated with the host vehicle that could affect the host vehicle's ability to navigate with respect to navigation restrictions associated with a navigation state of the host vehicle.

[0231] If the presence of a navigation restriction extension factor is identified (e.g., at step 1507), a second navigation restriction can be determined or developed in response to the restriction extension factor detection. This second navigation restriction may differ from the first navigation restriction and may include at least one feature that is extended compared to the first navigation restriction. The second navigation restriction may be more restrictive than the first navigation restriction because the detection of a restriction extension factor in or in association with the host vehicle suggests that the host vehicle has at least one navigation capability that is limited compared to normal operating conditions. Such limited capabilities may include reduced road grip (e.g., ice, snow, or water on the road surface; reduced tire pressure, etc.).), include limited visibility (e.g., rain, snow, dust, smoke, fog, etc., which impair the quality of the captured image), limited detection capability (e.g., sensor failure or partial failure, reduced sensor performance, etc.), or any other reduction in the host vehicle's ability to navigate in response to a detected navigation condition.

[0232] If at least one constraint extension factor is detected in step 1507 and at least one constraint has been extended in step 1509, a navigation action for the host vehicle can be determined in step 1511. The navigation action for the host vehicle can be based on the identified navigation state and can satisfy the second (i.e., extended) navigation constraint. The navigation action can be implemented in step 1513 by initiating at least one adjustment of a navigation actuator of the host vehicle in response to the determined navigation action.

[0233] As discussed, the use of navigation constraints and extended navigation constraints can be employed with navigation systems that are either trained (e.g., through machine learning) or untrained (e.g., systems programmed to respond with predetermined actions in response to specific navigation states). When trained navigation systems are used, the availability of extended navigation constraints for certain navigation situations can represent a mode switch from a trained system response to an untrained system response. For example, a trained navigation network might determine an initial navigation action for the host vehicle based on the first navigation constraint. However, the action taken by the vehicle might differ from the navigation action that satisfies the first navigation constraint.Instead, the action taken can satisfy the extended second navigation constraint and be an action developed by an untrained system (e.g., in response to the detection of a certain condition in the environment of the host vehicle, such as the presence of a navigation constraint extension factor).

[0234] There are many examples of navigation constraints that can be created, augmented, or extended in response to the detection of a constraint extension factor in the host vehicle's environment. For example, if a predefined navigation constraint includes a buffer zone associated with a detected pedestrian, object, vehicle, etc., and at least part of the buffer zone extends at a distance from the detected pedestrian / object / vehicle, an extended navigation constraint (either newly developed, retrieved from memory from a predetermined set, or created as an extended version of an existing constraint) can include a different or modified buffer zone.For example, the other or modified buffer zone may have a distance relative to the pedestrian / object / vehicle that is greater than the distance of the original or unmodified buffer zone relative to the detected pedestrian / object / vehicle. As a result, the host vehicle may be forced to navigate further away from the detected pedestrian / object / vehicle due to the extended constraint if a suitable constraint extension factor is detected in the vicinity of the host vehicle or relative to the host vehicle.

[0235] The at least one extended property can also include other types of modifications to the navigation restriction properties. For example, the extended property can include a reduction in the speed associated with at least one predefined navigation restriction. The extended property can also include a reduction in the maximum permissible deceleration / acceleration associated with at least one predefined navigation restriction. Navigation based on long-term planning

[0236] In some embodiments, the disclosed navigation system can not only react to a detected navigation state in the environment of the host vehicle, but also determine one or more navigation actions based on long-term planning. For example, the system can consider the potential impact of one or more navigation actions, available as navigation options with respect to a detected navigation state, on future navigation states. Considering the impact of available actions on future states allows the navigation system to determine navigation actions not only based on a currently detected navigation state, but also based on long-term planning.Navigation using long-term planning techniques can be particularly relevant where one or more reward functions are employed by the navigation system as a technique for selecting navigation actions from available options. Potential rewards can be analyzed in relation to the available navigation actions that can be taken in response to a detected, current navigation state of the host vehicle. Furthermore, potential rewards can also be analyzed relative to actions that can be taken in response to future navigation states that result from the available actions for a current navigation state.Consequently, in some cases, the disclosed navigation system may select a navigational action in response to a detected navigational state, even if the selected action might not provide the highest reward among the available actions that can be taken in response to the current navigational state. This may be particularly true if the system determines that the selected action could lead to a future navigational state that has one or more potential navigational actions offering higher rewards than the selected action or, in some cases, any of the actions available with respect to a current navigational state. The principle can be expressed more simply as taking a less favored action now in order to generate higher-reward options in the future.Therefore, the revealed navigation system, which is capable of long-term planning, may choose a suboptimal short-term action if a long-term forecast indicates that a short-term loss of reward may lead to long-term reward gains.

[0237] In general, autonomous driving applications can include a number of planning problems where the navigation system can decide on immediate actions to optimize a longer-term goal. For example, if a vehicle encounters a merging situation at a roundabout, the navigation system can issue an immediate acceleration or braking command to initiate navigation into the roundabout. While the immediate response to the detected navigation condition at the roundabout might include an acceleration or braking command in response to the detected condition, the long-term goal is a successful merge, and the long-term effect of the chosen command is the success or failure of the merge. The planning problem can be approached by decomposing the problem into two phases.First, supervised learning can be applied to predict the near future based on the present (assuming the predictor is differentiable with respect to its representation of the present). Second, a complete trajectory of the agent can be modeled using a recurrent neural network, with unexplained factors modeled as (additive) input nodes. This can allow solutions to the long-term planning problem to be determined using supervised learning and direct optimization techniques via the recurrent neural network. Such an approach can also enable the learning of robust policies by incorporating adversarial elements into the environment.

[0238] Two of the most fundamental elements of autonomous driving systems are sensor acquisition and planning. Sensor acquisition focuses on finding a concise representation of the current state of the environment, while planning deals with deciding which actions to take to optimize future objectives. Supervised machine learning techniques are useful for solving sensor acquisition problems. For the planning component, algorithmic machine learning frameworks, particularly reinforcement learning (RL) frameworks like the ones described above, can also be used.

[0239] RL can be performed in a sequence of successive rounds. In round t, the planner (also known as the agent or the Driving Guide Module 803) can define a state, s t∈ S, observe, representing both the agent and the environment. The agent should then choose an action a ∈ A. After performing the action, the agent receives an immediate reward, r. t ∈ ℝ, and is transformed into a new state, s t+1 , offset. For example, the host vehicle might include an adaptive cruise control (ACC) system where the vehicle should autonomously accelerate / brake to maintain an appropriate distance to a vehicle ahead while maintaining a smooth driving style. The state can be described as a pair s t = (s t , υ t ) ∈ ℝ 2 be modeled, where x t the distance to the vehicle in front is and v t The speed of the host vehicle relative to the speed of the vehicle in front. The action a t∈ ℝ is the acceleration command (where the host vehicle slows down when a t < 0). The reward can be a function that depends on |a t The process depends on the smoothness of the journey (reflecting the smoothness of the journey) and on the distance the host vehicle maintains from the vehicle ahead (reflecting that the host vehicle maintains a safe distance). The planner's goal is to maximize the cumulative reward (perhaps up to a certain time horizon or a discounted sum of future rewards). To achieve this, the planner can rely on a guideline π : S → A, which maps a state to an action.

[0240] Supervised learning (SL) can be seen as a special case of RL, where st is sampled from a certain distribution over S and the reward function has the form r t = -ℓ(a t , y t ) can exhibit, where is a loss function and the learner determines the value of y. tobserved, which is the (possibly noisy) value of the optimal action that would be taken when considering the state s t to grasp. There can be several differences between a general RL model and a specific case of SL, and these differences can make the general RL problem more difficult.

[0241] In some SL situations, the actions (or predictions) taken by the learner may have no effect on the environment. In other words, s t+1 and a t are independent. This can have two important implications.

[0242] Firstly, in SL, a sample (s1, y1 ..... s) can be taken. m , y m ) must be collected in advance, and only then can the search begin for a guideline (or predictor) that has good accuracy with respect to the sample. In contrast, with RL, the state s depends on t+1This generally depends on the action taken (and also on the previous state), which in turn depends on the policy used to generate the action. This links the data generation process with the policy learning process. Secondly, since actions in SL do not affect the environment, the contribution of choosing a t The power of π is locally affected. Specifically, a t only the value of the immediate reward. In contrast, with RL, actions taken in round t can have a long-term impact on reward values ​​in future rounds.

[0243] In SL, knowledge of the "correct" answer, y t , together with the form of the reward r t = -ℓ(a t , y t ) a complete knowledge of the reward for all possible choices of a t provide what is needed to calculate the derivative of the reward in relation to a tThis can enable [further explanation needed]. In contrast, with RL, a "one-off" reward value can be anything that can be observed for a specific choice of action taken. This can be described as "bandit" feedback. This is one of the most important reasons for the necessity of "exploration" as part of long-term navigation planning, because if only "bandit" feedback is available in RL-based systems, the system may not always know whether the action taken was the best one.

[0244] Many real-world algorithms rely, at least in part, on the mathematically elegant model of a Markov decision process (MDP). The Markov assumption states that the distribution of s t+1 is fully determined when s t and a tare given. This results in a closed-form expression for the cumulative reward of a given policy in the form of the stationary distribution over the states of the MDP. The stationary distribution of a policy can be expressed as the solution of a linear programming problem. This leads to two families of algorithms: 1) optimization with respect to the primal problem, which can also be called policy search, and 2) optimization with respect to a dual problem whose variables are value functions V. π The value function determines the expected cumulative reward when the MDP starts from the initial state s and chooses actions according to π. A related quantity is the state-action value function Q. π(s, a), which determines the cumulative reward if one starts from state s, performs an immediately chosen action a, and from then on chooses actions according to π. The Q-function can lead to a characterization of the optimal policy (using the Bellman equation). In particular, the Q-function can show that the optimal policy is a deterministic function from S to A (in fact, it can be characterized as a "greedy" policy with respect to the optimal Q-function).

[0245] A potential advantage of the MDP model is that it allows the coupling of the future with the present using the Q function. For example, if a host vehicle is currently in state s, the value of Q can be π(s,a) indicate the impact of performing action a on the future. Therefore, the Q function can provide a local measure of the quality of action a, making the RL problem more similar to an SL scenario.

[0246] Many RL algorithms approximate the V-function or the Q-function in one way or another. Value iteration algorithms, such as the Q-learning algorithm, may rely on the fact that the V- and Q-functions of the optimal policy can be fixed points of some operators derived from the Bellman equation. Actor-critique policy iteration algorithms aim to learn a policy iteratively, where the "critique" is Q. πt At iteration t, the value is estimated, and the "actor" improves the policy based on this estimate.

[0247] Despite the mathematical elegance of MDPs and the convenience of converting to Q-function representation, this approach can have several limitations. For example, in some cases, an approximate idea of ​​a Markovian behavioral state may be all that can be found. Furthermore, the transition between states may depend not only on the agent's actions but also on the actions of other actors in the environment. For instance, in the aforementioned ACC example, while the dynamics of the autonomous vehicle may be Markovian, the next state may depend on the behavior of the driver of the other car, which is not necessarily Markovian. One possible solution to this problem is to use partially observed MDPs, where it is assumed that a Markovian state exists, but an observation distributed according to the hidden state is what can be seen.

[0248] A more direct approach can consider game-theoretic generalizations of MDPs (e.g., the Stochastic Games framework). Indeed, algorithms for MDPs can be generalized to multi-agent games (e.g., minimax Q learning or Nash Q learning). Other approaches may involve explicit modeling of the other actors and vanishing-regret learning algorithms. Learning in a multi-agent environment can be more complex than in a single-agent environment.

[0249] A second limitation of the Q-function representation can arise when deviating from a tabular representation. A tabular representation is appropriate when the number of states and actions is small, and Q can therefore be expressed as a table with |S| rows and |A| columns. However, if the natural representation of S and A involves Euclidean spaces, and the state and action spaces are discretized, the number of states / actions can be exponential with dimension. In such cases, using a tabular representation can be impractical. Instead, the Q-function can be approximated by a certain function from a class of parametric hypotheses (e.g., neural networks of a particular architecture). For example, a deep Q network (DQN) learning algorithm can be used. In the DQN, the state space can be continuous, but the action space can remain a small discrete set.While there are approaches to dealing with continuous action spaces, they may rely on an approximation of the Q-function. In any case, the Q-function can be complex and susceptible to noise, and therefore difficult to learn.

[0250] Another approach might be to tackle the RL problem using a recurrent neural network (RNN). In some cases, RNNs can be combined with game-theoretic concepts such as multi-agent games and robustness to adverse environments. Furthermore, this approach can be one that does not explicitly rely on any arbitrary Markov assumptions.

[0251] The following section describes in more detail an approach for navigation by planning based on prediction. In this approach, it can be assumed that the state space S is a subset of ℝ. d is, and the action space A is a subset of ℝk This can be a natural representation in many applications. As mentioned above, there are two key differences between RL and SL: (1) because past actions affect future rewards, information from the future may need to be propagated back into the past; and (2) the "bandit" nature of rewards can obscure the dependency between (state, action) and reward, which can complicate the learning process.

[0252] As a first step in this approach, one could observe that there are interesting problems where the banditry nature of rewards is not an issue. For example, the reward value (as will be discussed in more detail below) for the ACC application can be differentiable with respect to the current state and action. Indeed, even if the reward is given in a "banditry" manner, the problem of learning a differentiable function r̂(s, a) such that r̂(s) t , a t ) ≈ r tThis is considered to be a relatively simple SL problem (e.g., a one-dimensional regression problem). Therefore, the first step of the approach can be either to define the reward as a function r̂(s, a) that is differentiable with respect to s and a, or to use a regression learning algorithm to learn a differentiable function r̂ that minimizes at least one regression loss over a sample, where the instance vector (s) t , a t ) ∈ ℝ d × ℝ k is and the target scale r t In some situations, elements of exploration can be used to create a training set.

[0253] To establish a connection between the past and the future, a similar idea can be used. For example, suppose that a differentiable function N̂(s, a) can be learned such that N̂(s t , a t ) ≈ s t+1Learning such a function can be called a SL problem. N̂ can be viewed as a predictor for the near future. Next, a guideline mapping S to A can be derived using a parametric function π. θ : S → A can be described. Expressing π θ As a neural network, it can enable the expression of an episode of the agent's execution for T rounds using a recurrent neural network (RNN), where the next state is represented as s t+1 = N̂(s t , a t ) + v t is defined here. t ∈ ℝ d It is defined by its environment and can express unpredictable aspects of the near future. The fact that s t+1 in a distinguishable way from s t and a t This can enable a connection between future reward values ​​and past actions. A parameter vector of the policy function π θLearning can be achieved through backpropagation via the resulting RNN. It should be noted that no explicit probabilistic assumptions are made. t must be imposed. In particular, it is not necessary for a Markovian relationship to exist. Instead, the recurrent network can be used to propagate "sufficient" information between past and future. Intuitively, N̂ (s t , a t ) describe the predictable part of the near future, while v t It can express the unpredictable aspects that may arise from the behavior of other actors in the environment. The learning system should learn a guideline that is robust against the behavior of the other actors. If ||v tIf the value of N̂ is large, the connection between past actions and future reward values ​​may be too noisy for learning a meaningful guideline. Explicitly representing the system dynamics in a transparent manner can facilitate the incorporation of prior knowledge. For example, prior knowledge can simplify the problem of defining N̂.

[0254] As discussed above, the learning system can benefit from robustness against an adversarial environment, such as the environment of a host vehicle, which may include several other drivers who might behave unexpectedly. In a model that makes no probabilistic assumptions on v t imposed, environments can be considered in which v t is chosen in an opposing way. In some cases, restrictions on µ may apply. tThis applies because otherwise the opponent could complicate or even prevent the planning problem. A natural limitation can be that ||µ t || is limited by a constant.

[0255] Robustness against adverse environments can be useful in autonomous driving applications. Choosing µ t In an adversarial way, it can even accelerate the learning process, as it can align the learning system with a robust, optimal guideline. This concept can be illustrated using a simple game. The state is s t ∈ ℝ, the action is a t ∈ ℝ, and the immediate loss function is 0.1 |a t | + [|s t | - 2] + , where [x] + The ReLU (Rectified Linear Unit) function is max{x.0}. The next state is s t+1 = s t + at + v t , where v t∈ [-0,5, 0,5] for the neighborhood is chosen in an adversarial way. Here, the optimal guideline can be a two-layer net with ReLU: a t = -[s t - 1.5] + [-s t - 1.5] + be written. It should be noted that with |s t | ∈ (1,5, 2], the optimal action may have a greater immediate loss than the action a = 0. Therefore, the system can plan for the future and does not have to rely solely on the immediate loss. It should be noted that the derivative of the loss with respect to a t 0,1 sign(a t ) is, and the derivative with respect to s t 1[|s t | > 2] sign(s t ) is. In a situation where s t ∈ (1,5. 2], would be the opponent's choice of v t , v t to set = 0.5, and therefore there can be a non-zero loss in round t + 1 whenever a t > 1.5 - s tIn such cases, the derivative of the loss can be directly applied to a t propagate back. Therefore, the opposing choice of v t to help the navigation system in cases where the choice of a t It is suboptimal to obtain a backpropagation that deviates from zero. Such a relationship can help the navigation system select present actions based on the expectation that such a present action (even if it would lead to a suboptimal reward or even a loss) will provide opportunities for more optimal actions in the future that will lead to higher rewards.

[0256] Such an approach can be applied to virtually any navigation situation that may arise. The following describes the approach using an example: adaptive cruise control (ACC). In the ACC problem, the host vehicle may attempt to maintain an appropriate distance from a target vehicle ahead (e.g., 1.5 seconds from the target vehicle). Another objective may be to drive as smoothly as possible while maintaining the desired distance. A model representing this situation can be defined as follows. The state space is ℝ 3, and the action space is ℝ. The first coordinate of the state is the velocity of the target vehicle, the second coordinate is the velocity of the host vehicle, and the last coordinate is the distance between the host vehicle and the target vehicle (e.g., the location of the host vehicle minus the location of the target vehicle along the curve of the road). The action to be taken by the host vehicle is acceleration, and can be expressed as a t The quantity τ can denote the time difference between successive laps. While τ can be set to any suitable value, in one example τ could be 0.1 seconds. The position s t can be st=(vtZiel,vtHost,xt) can be described as, and the (unknown) acceleration of the target vehicle can be described as atTarget be designated.

[0257] The entire dynamics of the system can be described by: vtTarget=[vt−1Target+τ at−1Target]+vtHost=[vt−1Host+τ at−1]+xt=[xt−1+τ(vt−1Target−vt−1Host)]+

[0258] This can be described as the sum of two vectors: st=([st−1[0]+τat−1target]+,[st−1[1]+τat−1]+,[st−1[2]+τ(st−1[0]−st−1[1])]+)=(st−1[0],[st−1[1 ]+τat−1]+,[st−1[2]+τ(st−1[0]+st−1[1])]+)︸N^(st−1,at)+([st−1[0]+τat−1target]+−st−1[0],0,0)︸vt

[0259] The first vector is the predictable part, and the second vector is the unpredictable part. The reward in round t is defined as follows: −rt=0.1|at|+[|xt / xt*−1|−0.3]+ where xt*=max{1,1,5 vtHost}

[0260] The first term can result in a penalty for non-zero accelerations, thus encouraging smooth driving. The second term depends on the ratio between the distance to the target vehicle x. t and the desired distance xt* from, which is defined as the maximum between a distance of 1 meter and a braking distance of 1.5 seconds. In some cases, this ratio can be exactly 1, but as long as this ratio is within [0,7, 1,3], the policy can waive penalties, giving the host vehicle some leeway in navigation – a feature that can be important for achieving a smooth ride.

[0261] By implementing the approach outlined above, the host vehicle's navigation system (e.g., by operating the driving guidance module 803 within the navigation system's processing unit 110) can select an action in response to an observed state. The selected action can be based not only on the analysis of the rewards associated with the available response actions relative to a sensor-detected navigation state, but also on the consideration and analysis of future states, potential actions in response to those future states, and the rewards associated with those potential actions.

[0262] Fig. Figure 16 illustrates an algorithmic approach to navigation based on detection and long-term planning. For example, in step 1601, the at least one processing device 110 of the navigation system for the host vehicle can receive a variety of images. These images can capture scenes representative of the host vehicle's environment and can be supplied by any of the image acquisition devices described above (e.g., cameras, sensors, etc.). Analyzing one or more of these images in step 1603 can enable the at least one processing device 110 to identify an existing navigation state associated with the host vehicle (as described above).

[0263] In steps 1605, 1607, and 1609, various potential navigation actions can be determined in response to the sensor-detected navigation state. These potential navigation actions (e.g., an initial navigation action up to an Nth available navigation action) can be determined based on the sensor-detected state and the long-term goals of the navigation system (e.g., to complete a merge, smoothly follow a vehicle ahead, overtake a target vehicle, avoid an object in the roadway, slow down for a detected stop sign, avoid a merging target vehicle, or any other navigation action that can advance the system's navigation goals).

[0264] For each of the identified potential navigation actions, the system can determine an expected reward. The expected reward can be determined according to any of the techniques described above and may include the analysis of a particular potential action relative to one or more reward functions. The expected rewards 1606, 1608, and 1610 can be determined for each of the potential navigation actions (e.g., the first, second, and Nth) identified in steps 1605, 1607, and 1609, respectively.

[0265] In some cases, the host vehicle's navigation system can select from available potential actions based on values ​​associated with expected rewards 1606, 1608, and 1610 (or any other type of expected reward indicator). For example, in some situations, the action that yields the highest expected reward can be selected.

[0266] In other cases, particularly when the navigation system is engaged in long-term planning to determine navigation actions for the host vehicle, the system may not select the potential action that yields the highest expected reward. Rather, the system may look into the future to analyze whether there are opportunities to achieve higher rewards later by choosing lower-reward actions in response to a current navigation state. For example, a future state may be determined for any or all of the potential actions identified at steps 1605, 1607, and 1609. Each future state identified at steps 1613, 1615, and 1617 may represent a future navigation state based on the current navigation state determined by a given potential action (e.g.,which was modified in steps 1605, 1607 and 1609 (the potential actions determined).

[0267] For each of the future states predicted in steps 1613, 1615, and 1617, one or more future actions (as navigation options available in response to the specific future state) can be determined and evaluated. For steps 1619, 1621, and 1623, for example, values ​​or any other type of indicator for expected rewards associated with one or more of the future actions can be developed (e.g., based on one or more reward functions). The expected rewards associated with one or more future actions can be evaluated by comparing values ​​of reward functions associated with each future action or by comparing other indicators associated with the expected rewards.

[0268] At step 1625, the host vehicle's navigation system can select a navigation action for the host vehicle based on a comparison of expected rewards. This selection is based not only on potential actions identified relative to a current navigation state (e.g., at steps 1605, 1607, and 1609), but also on expected rewards determined as a result of potential future actions available in response to predicted future states (e.g., determined at steps 1613, 1615, and 1617). The selection at step 1625 can be based on the options and reward analysis performed at steps 1619, 1621, and 1623.

[0269] The selection of a navigation action at step 1625 can be based on a comparison of expected rewards associated solely with future action options. In such a case, the navigation system can select an action for the current state based solely on a comparison of expected rewards resulting from actions for potential future navigation states. For example, the system can select the potential action identified at step 1605, 1607, or 1609 that is associated with the highest future reward value, as determined by the analysis at steps 1619, 1621, and 1623.

[0270] The selection of a navigation action at step 1625 can also be based solely on a comparison of the current action options (as noted above). In this situation, the navigation system can select the potential action identified at step 1605, 1607, or 1609 that is associated with the highest expected reward, 1606, 1608, or 1610, respectively. Such a selection can be made with little or no consideration of future navigation states or future expected rewards for navigation actions available in response to those future navigation states.

[0271] On the other hand, the selection of a navigation action at step 1625 may, in some cases, be based on a comparison of the expected rewards associated with both future and current action options. This may indeed be one of the principles of navigation based on long-term planning. For example, expected rewards for future actions may be analyzed to determine whether any can justify selecting a lower-reward action in response to the current navigation state in order to achieve a potentially higher reward in response to a subsequent navigation action expected to be available in response to future navigation states. As an example, a value or other indicator for expected reward 1606 may indicate the highest expected reward among rewards 1606, 1608, and 1610.On the other hand, the expected reward 1608 may indicate the lowest expected reward among rewards 1606, 1608, and 1610. Instead of simply selecting the potential action determined at step 1605 (i.e., the action that elicits the highest expected reward 1606), an analysis of future states, potential future actions, and future rewards can be used when making a navigation action selection at step 1625. For example, it may be determined that a reward identified at step 1621 (as a response to at least one future action leading to a future state determined at step 1615 based on the second potential action identified at step 1607) may be higher than the expected reward 1606.Based on this comparison, the second potential action determined at step 1607 can be selected instead of the first potential action determined at step 1605, even though the expected reward in step 1606 is higher than the expected reward in step 1608. For example, the potential navigation action determined at step 1605 might include merging in front of a detected target vehicle, while the potential navigation action determined at step 1607 might include merging behind the target vehicle.While the expected reward 1606 for merging in front of the target vehicle may be higher than the expected reward 1608 associated with merging behind the target vehicle, it can be determined that merging behind the target vehicle may lead to a future state for which there may be action options that provide even higher potential rewards than the expected reward 1606, 1608, or other rewards based on available actions in response to a current, sensor-detected navigation state.

[0272] The selection from the potential actions in step 1625 can be based on a suitable comparison of expected rewards (or any other metric or indicator of the utility associated with one potential action versus another). In some cases, as described above, a second potential action may be chosen over a first if the second potential action is expected to provide at least one future action with an expected reward higher than the reward associated with the first potential action. In other cases, more complex comparisons may be used. For example, rewards associated with action options in response to predicted future states may be compared with more than one expected reward associated with a particular potential action.

[0273] In some scenarios, actions and expected rewards based on predicted future states can influence the selection of a potential action for a current state if at least one of the future actions is expected to provide a higher reward than any of the rewards expected as a result of the potential actions for a current state (e.g., expected rewards 1606, 1608, 1610, etc.). In some cases, the future action option that provides the highest expected reward (e.g., among the expected rewards associated with potential actions for a sensor-detected current state, as well as among the expected rewards associated with potential future action options relative to potential future navigation states) can be used as a guide for selecting a potential action for a current navigation state.That is, after identifying a future action option that provides the highest expected reward (or a reward above a predetermined threshold, etc.), the potential action that would lead to the future state associated with the identified future action providing the highest expected reward can be selected at step 1625.

[0274] In other cases, the selection of available actions can be based on specific differences between expected rewards. For example, a second potential action, determined at step 1607, can be selected if the difference between an expected reward associated with a future action, determined at step 1621, and the expected reward in step 1606 is greater than the difference between the expected reward in step 1608 and the expected reward in step 1606 (assuming differences are denoted by a + sign).In another example, a second potential action, determined at step 1607, can be selected if the difference between an expected reward associated with a future action, determined at step 1621, and an expected reward associated with a future action, determined at step 1619, is greater than the difference between an expected reward in 1608 and an expected reward in 1606.

[0275] Several examples of selecting from potential actions for a current navigation state have been described. However, any other suitable comparison techniques or criteria can also be used to select an available action through long-term planning based on an action and reward analysis that extends to predicted future states. Fig. While the analysis of long-term planning may involve two layers (e.g., a first layer considering the rewards of potential actions for a current state, and a second layer considering the rewards of future action options in response to predicted future states), an analysis based on more layers is also possible. For example, instead of basing the long-term planning analysis on one or two layers, three, four, or more layers of analysis could be used when selecting from available potential actions in response to a current navigation state.

[0276] After a selection has been made from the potential actions in response to a sensor-detected navigation state, the at least one processor at step 1627 can cause at least one adjustment of a navigation actuator of the host vehicle in response to the selected potential navigation action. The navigation actuator can include any suitable device for controlling at least one aspect of the host vehicle. For example, the navigation actuator can include at least a steering mechanism, a brake, or an accelerator. Navigation based on presumed aggression from others

[0277] Target vehicles can be monitored by analyzing a captured image stream to determine indicators of aggressive driving behavior. Aggression is described here as a qualitative or quantitative parameter, but other characteristics can also be used: perceived level of attention (potential impairment of the driver, distracted – mobile phone, asleep, etc.). In some cases, a target vehicle is assumed to be adopting a defensive posture, and in other cases, it can be determined that the target vehicle is adopting a more aggressive posture. Navigation actions can be selected or developed based on indicators of aggression. For example, in some cases, relative speed, relative acceleration, increases in relative acceleration, following distance, etc., relative to a host vehicle can be tracked to determine whether the target vehicle is aggressive or defensive.If it is determined that the target vehicle exhibits a level of aggression that exceeds, for example, a certain threshold, the host vehicle may be inclined to yield the right-of-way to the target vehicle. A target vehicle's level of aggression can also be detected based on specific behavior of the target vehicle relative to one or more obstacles on a path or in the vicinity of the target vehicle (e.g., a vehicle ahead, an obstacle in the road, a traffic light, etc.).

[0278] As an introduction to this concept, an example experiment is described involving the host vehicle merging into a roundabout, where the navigation goal is to complete and exit the roundabout. The situation can begin with the host vehicle approaching an entrance to the roundabout and end with the host vehicle reaching an exit (e.g., the second exit). Success can be measured by whether the host vehicle maintains a safe distance from all other vehicles at all times, whether the host vehicle completes the route as quickly as possible, and whether the host vehicle adheres to a guideline for smooth acceleration. In this illustration, the N TTarget vehicles are randomly placed in the roundabout. To model a mixture of adversarial and typical behavior with a probability p, one target vehicle can be modeled with an "aggressive" driving policy, such that the aggressive target vehicle accelerates when the host vehicle attempts to merge in front of it. With a probability of 1-p, the target vehicle can be modeled with a "defensive" driving policy, i.e., the target vehicle brakes and allows the host vehicle to merge. In this experiment, p = 0.5, and the host vehicle's navigation system may have no information about the types of other drivers. The types of other drivers can be randomly selected at the beginning of the episode.

[0279] The navigation state can be represented as the speed and location of the host vehicle (the agent), as well as the locations, speeds, and accelerations of the target vehicles. Maintaining target acceleration observations can be important to differentiate between aggressive and defensive drivers based on the current state. All target vehicles can move along a one-dimensional curve that outlines the path of the roundabout. The host vehicle can move along its own one-dimensional curve that intersects the target vehicles' curve at the merging point, and this point is the origin of both curves. To model reasonable driving behavior, the absolute value of all vehicles' accelerations can be bounded by a constant. Speeds can also be guided by a ReLU (Reverse Lumen Entity), since reversing is not permitted.It should be noted that by not allowing reversing, long-term planning may be necessary, as the agent cannot regret his past actions.

[0280] As described above, the next state can be s t+1 into a sum of a predictable part N(s t , a t ) and an unpredictable part of v t be decomposed. The expression N(s t , a t ) can represent the dynamics of vehicle positions and speeds (which can be well-defined in a differentiable way), while v t This can represent the acceleration of the target vehicles. It can be verified that N(s t ,a t ) can be expressed as a combination of ReLU functions via an affine transformation, therefore it is in terms of s t and a t differentiable. The vector v tcan be defined by a simulator in an indistinguishable way and can implement aggressive behavior for some targets and defensive behavior for others. Two still images from such a simulator are in Fig. 17A and Fig. 17B. In this example experiment, a host vehicle, 1701, has learned to slow down as it approaches the roundabout entrance. It has also learned to yield to aggressive vehicles (e.g., vehicles 1703 and 1705) and to proceed safely ahead of defensive vehicles (e.g., vehicles 1706, 1708, and 1710) when merging. In the Fig. 17A and Fig. In the example shown in Figure 17B, the navigation system of the host vehicle 1701 is not equipped with the type of target vehicles. Rather, whether a particular vehicle is identified as aggressive or defensive is determined by inferences based, for example, on the observed position and acceleration of the target vehicles. Fig. 17A allows the host vehicle 1701 to determine, based on position, speed, and / or relative acceleration, that the target vehicle 1703 exhibits an aggressive tendency, and therefore the host vehicle 1701 can stop and wait until the target vehicle 1703 has passed, instead of attempting to merge in front of the target vehicle 1703. Fig. However, 17B, the host vehicle 1701, recognized that the target vehicle 1710, which was traveling behind vehicle 1703, was exhibiting defensive tendencies (again based on the observed position, speed and / or relative acceleration of vehicle 1710), and therefore successfully completed a merge maneuver in front of target vehicle 1710 and behind target vehicle 1703.

[0281] Fig. Section 18 provides a flowchart illustrating an example algorithm for navigating a host vehicle based on predicted aggression from other vehicles. In the example of Fig. Step 18 allows an inference to be made about the level of aggression associated with at least one target vehicle based on the observed behavior of the target vehicle relative to an object in the target vehicle's environment. For example, in step 1801, at least one processing device (e.g., processing device 110) of the host vehicle's navigation system can receive a multitude of images from a camera associated with the host vehicle, images that are representative of the host vehicle's environment. In step 1803, the analysis of one or more of the received images can enable the at least one processor to identify a target vehicle (e.g., vehicle 1703) in the environment of host vehicle 1701. In step 1805, the analysis of one or more of the received images can enable the at least one processing device to identify at least one obstacle for the target vehicle in the environment of the host vehicle.The object can include dirt on a roadway, a stop light / traffic light, a pedestrian, another vehicle (e.g., a vehicle traveling in front of the target vehicle, a parked vehicle, etc.), a piece of cardboard on the roadway, a road barrier, a curb, or any other type of object that can be found in the vicinity of the host vehicle. In step 1807, the analysis of one or more of the received images can enable the at least one processing device to determine at least one navigation property of the target vehicle relative to the at least one identified obstacle for the target vehicle.

[0282] Various navigational characteristics can be used to infer the aggression level of a detected target vehicle in order to develop an appropriate navigational response to the target vehicle. For example, such navigational characteristics can include the relative acceleration between the target vehicle and the one or more identified obstacles, the distance of the target vehicle from the obstacle (e.g., the following distance of the target vehicle behind another vehicle), and / or the relative speed between the target vehicle and the obstacle, etc.

[0283] In some embodiments, the navigation characteristics of the target vehicles can be determined based on the outputs of sensors associated with the host vehicle (e.g., radar, speed sensors, GPS, etc.). However, in some cases, the navigation characteristics of the target vehicles can be determined partially or completely based on the analysis of images of the host vehicle's environment. For example, the image analysis techniques described above and, for instance, in U.S. Patent No. 9,168,868, incorporated herein by reference, can be used to detect target vehicles within the host vehicle's environment. This includes monitoring the location of a target vehicle in the captured images over time and / or monitoring the locations in the captured images of one or more features associated with the target vehicle (e.g., taillights, headlights, bumper, wheels, etc.).) can enable the determination of the relative distances, speeds and / or accelerations between the target vehicles and the host vehicle, or between the target vehicles and one or more other objects in the environment of the host vehicle.

[0284] The aggressiveness level of an identified target vehicle can be inferred from any suitable observed navigation property of the target vehicle or any combination of observed navigation properties. For example, a determination of aggressiveness can be made based on any observed property and one or more predetermined threshold levels, or any other suitable qualitative or quantitative analysis. In some embodiments, a target vehicle may be considered aggressive if it is observed following the host vehicle or another vehicle at a distance less than a predetermined threshold for aggressive distance.On the other hand, a target vehicle observed following the host vehicle or another vehicle at a distance greater than a predetermined defensive distance threshold may be considered defensive. The predetermined aggressive distance threshold need not be the same as the predetermined defensive distance threshold. Additionally, one or both of the predetermined aggressive distance thresholds and the predetermined defensive distance thresholds may encompass a range of values ​​rather than a single, defined line value. Furthermore, neither the predetermined aggressive distance threshold nor the predetermined defensive distance threshold needs to be fixed.Rather, these values ​​or ranges of values ​​can shift over time, and different thresholds / threshold ranges can be applied based on observed characteristics of a target vehicle. For example, the applied thresholds may depend on one or more other characteristics of the target vehicle. Higher observed relative speeds and / or accelerations may justify the application of larger thresholds / threshold ranges. Conversely, lower relative speeds and / or accelerations, including zero relative speeds and / or accelerations, may justify the application of smaller distance thresholds / threshold ranges, indicating an aggressive / defensive driving pattern.

[0285] The aggressive / defensive inference can also be based on thresholds for relative speed and / or relative acceleration. A target vehicle may be considered aggressive if its observed relative speed and / or relative acceleration relative to another vehicle exceeds a predetermined value or range. A target vehicle may be considered defensive if its observed relative speed and / or relative acceleration relative to another vehicle falls below a predetermined value or range.

[0286] While the determination of aggressive / defensive behavior can be made solely based on a single observed navigational characteristic, it can also depend on any combination of observed characteristics. For example, as noted above, in some cases a target vehicle may be deemed aggressive solely based on the observation that it is following another vehicle at a distance below a certain threshold or range. In other cases, however, the target vehicle may be deemed aggressive if it is following another vehicle at a distance less than a predetermined value (which may be the same as or different from the threshold applied when the determination is based solely on distance) and exhibits a relative speed and / or relative acceleration greater than a predetermined value or range.Similarly, a target vehicle may be deemed defensive solely based on the observation that it is following another vehicle at a distance greater than a specified threshold or range. In other cases, however, the target vehicle may be deemed defensive if it is following another vehicle at a distance greater than a predetermined value (which may be the same or different from the threshold applied when the determination is based solely on distance) and has a relative speed and / or relative acceleration that is less than a predetermined value or range. System 100 may make an aggressive / defensive conclusion if, for example, a vehicle experiences an acceleration or deceleration exceeding 0.5 G (e.g.,a sudden movement of 5 m / s3), a vehicle exhibits a lateral acceleration of 0.5 G when changing lanes or cornering, a vehicle causes another vehicle to do any of the foregoing, a vehicle changes lanes and causes another vehicle to yield by deceleration of more than 0.3 G or a sudden movement of 3 m / s3, and / or a vehicle changes two lanes without stopping.

[0287] It should be noted that references to a quantity exceeding a range can indicate that the quantity either exceeds all values ​​associated with the range or falls within the range. Similarly, references to a quantity falling below a range can indicate that the quantity either falls below all values ​​associated with the range or falls within the range. Additionally, while the described examples of inferring an aggressive / defensive course of action are based on distance, relative acceleration, and relative velocity, any other suitable quantities can be used. For example, a time-to-collision calculation or any indirect indicator of the target vehicle's distance, acceleration, and / or velocity could be employed.It should also be noted that while the above examples focus on target vehicles relative to other vehicles, the aggressive / defensive inference can be made by observing the navigation characteristics of a target vehicle relative to any other type of obstacle (e.g., a pedestrian, a road barrier, a traffic light, debris, etc.).

[0288] With renewed reference to the in Fig. 17A and Fig. In the example shown in Figure 17B, as the host vehicle 1701 approaches the roundabout, the navigation system, including its at least one processing device, can receive a stream of images from a camera associated with the host vehicle. Based on the analysis of one or more of the received images, any one of the target vehicles 1703, 1705, 1706, 1708, and 1710 can be identified. Furthermore, the navigation system can analyze the navigation characteristics of one or more of the identified target vehicles. The navigation system can recognize that the gap between target vehicles 1703 and 1705 represents the first opportunity for a potential merge into the roundabout. The navigation system can analyze target vehicle 1703 to determine aggression indicators associated with target vehicle 1703.If the target vehicle 1703 is deemed aggressive, the host vehicle's navigation system may choose to yield the right-of-way to vehicle 1703 instead of merging in front of it. Conversely, if the target vehicle 1703 is deemed defensive, the host vehicle's navigation system may attempt to complete a merging maneuver in front of it.

[0289] As the host vehicle 1701 approaches the roundabout, the navigation system's at least one processing device can analyze the captured images to determine the navigation characteristics associated with the target vehicle 1703. For example, based on the images, it can be determined that vehicle 1703 is following vehicle 1705 at a distance that provides a sufficient gap for the host vehicle 1701 to enter safely. In fact, it can be determined that vehicle 1703 is following vehicle 1705 at a distance that exceeds a threshold for aggressive following distance, and therefore, based on this information, the host vehicle's navigation system may be inclined to identify the target vehicle 1703 as defensive. However, in some situations, more than one navigation characteristic of a target vehicle can be analyzed to make the aggressive / defensive determination, as discussed above.By deepening its analysis, the host vehicle's navigation system can determine that while target vehicle 1703 is following target vehicle 1705 at a non-aggressive distance, vehicle 1703 is exhibiting a relative speed and / or acceleration relative to vehicle 1705 that exceeds one or more thresholds associated with aggressive behavior. In fact, the host vehicle 1701 can determine that target vehicle 1703 is accelerating relative to vehicle 1705 and closing the gap between them. Based on further analysis of the relative speed, acceleration, and distance (and even the rate at which the gap between vehicles 1703 and 1705 is closing), the host vehicle 1701 can determine that target vehicle 1703 is behaving aggressively.Therefore, even if there is a sufficiently large gap in which the host vehicle 1701 can safely navigate, it can expect that merging in front of the target vehicle 1703 would result in an aggressively navigating vehicle directly behind the host vehicle. Furthermore, based on behavior observed through image analysis or other sensor output, the target vehicle 1703 can be expected to accelerate further toward the host vehicle 1701 or continue traveling toward the host vehicle 1701 at a non-zero relative speed if the host vehicle 1701 were to merge in front of the host vehicle 1703. Such a situation can be undesirable from a safety perspective and can also cause discomfort for the host vehicle's passengers. For these reasons, the host vehicle 1701 may choose to yield the right-of-way to the host vehicle 1703, as shown in [reference to relevant section]. Fig. 17B shown, and to merge into the roundabout behind vehicle 1703 and in front of vehicle 1710, which is considered defensive based on the analysis of one or more of its navigation characteristics.

[0290] With renewed reference to Fig. 18. At step 1809, the at least one processing device of the host vehicle's navigation system can determine a navigation action for the host vehicle based on the identified at least one navigation property of the target vehicle relative to the identified obstacle (e.g., merging in front of vehicle 1710 and behind vehicle 1703). To implement the navigation action (at step 1811), the at least one processing device can, in response to the determined navigation action, initiate at least one adjustment of a navigation actuator of the host vehicle. For example, a brake can be applied to allow vehicle 1703 to merge into the obstacle. Fig. 17A to yield the right-of-way, and an accelerator can be operated together with the steering of the host vehicle's wheels to cause the host vehicle to enter the roundabout behind vehicle 1703, as shown in Fig. 17B shown.

[0291] As described in the preceding examples, the host vehicle's navigation can be based on the navigation properties of a target vehicle relative to another vehicle or object. Additionally, the host vehicle's navigation can be based solely on the navigation properties of the target vehicle without any specific reference to another vehicle or object. For example, in step 1807 of Fig. 18. The analysis of a large number of images taken from the environment of a host vehicle enables the determination of at least one navigational property of an identified target vehicle that indicates a level of aggression associated with the target vehicle. The navigational property may include a speed, acceleration, etc., which need not be referenced to another object or target vehicle to make an aggressive / defensive determination. For example, observed accelerations and / or speeds associated with a target vehicle that exceed a predetermined threshold or fall within or exceed a range of values ​​may indicate aggressive behavior.Conversely, observed accelerations and / or speeds associated with a target vehicle that fall below a predetermined threshold, or fall within or exceed a range of values, may indicate defensive behavior.

[0292] Of course, in some cases, the observed navigational characteristic (e.g., a position, distance, acceleration, etc.) relative to the host vehicle can be used to determine aggressive / defensive behavior. For example, an observed navigational characteristic of the target vehicle that indicates a level of aggression associated with the target vehicle could include an increase in the relative acceleration between the target vehicle and the host vehicle, a following distance of the target vehicle behind the host vehicle, a relative speed between the target vehicle and the host vehicle, etc. Navigation based on the limitation of liability for accidents

[0293] As described in the preceding sections, planned navigation actions can be checked against predetermined constraints to ensure compliance with certain rules. In some embodiments, this concept can be extended to considerations of potential accident liability. As discussed below, a primary goal of autonomous navigation is safety. Since absolute safety may be impossible (for example, because a given host vehicle under autonomous control cannot control the other vehicles in its vicinity—it can only control its own actions), using potential accident liability as a consideration in autonomous navigation, and indeed as a constraint on planned actions, can help ensure that a given autonomous vehicle does not take actions deemed unsafe—for example,those for which the host vehicle could potentially be liable for an accident. If the host vehicle only takes actions that are safe and that are determined not to lead to an accident caused by the host vehicle's own fault or responsibility, then the desired levels of accident avoidance (e.g., less than 10) can be achieved. -9 per driving lesson).

[0294] The challenges posed by most current approaches to autonomous driving include a lack of safety guarantees (or at least the inability to provide the desired levels of safety) and also a lack of scalability. Consider the problem of guaranteeing safe driving with multiple agents. Since society is unlikely to tolerate fatalities from road traffic accidents caused by machines, an acceptable level of safety is crucial for the acceptance of autonomous vehicles. While one goal may be to reduce the number of accidents to zero, this may not be achievable, as accidents typically involve multiple agents, and situations can be imagined where an accident is caused solely by the fault of other agents. For example, the host vehicle 1901, as in Fig. Figure 19 shows a scenario on a multi-lane highway. While the host vehicle 1901 can control its own actions relative to the target vehicles 1903, 1905, 1907, and 1909, it cannot control the actions of the surrounding target vehicles. Consequently, the host vehicle 1901 may be unable to avoid a collision with at least one of the target vehicles if, for example, vehicle 1905, on a collision course with the host vehicle, suddenly cuts into the host vehicle's lane. To address this difficulty, autonomous vehicle practitioners typically employ a statistical, data-driven approach, where the safety validation becomes more rigorous as more data is collected over extended driving time.

[0295] To understand the problems with a data-driven safety concept, one must bear in mind that the probability of a fatality caused by an accident per hour of (human) driving is known to be 10 -6 It is reasonable to assume that the death rate would have to be reduced by three orders of magnitude, namely to a probability of 10 -9 per hour, so that society accepts machines replacing humans in the task of driving. This estimate is similar to the assumed fatality rate of airbags and the standards in aviation. For example, 10 -9The probability of a wing spontaneously detaching from an aircraft in mid-air. However, attempts to guarantee safety using a data-driven statistical approach that builds additional confidence through the accumulation of miles flown are not practical. The amount of data required to calculate a probability of 10 -9 The number of deaths per driving lesson to guarantee is proportional to its inverse (i.e., 10). 9Hours of data), which is roughly on the order of thirty billion miles. Furthermore, a multi-agent system interacts with its environment and is unlikely to be validated offline (unless a realistic simulator is available that replicates real-world human driving with all its facets and complexities, such as reckless driving—but validating the simulator would be even more difficult than creating a safe autonomous vehicle agent). And any change to the planning and control software requires a new data collection of the same order of magnitude, which is clearly cumbersome and impractical. Moreover, developing a data-driven system inevitably suffers from a lack of interpretability and explainability of the actions taken—if an autonomous vehicle (AV) has a fatal accident, we need to know why.Consequently, a model-based safety approach is required, but the existing requirements for "functional safety" and ASIL in the automotive industry are not designed to deal with multi-agent environments.

[0296] A second primary challenge in developing a safe driving model for autonomous vehicles is the need for scalability. The premise underlying AV goes beyond “building a better world” and is instead based on the premise that driverless mobility can be maintained at a lower cost than with drivers. This premise is inextricably linked to the concept of scalability—in the sense of supporting the mass production of AVs (in the millions) and, more importantly, supporting negligible additional costs to enable driving in a new city.Therefore, the costs of data processing and sensor acquisition are important if AV is to be mass-produced, while the costs of validation and the ability to drive “everywhere” and not just in a few selected cities are necessary prerequisites for maintaining a business.

[0297] The problem with most current approaches lies in a brute-force mindset along three axes: (i) the required computing density, (ii) how high-resolution maps are defined and built, and (iii) the required sensor specifications. A brute-force approach undermines scalability and shifts the emphasis toward a future where unlimited on-board computing is ubiquitous, where the cost of building and maintaining HD maps becomes negligible and scalable, and where exotic, highly advanced sensors are developed and manufactured to automotive-grade standards at negligible cost. A future in which one of the above occurs is certainly plausible, but all of them occurring is likely a low-probability event.Therefore, there is a need to provide a formal model that combines security and scalability in an AV program that can be accepted by society and is scalable in the sense that it supports millions of cars driving throughout developed countries.

[0298] The disclosed embodiments represent a solution capable of providing (or even exceeding) the targeted levels of safety and scalable to systems encompassing millions (or more) of autonomous vehicles. In the area of ​​safety, a model called Responsibility-Sensitive Safety (RSS) is introduced, which formalizes the concept of "accident fault," making it interpretable and explainable, and instills a sense of "responsibility" in the actions of a robotic agent. The definition of RSS is agnostic to how it is implemented—a crucial feature for facilitating the goal of creating a compelling global safety model. RSS is motivated by observation (as in Fig. 19) that the agents play an asymmetrical role in an accident, with usually only one of the agents being responsible for the accident and therefore having to be held accountable. The RSS model also includes a formal treatment of “cautious driving” under restricted perceptual conditions, where not all agents are always visible (for example, due to obstructions). A key objective of the RSS model is to guarantee that no agent ever causes an accident for which it is “at fault” or responsible. A model can only be useful if it is accompanied by an efficient policy (e.g., a function that maps the “sensor detection state” to an action) that is RSS-compliant. For example, an action that appears harmless at the moment could lead to a catastrophic event in the distant future (“butterfly effect”).RSS can be useful for creating a set of local restrictions for the short-term future that can guarantee (or at least practically guarantee) that no accidents will occur in the future as a result of the actions of the host vehicle.

[0299] Another contribution focuses on the introduction of a “semantic” language consisting of units, measurements, and action spaces, as well as specifications for how these are incorporated into the planning, sensor acquisition, and operation of the AV. To get a sense of the semantics in this context, consider how a person taking driving lessons is instructed to think about a “driving guideline.” These instructions are not geometric—they don’t take the form of “drive 13.7 meters at the current speed and then accelerate at a rate of 0.8 m / s².” 2Instead, the instructions are semantic in nature—"follow the car in front of you" or "overtake the car on your left." The typical language of human driving direction deals with longitudinal and lateral objectives rather than geometric units of acceleration vectors. A formal semantic language can be useful in several ways related to the computational complexity of planning, which does not scale exponentially with time and the number of agents; the way safety and comfort interact; how the calculation of sensor acquisition is defined; and the specification of sensor modalities and how they interact in a fusion methodology. Using a fusion methodology (based on the semantic language) can ensure that the RSS model achieves the required probability of 10 -9Deaths per driving hour were achieved, using only offline validation over a dataset of approximately 10 5 Hours of driving data are carried out.

[0300] For example, in a reinforcement learning environment, a Q-function (e.g., a function that evaluates the long-term quality of performing an action a ∈ A when the agent is in state s ∈ S) may be used; given such a Q-function, the natural choice of action may be to take the one with the highest quality π(s) = argmax a Q(s, a) to select) is defined over a semantic space in which the number of trajectories to be examined at a given time is divided by 10, independent of the time horizon used for planning. 4The space is limited. The signal-to-noise ratio in this space can be high, so effective machine learning approaches can be successful in modeling the Q-function. In the case of sensor acquisition calculations, the semantics can allow a distinction between errors that affect safety and errors that affect ride comfort. We define a PAC (Probably Approximate Correct) model (following Valiant's PAC learning terminology) for sensor acquisition linked to the Q-function and show how measurement errors are incorporated into the design in a way that complies with RSS while still allowing for ride comfort optimization.The language of semantics can be important for the success of certain aspects of this model, as other standard error measurements, such as errors with respect to a global coordinate system, may not align with the PAC sensor acquisition model. Additionally, semantic language can be a crucial enabler for defining high-definition maps that can be created using low-bandwidth sensor data, thus facilitating crowdsourcing and supporting scalability.

[0301] In summary, the disclosed embodiments can include a formal model that covers key components of an AV: perception, planning, and action. This model can help ensure that, from a planning perspective, no accident falls under the AV's own responsibility. Furthermore, even with a PAC sensor acquisition model, the described fusion methodology, even in the event of sensor acquisition errors, can require only a very reasonable amount of offline data collection to comply with the described safety model. Moreover, the model can combine safety and scalability through semantic language, thus providing a complete methodology for a safe and scalable AV. Finally, it is worth noting that developing an accepted safety model that gains adoption by industry and regulatory bodies could be a necessary condition for the success of AVs.

[0302] The RSS model can generally follow a classic robot control methodology based on the "sense-plan-act" principle. The sensor acquisition system can be responsible for understanding the current state of the environment surrounding a host vehicle. The planning part, which can be referred to as the "drive policy" and can be implemented through a set of hard-coded instructions, a trained system (e.g., a neural network), or a combination of both, can be responsible for determining the best next step in terms of the available options for achieving a driving destination (e.g., how to change from the left lane to a right lane to exit a highway). The acting part is responsible for implementing the plan (e.g., the system of actuators and one or more controllers for steering, accelerating, and / or braking, etc.).(of a vehicle, in order to implement a selected navigation action). The embodiments described below focus mainly on the sensor acquisition and planning aspects.

[0303] Accidents can be caused by sensor detection errors or planning errors. Planning is a multi-agent endeavor, as other road users (humans and machines) react to the actions of an automated guided vehicle (AGV). The described RSS model is designed, among other things, to address the safety of the planning component. This can be referred to as multi-agent safety. With a statistical approach, the probability of planning errors can be estimated "online." After each software update, billions of miles would have to be driven with the new version to provide an acceptable level for estimating the frequency of planning errors. This is clearly not feasible. Alternatively, the RSS model can provide a 100% (or near 100%) guarantee that the planning module will not make errors attributable to the AGV (the term "fault" is formally defined).The RSS model can also provide an efficient means for its validation that does not rely on online testing.

[0304] Faults in a sensor acquisition system can potentially be validated more easily because sensor acquisition can occur independently of vehicle actions, and therefore we can validate the probability of a serious sensor acquisition fault using "offline" data. But even collecting offline data from more than 10 9 Driving lessons are a challenge. As part of the description of a disclosed sensor system, a fusion approach is described that can be validated using a significantly smaller amount of data.

[0305] The described RSS system can also be scalable for millions of cars. For example, the described semantic driving policy and the applied safety restrictions can be compatible with sensor and mapping requirements that, even with today's technology, can be scaled to millions of cars.

[0306] A fundamental building block of such a system is a thorough safety definition, i.e., a minimum standard that AV systems must adhere to. The following technical lemma demonstrates that a statistical approach to validating an AV system is feasible, even for validating a simple claim such as "the system causes N accidents per hour." This implies that a model-based safety definition is the only practical tool for validating an AV system.

[0307] Lemma 1: Let X be a probability space and A an event for which Pr(A) = p1 < 0,1. Suppose we take m=1p1 uiv samples from X as a sample, and let it be Z=∑i=1m1[x∈A]. Then the following applies Pr(Z=0)≥e−2.

[0308] Proof: We use the inequality 1 - x ≥ e -2x (proven for the sake of completeness in Appendix A.1) and obtained Pr(Z=0)=(1−p1)m≥e−2p1m=e−2

[0309] Corollary 1 Suppose an AV system AV1 causes an accident with a small but insufficient probability p1. Any deterministic validation procedure that obtains 1 / p1 samples will, with constant probability, not distinguish between AV1 and another AV system AV0 that never causes accidents.

[0310] To get an overview of the typical values ​​for such probabilities, let's assume an accident probability of 10 -9 per hour and a specific AV system only has a probability of 10 -8 provides. Even if the system is 10 8 Once driving lessons have been completed, there is a constant probability that the validation process will be unable to indicate that the system is dangerous.

[0311] Finally, it should be noted that the difficulty lies in invalidating a single, specific, dangerous AV system. A complete solution cannot be considered a single system, as new versions, bug fixes, and updates will be required. Every change, even a single line of code, creates a new system from the validator's perspective. Therefore, a solution that is statistically validated must do so online, with new samples after every minor fix or change to account for the shift in the distribution of observed states and those achieved by the new system. Repeatedly and systematically obtaining such a large number of samples (and even then, with a constant probability of failing to validate the system) is not feasible.

[0312] Furthermore, every statistical claim must be formalized in order to be measurable. A claim about a statistical property regarding the number of accidents caused by a system is significantly weaker than the claim "it drives safely." To make this statement, one must formally define what safety is. Absolute security is impossible

[0313] An action a taken by a car c can be considered absolutely safe if the action cannot lead to an accident at any later time. It is evident that achieving absolute safety is impossible when considering simple driving scenarios, such as those found in... Fig. Figure 19 illustrates this. From the perspective of vehicle 1901, no action can guarantee that none of the surrounding cars will collide with the vehicle. Solving this problem by prohibiting the autonomous car from being in such situations is also impossible. Since any highway with more than two lanes will eventually lead to this, prohibiting this scenario is tantamount to forcing it to stay in the garage. The implications may seem disappointing at first glance. Nothing is absolutely safe. However, such a requirement for absolute safety, as defined above, can be too strict, as demonstrated by the fact that human drivers do not adhere to the requirement of absolute safety. Instead, humans behave according to a safety concept based on responsibility. Responsibility-Sensitive Safety (RSS)

[0314] An important aspect missing from the absolute safety concept is the asymmetry of most accidents – usually, one of the drivers is responsible for an accident and is held accountable. For example, in Fig. In equation 19, the middle car, 1901, is not at fault if, for example, the left car, 1909, suddenly crashes into it. To formalize the fact of taking its lack of responsibility into account, the behavior of AV 1901, remaining in its own lane, can be considered safe. This involves describing a formal concept of "accident fault" or accident responsibility, which can serve as a premise for a safe driving approach.

[0315] As an example, consider the simple case of two cars c f , c r , driving one behind the other at the same speed on a straight road. Assume c fThe car in front brakes suddenly due to an obstacle appearing in the road and manages to avoid it. Unfortunately, c r not sufficient distance to c f , cannot react in time and drives c f into the rear. It is clear that c r who is at fault; it is the responsibility of the car behind to maintain a safe distance from the car in front and to be prepared for unexpected but appropriate braking.

[0316] Next, consider a much broader family of scenarios: driving on a multi-lane road where cars can freely change lanes, merge into other cars' lanes, travel at different speeds, and so on. To simplify the following discussion, assume a straight road on a flat surface, where the lateral and longitudinal axes are the x- and y-axes, respectively. This can be achieved under mild conditions by defining a homomorphism between the actual curved road and a straight road. Additionally, consider a discrete time period. Definitions can help distinguish between two intuitively different sets of cases: simple cases where no significant lateral maneuver is performed, and more complex cases that involve lateral movement.

[0317] Definition 1 (Car corridor) The corridor of a car c is the area [c x,links ,c x,rechts ] × [±∞], where c x,links ,c x,rechts The positions of the outermost left and outermost right corners of c are.

[0318] Definition 2 (Merging): A car c1 (e.g., car 2003 in Fig. 20A and Fig. 20B) shears into the corridor of car c0 at time t (e.g., car 2001 in Fig. 20A and Fig. 20B) if it has not crossed the corridor of c0 at time t-1, but crosses it at time t.

[0319] A further distinction can be made between the front and rear sections of the corridor. The term "direction of merging" can describe movement towards the respective corridor boundary. These definitions can address cases involving lateral movement. For the simple case where no such occurrence exists, such as when one car follows another, the safe longitudinal distance is defined:

[0320] Definition 3 (Safe longitudinal distance): A longitudinal distance 2101 ( Fig. 21) between a car c r (Car 2103) and another car c f (Car 2105), which is in the frontal corridor of c r is certain with respect to a reaction time p if, at any of c f executed braking command a, |a| < a max,bremsen , if c r its maximum brake is applied from time p until it comes to a standstill, not with c fwill collide.

[0321] The following Lemma 2 calculates d as a function of the velocities of c. r , c f , the reaction time ρ and the maximum acceleration a max,bremsen Both ρ and a max,bremsen These are constants that should be fixed to some reasonable values ​​through regulation. In further examples, either one of the reaction times ρ and the maximum acceleration a can be used. max,bremsen can be set for a specific vehicle or vehicle type, or can be adjusted / set according to measurements or other input parameters relating to the vehicle condition, road condition, user preferences (e.g., of a driver or passenger), etc.

[0322] Lemma 2 Let c r a vehicle that is positioned behind c along its longitudinal axis f is located. There are a max,bremsen , a max,beschlthe maximum braking and acceleration commands, and ρ is the reaction time of c r '. There were u r , u f the longitudinal speeds of the car and let it be I f , I r their lengths. Define u p,max = u r + ρ · a max,beschl and define Tr=p+vpmaxamax,brakes and Tf=vfamax,brakes. Let L = (I r +I f ) / 2. Then the minimum safe longitudinal distance for c is r : dmin={Lif Tr≤TfL+Tf[(vp,max−vf)+ρamax,brake,]−ρ2amax,brake2+(Tr−Tf)(vp,max−(Tf−ρ)cmax,brake)2otherwise

[0323] Proof: Let it be d t The distance at time t. To prevent an accident, dt > L must hold for every t. To ensure d minTo construct this, we need to find the tightest required lower bound for d0. Obviously, d0 must be at least L. As long as the two cars do not stop after T ≥ ρ seconds, the speed of the car in front is u. f - T a max,bremsen while the speed of c r through u ρ,max - (T - ρ) a max,beschl The distance between the cars after T seconds is therefore limited from above by: dT:=d0+T2(2vfT amax,braking)−[ρvρ,max+T−ρ2(2vρ,max−(T−ρ)amax,braking)] =d0+T[(vf−vρ,max)−ρamax,braking]+ρ2amax,braking2

[0324] Note that T r The time is when c r comes to a standstill (a speed of 0) and T f the time at which the other vehicle comes to a complete stop. Note that a) holds true. max,bremsen (T r - T f ) = u ρ,max - u f + ρa max,bremsen If T r≤ T f If T r > T f then applies dTr=d0+Tf[(vf−vρ,max)−ρamax,brake]+ρ2amax,brake2−(Tr−Tf)(vρ,max−(Tf−ρ)amax,brake)2.

[0325] The demand d Tr L and the transformation of the terms complete the proof.

[0326] Finally, a comparison operator is defined that allows comparisons with a certain notion of "margin": When comparing lengths, speeds, etc., it is necessary to accept very similar quantities as "equal".

[0327] Definition 4 (µ-comparison) The µ-comparison of two numbers a, b is a > µ b, if a > b + µ, a < µ b, if a < b - µ and a = µ b, if |a - b| ≤ µ.

[0328] The comparisons listed below (argmin, argmax, etc.) are µ-comparisons for some suitable µ. Suppose an accident occurs between vehicles c1 and c2. To determine who is responsible for the accident, the relevant moment that needs to be investigated is defined. This is a point in time that preceded the accident and was intuitively the "turning point"; after which nothing could be done to prevent the accident.

[0329] Definition 5 (Time of Fault) The time of fault in an accident is the earliest point in time before the accident at which: • there was an overlap between one of the cars and the corridor of the other and • the longitudinal distance was not certain.

[0330] Such a point in time clearly exists, since both con...

Claims

[1] Non-transitory computer-readable device with instructions stored thereon which, when executed by at least one computing device, cause at least one computing device to perform operations, comprising: Receiving an image of the vehicle's driver; Determining the driver's viewing direction based on the image; Predictions of the vehicle's path, at least partially, based on the direction of view; and [2] Device according to claim 1, wherein the image is a first image, further comprising: Receiving a second image of the vehicle's surroundings; Identifying an object in the second image; and Determine whether the viewing direction is directed towards the object, where the prediction includes predicting the path based on whether the viewing direction is directed towards the object. [3] Device according to claim 2, further comprising classifying the object as either of a type over which vehicles drive or of a type over which vehicles do not drive; wherein, based on the classification of the object as of a type over which vehicles drive, the path prediction comprises predicting the path directed towards the object. [4] Device according to claim 2, further comprising classifying the object as either of a type over which vehicles drive or of a type over which vehicles do not drive; wherein, on the basis of classifying the object as of a type over which vehicles do not drive, the path prediction comprises predicting the path directed away from the object. [5] Device according to claim 1, wherein the operations further comprise capturing a sequence of photographic images taken by at least one camera attached to the vehicle and positioned to capture the vehicle's surroundings, wherein predicting the path comprises predicting the path based on an inherent motion estimated from the sequence of photographic images. [6] Device according to claim 1, wherein the path prediction comprises predicting the path based on an input from a sensor on the vehicle. [7] System according to claim 1, further comprising the operations: Adapting the operation of a vehicle's advanced driver assistance system (ADAS) based on the predicted path. [8] Device according to claim 1, wherein the control comprises: Adjusting the degree of automatic braking based on the fact that the object is not on the predicted path. [9] Device according to claim 1, wherein the control comprises: Providing a notification to the driver based on the predicted path. [10] System for predicting the driving path of a vehicle, comprising: one or more inward-facing cameras that monitor a driver of the vehicle; a storage facility; and at least one processor that is coupled with and configured to the memory for: Receiving an image of the vehicle's driver from an inward-facing camera that monitors the vehicle's driver; Determining the driver's viewing direction based on the image; Predictions of the vehicle's path, at least partially, based on the direction of view; and Steering the vehicle based on the predicted path. [11] System according to claim 10, wherein the image is a first image, the system further comprises one or more outwardly facing sensors that monitor an external environment of the vehicle, and the at least one processor is further configured to: Receiving a second image of the vehicle's surroundings from the one or more outward-facing sensors; Identifying an object in the second image; and Determine whether the viewing direction is directed towards the object, where the prediction includes predicting the path based on whether the viewing direction is directed towards the object. [12] System according to claim 11, wherein the at least one processor is further configured to classify the object either as being of a type over which vehicles drive or as being of a type over which vehicles do not drive; wherein, based on the classification of the object as being of a type over which vehicles drive, the path prediction includes predicting the path directed towards the object. [13] System according to claim 11, wherein the at least one processor is further configured to capture a sequence of photographic images taken by at least one camera attached to the vehicle and positioned to capture the vehicle's surroundings, wherein the path prediction comprises predicting the path based on an inherent motion estimated from the sequence of photographic images. [14] System according to claim 10, wherein the path prediction comprises predicting the path based on an input from a sensor on the vehicle. [15] System according to claim 10, wherein the at least one processor is further configured to: Adapting the operation of a vehicle's ADAS system based on the predicted path. [16] System according to claim 10, wherein the control comprises: Adjusting the degree of automatic braking based on the fact that an object is not on the predicted path. [17] System according to claim 10, wherein the control comprises: Providing a notification to the driver based on the predicted path. [18] Computer-implemented method for predicting the path of a vehicle, comprising: Receiving an image of the vehicle's driver; Determining the driver's viewing direction based on the image; Predictions of the vehicle's path, at least partially, based on the direction of view; and Steering the vehicle based on the predicted path. [19] The method of claim 18, wherein the image is a first image, further comprising: Receiving a second image of the vehicle's surroundings; Identifying an object in the second image; and Determine whether the viewing direction is directed towards the object, where the prediction includes predicting the path based on whether the viewing direction is directed towards the object. [20] Device according to claim 19, further comprising classifying the object as either of a type over which vehicles drive or of a type over which vehicles do not drive; wherein, based on the classification of the object as of a type over which vehicles drive, the path prediction comprises predicting the path directed towards the object.