Method and system for conditional operation of autonomous agents
A multi-step conditional policy with trigger conditions optimizes autonomous vehicle decision-making, enhancing prediction accuracy and reducing computational load, enabling safe and efficient navigation in complex scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MAY MOBILITY INC
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-29
AI Technical Summary
Existing autonomous vehicle systems struggle to optimize decision-making to minimize risks while mimicking human behavior and handling complex scenarios on the road, leading to unpredictable and disruptive operations.
Implementing a multi-step conditional policy with trigger conditions for autonomous agents, allowing for forward simulation and optimized policy selection based on the vehicle's environment, reducing computational load, and enhancing prediction accuracy.
This approach enables more predictable and accurate vehicle behavior, reliably handling complex operations, and reduces the occurrence of unexpected actions, ensuring safe and efficient navigation.
Smart Images

Figure 2026122968000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 309,945, filed on February 14, 2022, the entire disclosure of which is incorporated herein by reference.
[0002]
[0002] The present invention generally relates to the field of autonomous vehicles, and more specifically, to novel and useful systems and methods for the conditional operation of autonomous agents in the field of autonomous vehicles.
Background Art
[0003]
[0003] There are numerous challenges in optimizing the decision - making of autonomous vehicles. One important challenge is to be configured to carefully drive an autonomous vehicle to minimize risks while mimicking human behavior to minimize disruption to other drivers on the road and be prepared to handle complex scenarios. Conventional systems and methods have approached this in various ways, but it has not yet been achieved and is not reliably implemented.
[0004]
[0004] Therefore, in the field of autonomous vehicles, there is a need to create improved and useful systems and methods for the operation of autonomous agents.
Brief Description of the Drawings
[0005] [Figure 1]
[0005] FIG. 1 is a schematic diagram of a method for the conditional operation of an autonomous agent. [Figure 2]
[0006] FIG. 2 is a schematic diagram of a system for the conditional operation of an autonomous agent. [Figure 3]
[0007] FIG. 3 shows a schematic variant of a method for the conditional operation of an autonomous agent. [Figure 4]
[0008] Figure 4 shows a first example of a method for conditional behavior of an autonomous agent. [Figure 5]
[0009] Figure 5 shows a second example of a method for conditional behavior of an autonomous agent. [Figure 6]
[0010] Figure 6 shows a third example of a method for conditional behavior of an autonomous agent. [Figure 7]
[0011] Figure 7 shows a modified system for the conditional operation of an autonomous agent and information exchange within the system. [Figure 8A]
[0012] Figure 8A shows an exemplary example of a method for conditional behavior of an autonomous agent. [Figure 8B]
[0012] Figure 8B shows an exemplary example of a method for conditional operation of an autonomous agent. [Figure 8C]
[0012] Figure 8C shows an exemplary example of a method for conditional operation of an autonomous agent. [Figure 8D]
[0012] Figure 8D shows an exemplary example of a method for conditional operation of an autonomous agent. [Figure 8E]
[0012] Figure 8E shows an exemplary example of a method for conditional operation of an autonomous agent. [Figure 8F]
[0012] 8F illustrates an exemplary method for conditional operation of an autonomous agent. [Modes for carrying out the invention]
[0006]
[0013] The following description of preferred embodiments of the present invention is not intended to limit the invention to these preferred embodiments, but rather to enable those skilled in the art to manufacture and use the invention.
[0007] 1. Overview
[0014] As shown in Figure 1, the method 100 for the conditional operation of an autonomous agent includes the steps of: collecting a set of inputs S110; processing the set of inputs S120; determining a set of policies for the agent S130; evaluating the set of policies S140; and operating the agent S150. As an addition or alternative, Method 100 is a step that repeats any or all of the process; each of the entire applications is incorporated by this reference: U.S. Application No. 16 / 514,624 filed July 17, 2019; U.S. Application No. 16 / 505,372 filed July 8, 2019; U.S. Application No. 16 / 540,836 filed August 14, 2019; U.S. Application No. 16 / 792,780 filed February 17, 2020; U.S. Application No. 17 / 365,538 filed July 1, 2021; U.S. Application No. 17 / 365,538 filed December 14, 2021 Method 100 may include a step of repeating any or all of the processes described in any or all of the following: U.S. Application No. 550,461; U.S. Application No. 17 / 554,619 filed December 17, 2021; U.S. Application No. 17 / 712,757 filed April 4, 2022; U.S. Application No. 17 / 826,655 filed May 27, 2022; U.S. Application No. 18 / 073,209 filed December 1, 2022; and U.S. Application No. 18 / 072,939 filed December 1, 2022; or any other suitable processes performed in any suitable order. Method 100 may be performed using System 200 and / or any other suitable system described later.
[0008]
[0015] As shown in Figure 2, System 200 for the conditional operation of an autonomous agent (hereinafter also referred to as the self-agent and autonomous vehicle) includes a set of computing subsystems (hereinafter also referred to as a set of computers) and / or processing subsystems (hereinafter also referred to as a set of processors) that function to implement any or all of the processes of Method 100. Additionally or alternatively, System 200 may include or interface with any or all of the following components: an autonomous agent, one or more sets of sensors (e.g., sensors mounted on the self-agent, sensors mounted on a set of infrastructure devices, etc.), memory associated with the computing subsystem (e.g., memory storing a set of maps and / or a set of databases, as shown in Figure 7), a simulator module, a control subsystem, a set of infrastructure devices, a teleoperator platform, a tracker, a positioning system, a guidance system, a communication interface and / or any other components. In addition or alternatively, the system incorporates, by this reference, the entirety of each of the following applications: U.S. Application No. 16 / 514,624 filed July 17, 2019; U.S. Application No. 16 / 505,372 filed July 8, 2019; U.S. Application No. 16 / 540,836 filed August 14, 2019; U.S. Application No. 16 / 792,780 filed February 17, 2020; U.S. Application No. 17 / 365,538 filed July 1, 2021; and U.S. Application No. 17 / 365,538 filed December 14, 2021. U.S. applications No. 17 / 550,461; U.S. application No. 17 / 554,619 filed on 17 December 2021; U.S. application No. 17 / 712,757 filed on 4 April 2022; U.S. application No. 17 / 826,655 filed on 27 May 2022; U.S. application No. 18 / 073,209 filed on 1 December 2022; and U.S. application No. 18 / 072,939 filed on 1 December 2022 may include any or all of the components described in any or all of these applications.
[0009]
[0016] Alternatively, method 100 may be performed and / or implemented by any other suitable system.
[0010] 2. Advantages
[0017] Systems and methods for the conditional operation of autonomous agents can offer several advantages over current systems and methods.
[0011]
[0018] In the first modification, this technology offers the advantage of optimizing the behavior of an autonomous vehicle to its current environment (e.g., scene, scenario, context, etc.) by using a multi-step conditional policy (also referred to herein as behavior) that implements conditional logic (e.g., trigger conditions) optimized for navigating in relation to the vehicle's current specific environment. This enables a forward simulation that can be used when selecting actions to implement on the vehicle, representing what the vehicle will actually do over the length of time represented by the forward simulation, thereby enabling the selection of the optimal policy. For example, if the simulation is limited to simulating a single action over a relatively long simulation timescale, the results may not adequately reflect what the vehicle is actually trying to do, and a suboptimal action may be selected for the vehicle to implement (e.g., the vehicle may become stuck in a certain location). By using a multi-step policy with simulable trigger conditions, more natural behavior can be implemented on the vehicle even if not all actions of the multi-step policy are fully implemented.
[0012]
[0019] In a series of examples, using a multi-step conditional policy eliminates the need to prepare different types of the same policy for different types of situations, where different logics are hard-coded into the metrics (e.g., reward functions, risk metrics, cost functions, loss functions, etc.) used to evaluate these individual policies and / or policies. Instead, single-step policies are combinable in a modular fashion and can be implemented using trigger conditions that initiate transitions between them within a simulation and optionally in the actual operation of the vehicle.
[0013]
[0020] In another series of examples additional or alternative to the foregoing examples, using a multi-step conditional policy enables the simulated behavior of a vehicle to match the behavior that actually occurs during operation of the vehicle even when the simulated multi-step policy is not actually executed as simulated. This makes the simulation very predictable and accurate with respect to the actual behavior of the vehicle, resulting in improved prediction accuracy of the simulation, preventing unexpected and / or unanticipated behavior from being executed by the vehicle, and / or otherwise providing an advantage to the operation of the vehicle. For example, in certain examples, many single-step policies are not relevant for the entire planning period of a simulation (e.g., 5 - 10 seconds ahead), so the simulated vehicle can be made to behave in a way that does not match the behavior the vehicle would actually perform in its actual environment. Additionally or alternatively, this enables selection of policies that would result in disadvantageous outcomes in reality for the vehicle (e.g., stopping the vehicle, not advancing towards a target / destination, etc.).
[0014]
[0021] In the second modification, as an addition or alternative to the first modification, this technique provides the advantage of reducing the computational load and / or computational time required to evaluate and select a policy for an autonomous vehicle to implement over one or more selected cycles of the vehicle. For example, in a specific example, implementing a multi-step conditional policy can reduce the computational load and / or shorten the computational time compared to the case of individually determining each of the steps (e.g., by equally considering all possible policies in each selection cycle).
[0015]
[0022] Implementing a multi-step conditional policy can optionally additionally provide the advantage of reducing the occurrence of the self-agent being in an idle state, completely stopping, and / or not advancing towards the target in another way (e.g., not reaching the destination). For example, in a specific example, by defining a plurality of steps based on defined trigger conditions, the vehicle can be made not to wait for the next selection cycle (and / or processing a plurality of policies in the selection cycle) to determine and implement the next action.
[0016]
[0023] In the third modification, as an addition or alternative to the foregoing, this technique provides the advantage of improving the ability to accurately predict the actions taken by the vehicle. For example, in a specific example, the transition between steps of a conditional policy is not only implemented at the start of a selection cycle but is triggerable within the selection cycle and the policy is completed before the selection cycle ends. In these latter cases, the vehicle may complete the execution of the policy before the selection cycle ends and use the remaining time to perform unexpected and / or inconvenient and / or dangerous actions such as stopping. Instead, the system and / or method can provide the advantage of performing more predictable and acceptable actions throughout its operation.
[0017]
[0024] In the fourth modification, as an addition to or alternative to those described above, this technology offers the advantage of reliably handling complex operations, such as those related to the custom of right-of-way over other vehicles on the road. For example, in the specific example, the trigger conditions associated with the multi-step conditional policy are determined, where applicable, according to the status of other vehicles on the road and in the order in which the custom of right-of-way over those vehicles should be handled. By reliably handling complex operations, it can function to allow the vehicle to operate even when a safety operator (and / or a teleoperator located remotely from the vehicle) is not on board, and / or with minimal possibility of intervention occurring while the vehicle is in motion. In the specific example, for example, when the vehicle is approaching another vehicle, and / or when the vehicle is at an intersection with another vehicle, the system and / or method can prevent cases of stop-and-go behavior in which the vehicle attempts to move forward without right-of-way over, causing a safe stop by the operator (or any other stop, such as an emergency stop by the vehicle).
[0018]
[0025] In addition or alternatively, the method and system may also provide any other advantages.
[0019] 3. System
[0026] As shown in Figure 2, the system 200 for the conditional operation of an autonomous agent (hereinafter also referred to herein as an autonomous vehicle, self-vehicle and / or self-agent) includes and / or interfaces with a computing subsystem (hereinafter also referred to herein as a computer), the computing subsystem includes and / or interfaces with a simulator subsystem (e.g., a simulator module, a simulation program, a simulator, etc.) (e.g., communicates with the simulator subsystem, implements the simulator subsystem, runs the simulator subsystem, etc.), and / or is configured to execute and / or trigger a series of simulations in another manner (e.g., as described later). The system 200 preferably further includes and / or interfaces with an autonomous agent (hereinafter also referred to herein as a self-agent and / or autonomous vehicle and / or self-vehicle), a set of one or more sensors (e.g., mounted on the self-agent, mounted on a set of infrastructure devices, etc.), and / or any other components. In addition or alternatively, the system incorporates, by this reference, the entirety of each of the following applications: U.S. Patent Application No. 16 / 514,624 filed July 17, 2019; U.S. Patent Application No. 16 / 505,372 filed July 8, 2019; U.S. Patent Application No. 16 / 540,836 filed August 14, 2019; U.S. Patent Application No. 16 / 792,780 filed February 17, 2020; U.S. Patent Application No. 17 / 365,538 filed July 1, 2021; and U.S. Patent Application No. 17 / 365,538 filed December 14, 2021. U.S. Application No. 17 / 550,461; U.S. Application No. 17 / 554,619 filed December 17, 2021; U.S. Application No. 17 / 712,757 filed April 4, 2022; U.S. Application No. 17 / 826,655 filed May 27, 2022; U.S. Application No. 18 / 073,209 filed December 1, 2022; and U.S. Application No. 18 / 072,939 filed December 1, 2022 may include any or all of the components described in any or all of these.
[0020]
[0027] System 200 preferably includes and / or interfaces with an autonomous vehicle (hereinafter also referred to as an autonomous agent, agent and / or self-agent) (for example, integrated within the autonomous vehicle). The autonomous agent is preferably an autonomous vehicle, more preferably a fully autonomous vehicle and / or a vehicle capable of operating as a fully autonomous vehicle, but may also be a semi-autonomous vehicle and / or any other vehicle.
[0021]
[0028] In a preferred variation, the autonomous vehicle is an automobile (e.g., a car, driverless vehicle, bus, shuttle, taxi, rideshare vehicle, truck, trailer truck, etc.). In addition or alternative, the autonomous vehicle may include any or all of the following: water vehicles (e.g., boats, water taxis, etc.), aircraft (e.g., airplanes, helicopters, drones, etc.), land vehicles (e.g., two-wheeled vehicles, bicycles, motorcycles, scooters, etc.), and / or any other suitable vehicles and / or transport devices, autonomous machines, autonomous devices, autonomous robots and / or any other suitable devices.
[0022]
[0029] The autonomous agent preferably includes and / or interfaces with a computing subsystem, which processes information (e.g., sensor inputs) and performs processing and decision-making for the agent's actions. This may include, for example, determining a set of policies to be executed by the agent (e.g., behavior, actions, high-level behavior and / or planning), behavior and / or actions to be performed by the vehicle, a trajectory to be performed by the agent, a set of control commands to be performed by the vehicle (e.g., actuation subsystem, steering subsystem, brake subsystem, acceleration subsystem, etc.), and / or any other information, either or all of it. Additionally or alternatively, the computing subsystem may also function to perform any or all processes related to perception, prediction, localization, planning, and / or any other processes related to the autonomous agent's actions.
[0023]
[0030] The computing system preferably includes an onboard computing subsystem located within the agent (e.g., integrated within the agent). Additionally or alternatively, the computing system may include any or all of any other suitable computing subsystems and devices, such as a remote computing subsystem (e.g., a cloud computing system, remote computing communicating with the onboard computing system, or a substitute for the onboard computing system), a computing subsystem integrated into an auxiliary device (e.g., a mobile device, a user device, etc.), an edge device including a mobile computing device, and / or any other suitable computing subsystems and devices. For example, in one modification, the agent may communicate with and operate from remote or different computing systems, which may include a user device (e.g., a mobile phone, a laptop, etc.), a remote server, a cloud server, or any other suitable local and / or distributed computing system located remotely from the vehicle. The remote computing subsystem may be connected to one or more systems of the autonomous agent through one or more data connections (e.g., channels), but alternatively, it may communicate with the vehicle system in any suitable manner.
[0024]
[0031] The computing subsystem may include and / or interface with processing subsystems (e.g., a processor or set of processors, a graphical processing unit i.e., a GPU, a central processing unit i.e., a CPU, or any suitable processing circuit) and memory, but may also include any other suitable components as additions or alternatives. The memory may be short-term memory (e.g., volatile, non-volatile, random-access memory, or RAM, etc.) and / or long-term memory (e.g., flash memory, hard disk, etc.). Preferably, the memory functions to store a database (e.g., a lookup table) and / or a set of maps, which may be used when the agent selects any or all of the policies to consider (e.g., as described below), and optionally when selecting any or all of the simulated policies for objects in the vehicle's environment (e.g., during the intent estimation process). In a set of preferred modifications, for example, one or more maps may be referenced and used to check and determine location-related policies (e.g., those to be simulated in the simulation subsystem) for the vehicle to consider. These location-related policies may include multi-step policies (e.g., including at least two actions on a vehicle and a set of trigger conditions to initiate the transition between them), single-step policies (e.g., a single action, a policy without trigger conditions, etc.), or any combination of multi-step and single-step policies. Location-related policies are preferably added to a basic set of policies (e.g., a default set, location-independent policies, etc.) for consideration by the vehicle, but may also be the only policy considered by the vehicle, supplement other dynamically determined policies, supplement default dynamic policies for consideration, and / or may be considered by the vehicle in any other way along with any other policies.As an addition or alternative, any or all of the components and / or processes described in U.S. Application No. 17 / 365,538 filed July 1, 2021, may be used in determining the policy for which the vehicle should consider.
[0025]
[0032] In one modification, for example, the onboard computing subsystem functions to interact with and / or control one or more of the identified components or modules described herein in an operable manner. In a preferred modification, for example, the onboard computing subsystem executes computer instructions for implementing a multi-policy decision module. In a specific example, the processing system and memory together function to dynamically manage a set of policies available to the autonomous agent within the framework of a multi-policy decision framework, such as those described in either or all of U.S. Application No. 16 / 514,624 of July 17, 2019 and U.S. Application No. 17 / 365,538 of July 1, 2021, each of which is incorporated herein by this reference in its entirety. In addition or alternatively, the processing system and memory, and / or any other suitable components may be used for any other suitable function.
[0026]
[0033] The computing subsystem preferably includes, interfaces with, and / or is configured to perform processes in conjunction with the system's simulator subsystem (also referred to herein as the simulation subsystem), the simulator subsystem functions to perform a set of simulations (e.g., as described below), the set of simulations functions to predict future scenarios associated with the agent itself and its environment (e.g., around the agent, within the field of view of the agent's sensors, within a predetermined radius relative to the agent, etc.) and environment agents (e.g., other vehicles, pedestrians, dynamic and / or static objects, etc.). In addition or alternatively, the simulator subsystem may also perform any other functions.
[0027]
[0034] The simulator subsystem preferably includes a simulation program executable by the computing subsystem (e.g., a simulation module, simulation software, a programming language, a software script, and / or program commands), but may include and / or be executable by any other components as an addition or alternative.
[0028]
[0035] The simulator subsystem is preferably configured to perform a forward simulation that functions to predict and analyze how the agent and its environment will evolve in the future (e.g., up to a predetermined future point in time) based on the agent's current and / or past understanding of its environment (e.g., the current positions of the agent and environment agents, past positions of the agent and environment agents, movement information of current and / or past information associated with the agent and / or environment agents, etc.). In a preferred set of modifications, for example, throughout the entire operation of the autonomous vehicle, a set of simulations is performed, for example, continuously, at a predetermined frequency (also known herein as a selection cycle) (e.g., every tenth of a second to every second, at least every second, at least every 5 seconds, every millisecond to every second, 5 to 15 times per second, 10 times per second, 1 to 100 times per second, 1 to 20 times per second, 1 to 50 times per second, etc.) at predetermined intervals, such as when new sensor information is collected, to forward simulate the vehicle's environment over a planned period associated with the simulation (e.g., up to a predetermined time in the future, at each of the set of predetermined time intervals for the predetermined time in the future, such as the next 1 to 10 seconds, the next less than 1 second, the next more than 10 seconds, the next 0.1 to 30 seconds, the next 2 to 8 seconds, the next 5 to 10 seconds, the next 8 seconds, etc.).
[0029]
[0036] In a preferred variation, for example, the future period to be simulated—referred herein to as the planning period—is longer than the period between consecutive sets of simulations for policy selection (specified by the selection cycle). In the example, the planning period is at least an order of magnitude larger than the time between consecutive simulations. In a particular example, simulations are run multiple times per second (e.g., 1-50 times per second, 1-40 times per second, 1-30 times per second, 1-20 times per second, 1-10 times per second, 5-10 times per second, etc.), and each simulation looks ahead by several seconds (e.g., 1-10 seconds, 5-10 seconds, 1-20 seconds, 1-30 seconds, etc.). Additionally or alternatively, the planning period may be equal to the period between consecutive sets of simulations for policy selection, or shorter than the period between consecutive sets of simulations for policy selection, and the planning period and / or the period between consecutive sets of simulations for policy selection may be variable (e.g., dynamically determined) and / or may be determined appropriately in another way.
[0030]
[0037] As an addition or alternative, the simulator subsystem may perform any other simulation and / or type of simulation.
[0031]
[0038] In a concrete example, a multi-policy decision module includes or implements a simulator module or similar machine or system that functions to estimate future (e.g., forward-moving in time) behavior policies (movements or actions) for each of the environmental agents (e.g., other vehicles in the agent's environment) and / or objects (e.g., pedestrians) identified within the autonomous agent's operating environment (real or virtual), including potential behavior policies that may be performed by the agent itself. The simulation may be based on each agent's past actions or past behaviors derived from the agent's current state (e.g., current hypothesis) and / or historical data buffer (preferably including data up to the present). The simulation may provide data relating to the interaction (e.g., relative position, relative velocity, relative acceleration, etc.) between each agent's predicted behavior policy and one or more potential behavior policies that may be performed by the autonomous agent.
[0032]
[0039] As an addition or alternative, the simulation subsystem may operate independently of and / or outside of the multi-policy decision module.
[0033]
[0040] System 100 may optionally include a communication interface for communicating with a computing system, which functions to enable information to be received by the computing system (e.g., from an infrastructure device, from a remote computing system and / or remote server, from a teleoperator platform, from another autonomous agent or other vehicle, etc.) and information to be transmitted from the computing system (e.g., to a remote computing system and / or remote server, teleoperator platform, infrastructure device, another autonomous agent or other vehicle, etc.). The communication interface preferably includes a wireless communication system (e.g., Wi-Fi, Bluetooth, cellular 3G, cellular 4G, cellular 5G, multiple input multiple output (MIMO), one or more radios, or any other suitable wireless communication system or protocol), but may also include, in addition or alternatively, a wired communication system (e.g., modulated power line data transfer, Ethernet, or any other suitable wired data communication system or protocol), a data transfer bus (e.g., CAN, FlexRay), and / or any other suitable components.
[0034]
[0041] System 100 may optionally include and / or interface with a set of infrastructure devices (e.g., as shown in Figure 2) (e.g., to receive information from there), these infrastructure devices, also referred to herein as roadside units, function individually or collectively to observe one or more aspects and / or features of the environment and to collect observational data related to one or more aspects and / or features of the environment. The set of infrastructure devices preferably communicates with the onboard computing system of the autonomous agent, but may also communicate with a teleassist platform, any other components and / or any combination thereof, either additionally or alternatively.
[0035]
[0042] The infrastructure device preferably includes a device located very close to or near the operating position of the autonomous agent, or a device within short-range communication range, and may function to collect data on the circumstances surrounding the autonomous agent and the area adjacent to the autonomous agent's operating zone. In one embodiment, the roadside unit includes one or more offboard sensing devices, including flash LiDAR, thermal imaging devices (thermal cameras), still image or video capture devices (e.g., image cameras and / or video cameras), global positioning systems, radar systems, microwave systems, inertial measurement units (IMUs), and / or any other suitable sensing devices or combinations of sensing devices.
[0036]
[0043] The system preferably includes and / or interfaces with a sensor suite (e.g., a computer vision system, LiDAR, Radar, wheel speed sensors, GPS, cameras, etc.), the sensor suite (also referred to herein as a sensor system) communicates with an onboard computing system and functions to collect information for determining one or more trajectories for the autonomous agent. Additionally or alternatively, the sensor suite may enable the operation of the autonomous agent (e.g., autonomous driving), data capture regarding the circumstances surrounding the autonomous agent, detection of maintenance needs for the autonomous agent (e.g., through engine diagnostic sensors, external pressure sensor strips, sensor state sensors, etc.), detection of cleanliness criteria inside the autonomous agent (e.g., internal cameras, ammonia sensors, methane sensors, alcohol vapor sensors), and / or perform any other appropriate functions.
[0037]
[0044] The sensor suite preferably includes sensors mounted on the autonomous vehicle (e.g., RADAR sensors and / or LIDAR sensors and / or cameras coupled to the exterior of the agent, IMUs and / or encoders coupled to and / or located within the agent, audio sensors, proximity sensors, temperature sensors, etc.), but may also include, as an addition or alternative, sensors located remotely from the agent (e.g., sensors communicating with the agent as part of one or more infrastructure devices), and / or any suitable sensors located at any suitable location.
[0038]
[0045] The sensors may include any or all of the following types of sensors: cameras (e.g., field of view, multispectral, hyperspectral, IR, stereoscopic, etc.), LiDAR sensors, RADAR sensors, compass sensors (e.g., accelerometers, gyroscopes, altimeters), acoustic sensors (e.g., microphones), other optical sensors (e.g., photodiodes, etc.), temperature sensors, pressure sensors, flow sensors, vibration sensors, proximity sensors, chemical sensors, electromagnetic sensors, force sensors, or any other suitable type of sensor.
[0039]
[0046] In a preferred set of modifications, the sensor includes at least a set of optical sensors (e.g., cameras, LiDAR, etc.), and optionally includes any or all of a RADAR sensor, vehicle sensors (e.g., speedometer, compass sensor, accelerometer, etc.), and / or any other sensors.
[0040]
[0047] The system may optionally include and / or interface with one or more vehicle control subsystems, which include one or more controllers and / or control systems, and which include any suitable software and / or hardware components (e.g., a processor and a computer-readable memory device) used to generate control signals for controlling the autonomous agent according to the autonomous agent's routing goals, a selected behavior policy, and / or a selected trajectory.
[0041]
[0048] In a preferred modification, the vehicle control system includes, interfaces with, and / or implements the vehicle's drive-by-wire system. Additionally or alternatively, the vehicle may be operated in accordance with the operation of one or more mechanical components, and / or otherwise implemented.
[0042]
[0049] Additionally or as an alternative, system 100 may include and / or interface with any other suitable components.
[0043] 4. Method
[0050] As shown in Figure 1, the method 100 for the conditional operation of an autonomous agent may include any or all of the following steps: collecting a set of inputs S110; processing the set of inputs S120; determining a set of policies for the agent S130; evaluating the set of policies S140; and operating the agent S150. As an addition or alternative, Method 100 is a step that repeats any or all of the process; the entirety of each of the applications is incorporated by this reference: U.S. Patent Application No. 16 / 514,624 filed July 17, 2019; U.S. Patent Application No. 16 / 505,372 filed July 8, 2019; U.S. Patent Application No. 16 / 540,836 filed August 14, 2019; U.S. Patent Application No. 16 / 792,780 filed February 17, 2020; U.S. Patent Application No. 17 / 365,538 filed July 1, 2021; U.S. Patent Application No. 17 / 365,538 filed December 14, 2021 Method 100 may include any or all of the processes described in any or all of the following: U.S. Application No. 17 / 550,461; U.S. Application No. 17 / 554,619 filed December 17, 2021; U.S. Application No. 17 / 712,757 filed April 4, 2022; U.S. Application No. 17 / 826,655 filed May 27, 2022; U.S. Application No. 18 / 073,209 filed December 1, 2022; and U.S. Application No. 18 / 072,939 filed December 1, 2022, or any other suitable processes performed in any suitable order. Method 100 may be performed using the aforementioned System 200 and / or any other suitable system.
[0044]
[0051] Method 100 is preferably configured to interface with the agent's multi-policy decision-making process (e.g., a multi-policy decision-making task block on a computer-readable medium) and any associated components (e.g., a computer, processor, software module, etc.), but may also interface with any other decision-making process, either additionally or alternatively. For example, in a preferred set of modifications, the multi-policy decision-making module of a computing system (e.g., an onboard computing system) includes a simulator module (or similar machine or system) (e.g., a simulator task block on a computer-readable medium) that functions to predict (e.g., estimate) the effects of future (i.e., temporally forward) behavioral policies (actions or behaviors) implemented by the agent, and optionally the effects of behavioral policies (actions or behaviors) performed on each of the configured environment agents (e.g., other vehicles in the agent's environment) and / or objects identified in the agent's operating environment (e.g., pedestrians). The simulation may be based on each agent's historical actions or historical behavior derived from the agent's current state (e.g., current hypothesis) and / or historical data buffer (preferably including data up to the present moment). The simulation may provide data on the interaction between the predicted behavioral policies of each environment agent and one or more potential behavioral policies that can be implemented by the autonomous agent (e.g., relative position, relative velocity, relative acceleration, etc.).
[0045]
[0052] The data obtained from the simulation may be used to determine (e.g., calculate) any number of metrics, which may individually and / or collectively function to evaluate any or all of the following: the potential impact of the agent on any or all of the environment agents when implementing a particular policy; the risks of implementing a particular policy (e.g., collision risk); the extent to which the agent makes progress toward a particular objective by implementing a particular policy; and / or any other metrics related to evaluating, comparing, and / or ultimately selecting the policies that the agent will implement in actual operation.
[0046]
[0053] The set of metrics may optionally include, and / or be determined collaboratively (e.g., by aggregating any or all of the sets of metrics described below) the cost function (and / or loss function) associated with each proposed agent policy, based on a set of simulations performed against the proposed policy. Additionally or alternatively, the sets of metrics described below may be determined and / or analyzed individually, other metrics may be determined, metrics may be aggregated in other appropriate ways, and / or metrics may be constructed in other ways. Using these metrics (e.g., scores) and / or functions, the best policy from the set of policies may be selected (e.g., the policy with the highest metric value, the policy with the lowest metric value, the policy with the lowest cost / loss function, the policy that optimizes the objective function (e.g., maximizing, minimizing, etc.), the policy with the best risk-normalized reward function, etc.) by comparing the metrics and / or functions among the various proposed policies and selecting a policy based on such comparison.
[0047]
[0054] The multi-policy decision-making process may, in addition or alternative to, include any other processes, such as any or all of the processes described in U.S. Application No. 16 / 514,624 filed July 17, 2019 and U.S. Application No. 17 / 365,538 filed July 1, 2021, each of which is incorporated in whole by this reference, or any other suitable processes that are performed in any suitable order, and / or interface with those processes.
[0048]
[0055] Additionally or alternatively, method 100 may include and / or interface with any other decision-making process.
[0049] 4.1 Method - Step S110 for collecting the set of inputs
[0056] Method 100 may include a step S110 that collects a set of inputs which functions to receive information for performing and / or initiating any or all of the remaining processes of Method 100. In preferred modifications, for example, S110 may function to receive information for performing any or all of the following steps: (e.g., S120) reviewing and / or characterizing a scene associated with its agent; (e.g., S130) selecting a set of policies for consideration by its agent; (e.g., S140) running a set of forward simulations and / or evaluating policies for consideration; (e.g., S150) operating its agent; (e.g., S150) triggering transitions of actions within a multistep policy; and / or functioning to perform any other purpose.
[0050]
[0057] S110 is preferably executed continuously throughout the entire operation of the agent (e.g., at a default frequency, at irregular intervals, etc.), but may also be executed in accordance with any or all of the cycles associated with the agent, such as the selection cycle associated with the agent (e.g., a 10Hz cycle, a 5-20Hz cycle, etc.) (e.g., a cycle for selecting a policy to implement, a cycle for selecting a new policy, etc.), the perception cycle associated with the agent, the planning cycle associated with the agent (e.g., a 30Hz cycle, a 20-40Hz cycle, a cycle that occurs more frequently than the selection cycle, etc.) (e.g., at the start of each cycle, during each cycle, etc.); in response to a trigger (e.g., a request, the start of a new cycle, etc.); and / or at any other point in method 100.
[0051]
[0058] The inputs preferably include sensor inputs received from a sensor suite mounted on the agent (e.g., a camera, Lidar, Radar, motion sensors (e.g., accelerometer, gyroscope, etc.), OBD port output, etc.), position sensors (e.g., GPS sensor, etc.), but may also include, additionally or alternatively, historical information related to the agent (e.g., estimated historical state of the agent) and / or environmental agents (e.g., estimated historical state of the environmental agent), sensor inputs, information, and / or any other inputs from sensor systems outside the agent (e.g., mounted on other agents or environmental agents, on infrastructure devices and / or roadside units, etc.).
[0052]
[0059] The input preferably includes information associated with the self-agent, which refers to the vehicle operating in Method 100. This may include information characterizing the self-agent's position (e.g., relative to the world, to one or more maps, to other objects, etc.), the self-agent's motion (e.g., velocity, acceleration, etc.), the self-agent's orientation (e.g., angle of direction of travel), the performance and / or health of the self-agent and any of its subsystems (e.g., sensor health, computing system health, etc.), and / or any other information.
[0053]
[0060] The input may more preferably include and / or be used to determine (e.g., pre-process, process, etc.) information associated with (e.g., characterizing) the agent's environment, which may include other objects (e.g., vehicles, pedestrians, stationary objects, etc.) near the agent (e.g., within the field of view of its sensors, within a predetermined distance, etc.); the presence of objects not directly detected in the agent's environment (e.g., by obstacles in the agent's environment that may conceal the presence of such objects); environmental features around the agent (e.g., for reference in a map to determine the agent's location, etc.); and / or any other information. In one variation, for example, the set of inputs may include information characterizing any or all of the location, type / classification (e.g., vehicle vs. pedestrian), direction (e.g., angle of travel), and / or motion of environment agents being tracked by System 200 (e.g., from sensors mounted on the agent, from sensors in the agent's environment, from sensors mounted on objects, etc.), where environment agents refer to other vehicles in the agent's environment (e.g., manually driven vehicles, autonomous vehicles, semi-autonomous vehicles, etc.). Additional or alternative, the set of inputs may include information characterizing roads and / or other landmarks / infrastructure features (e.g., to locate, identify, etc.) so that the agent can locate its own position in its environment (e.g., to refer to a map) (e.g., lane location, road edge location, traffic signal location and type, where the agent is relative to these landmarks, etc.), and / or any other information.
[0054]
[0061] S110 may optionally include a step of preprocessing any or all of the set of inputs, which serves to prepare the set of inputs for analysis in a subsequent process of the method 100. The step of preprocessing the set of inputs may optionally include a step of calculating state estimates of the agent and / or the environment agent based on the set of inputs. The state estimates preferably include at least the position and velocity associated with the agent, but may also include, additionally or alternatively, other motion / movement information such as directional information (e.g., azimuth angle), acceleration and / or angular motion parameters (e.g., angular velocity, angular acceleration, etc.), and / or any other parameters.
[0055]
[0062] The step of preprocessing the input set may optionally include, as an addition or alternative, a step of determining one or more geometric properties / features associated with the environment agent / object (e.g., using a computer vision module of a computing system), such as defining a 2D geometric shape associated with the environment agent (e.g., a 2D geometric hull, 2D profile, agent outline, etc.), a 3D geometric shape associated with the environment agent, and / or any other geometric shape. This may be used to determine, for example, one or more lanes in which the environment agent may reside (e.g., associated with probability / confidence values); the width of the lane obstructed by the object (e.g., for selecting the turning behavior described later); parameter values for implementation in trigger conditions associated with a multistep policy; the size of the agent or object (e.g., an obstacle); and / or any other information.
[0056]
[0063] The step of preprocessing the input set may optionally include, as an additional or alternative step, determining one or more classification labels associated with any or all of the set of environmental objects / agents, and further optionally, determining the probability and / or confidence (expressed as probability) associated with the classification labels. The classification labels preferably correspond to agent types, such as vehicles (e.g., binary classification of vehicles) and / or vehicle types (e.g., sedans, trucks, shuttles, buses, emergency vehicles, etc.), pedestrians, animals, inanimate objects (e.g., road obstacles, construction machinery, traffic cones, etc.), and / or any other types of agents. The classification labels are preferably determined at least in part based on any or all of the agent's geometric characteristics (e.g., size, profile, 2D hull, etc.) and state estimates (e.g., speed, position, etc.), but may be determined in other ways, as an additional or alternative step.
[0057]
[0064] As an addition or alternative, S110 may include any other process.
[0058] 4.2 Method - Step S120 for processing the set of inputs
[0065] Method 100 includes step S120 which processes a set of inputs that can function to detect and / or characterize scenarios (hereinafter also referred to as scenes) and / or contexts associated with an agent, the scenarios and / or contexts which may be used in subsequent processes of the Method to determine the optimal set of policies for the agent to consider. Additionally or alternatively, S120 may enable the execution of any other processes of Method 100, function to refer to a set of maps and / or databases based on the set of inputs to inform the agent of the set of policies that will be considered by the vehicle, and / or function to perform any other functions.
[0059]
[0066] S120 is preferably executed in response to S110, but additionally or alternatively, it may be executed in response to any other process of method 100, continuously (e.g., at a predetermined frequency, according to the planner cycle frequency, etc.), in response to a trigger, and / or at any other point in time. Additionally or alternatively, method 100 may be executed without S120, and / or the method may be adequately executed in another manner.
[0060]
[0067] S120 is preferably executed in the agent's computing system (e.g., onboard computing system), for example, in or in conjunction with the agent's planner. Additionally or alternatively, S120 may be executed in / with the agent's perception module (e.g., perception processor, perception computing subsystem), in / with the agent's prediction module, and / or with any other component.
[0061]
[0068] S120 may optionally include a step of characterizing the environment surrounding the vehicle (e.g., the current environment, the expected environment, and / or future environment) as a scenario (and / or context described later), preferably based on processing perceptual information (e.g., sensor data) to determine geometric features associated with the environment of the agent. Optionally, the step of determining the scenario (and / or context) may include a step of comparing any or all of these geometric features with a set of maps (e.g., custom maps reflecting local static objects) and / or databases, which may function to identify the scenario (and / or context) based on the positioning of the agent within one or more maps (e.g., features / geometry detected within the agent's field of view) and / or based on identifying geometric features in the database. Additionally or alternatively, scenarios and / or contexts may be characterized based on dynamic features detected and / or determined by the vehicle (e.g., based on sensor data, based on processing of sensor data, etc.) (e.g., the presence of pedestrians, the presence of pedestrians in a crosswalk, the presence of dynamic obstacles in the vehicle's direction of travel, etc.), combinations of static and dynamic features, and / or any other information.
[0062]
[0069] The step of characterizing the environment surrounding the vehicle may include the step of determining (e.g., detecting, characterizing, identifying, etc.) scenarios (also referred to herein as scenes) associated with the agent and its location. Scenarios preferably define driving conditions and / or road features in which the agent is and / or approaching, and optionally include scenarios that are or could be complex for the vehicle to navigate (e.g., involving pedestrians, involving right-of-way customs, involving the vehicle violating general road rules, involving the risk of the vehicle being stopped for an extended period, involving the risk of causing confusion to other drivers, etc.). Examples of scenarios, but not limited to, include pedestrian crossings (e.g., as shown in Figure 4); four-way intersections (e.g., as shown in Figure 5); other intersections (e.g., three-way intersections); obstacles on the road (e.g., as shown in Figure 6); transitions to one-way streets; construction zones; merging zones and / or lane ends; parking lots; and / or any other scenarios.
[0063]
[0070] S120 optionally further includes the step of determining a set of contexts associated with a scenario, the contexts further characterizing the scenario. The contexts may optionally include and / or define one or more features associated with the scenario (e.g., geometric features, parameters, etc.), such as the size and / or type of obstacles ahead; the type of intersection (e.g., a four-way intersection, a three-way intersection, whether there is a pedestrian crossing at the intersection, etc.); and / or any other information. The scenario may be associated with a single context, multiple contexts, no context, and / or any other information. Alternatively, the scene and contexts may be the same, only the scene may be determined, only a set of one or more contexts may be determined, and / or S120 may be appropriately performed in another way.
[0064]
[0071] In any iteration of S120, the agent may be associated with one or all of the following: a single scenario, multiple scenarios, no scenario, additional features, and / or any other information. Alternatively, the agent may always be associated with one or more scenarios. Additionally or alternatively, in any iteration of S210, the agent may be associated with one or more contexts (for example, as described below). Further additional or alternatively, the method may be performed without characterizing a set of scenarios and / or contexts.
[0065]
[0072] S120 may also include, as an addition or alternative, a step of referencing a set of maps and / or databases based on sensor data (e.g., raw sensor data, processed sensor data, aggregated sensor data, etc.), which may function to check and / or retrieve policies that will be considered in S130. For example, in a series of modifications, sensor data representing the vehicle's position (e.g., determined by a position sensor such as a GPS sensor) may be compared with maps and / or databases to retrieve a set of policies (e.g., multi-step policies, single-step policies, etc.) that the vehicle will consider based on its position. Any other sensor data may be used as an addition or alternative. For example, to retrieve one or more policies, a determination may be used that the vehicle is located in close proximity (e.g., within its default distance threshold) and / or approaching a default position based on a map reference (e.g., based on the angle of the vehicle's direction of travel, based on the angle of the direction of travel and its position, etc.).
[0066]
[0073] For example, a series of examples may refer to a custom map that assigns a default scenario and / or context to any or all locations on the map. In other examples, the agent's location (e.g., GPS coordinates) may be identified and compared to the map to determine the default scenario associated with that location.
[0067]
[0074] The step of determining the scenario and / or context may optionally be performed using one or more computer vision processes, machine learning models and / or algorithms, and / or any other tools, as an additional or alternative step.
[0068]
[0075] In a specific example, a set of policies specific to the parking environment (e.g., a multi-step policy involving the vehicle finding and parking in an available parking space) may be added to the set of policies the vehicle considers in response to the determination that the vehicle is located in and / or approaching a parking space, based on the vehicle's location and referencing a default labeled map. Alternatively, policies specific to the parking environment may be added for consideration based on the detection of features corresponding to the parking environment (e.g., processing camera data to identify rows of parking spaces). Further alternatives include the fact that policies specific to the parking environment (and / or other scenarios) may always be considered by the vehicle (e.g., they may be included in a default set of policies the vehicle considers), and the simulation output (e.g., metrics) reflects that these are not relevant and / or optimal policies to implement when the vehicle is not in a parking space.
[0069]
[0076] For example, in a particular example (such as shown in Figure 4), a multi-step policy specific to a vehicle approaching a crosswalk may be considered in response to any or all of the following: identification of crosswalk-specific features (e.g., road markings indicating crosswalk features, pedestrians and / or pedestrian-sized objects moving perpendicular to the direction of travel in the lane, and / or crosswalk signs); or reference to a map containing default crosswalk labels based on the vehicle's location, or based on any other information and / or any combination of information. Alternatively, a crosswalk-specific policy (e.g., the multi-step policy shown in Figure 4) may always be considered by the vehicle itself (e.g., as part of a default set of policies to be considered).
[0070]
[0077] For example, in a specific example (such as shown in Figure 5), a multi-step policy specific to a vehicle approaching a four-way intersection may be considered in response to any or all of the following: identification of intersection-specific features (e.g., road markings indicating merging lanes in orthogonal directions, detection of vehicles dynamically moving in orthogonal directions); referencing a map containing a default intersection label based on the vehicle's location; and based on any other information and / or any combination of information. Alternatively, an intersection-specific policy (e.g., the multi-step policy shown in Figure 5) may always be considered by the vehicle itself (e.g., as part of a default set of policies to be considered).
[0071]
[0078] For example, in a specific example (such as shown in Figure 6), a multi-step policy specific to a vehicle encountering an obstacle (also referred to herein as an obstacle) may be considered based on any other information and / or any combination of information in response to either or all of the following: identifying the obstacle (e.g., detecting an object overlapping the vehicle's lane); or referring to a map containing default obstacle labels based on the vehicle's location (e.g., for static, long-lasting obstacles such as construction machinery, or potholes in the road). Alternatively, an obstacle-specific policy (e.g., the multi-step policy shown in Figure 6) may always be considered by the vehicle (e.g., as part of a default set of policies to be considered).
[0072]
[0079] For example, in a preferred implementation of this particular example, since obstacles can occur at any location, an obstacle-specific multi-step policy is included in the default set of policies for consistent consideration by the vehicle.
[0073]
[0080] As an addition or alternative, S120 may include any other subprocesses, such as determining a set of features associated with the environment agent, and / or any other processes.
[0074] 4.3 Method - Step S130 for determining the set of policies for the agent
[0081] Method 100 preferably includes step S130 of determining (e.g., selecting, aggregating, compiling, etc.) a set of policies for the agent, which functions to determine a set of policies that the agent will consider implementing in subsequent processes of Method 100. Additionally or alternatively, S130 may function to identify an optimal set of policies for the vehicle to consider (e.g., based on the specific environment surrounding the vehicle), a minimal set of policies (e.g., to reduce the computational load associated with the relevant simulations), a prioritized set of policies (e.g., so that the simulations run in the optimal order if there is no time to select policies), a comprehensive set of policies (e.g., all policies that may be implemented and / or applied to the vehicle), any other set of policies, and / or S130 may perform any other function.
[0075]
[0082] S130 is preferably executed in response to S120, but additionally or alternatively, it may be executed in response to any other process of method 100, sequentially (e.g., at a default frequency, according to a selected cycle frequency, etc.), in response to a trigger, and / or at any other time. Additionally or alternatively, S130 may be executed without S120, and / or at any other time.
[0076]
[0083] For example, in a preferred set of modifications, S130 is executed according to a selection cycle associated with the vehicle, which defines the frequency at which a set of policies is determined and evaluated for its own agent. In a set of specific examples, the selection cycle is associated with frequencies between 1 and 50 Hz (e.g., 10 Hz, 20 Hz, 30 Hz, 5 to 15 Hz, 1 to 20 Hz, etc.), but alternatively, it may be associated with frequencies less than 1 Hz, frequencies greater than 50 Hz, a set of irregular intervals, and / or any other time.
[0077]
[0084] The set of policies preferably includes multiple policies so that multiple policies are evaluated in S140 (for example, in a simulation) and the best policy for the vehicle is selected and implemented in the vehicle (for example, according to a scoring system). Alternatively, any iteration of S130 may include a step to determine a single policy for evaluation, such as a single multi-step policy determined based on the scenario and / or context determined in S120.
[0078]
[0085] The number of policies to be determined in each iteration of S130 may be a default number (e.g., a fixed number, a constant, etc.), a variable number, and / or any other number, or all of them.
[0079]
[0086] The set of policies may include single-step policies, multi-step policies, combinations of single-step and multi-step policies, and / or any other policies, either or all of them. In this specification, a step preferably refers to a single action of the vehicle (e.g., a task, behavior, etc.), but may also, in addition or alternatively, refer to a grouping of actions, sub-actions, and / or any other behavior of the vehicle. For example, in a preferred set of modifications, each action in a multi-step policy is an action specified by a single-step policy, along with a set of trigger conditions associated with transitions between these actions. In addition or alternatively, any or all actions in a multi-step policy may differ from any single-step policy.
[0080]
[0087] Any or all of the multi-step policies may be dynamically determined based on S120, for example, by detecting the scenarios and / or contexts associated with the agent, and the detected scenarios and / or contexts define the associated multi-step policies (for example, according to the default mapping). Additionally or alternatively, the detected scenarios and / or contexts may define multiple multi-step policies, one or more single-step policies, and / or any other policies that are considered in subsequent processes of method 100.
[0081]
[0088] In addition or alternatively, any or all of the multistep policies may be determined by referencing maps and / or databases, or by prior determination, and / or by any other method and / or any combination of methods.
[0082]
[0089] As an additional or alternative, any or all of the multi-step policies may be dynamically constructed (e.g., modularly) and / or adjusted (e.g., based on the completion of one or more actions of the multi-step policy) before consideration (e.g., simulation). For example, in one variation, the multi-step policy may be constructed depending on the environment surrounding the vehicle. This may include, for example, aggregating some or all of different multi-step policies (e.g., detecting an obstacle as the vehicle approaches a crosswalk), linking multiple single-step policies (e.g., if it is determined that the vehicle has already performed one or more of the initial actions of the multi-step policy (e.g., from time step t3 to t) 10 For the selected multistep policy, as shown in the table in Figure 8F, this may include removing actions and / or trigger conditions from the multistep policy, modifying trigger conditions associated with the multistep policy, and / or constructing and / or modifying the multistep policy in a different way.
[0083]
[0090] Each multi-step policy is preferably configured to enable its agent to behave in a human-like manner and / or to enable its agent to move forward toward its goals (e.g., to reach a destination, prevent undesirable / unnecessary stops, prevent handover by a human operator, prevent intervention from a teleoperator, etc.), which may include any or all of the following: reducing and / or eliminating the time the agent is idle (even though it can / is allowed to move because only a single-step policy is being simulated over a long planning cycle); reducing the number of stops and / or occurrences of roadside stops; reducing behavior that interferes with other vehicles on the road; reducing dangerous behavior; reducing overly conservative behavior; and / or the multi-step policy may be configured in a different way.
[0084]
[0091] In one variation, one or more multistep policies are associated with a single scenario and one or more contexts. For example, in a set of specific examples, a scenario in the form of a four-way intersection is determined, which is associated with two (or more) contexts, including that a vehicle is approaching the intersection and that there is a crosswalk at the intersection. In another set of specific examples (e.g., shown in Figure 5), a scenario in the form of a four-way intersection is determined, which is associated with two contexts, the two contexts including that a vehicle is approaching the intersection and that there is no crosswalk at the intersection. In variations involving multiple contexts, it is preferable that the multistep policy defines an action aggregated from the multiple contexts (e.g., in a non-repeating manner). As an addition or alternative, the multistep policy associated with each context may be evaluated individually in S140. Further addition or alternative, each scenario may be associated with a single context, and / or the multistep policy may be determined in a different way.
[0085]
[0092] Each set of multi-step policies is preferably associated with a set of conditions (also referred to herein as trigger conditions), e.g., specifies, defines, includes, etc., where the conditions function to initiate transitions between steps of the multi-step policy. The set of conditions is preferably defined and / or determined using a set of parameters, which may include any or all of the following: distance parameters (e.g., distance to a crosswalk, distance to an intersection, distance to a dynamic object, etc.), size parameters (e.g., size of an obstacle, width of a lane, etc.), presence and / or proximity of an object (e.g., a pedestrian), time parameters, motion parameters (e.g., motion of the agent itself, motion of an environmental object, speed threshold, acceleration threshold, etc.), and / or any other parameters. Further or alternatively, any or all of the conditions may be associated with a set of driving rules, such as a custom of priority right of way between multiple vehicles (e.g., a four-way stop, at an intersection, etc.). Further or alternatively, any or all of the conditions may be associated with any other information.
[0086]
[0093] The parameters may be predetermined, dynamically determined, and / or any combination of any of them.
[0087]
[0094] The steps that trigger transitions between steps in a multi-step policy (e.g., in a simulation, in the actual operation of the vehicle itself) preferably include steps to check whether conditions are met, and these steps may include any or all of the following: comparing a parameter value with one or more thresholds; comparing a parameter value with a defined set of values (e.g., optimal values); evaluating a decision tree and / or a set of algorithms; referring to a lookup table and / or database; evaluating a set of models (e.g., a training model, a machine learning model, etc.); checking whether the agent has priority (or whether any other driving / traffic / courtesy customs are met); and / or evaluating a set of parameters in another way.
[0088]
[0095] The single-step policy is preferably determined in advance, such as remaining constant across all iterations of S130. This may include, but is not limited to, a set of standard (common) policies considered in each selection cycle, such as stop, decelerate, accelerate, drive straight (e.g., lane keeping), change lanes, merge, and / or any other policies. Alternatively, any or all of the single-step policies may be determined dynamically based on any or all of the following: the location associated with the agent (e.g., comparison with a map), any other inputs collected in S110 (e.g., sensor information), information associated with the agent's environment (e.g., other vehicles, adjacent objects, etc.), and / or any other information. Further or alternatively, any or all of the single-step policies may be determined based on S120, for example, based on the scenario and / or context associated with the agent.
[0089]
[0096] In the first set of variations, the set of policies determined in S130 includes a default set of single-step policies, a default set of multi-step policies, and optionally, if a scenario and / or context and / or a specific location of the vehicle are determined in S120, one or more environment-specific multi-step policies. In a specific example, if no scenario and / or context are determined in S120, the set of policies determined in S130 includes only single-step policies. In another specific example, if no scenario and / or context are determined in S120, the set of policies determined in S130 includes both a default single-step policy and a default multi-step policy.
[0090]
[0097] In additional or alternative modifications, S120 may define a single-step policy, the set of policies may include a default multi-step policy, and / or any other policy may be determined.
[0091]
[0098] S130 may optionally include a step that leverages policy determinations from previous iterations of S130 (e.g., previous selection cycles) when determining which policies to consider in the current iteration of S130 (e.g., the current selection cycle). This may include, for example, a step that transfers the best policy from the previous selection cycle to the new / current cycle. This may also include, additionally or alternatively, a step that reduces the number of policies considered in the current / future selection cycles. For example, if the agent is still in the middle of a multi-step policy when the next selection cycle occurs, consideration of other policies may be optionally excluded and / or minimized (e.g., only policies that may naturally occur in a particular scenario may be considered, or only policies that reflect emergencies may be considered alongside the current multi-step policy). Alternatively, a standard set of policies may be consistently considered in each selection cycle, regardless of whether the agent implements a multi-step policy or not.
[0092]
[0099] As an addition or alternative, S130 may include a step to modify the multistep policy (for example, as described above) to remove the actions and / or trigger conditions from the multistep policy when these actions and / or trigger conditions have already been initiated by the vehicle and / or are no longer relevant to the implementation.
[0093]
[0100] In a first variant of the series of S130, the step of determining the set of policies includes determining one or more multi-step policies in response to the detection of a set of scenarios and / or contexts associated with the agent in S120. The set of policies may optionally include one or more default policies, such as a set of standard single-step policies.
[0094]
[0101] In the first specific example (for example, shown in Figure 4), a scenario corresponding to a pedestrian crossing is detected in the agent's environment (e.g., future environment) (e.g., based on dynamic processing of sensor data, based on camera data, based on location data, and based on a map reference based on location data, etc.), and this scenario may optionally be associated with the context of a vehicle approaching the crossing. In response, a multi-step policy is added to the set of policies to be evaluated for consideration (e.g., in S140), and the multi-step policy defines three actions. The first action involves the agent stopping, which is preferably defined based on a fixed distance parameter to the crossing (e.g., distance to the nearest crossing boundary to the agent, distance to the wide end of the crossing, etc.), such that the first action stipulates that the agent stops at a predetermined distance from the crossing boundary. Alternatively, the fixed distance parameter may instead include a range of acceptable distances and / or any other parameters. For example, in one example, instead of a distance parameter, the first action may define a vehicle deceleration value, a set of decreasing speeds at which the vehicle moves forward, and / or any other parameters. The second action includes waiting until the crosswalk is clear. The second action is triggered in response to a trigger condition indicating that the first action has been completed / fulfilled (e.g., in a simulation, in actual operation, etc.). In this example, the trigger condition from the first action to the second action preferably includes detecting that the vehicle has stopped (e.g., that its speed has become zero), but may also include, additionally or alternatively, detecting that the vehicle is within a predetermined distance from the crosswalk, that the vehicle has decelerated by a certain amount, that the vehicle's speed is below a predetermined threshold, or that the vehicle is in a predetermined position, and / or the trigger condition may be appropriately defined in another way.The third action preferably includes passing through a crosswalk (and optionally continuing to move forward, such as in lane-keeping behavior), and is triggered when it is detected that there are no pedestrians at the crosswalk (e.g., in a simulation, in actual operation, etc.) through determination of whether there are no pedestrians within the boundaries of the crosswalk, whether all pedestrians are within a predetermined distance from the vehicle, and / or determination of any other conditions. Additionally or alternatively, the trigger conditions may include determination of whether there are no objects within the crosswalk, whether there are no vehicles on the road ahead of the crosswalk so as not to block the crosswalk (e.g., there is at least a distance of a car ahead of the crosswalk), and / or determination of any other conditions being met. This multistep policy is preferably simulated (along with other policies considered by the vehicle), and the applicable actions and trigger conditions of the multistep policy are simulated over the planned duration of the simulation, and as a result, a score for the multistep policy is calculated and may be used to determine whether to implement the multistep policy (e.g., the initial action of the multistep policy, or a portion of the multistep policy that can be implemented by the vehicle before the next selection cycle) during the operation of the vehicle. In addition or alternatively, any other multistep policy (having any appropriate actions and / or conditions) may be evaluated, and / or, there may be no multistep policy in the set of policies for the scenario.
[0095]
[0102] In the second specific example (for example, shown in Figure 5), a scenario corresponding to a four-way intersection is detected within the agent's environment, and this intersection is further characterized (for example, through context determination) as an intersection without a pedestrian crossing. The scene evaluation (for example, in S120) may optionally determine the order in which agents in the environment (for example, agents waiting at the intersection and / or agents approaching the intersection) have priority (for example, if it is detected that the four-way intersection has a stop sign and / or yield sign instead of a traffic light). In response to this scenario and / or context, a multi-step policy is added to the set of policies to be evaluated for consideration (for example, in S140), and the multi-step policy defines three actions. The first action involves the agent stopping at a default position or similar location relative to one or more features of the intersection (for example, a stop sign, the edge of a crossing lane, etc.). The second action includes the agent waiting, which is preferably triggered in response to any or all of the following conditions: the agent has stopped, the agent is positioned at a specific location relative to the intersection, and / or any other conditions. The third action includes the agent passing through the intersection (e.g., moving forward, turning left, turning right, etc.), which is preferably triggered in response to the detection that the intersection is clear and that all other vehicles with priority have passed through the intersection before the agent. This multi-step policy is preferably simulated (along with other policies considered by the vehicle) and the applicable actions and trigger conditions of the multi-step policy are simulated over the planned duration of the simulation, and as a result, a score for the multi-step policy is calculated and may be used in determining whether to implement the multi-step policy (e.g., the initial actions of the multi-step policy, parts of the multi-step policy that the vehicle can execute before the next selection cycle, etc.) during the operation of the vehicle.Additionally or alternatively, any other multistep policy (with any appropriate action and / or condition) may be evaluated, and / or, there may be no multistep policies in the set of policies for the scenario.
[0096]
[0103] In the third specific example (shown, for example, in Figure 6), a scenario is detected that corresponds to an obstacle on the road ahead of the local agent. S120 may optionally further include a step of detecting the type of obstacle, the size of the obstacle (e.g., length, width, amount of lane obstructed by the object, etc.), and / or any other features associated with the obstacle and / or the local agent's environment. In response to this scenario and / or context, a multi-step policy is added to the set of policies to be evaluated for consideration (e.g., in S140), and the multi-step policy defines two actions. The first action defines that the vehicle turns left (or, optionally, right depending on the direction of traffic and / or the location of the obstacle) according to one or more parameters associated with the obstacle (e.g., width, length, etc.) and / or the placement of the obstacle relative to the road (e.g., the percentage of lane obstructed by the object). Additional or alternative, the turning action may define that the vehicle follows lane boundaries and / or any other infrastructure markers (e.g., centered). The second action specifies that the vehicle resumes normal driving within its lane (e.g., returning to the center line of the lane and moving forward) and is preferably triggered in response to parameters associated with the obstacle, such as detecting that the vehicle has reached a position within a predetermined distance from the far end of the obstacle (e.g., ensuring the vehicle has adequate clearance from the obstacle). This multistep policy is preferably simulated (along with other policies considered by the vehicle) and the applicable actions and trigger conditions of the multistep policy are simulated over the planned duration of the simulation, and as a result, a score for the multistep policy may be calculated and used to determine whether the vehicle implements the multistep policy (e.g., the initial action of the multistep policy, or any part of the multistep policy that the vehicle can execute before the next selection cycle) during its operation. In addition or alternatively, any other multistep policy may be considered (along with any appropriate actions and / or conditions), and / or the set of policies for a scenario may not include any multistep policy.
[0097]
[0104] As an addition or alternative, S130 may include any other suitable process.
[0098] 4-4 Method - Step S140 to evaluate the set of policies
[0105] The method may include step S140, which evaluates a set of policies that function to select the best policy for its agent to implement.
[0099]
[0106] S140 is preferably performed in response to S130, but additionally or alternatively, it may be performed in response to any other process of method 100, in response to a trigger, according to a cycle and / or frequency, and / or may be performed appropriately in another way.
[0100]
[0107] In a preferred set of variations, S140 may be performed according to a multi-policy decision-making process (e.g., one described above), which may include simulating each of the proposed set of policies (e.g., forward simulation, forward simulation 5-10 seconds ahead, etc.), determining a score of 1 or more for each of the proposed set of policies, comparing scores between different policies, and selecting a policy for the agent to implement based on the comparison. Any or all of the scores may be determined selectively based on detected scenarios associated with the agent, and as a result, policies that are better configured for a particular scenario may be appropriately given higher scores (and / or have a lower cost / loss function). Additionally or alternatively, any or all of the scores may be determined without detected scenarios, based on other information (e.g., one or more contexts), and / or in other ways.
[0101]
[0108] As an addition or alternative, S140 may include any other processes that are executed in any appropriate order.
[0102] 4.5 Method - Step S150 to activate the agent
[0109] The method may include step S150, which causes the agent to operate (for example, through a set of control commands, such as those shown in Figure 3, configured to implement a selected policy) in order to control the agent itself.
[0103]
[0110] S150 is preferably performed using the agent's own planner and controller, but may be performed using any other suitable subsystem, either additionally or alternatively. In a series of preferred modifications, the planner is associated with a planning frequency higher than the selected cycle frequency, but may be associated with a lower frequency and / or any other frequency, either additionally or alternatively.
[0104]
[0111] If the agent implements a multi-step policy, it is preferable that S150 includes a step to check whether a set of trigger conditions associated with the multi-step policy (e.g., based on sensor data collected in S110 using a set of computer vision processes) are met so that the transitions between steps of the multi-step policy can be properly triggered. In addition or alternatively, S150 may include any other suitable process.
[0105]
[0112] If a multi-step policy is selected for implementation in S140, S150 preferably includes a step of implementing a portion of the multi-step policy that applies to the duration of that selection cycle and any future selection cycles in which that particular multi-step policy is subsequently selected. This may mean that only a portion of the selected multi-step policy is actually implemented in the operation of the vehicle, for example, in an implementation where the planning duration of the simulation is longer than the time between consecutive selection cycles. Additional or alternative, in an implementation, a multi-step policy may actually be implemented at least partially as one or more single-step policies, such as when any or all of the multi-step policies are already implemented and / or are no longer relevant. Furthermore, additional or alternative, in an implementation, a multi-step policy may actually be implemented at least partially initially as a multi-step policy, then as a modified version of the multi-step policy, and then as one or more single-step policies.
[0106]
[0113] Any or all of the trigger conditions associated with the selected multistep policy may optionally be triggered only in the simulation of the multistep policy, based on factors such as the rate at which the new policy is evaluated in S140. Additionally or alternatively, if the trigger condition is met within the selected cycle (e.g., during a continuous period of time in which the policy is evaluated), the transition between the trigger condition and the multistep policy action may be implemented in the actual operation of the vehicle.
[0107]
[0114] In one modified example shown in Figures 8A to 8F, the vehicle is shown at various times and in its relevant positions within its environment during the actual operation in Figures 8A to 8C (represented as the vehicle with a solid outline). When the vehicle approaches a crosswalk, a multi-step policy including actions 1, 2, and 3, and related trigger conditions between action 1 and action 2, and between action 2 and action 3, are considered to be implemented through a set of simulations performed at t1 (for example, as shown in Figure 8D, the simulated vehicle is shown with a dashed outline). During the simulation of this multi-step policy, the vehicle is shown to execute parts of the multi-step policy that occur during the simulation planning period, and this multi-step policy may be evaluated (e.g., through a set of metrics, through a score, etc.) and compared with the evaluation of other potential policies. This process is then performed from time t2 to t 14 This process is repeated (for example, according to a default selection cycle), and this multi-step policy may be optionally considered at each of these times. Additional or alternative options may be considered, such as single-step actions representing individual actions within a multi-step policy; modified versions of the multi-step policy (e.g., including only actions 2 and 3 when action 1 has already been completed and / or is no longer relevant, or always including actions 2 and 3, etc.); other default (e.g., baseline) policies may be considered; and / or any other policies may be considered.
[0108]
[0115] In a concrete example of the selected (e.g., consistent with the simulated) policy actually implemented as a result of these simulations, as shown in Figure 8E, the vehicle first implements the relevant parts of the multi-step policy, including actions 1, 2, and 3, at times t1 and t2; and then from time t3 to t 10 Then implement a single-step policy corresponding to action 2; and finally, time t 11 from t 14Then implement a single-step policy that corresponds to action 3.
[0109]
[0116] In a concrete example of the selected policy (e.g., one that matches the simulated one) that was actually implemented as a result of these simulations, as shown in Figure 8F, the vehicle first implements the relevant parts of the multi-step policy, including actions 1, 2, and 3, at times t1 and t2; and then from time t3 to t 10 Implement a multi-step policy including actions 2 and 3; and finally, time t 11 from t 14 Then implement a single-step policy that corresponds to action 3.
[0110]
[0117] As an addition or alternative, the vehicle may appropriately implement any other policies in a different manner.
[0111] 4.6 Method - Steps to repeat one or all of the processes
[0118] The method may optionally include steps of repeating any or all of the above processes, such as: continuously collecting new inputs in S110; checking for new scenarios and / or determining whether the current scenario has changed based on repeated iterations of S120; repeating the determination of a new set of policy options in S130 (e.g., according to a selection cycle); selecting a new policy in S140 (e.g., according to a selection cycle); keeping the agent running continuously in S150; and / or any or all of any other processes repeated in any appropriate manner.
[0112]
[0119] In one variation, the method includes a step of using the results of previously performed simulations and / or previously implemented policies to notify the execution of future simulations and / or the creation and selection of future policies.
[0113]
[0120] Alternatively, the method may be carried out appropriately in another manner.
[0114]
[0121] For the sake of brevity, some details have been omitted, but preferred embodiments include all combinations and permutations of various system components and various method processes, the method processes may be executed sequentially or simultaneously in any suitable order.
[0115]
[0122] Embodiments of a system and / or method may include any combination and permutation of various system components and various method processes, and one or more examples of the methods and / or processes described herein may be executed asynchronously (e.g., sequentially), concurrently (e.g., simultaneously, in parallel, etc.) or in any other suitable order by and / or using one or more examples of the systems, elements and / or entities described herein. The following system and / or method components and / or processes may be used in addition to, in place of, or in other ways integrated with, all or part of the systems and / or methods disclosed in the above application, each of which is incorporated by this reference.
[0116]
[0123] Additional or alternative embodiments implement the above method and / or processing module on a confidential, temporary, computer-readable medium that stores computer-readable instructions. Instructions may be executed by a computer-executable component integrated into the computer-readable medium and / or processing system. The computer-readable medium may include any suitable computer-readable medium, such as RAM, ROM, flash memory, EEPROM, optical devices (CD or DVD), hard drives, floppy drives, confidential, temporary, computer-readable medium, or any suitable device. The computer-executable component may include a computing system and / or processing system (e.g., including one or more coexisting or distributed remote or local processors) connected to a confidential, temporary, computer-readable medium, such as a CPU, GPU, TPUS, microprocessor, or ASIC, but instructions may, as an alternative or addition, be executed by any suitable dedicated hardware device.
[0117]
[0124] As those skilled in the art will understand from the above detailed description, drawings, and claims, preferred embodiments of the present invention may be modified and altered without departing from the scope of the invention as defined in the following claims.
Claims
1. A method for operating an autonomous vehicle, the method being A step of selecting a first set of policies for evaluation by the autonomous vehicle, wherein the first set of policies is: A set of single-action policies, A set of multiple action policies, each of which is: A set of multiple actions, The steps include selecting a first set of policies, each comprising a set of trigger conditions, each associated with a transition between consecutive actions in the set of multiple actions, and a set of multiple action policies, A step of evaluating the first set of policies, wherein the step of evaluating the first set of policies is: For each policy in the first set of the aforementioned policies, A step of simulating the behavior of the autonomous vehicle and the behavior of each set of tracked agents in the autonomous vehicle's environment over a predetermined future simulation period, in response to the autonomous vehicle executing the policy, A step of determining the quantitative metrics of the policy based on the behavior of the autonomous vehicle and the behavior of the tracking agent, A step of evaluating the first set of policies, including the step of selecting a policy from the first set of policies based on the set of quantitative metrics, A step of operating the autonomous vehicle in accordance with the selected policy, wherein the selected policy includes multiple action policies from the set of multiple action policies, A step of operating the autonomous vehicle, which includes the step of implementing the first action of a set of multiple actions from the selected multiple action policy, The steps include selecting a second set of policies for evaluation by the autonomous vehicle while the first action is implemented and according to a default selection cycle period having a shorter duration than the default simulation period, The steps include evaluating a second set of the aforementioned policies and selecting a second policy based on the evaluation, A step of refraining from completing the remaining parts of the selected multiple action policy, A method comprising the step of operating the autonomous vehicle in accordance with the selected second policy.
2. The method according to claim 1, wherein at least a portion of the set of trigger conditions depends on the progress of the set of agents being tracked.
3. The method according to claim 1, wherein at least a portion of the set of multiple action policies in the first set of policies is selected based on the location of the autonomous vehicle.
4. The method according to claim 3, wherein at least a second portion of the multiple action policy set of the first set of policies is determined and selected in advance, independently of its position.
5. The method according to claim 4, wherein the second part includes multiple action policies configured to bypass obstacles.
6. The method according to claim 3, wherein the position is determined based on sensor data collected by a set of sensors mounted on the autonomous vehicle.
7. The method according to claim 6, wherein the portion of the set of multiple action policies is further determined based on referencing a map labeled based on the location, the labeled map includes a set of default label assignments.
8. The method according to claim 7, wherein the position overlaps with a specific label assignment from the set of default label assignments, and the specific label assignment corresponds to a specific scenario in the environment.
9. The method according to claim 8, wherein the scenario includes at least one of a pedestrian crossing, an intersection, or a parking lot.
10. The method according to claim 1, wherein at least a portion of the set of trigger conditions is implemented in response to the set of tracked agents following a set of right-of-way driving practices during a relevant simulation performed while evaluating a first set of policies.
11. The method according to claim 1, wherein the default selection cycle period is less than one-tenth of the default simulation period.
12. The method according to claim 1, wherein the selected second policy includes a specific single action policy from the set of single action policies.
13. The method according to claim 12, wherein the specific single action policy includes a second action of the selected multiple action policy.
14. A method for operating an autonomous vehicle, the method being A step of selecting a set of policies for evaluation by the autonomous vehicle, wherein the set of policies is: A set of single-action policies, A set of multiple action policies, each of which is: A set of multiple actions, The steps include selecting a set of policies that define a set of trigger conditions associated with the set of multiple actions, A step of evaluating the set of policies, wherein the step of evaluating the set of policies is: For each policy in the aforementioned set of policies, A step of simulating the movement of the autonomous vehicle and the movement of each set of tracked agents in the environment of the autonomous vehicle over a predetermined future simulation period, A step of determining the quantitative metrics of the policy based on the simulation, A step of evaluating the set of policies, including the step of selecting a policy from the set of policies based on the set of quantitative metrics, A step of operating the autonomous vehicle in accordance with the selected policy, wherein the selected policy includes multiple action policies from the set of multiple action policies, and the step of operating the autonomous vehicle is A step of implementing the first action of the set of multiple actions of the selected multiple action policy, A step of checking whether the first trigger condition of the set of trigger conditions is met, A method comprising the steps of operating the autonomous vehicle, which includes the step of moving the operation of the autonomous vehicle to a second action of the set of multiple actions when the first trigger condition is met.
15. The method according to claim 14, wherein at least a portion of the set of trigger conditions depends on the progress of the set of agents being tracked.
16. The method according to claim 14, wherein at least a portion of the set of multiple action policies among the set of policies is selected based on the location of the autonomous vehicle.
17. The method according to claim 16, wherein at least a second portion of the set of multiple action policies in the set of policies is predetermined and selected independently of the location.
18. The method according to claim 17, wherein the second part includes multiple action policies configured to bypass obstacles.
19. The method according to claim 17, wherein the portion of the set of multiple action policies is further determined based on referring to a map labeled based on the location and determining a default scenario label based on referring to the labeled map.
20. The method according to claim 19, wherein the default scenario label includes at least one of a pedestrian crossing, an intersection, or a parking lot.