Online State Space Refinement for Competence Recognition Systems
The Competence Awareness System in autonomous vehicles addresses the reliance on human intervention by adaptively determining autonomy levels, enhancing operational efficiency and safety through iterative learning.
Patent Information
- Application Number
- JP2024570291
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-31
- Filing Date
- 2023-05-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-05-03
AI Technical Summary
Autonomous vehicles often rely on human intervention due to inadequate modeling of complex and unstructured environments, leading to inefficient operation and potential safety risks.
A Competence Awareness System (CAS) that uses an Autonomous Cognitive Agent (ACA) to determine the level of autonomy based on an autonomy model and feedback model, enabling the vehicle to learn and adapt its competence over time through iterative state space refinement, reducing dependence on human assistance.
The CAS allows autonomous vehicles to operate more reliably and efficiently by proactively selecting the appropriate level of autonomy, minimizing the need for human intervention and optimizing performance in dynamic environments.
Smart Images

Figure 2025520096000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the operation management and autonomous driving of autonomous vehicles, and more specifically to determining the level of autonomy of the operation of an autonomous vehicle according to a competence model.
Background Art
[0002] Vehicles such as autonomous vehicles can cross a part of a vehicle traffic network (for example, a road). Driving on a part of a vehicle traffic network may include generating or capturing data such as data representing the operating environment of the vehicle or a part thereof by means of a sensor of the vehicle or the like. In some cases, some data may become unavailable due to obstacles.
Summary of the Invention
[0003] A first aspect of the disclosed embodiments is a method of autonomous driving by an autonomous vehicle (AV). The method includes detecting an environmental state based on sensor data, selecting an action based on the environmental state, identifying a set of currently indistinguishable states, identifying a discriminator from the set of currently indistinguishable states, training a feedback model for the discriminator, determining an autonomy level associated with the environmental state and the action, and executing the action according to the autonomy level. The autonomy level can be selected based at least on an autonomy model and a feedback model. The feedback model can be trained regardless of the presence or absence of a discriminator and an iterative state space refinement technique.
[0004] A second aspect of the disclosed embodiments is a system for autonomy that includes a memory and a processor. The processor is configured to execute instructions stored in the memory to calculate a policy for solving a task by solving an extended stochastic shortest path (SSP) problem, identify discriminators from a current set of discrimination states, and train a feedback model for the discriminators. The policy can map environmental states and autonomy levels to actions and autonomy levels. Calculating the policy can include generating a plan that operates across multiple levels of autonomy. The feedback model can be trained regardless of the presence or absence of discriminators and iterative state space refinement techniques.
[0005] A third aspect of the disclosed embodiments is a method for autonomous driving. The method includes calculating a policy for solving a task by solving an extended stochastic shortest path (SSP) problem, identifying discriminators from a current set of discrimination states, and training a feedback model for the discriminators. The policy can map environmental states and autonomy levels to actions and autonomy levels. Calculating the policy can include generating a plan that operates across multiple levels of autonomy. The feedback model can be trained regardless of the presence or absence of discriminators and iterative state space refinement techniques.
[0006] These and other aspects, features, elements, implementations, and variations of the methods, apparatuses, procedures, and algorithms disclosed herein will be described in further detail below.
[0007] Various aspects of the methods and apparatuses disclosed herein will become more apparent by reference to the examples provided in the following description and drawings in which like reference numerals refer to like elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
[0012]
Figure 5
[0013]
Figure 6
[0014]
Figure 7
[0015]
Figure 8
[0016]
Figure 9
[0017]
Figure 10
[0018]
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0019] Vehicles such as AVs or semi-autonomous vehicles may cross a portion of a vehicle transportation network. The vehicle may include one or more sensors, and crossing the vehicle transportation network may include the sensors generating or capturing sensor data such as sensor data corresponding to the operating environment of the vehicle or a portion thereof. For example, the sensor data may include information corresponding to one or more external objects such as pedestrians, remote vehicles, other objects within the vehicle operating environment, vehicle transportation network geometry, or combinations thereof.
[0020] The AV may include an autonomous vehicle operation management system that may include one or more operating environment monitors capable of processing operating environment information such as sensor data for an autonomous vehicle.
[0021] The AV operation management system may include an autonomous vehicle operation management controller, and the autonomous vehicle operation management controller may detect one or more operation scenarios such as a pedestrian scenario, an intersection scenario, a lane change scenario, or any other vehicle operation scenario or combination of vehicle operation scenarios corresponding to an external object. An operation scenario or a set associated with the operation scenario may be referred to herein as an environmental state.
[0022] The AV operation management system may include one or more scenario-specific operation control evaluation modules (SSOCEMs). Each scenario-specific operation control evaluation module may be a model such as a partially observable Markov decision process (POMDP) model of the respective operation scenario. That is, each model is configured to handle a specific scenario. The AV operation management controller may instantiate respective instances of the scenario-specific operation control evaluation modules in response to detecting the corresponding operation scenario.
[0023] The AV motion management controller may receive candidate vehicle control actions from each instantiated scenario-specific motion control evaluation module (SSOCEM) instance, may identify vehicle control actions from the candidate vehicle control actions, and may control the AV to cross a portion of the vehicle traffic network according to the identified vehicle control actions.
[0024] In some implementations, the SSOCEM may be configured to autonomously complete some tasks while requiring human intervention to complete other tasks. That is, the SSOCEM can operate autonomously under certain conditions, but may require human intervention or assistance to achieve its goals (e.g., crossing an intersection). Thus, the SSOCEM can be in one of two binary autonomous states or levels.
[0025] For example, in response to detecting an obstacle on the road (i.e., on a portion of the vehicle traffic network), the SSOCEM may issue a request for assistance to a remote operator. The remote operator can be a human operator responsible for remotely monitoring and assisting one or more autonomous vehicles. The remote operator can stream sensor data (e.g., camera images and / or video) to the remote operator, as a result of which the remote operator can obtain a situation awareness, plot a navigation route for the AV around the obstacle, and / or remotely control the AV's actions.
[0026] In another example, the lane-crossing SSOCEM may use lane markings to keep the AV within the lane. In some situations, the lane-crossing SSOCEM may no longer be able to define the lane boundary, such as when the sensor may be dirty or malfunctioning, or when the lane markings are covered with snow or mud. In such situations, the vehicle-crossing SSOCEM may request that a human (e.g., the AV's driver passenger or a remote operator) take over control of the AV.
[0027] Dependence on human assistance (i.e., intervention) can indicate the limited competence of SSOCEM in its autonomous model. Human intervention can be costly. For example, it may take a relatively long time for a remote operator to respond to a request for assistance from an AV. During that time, the AV may be obstructing traffic. For example, as the number of requests for assistance from AVs to remote operators increases, the number of available remote operators must necessarily also increase.
[0028] In an implementation according to the present disclosure, SSOCEM can be or include an Autonomous Cognitive Agent (ACA), and the Autonomous Cognitive Agent (ACA) selects the next action to execute and the level of autonomy for executing the action based on an autonomous model that the ACA maintains and evolves. For example, initially, the autonomous model can indicate that the ACA requests human assistance (e.g., feedback) for the action in view of the detected environmental state. As the ACA receives more feedback from humans or the like, the ACA can gradually become less dependent on human assistance, and the ACA learns the appropriate timing for executing actions under a lower level of assistance, which means higher autonomy (i.e., competence). Thus, it can be said or considered that the ACA is aware of its level of competence.
[0029] In an implementation according to the present disclosure, the ACA can consider all levels of autonomy available during plan generation (e.g., as opposed to adjusting the level of autonomy during plan execution). Thus, the ACA can create a plan that more effectively utilizes the ACA's knowledge about its own level of autonomy.
[0030] Furthermore, in implementations according to the present disclosure, the ACA can model feedback from multiple forms of humans, thereby enabling the ACA to plan ahead in a way that also takes into account the likelihood of feedback from each form of human. Thus, the ACA can proactively avoid situations where negative feedback is possible.
[0031] Furthermore, the ACA not only avoids situations that are more likely to require human intervention through experiences that enable the ACA to reduce its dependence on humans over time, but also maintains a predictive model of human feedback and intervention by adjusting the level of autonomy of the ACA over time. Thus, the ACA can execute at a minimum-cost level of autonomy for any situation (i.e., environmental state) that the ACA encounters (i.e., is detected based on sensor data).
[0032] Thereafter, the ACA can use the autonomy model to cross the vehicle transportation network.
[0033] Autonomous systems are increasingly deployed in an open world that includes very complex and unstructured domains. Examples of these systems include space exploration rovers, autonomous underwater vehicles, service robots, and AVs. Since it is difficult to fully model the open world, these systems must rely on approximate models of those domains that may not be sufficient to handle all situations, and as a result, dangerous behavior may occur. Nevertheless, these systems are expected to potentially maintain safe and reliable operation over a long-term deployment process. To achieve this, these systems rely on various forms of human supervision, support, and intervention. In that sense, even the most sophisticated AI systems currently under development are at best semi-autonomous in that they operate autonomously only under certain conditions and require human intervention to complete the assigned tasks otherwise.
[0034] For example, a space exploration rover may pause its operation and wait for a new plan from ground control when it encounters an unexpected obstacle or when system parameters such as wheel resistance torque go out of the allowable range. Similarly, in case of bad weather or when there is high-density intersecting traffic and no safe gap exists for a long time when attempting to merge, the AV may request to return control to a human driver. Human assistance and feedback may be available in different forms or modalities corresponding to the behavior. For example, enabling the system to operate autonomously under human supervision represents a higher level of competence compared to requiring explicit approval from a human before executing each action, while representing a lower level of competence compared to unsupervised autonomous operation.
[0035] The embodiments disclosed herein are described in the context of semi-autonomous systems, where competence may be a measure of the optimal range of autonomous operation of the system in any given situation, taking into account the available human involvement, i.e., the various forms of assistance that a human can provide and the knowledge related to the ability to communicate with the semi-autonomous system. Competence may be a measure of the performance of the system or algorithm. The competence of automated decision-making in a semi-autonomous system may be based on the level of autonomy. The level of autonomy may be a paradigm for modeling the stages of autonomous behavior in domains where safety is important, and each level of autonomy may correspond to some set of constraints, limitations, or requirements for autonomous operation. Thus, competence recognition may be defined as the ability of an agent to predict the correct level of autonomy to operate in any given situation.
[0036] The embodiments disclosed herein include a planning model called a Competence Awareness System (CAS) that operates at multiple levels of autonomy, with each level associated with different forms of human assistance that compensate for the system's constrained capabilities. The system may associate each type of human assistance with a unique set of feedback signals that the system can receive from humans, and the likelihood thereof can be learned over time. This model enables the system to operate more reliably in an open world, reduce inappropriate dependence on humans, and ultimately optimize the autonomous behavior of the system. To address situations where the initial domain model has only inadequate representations for correctly modeling human feedback, an iterative approach is introduced to refine the state space of the system to better discriminate human feedback, generating a finer partitioning of the state-action space with different levels of competence, enabling the system to learn and act better with its true competence.
[0037] The embodiments disclosed herein may include a method for iterative state space refinement. Iterative state space refinement may enable the CAS to refine the granularity of its state representation online. The CAS models multiple forms of human feedback and uses this feedback to enable a semi-autonomous system to learn its competence over time. Further, the CAS enables the semi-autonomous system to learn a predictive model of human feedback, thereby enabling the semi-autonomous system to converge to an optimal level of autonomy over time. The CAS may be configured to learn from human interactions, and in particular, to learn from human input and learn from demonstrations.
[0038] The decision-making process associated with the CAS may be based on an Uncertain Markov Decision Process (UMDP). The embodiments disclosed herein may use two sources of prior models that are underspecified or incompletely specified. The first may be used in the transition dynamics in the form of human feedback and sub-competencies, and the second may be used in the feature space in the form of missing features.
[0039] FIG. 1 is a diagram of an example of a vehicle in which the aspects, features, and elements disclosed herein can be implemented. In the illustrated embodiment, vehicle 1000 includes various vehicle systems. The vehicle systems include chassis 1100, powertrain 1200, controller 1300, and wheels 1400. Additional or different combinations of vehicle systems may be used. For simplicity, vehicle 1000 is shown as including four wheels 1400, although one or more any other propulsion devices such as a propeller or tread may be used. In FIG. 1, the lines interconnecting elements such as powertrain 1200, controller 1300, and wheels 1400 indicate that information such as data or control signals, forces such as power or torque, or both information and power can be communicated between the respective elements. For example, controller 1300 may receive power from powertrain 1200 and communicate with powertrain 1200, wheels 1400, or both to control vehicle 1000, which may include accelerating, decelerating, steering, or otherwise controlling vehicle 1000.
[0040] The powertrain 1200 shown by the example of FIG. 1 includes power source 1210, transmission 1220, steering unit 1230, and actuator 1240. Any other elements or combinations of elements of the powertrain, such as a suspension, drive shaft, axle, or exhaust system, may be included. Although shown separately, wheels 1400 may be included in powertrain 1200.
[0041] The power source 1210 includes an engine, a battery, or a combination thereof. The power source 1210 can be any device or combination of devices that operates to supply energy such as electrical energy, thermal energy, or kinetic energy. In one example, the power source 1210 includes an engine such as an internal combustion engine, an electric motor, or a combination of an internal combustion engine and an electric motor, and operates to supply kinetic energy as power to one or more of the wheels 1400. Alternatively or additionally, the power source 1210 includes one or more dry batteries such as nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li ion), a solar cell, a fuel cell, or any other device capable of supplying energy, such as a potential energy unit.
[0042] The transmission 1220 receives energy such as kinetic energy from the power source 1210 and transmits that energy to the wheels 1400 to supply power. The transmission 1220 can be controlled by the controller 1300, the actuator 1240, or both. The steering unit 1230 can be controlled by the controller 1300, the actuator 1240, or both, and can control the wheels 1400 to steer the vehicle. The actuator 1240 can receive a signal from the controller 1300 and can activate or control the power source 1210, the transmission 1220, the steering unit 1230, or any combination thereof to operate the vehicle 1000.
[0043] In the illustrated embodiment, the controller 1300 includes a location identification unit 1310, an electronic communication unit 1320, a processor 1330, a memory 1340, a user interface 1350, a sensor 1360, and an electronic communication interface 1370. There may be fewer of those that exist as part of the controller 1300 among these elements. Although shown as a single unit, any one or more of the elements of the controller 1300 may be integrated into any number of separate physical units. For example, the user interface 1350 and the processor 1330 may be integrated into a first physical unit, and the memory 1340 may be integrated into a second physical unit. Although not shown in FIG. 1, the controller 1300 may include a power source such as a battery. Although shown as separate elements, the location identification unit 1310, the electronic communication unit 1320, the processor 1330, the memory 1340, the user interface 1350, the sensor 1360, the electronic communication interface 1370, or any combination thereof may be integrated into one or more electronic units, circuits, or chips.
[0044] Processor 1330 can include any device or combination of devices that can operate on or process signals or other information, present or future, including an optical processor, a quantum processor, a molecular processor, or combinations thereof. For example, processor 1330 can include one or more dedicated processors, one or more digital signal processors, one or more microprocessors, one or more controllers, one or more microcontrollers, one or more integrated circuits, one or more application specific integrated circuits, one or more field programmable gate arrays, one or more programmable logic arrays, one or more programmable logic controllers, one or more state machines, or any combination thereof. Processor 1330 is operably coupled to one or more of location unit 1310, memory 1340, electronic communication interface 1370, electronic communication unit 1320, user interface 1350, sensor 1360, and power train 1200. For example, the processor may be operably coupled to memory 1340 via communication bus 1380.
[0045] Memory 1340 includes machine-readable instructions or any information associated therewith for use by or in connection with any processor such as processor 1330, and includes any tangible non-transitory computer-usable or computer-readable medium that can, for example, store, retain, communicate, or transport. Memory 1340 can be, for example, one or more solid state drives, one or more memory cards, one or more removable media, one or more read-only memories, one or more random access memories, one or more disks including hard disks, floppy disks, optical disks, magnetic or optical cards, or any other type of non-transitory medium suitable for storing electronic information, or any combination thereof. For example, the memory can be one or more read-only memories (ROMs), one or more random access memories (RAMs), one or more registers, low power double data rate (LPDDR) memories, one or more cache memories, one or more semiconductor memory devices, one or more magnetic media, one or more optical media, one or more magneto-optical media, or any combination thereof.
[0046] Communication interface 1370 can be a wireless antenna, a wired communication port, an optical communication port, or any other wired or wireless unit that can interface with a wired or wireless electronic communication medium 1500 as shown. FIG. 1 shows a communication interface 1370 that communicates via a single communication link, but the communication interface may be configured to communicate via multiple communication links. FIG. 1 shows a single communication interface 1370, but the vehicle may include any number of communication interfaces.
[0047] The communication unit 1320 is configured to transmit or receive signals via a wired or wireless electronic communication medium 1500, such as via a communication interface 1370. Although not explicitly shown in FIG. 1, the communication unit 1320 may be configured to transmit, receive, or both via any wired or wireless communication medium, such as radio frequency (RF), ultraviolet (UV), visible light, optical fiber, wireline, or a combination thereof. FIG. 1 shows a single communication unit 1320 and a single communication interface 1370, but any number of communication units and any number of communication interfaces may be used. In some embodiments, the communication unit 1320 includes a dedicated short range communication (DSRC) unit, an on-board unit (OBU), or a combination thereof.
[0048] The location unit 1310 may determine geographical location information such as the longitude, latitude, altitude, driving direction, or speed of the vehicle 1000. In one example, the location unit 1310 includes a GPS unit such as a wide area augmentation system (WAAS)-compatible National Marine Electronics Association (NMEA) unit, a wireless triangulation unit, or a combination thereof. The location unit 1310 may be used, for example, to obtain information representing the current direction of travel of the vehicle 1000, the current position of the vehicle 1000 in two or three dimensions, the current angular direction of the vehicle 1000, or a combination thereof.
[0049] The user interface 1350 includes any unit that can interface with a person, such as a virtual or physical keypad, touchpad, display, touch display, head-up display, virtual display, augmented reality display, tactile display, a feature tracking device such as an eye tracking device, speaker, microphone, video camera, sensor, printer, or any combination thereof. The user interface 1350 can be operably coupled to the processor 1330 as shown, or to any other element of the controller 1300. Although shown as a single unit, the user interface 1350 may include one or more physical units. For example, the user interface 1350 may include both an audio interface for performing audio communication with a person and a touch display for performing visual and touch-based communication with a person. The user interface 1350 may include multiple displays, such as multiple physically separate units, multiple defined portions within a single physical unit, or a combination thereof.
[0050] The sensor 1360 is operable to provide information that can be used to control the vehicle. The sensor 1360 may be an array of sensors. The sensor 1360 may provide information regarding the current operating characteristics of the vehicle 1000, including vehicle motion information. The sensor 1360 can include, for example, a speed sensor, an acceleration sensor, a steering angle sensor, a traction-related sensor, a braking-related sensor, a steering wheel position sensor, an eye tracking sensor, a seat position sensor, or any sensor, or combination of sensors, that are operable to report information regarding some aspect of the current dynamic situation of the vehicle 1000.
[0051] Sensor 1360 includes one or more sensors 1360 operable to obtain information regarding the physical environment surrounding vehicle 1000, such as operating environment information. For example, the one or more sensors may detect road geometry such as lane boundaries, and obstacles such as fixed obstacles, vehicles, and pedestrians. Sensor 1360 can be, or include, one or more video cameras, laser detection systems, infrared detection systems, acoustic detection systems, or any other suitable type of in-vehicle environment detection device, or combination of devices, known currently or developed later. In some embodiments, sensor 1360 and location unit 1310 are combined.
[0052] Although not shown separately, vehicle 1000 may include a trajectory controller. For example, controller 1300 may include a trajectory controller. The trajectory controller can be operable to obtain information describing the current state of vehicle 1000 and the route planned for vehicle 1000, and based on this information, determine and optimize the trajectory of vehicle 1000. In some embodiments, the trajectory controller can output a signal operable to control vehicle 1000 to follow the trajectory determined by the trajectory controller. For example, the output of the trajectory controller can be an optimized trajectory that can be supplied to power train 1200, wheels 1400, or both. In some embodiments, the optimized trajectory can be control inputs such as a set of steering angles, each steering angle corresponding to a point in time or a position. In some embodiments, the optimized trajectory can be one or more paths, lines, curves, or combinations thereof.
[0053] One or more of wheels 1400 can be steering wheels pivoted to a steering angle under the control of steering unit 1230, drive wheels given torque to propel vehicle 1000 under the control of transmission 1220, or steering and drive wheels that can steer and propel vehicle 1000.
[0054] Although not shown in FIG. 1, the vehicle may include additional units or elements not shown in FIG. 1, such as an enclosure, a Bluetooth (registered trademark) module, a frequency modulation (FM) radio unit, a near field communication (NFC) module, a liquid crystal display (LCD) display unit, an organic light emitting diode (OLED) display unit, a speaker, or any combination thereof.
[0055] Vehicle 1000 may be an autonomous vehicle that is autonomously controlled without direct human intervention to travel as part of a vehicle transportation network. Although not separately shown in FIG. 1, the autonomous vehicle may include an autonomous vehicle control unit that executes routing, navigation, and control of the autonomous vehicle. The autonomous vehicle control unit may be integrated with another unit of the vehicle. For example, controller 1300 may include an autonomous vehicle control unit.
[0056] When present, the autonomous vehicle control unit may control or operate vehicle 1000 to travel as part of a vehicle transportation network according to current vehicle operation parameters. The autonomous vehicle control unit may control or operate vehicle 1000 to perform defined operations or actions such as parking the vehicle. The autonomous vehicle control unit may generate a travel route from a starting point to a destination, such as the current position of vehicle 1000, based on vehicle information, environmental information, vehicle transportation network information representing the vehicle transportation network, or a combination thereof, and may control or operate vehicle 1000 to travel the vehicle transportation network according to the route. For example, the autonomous vehicle control unit may output the travel route to an orbit controller and operate vehicle 1000 to travel from the starting point to the destination using the generated route.
[0057] Figure 2 is a diagram of an example of a portion of a vehicle traffic and communication system in which the aspects, features, and elements disclosed herein may be implemented. The vehicle traffic and communication system 2000 may include one or more vehicles 2100 / 2110, such as vehicle 1000 shown in FIG. 1, that travel through one or more portions of a vehicle traffic network 2200 and communicate via one or more electronic communication networks 2300. Although not explicitly shown in FIG. 2, the vehicles may travel in off-road areas.
[0058] The electronic communication network 2300 may be a multiple access system that provides communication, such as voice communication, data communication, video communication, messaging communication, or a combination thereof, between the vehicles 2100 / 2110 and one or more communication devices 2400. For example, the vehicles 2100 / 2110 may receive information, such as information representing the vehicle traffic network 2200, from the communication devices 2400 via the network 2300.
[0059] In some embodiments, the vehicles 2100 / 2110 may communicate via a wired communication link (not shown), a wireless communication link 2310 / 2320 / 2370, or any number of combinations of wired or wireless communication links. As shown, the vehicles 2100 / 2110 communicate via a terrestrial wireless communication link 2310, via a non-terrestrial wireless communication link 2320, or a combination thereof. The terrestrial wireless communication link 2310 may include an Ethernet link, a serial link, a Bluetooth link, an infrared (IR) link, an ultraviolet (UV) link, or any link capable of providing electronic communication.
[0060] Vehicle 2100 / 2110 can communicate with another vehicle 2100 / 2110. For example, the host vehicle, i.e., the target vehicle (HV) 2100, can receive one or more automated inter-vehicle messages, such as a basic safety message (BSM), from a remote vehicle, i.e., the target vehicle (RV) 2110, via a direct communication link 2370 or via a network 2300. The remote vehicle 2110 can broadcast a message to the host vehicle within a defined broadcast range, such as 300 meters. In some embodiments, the host vehicle 2100 can receive a message via a third party, such as a signal repeater (not shown) or another remote vehicle (not shown). Vehicles 2100 / 2110 can periodically transmit one or more automated inter-vehicle messages based on a defined interval, such as 100 milliseconds.
[0061] Automated inter-vehicle messages can include vehicle identification information, geospatial state information such as longitude, latitude, or altitude information, geospatial position accuracy information, vehicle acceleration information, yaw rate information, speed information, vehicle travel direction information, braking system status information, throttle information, steering wheel angle information, or kinematic state information such as vehicle routing information, or vehicle operation state information such as vehicle size information, headlight state information, turn indicator information, wiper status information, transmission information, or any other information, or a combination of information related to the transmitting vehicle state. For example, transmission state information can indicate whether the transmission of the transmitting vehicle is in a neutral state, a parked state, a forward state, or a reverse state.
[0062] Vehicle 2100 may communicate with communication network 2300 via access point 2330. Access point 2330, which may include a computing device, is configured to communicate with vehicle 2100, communication network 2300, one or more communication devices 2400, or a combination thereof via wired or wireless communication links 2310 / 2340. For example, access point 2330 may be a base station, a base transceiver station (BTS), a Node B, an evolved Node B (eNode-B), a Home Node B (HNode-B), a wireless router, a wired router, a hub, a relay, a switch, or any similar wired or wireless device. Although shown here as a single unit, the access point may include any number of interconnected elements.
[0063] Vehicle 2100 may communicate with communication network 2300 via satellite 2350 or other non-terrestrial communication devices. Satellite 2350, which may include a computing device, is configured to communicate with vehicle 2100, communication network 2300, one or more communication devices 2400, or a combination thereof via one or more communication links 2320 / 2360. Although shown here as a single unit, the satellite may include any number of interconnected elements.
[0064] The electronic communication network 2300 is any type of network configured to provide voice, data, or any other type of electronic communication. For example, the electronic communication network 2300 may include a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a mobile or cellular phone network, the Internet, or any other electronic communication system. The electronic communication network 2300 uses communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), the Internet Protocol (IP), the Real-Time Transport Protocol (RTP), the Hypertext Transport Protocol (HTTP), or combinations thereof. Although shown here as a single unit, the electronic communication network may include any number of interconnected elements.
[0065] The vehicle 2100 may identify a portion or situation of the vehicle traffic network 2200. For example, the vehicle includes at least one in-vehicle sensor 2105, such as the sensor 1360 shown in FIG. 1, which can be a speed sensor, a wheel speed sensor, a camera, a gyroscope, an optical sensor, a laser sensor, a radar sensor, an acoustic sensor, or any other sensor or device or combination thereof that can determine or identify a portion or situation of the vehicle traffic network 2200, or may include them.
[0066] The vehicle 2100 may traverse a portion or portions of the vehicle traffic network 2200 using information communicated via the network 2300, such as information representing the vehicle traffic network 2200, information identified by one or more in-vehicle sensors 2105, or a combination thereof.
[0067] FIG. 2 shows, for simplicity, one vehicle transportation network 2200, one electronic communication network 2300, and one communication device 2400, although any number of networks or communication devices may be used. The vehicle transportation and communication system 2000 may include devices, units, or elements not shown in FIG. 2. Vehicle 2100 is shown as a single unit, although the vehicle may include any number of interconnected elements.
[0068] Vehicle 2100 is shown as communicating with communication device 2400 via network 2300, although vehicle 2100 may communicate with communication device 2400 via any number of direct or indirect communication links. For example, vehicle 2100 may communicate with communication device 2400 via a direct communication link such as a Bluetooth communication link.
[0069] FIG. 3 is a diagram of a portion of a vehicle transportation network according to the present disclosure. The vehicle transportation network 3000 may include one or more non-navigable areas 3100 such as a building, one or more partially navigable areas such as a parking area 3200, one or more navigable areas such as roads 3300 / 3400, or combinations thereof. In some embodiments, an autonomous vehicle such as vehicle 1000 shown in FIG. 1, one of vehicles 2100 / 2110 shown in FIG. 2, a semi-autonomous vehicle, or any other vehicle implementing autonomous driving may traverse a portion or portions of the vehicle transportation network 3000.
[0070] The vehicle transportation network 3000 may include one or more interchanges 3210 between one or more navigable or partially navigable areas 3200 / 3300 / 3400. For example, a portion of the vehicle transportation network 3000 shown in FIG. 3 includes an interchange 3210 between parking area 3200 and road 3400.
[0071] A portion of the vehicle transportation network 3000, such as roads 3300 / 3400, may include one or more lanes 3320 / 3340 / 3360 / 3420 / 3440 and may be associated with one or more driving directions indicated by the arrows in FIG. 3.
[0072] A vehicle transportation network or a portion thereof, such as a portion of the vehicle transportation network 3000 shown in FIG. 3, may be represented as vehicle transportation network information. For example, the vehicle transportation network information may be represented as a hierarchy of elements such as markup language elements that can be stored in a database or file. For simplicity, the figures in this specification depict vehicle transportation network information representing a portion of the vehicle transportation network as a diagram or map, but the vehicle transportation network information may be represented in any computer-usable form that can represent the vehicle transportation network or a portion thereof. In some embodiments, the vehicle transportation network information may include vehicle transportation network control information such as driving direction information, speed limit information, toll information, gradient information such as slope or angle information, surface material information, landscape information, the number of lanes, known hazards, or combinations thereof.
[0073] The vehicle transportation network may be associated with or may include a pedestrian transportation network. For example, FIG. 3 includes a portion 3600 of the pedestrian transportation network that may be a pedestrian passageway. Although not shown separately in FIG. 3, pedestrian-navigable areas such as pedestrian passageways or crosswalks may correspond to navigable areas or partially navigable areas of the vehicle transportation network.
[0074] In some embodiments, a portion of the vehicle traffic network, or a combination of portions, may be identified as an area of interest or a destination. For example, the vehicle traffic network information may identify buildings such as the non-navigable area 3100 and the adjacent partially navigable parking area 3200 as areas of interest, the vehicle may identify the area of interest as a destination, and the vehicle may travel from a starting point to the destination by traversing the vehicle traffic network. In FIG. 3, the parking area 3200 associated with the non-navigable area 3100 is shown as adjacent to the non-navigable area 3100, but the destination may include, for example, a building and a parking area that is not physically or geographically adjacent to the building.
[0075] Traversing a portion of the vehicle traffic network can proceed from the vehicle's topological position estimation to the destination. The destination can be an individually uniquely identifiable geographical location. For example, the vehicle traffic network may include defined locations such as the location address, postal address, vehicle traffic network address, GPS address, or a combination thereof of the destination. The destination may be associated with one or more entrances such as the entrance 3500 shown in FIG. 3. The destination may be associated with one or more docking locations such as the docking location 3700 shown in FIG. 3. The docking location 3700 can be a designated or undesignated location or area close to the destination where the vehicle can stop, wait, or park so that docking operations such as passenger boarding and alighting can be performed.
[0076] FIG. 4 is a diagram of an example of an autonomous vehicle operation management system 4000 according to an embodiment of the present disclosure. The autonomous vehicle operation management system 4000 can be implemented in an autonomous vehicle such as the vehicle 1000 shown in FIG. 1, one of the vehicles 2100 / 2110 shown in FIG. 2, a semi-autonomous vehicle, or any other vehicle implementing autonomous driving.
[0077] The self-driving vehicle can cross a vehicle transportation network or a part thereof, which may include crossing individual vehicle operation scenarios. An individual vehicle operation scenario (also referred to as a scenario in this specification) may include any clearly distinguishable set of operating situations that can affect the operation of the self-driving vehicle within a defined spatio-temporal area or operating environment of the self-driving vehicle. The individual vehicle operation scenario may be based on the number or density of roads, road sections, or lanes that the self-driving vehicle can cross within a defined spatio-temporal distance. The individual vehicle operation scenario may be based on one or more traffic control devices that can affect the operation of the self-driving vehicle within a defined spatio-temporal area or operating environment of the self-driving vehicle. The individual vehicle operation scenario may be based on one or more distinguishable rules, regulations, or laws that can affect the operation of the self-driving vehicle within a defined spatio-temporal area or operating environment of the self-driving vehicle. The individual vehicle operation scenario may be based on one or more distinguishable external objects that can affect the operation of the self-driving vehicle within a defined spatio-temporal area or operating environment of the self-driving vehicle.
[0078] Examples of individual vehicle operation scenarios include an individual vehicle operation scenario where the self-driving vehicle is crossing an intersection, an individual vehicle operation scenario where a pedestrian is crossing or approaching the predicted path of the self-driving vehicle, and an individual vehicle operation scenario where the self-driving vehicle is changing lanes. The individual vehicle operation scenario may separately include a merging lane, or an individual vehicle operation scenario where the self-driving vehicle is changing lanes may also include a merging lane.
[0079] For the sake of simplicity and clarity, similar vehicle operation scenarios may be described herein with reference to the type or class of vehicle operation scenarios. The type or class of vehicle operation scenarios may refer to a specific pattern or set of patterns on the scenario. For example, a vehicle operation scenario including a pedestrian may be referred to herein as a pedestrian scenario, which refers to the type or class of vehicle operation scenarios including a pedestrian. As an example, a first vehicle operation scenario of a pedestrian may include a pedestrian crossing a road at a crosswalk, and a second vehicle operation scenario of a pedestrian may include a pedestrian crossing a road in violation of traffic regulations. Although pedestrian vehicle operation scenarios, intersection vehicle operation scenarios, and lane change vehicle operation scenarios are described herein, any other vehicle operation scenario or vehicle operation scenario type may be used.
[0080] Aspects of the operating environment of an autonomous vehicle may be represented within each individual vehicle operation scenario. For example, the relative orientation, trajectory, and predicted path of an external object may be represented within each individual vehicle operation scenario. In another example, the relative geometry of a vehicle traffic network may be represented within each individual vehicle operation scenario.
[0081] As an example, a first individual vehicle operation scenario may correspond to a pedestrian crossing a road at a crosswalk, and the relative orientation and predicted path of the pedestrian, such as crossing from right to left as opposed to crossing from left to right, may be represented within the first individual vehicle operation scenario. A second individual vehicle operation scenario may correspond to a pedestrian crossing a road in violation of traffic regulations, and the relative orientation and predicted path of the pedestrian, such as crossing from right to left as opposed to crossing from left to right, may be represented within the second individual vehicle operation scenario.
[0082] An autonomous vehicle may traverse multiple different individual vehicle operation scenarios within an operating environment that may be an aspect of a composite vehicle operation scenario. For example, a pedestrian may approach the predicted path of an autonomous vehicle crossing an intersection.
[0083] The autonomous driving vehicle operation management system 4000 may operate or control an autonomous driving vehicle to cross individual vehicle operation scenarios according to defined constraints such as safety constraints, legal constraints, physical constraints, user acceptance constraints, or any other constraints or combinations of constraints that may be defined or derived for the operation of the autonomous driving vehicle.
[0084] Controlling an autonomous driving vehicle to cross individual vehicle operation scenarios may include identifying or detecting individual vehicle operation scenarios, identifying candidate vehicle control actions based on the individual vehicle operation scenarios, controlling the autonomous driving vehicle to cross a portion of the vehicle traffic network according to one or more of the candidate vehicle control actions, or combinations thereof.
[0085] Vehicle control actions may indicate vehicle control actions or maneuvers such as accelerating, decelerating, changing direction, stopping, creeping, or any other vehicle actions or combinations of vehicle actions that may be performed by the autonomous driving vehicle in relation to crossing a portion of the vehicle traffic network.
[0086] The autonomous driving vehicle motion management controller 4100 or another unit of the autonomous driving vehicle may control the autonomous driving vehicle to cross a vehicle transportation network or a part thereof according to vehicle control actions. Examples of vehicle control actions include a "stop" vehicle control action to stop or otherwise control the autonomous driving vehicle so that it stops or remains stationary, a "forward" vehicle control action to slowly move the autonomous driving vehicle forward a short distance such as a few inches or one foot, an "accelerate" vehicle control action to accelerate the autonomous driving vehicle (e.g., at a defined acceleration or within a defined range), a "decelerate" vehicle control action to decelerate the autonomous driving vehicle (e.g., at a defined deceleration rate or within a defined range), a "maintain" vehicle control action to maintain current operating parameters (e.g., current speed, current route or path, current lane orientation, etc.), a "turn" vehicle control action (which may include the angle of the turn), a "proceed" vehicle control action to initiate or resume a previously identified set of operating parameters, or any other standard vehicle motion.
[0087] The vehicle control action may be a composite vehicle control action that includes a sequence, combination, or both of vehicle control actions. For example, a "forward" or "inch forward" vehicle control action may indicate a "stop" vehicle control action, a subsequent "accelerate" vehicle control action associated with a defined acceleration, and a subsequent "stop" vehicle control action associated with a defined deceleration rate, such that controlling the autonomous driving vehicle according to the "forward" vehicle control action includes slowly moving the autonomous driving vehicle forward a short distance such as a few inches or one foot little by little.
[0088] The autonomous driving vehicle motion management system 4000 may include the autonomous driving vehicle motion management controller 4100, the blocking monitor 4200, the operating environment monitor 4300, the SSOCEM 4400, or a combination thereof. As will be described separately, the blocking monitor 4200 may be one or more instances of the operating environment monitor 4300.
[0089] The autonomous vehicle motion management controller 4100 can receive, identify, or otherwise access motion environment information representing the motion environment of the autonomous vehicle, such as the current motion environment or a predicted motion environment, or one or more aspects thereof. The motion environment of the autonomous vehicle can include a clearly distinguishable set of motion situations that can affect the motion of the vehicle within a defined spatio-temporal area of the vehicle.
[0090] The motion environment information may include vehicle information for the autonomous vehicle, such as information indicating the geographical spatial position of the vehicle, information correlating the geographical spatial position with information representing the vehicle traffic network, the route of the vehicle, the speed of the vehicle, the acceleration state of the vehicle, the passenger information of the vehicle, or any other information related to the vehicle or the motion of the vehicle.
[0091] The motion environment information may include information representing the vehicle traffic network in the vicinity of the autonomous vehicle, such as within a defined spatial distance (e.g., 300 meters) of the vehicle, information indicating the geometry of one or more aspects of the vehicle traffic network, information indicating the state such as the surface state of the vehicle traffic network, or any combination thereof.
[0092] The motion environment information may include information representing external objects within the motion environment of the autonomous vehicle, such as pedestrians, non-motivated means of transportation such as animals other than humans, bicycles or skateboards, motorized means of transportation such as remote vehicles, or any other external object or entity that can affect the motion of the vehicle.
[0093] The autonomous vehicle motion management controller 4100 can monitor the motion environment of the autonomous vehicle, or a defined aspect thereof. Monitoring the motion environment can include identifying and tracking external objects, identifying individual vehicle motion scenarios, or a combination thereof.
[0094] For example, the autonomous driving vehicle operation management controller 4100 can identify and track external objects using the operating environment of the autonomous driving vehicle. Identifying and tracking external objects can include identifying the spatio-temporal positions of each external object that can be relative to the vehicle, and identifying one or more predicted paths of each external object that can include identifying the speed, trajectory, or both of the external object. Explanations such as position, predicted position, path, and predicted path in this specification may omit explicit indication that the corresponding position and path refer to geospatial and temporal components, but unless explicitly indicated in this specification or otherwise unambiguously clear from the context, the position, predicted position, path, and predicted path described in this specification can include geospatial components, temporal components, or both.
[0095] The operation environment monitor 4300 can include a pedestrian operation environment monitor 4310, an intersection operation environment monitor 4320, a lane change operation environment monitor 4330, or a combination thereof. The operation environment monitor 4340 is shown using a dashed line to indicate that the autonomous driving vehicle operation management system 4000 can include any number of operation environment monitors 4300.
[0096] One or more individual vehicle operation scenarios can be monitored by each operation environment monitor 4300. For example, the pedestrian operation environment monitor 4310 may monitor operation environment information corresponding to vehicle operation scenarios of a plurality of pedestrians, the intersection operation environment monitor 4320 may monitor operation environment information corresponding to vehicle operation scenarios of a plurality of intersections, and the lane change operation environment monitor 4330 may monitor operation environment information corresponding to vehicle operation scenarios of a plurality of lane changes.
[0097] The operating environment monitor 4300 may receive or otherwise access operating environment information such as operating environment information generated or captured by one or more sensors of the autonomous vehicle, vehicle traffic network information, vehicle traffic network geometry information, or a combination thereof. For example, the pedestrian operating environment monitor 4310 may receive or otherwise access information such as sensor data that may indicate, correspond to, or otherwise be associated with one or more pedestrians within the operating environment of the autonomous vehicle.
[0098] The operating environment monitor 4300 may associate the operating environment information or a portion thereof with an external object such as a pedestrian, a remote vehicle, or an aspect of the vehicle traffic network geometry, etc., related to the operating environment or an aspect thereof.
[0099] The operating environment monitor 4300 may generate or otherwise identify information representing one or more aspects of the operating environment, including filtering, abstracting, or otherwise processing the operating environment information, of an external object such as a pedestrian, a remote vehicle, or an aspect of the vehicle traffic network geometry.
[0100] The operating environment monitor 4300 may output information representing one or more aspects of the operating environment to the autonomous vehicle motion management controller 4100 or for access by the autonomous vehicle motion management controller 4100, such as by storing the information representing one or more aspects of the operating environment in a memory such as the memory 1340 shown in FIG. 1 of the autonomous vehicle, which is accessible by the autonomous vehicle motion management controller 4100, or transmitting the information representing one or more aspects of the operating environment to the autonomous vehicle motion management controller 4100, or a combination thereof. The operating environment monitor 4300 may output information representing one or more aspects of the operating environment to one or more elements of the autonomous vehicle motion management system 4000, such as the blocking monitor 4200.
[0101] The pedestrian motion environment monitor 4310 can correlate, associate, or otherwise process motion environment information to identify, track, or predict the actions of one or more pedestrians. For example, the pedestrian motion environment monitor 4310 may receive information such as sensor data corresponding to one or more pedestrians from one or more sensors. The pedestrian motion environment monitor 4310 may associate the sensor data with one or more identified pedestrians, which may include identifying, for one or more of the respective identified pedestrians, a travel direction, a path such as a predicted route, a current or predicted speed, a current or predicted acceleration, or a combination thereof. The pedestrian motion environment monitor 4310 may output the identified, associated, or generated pedestrian information to the autonomous vehicle motion management controller 4100 or for access by the autonomous vehicle motion management controller 4100.
[0102] The intersection operation environment monitor 4320 correlates, associates, or otherwise processes operation environment information to identify, track, or predict the actions of one or more remote vehicles within the operation environment of the autonomous vehicle, thereby identifying an intersection or its aspects within the operation environment and potentially identifying vehicle traffic network geometries or combinations thereof. For example, the intersection operation environment monitor 4320 may receive information such as sensor data corresponding to one or more remote vehicles within the operation environment, an intersection, or one or more aspects thereof within the operation environment, vehicle traffic network geometries, or combinations thereof from one or more sensors. The intersection operation environment monitor 4320 may associate the sensor data with one or more identified remote vehicles within the operation environment, an intersection, or one or more aspects thereof within the operation environment, vehicle traffic network geometries, or combinations thereof, which may include identifying a current or predicted driving direction, a route such as a predicted route, a current or predicted speed, a current or predicted acceleration, or combinations thereof for one or more of the respective identified remote vehicles. The intersection operation environment monitor 4320 may output the identified, associated, or generated intersection information to the autonomous vehicle operation management controller 4100 or for access by the autonomous vehicle operation management controller 4100.
[0103] The lane change operation environment monitor 4330 correlates, associates, or otherwise processes operation environment information to identify, track, or predict the actions of one or more remote vehicles within the operation environment of the autonomous vehicle, such as information indicating a slow or stationary remote vehicle along the predicted path of the vehicle, thereby identifying one or more aspects of the operation environment, such as the vehicle traffic network geometry in the operation environment, or combinations thereof that geospatially correspond to the current or predicted lane change operation. For example, the lane change operation environment monitor 4330 may receive information such as sensor data corresponding to one or more remote vehicles within the operation environment of the autonomous vehicle, one or more aspects of the operation environment, or combinations thereof that geospatially correspond to the current or predicted lane change operation, from one or more sensors. The lane change operation environment monitor 4330 may associate the sensor data with one or more identified remote vehicles within the operation environment of the autonomous vehicle, one or more aspects of the operation environment, or combinations thereof that geospatially correspond to the current or predicted lane change operation, which may include identifying, for one or more of each identified remote vehicle, the current or predicted driving direction, a path such as the predicted path, the current speed or predicted speed, the current or predicted acceleration, or combinations thereof. The lane change operation environment monitor 4330 may output the identified, associated, or generated lane change information to the autonomous vehicle operation management controller 4100 or for access by the autonomous vehicle operation management controller 4100.
[0104] The autonomous driving vehicle operation management controller 4100 may identify one or more individual vehicle operation scenarios based on one or more aspects of the operation environment represented by the operation environment information. The autonomous driving vehicle operation management controller 4100 may identify an individual vehicle operation scenario in response to identifying the operation environment information indicated by one or more of the operation environment monitors 4300 or based on the operation environment information. For example, the operation environment information may include information representing a pedestrian approaching an intersection along the predicted route of the autonomous driving vehicle, and the autonomous driving vehicle operation management controller 4100 may identify the vehicle operation scenario of the pedestrian, the vehicle operation scenario of the intersection, or both.
[0105] The autonomous driving vehicle operation management controller 4100 may instantiate one or more respective instances of the SSOCEM 4400 based on one or more aspects of the operation environment represented by the operation environment information. For example, the autonomous driving vehicle operation management controller 4100 may instantiate each respective instance of the SSOCEM 4400 in response to identifying an upcoming scenario. The upcoming scenario may be an individual vehicle operation scenario that the autonomous driving vehicle operation management controller 4100 determines is likely to be encountered if the autonomous driving vehicle continues on its route. The upcoming scenario may be predictable (e.g., determinable from the route of the autonomous driving vehicle) or unexpected. An unexpected upcoming scenario may be detectable by the vehicle's sensors and may be a scenario that cannot be determined without sensor data.
[0106] When instantiated, the SSOCEM4400 can receive operation environment information including sensor data and determine and output candidate vehicle control actions, which are also referred to as candidate actions in this specification. A candidate action is a vehicle control action identified by a specific SSOCEM4400 as an optimal action that the vehicle is likely to execute to handle a specific scenario. For example, an SSOCEM4400 configured to handle intersections (e.g., intersection SSOCEM4420) can output "proceed", which is a candidate action proposing to pass through the intersection. At the same time, an SSOCEM4400 for handling lane changes (e.g., lane change SSOCEM4430) can output a candidate action of "turn left" indicating that the vehicle should merge two steps to the left. In some implementations, each SSOCEM4400 outputs a confidence score indicating the confidence level of the candidate action determined by the SSOCEM4400. For example, a confidence score greater than 0.95 can indicate very high confidence in the candidate action, while a confidence score less than 0.5 can indicate relatively low confidence in the candidate action. Further details of the SSOCEM4400 will be described later.
[0107] The autonomous driving vehicle operation management controller 4100 receives candidate actions and determines vehicle control actions based on the received candidate actions. In some implementations, the autonomous driving vehicle operation management controller 4100 uses hard-coded logic to determine vehicle control actions. For example, the autonomous driving vehicle operation management controller 4100 may select the candidate action with the highest confidence score. In other implementations, the autonomous driving vehicle operation management controller 4100 may select the candidate action with the lowest likelihood of resulting in a collision. In other implementations, the autonomous driving vehicle operation management controller 4100 may generate a composite action based on two or more non-conflicting candidate actions (e.g., combining "proceed" and "turn left two steps" to result in a vehicle control action that causes the vehicle to change its course to the left and pass through the intersection). In some implementations, the autonomous driving vehicle operation management controller 4100 may use a machine learning algorithm to determine vehicle control actions based on two or more different candidate actions.
[0108] For example, identifying a vehicle control action from a candidate action may include implementing a machine learning component such as supervised learning of a classification problem, and training the machine learning component using examples such as 1000 examples of corresponding vehicle motion scenarios. In another example, identifying a vehicle control action from a candidate action may include implementing a Markov decision process (MDP) or a partially observable Markov decision process (POMDP), which may describe how each candidate action affects subsequent candidate actions and may include a reward function that outputs a positive or negative reward for each vehicle control action.
[0109] The autonomous driving vehicle motion management controller 4100 may de-instantiate an instance of the SSOCEM 4400. For example, the autonomous driving vehicle motion management controller 4100 may identify an individual set of operating conditions as representing an individual vehicle motion scenario of the autonomous driving vehicle, instantiate an instance of the SSOCEM 4400 for the individual vehicle motion scenario, monitor the operating conditions, and then determine whether one or more of the operating conditions are expired or have a probability of affecting the operation of the autonomous driving vehicle below a defined threshold. The autonomous driving vehicle motion management controller 4100 may de-instantiate the instance of the SSOCEM 4400.
[0110] The blocking monitor 4200 may receive operating condition information representing the operating environment of the vehicle or its aspects. For example, the blocking monitor 4200 may receive operating condition information from the autonomous driving vehicle motion management controller 4100, from a sensor of the vehicle, from an external device such as a remote vehicle or an infrastructure device, or from a combination thereof. The blocking monitor 4200 may read the operating condition information or a part thereof from a memory such as a memory of the autonomous driving vehicle such as the memory 1340 shown in FIG. 1.
[0111] The blocking monitor 4200 may determine, for one or more portions of the vehicle traffic network, the respective probabilities of usability, or the corresponding blocking probabilities. The portions may include those portions of the vehicle traffic network corresponding to the predicted routes of the autonomous vehicles.
[0112] The probability of usability, or the corresponding blocking probability, may indicate the probability or likelihood that an autonomous vehicle safely traverses a portion of the vehicle traffic network or a spatial position therein, such as not being obstructed by external objects such as remote vehicles or pedestrians. For example, a portion of the vehicle traffic network may include obstacles such as stationary objects, and the probability of usability of a portion of the vehicle traffic network may be low, such as 0%, which may be expressed as a high blocking probability, such as 100%, of a portion of the vehicle traffic network. The blocking monitor 4200 may identify the respective probabilities of usability of each of a plurality of portions of the vehicle traffic network within an operating environment, such as within 300 meters of the autonomous vehicle.
[0113] The probability of usability may be indicated by the blocking monitor 4200 corresponding to each external object within the operating environment of the autonomous vehicle, and the geospatial area may be associated with a plurality of probabilities of usability corresponding to a plurality of external objects. The total probability of usability may be indicated by the blocking monitor 4200 corresponding to each type of external object within the operating environment of the autonomous vehicle, such as the probability of usability of pedestrians and the probability of usability of remote vehicles, and the geospatial area may be associated with a plurality of probabilities of usability corresponding to a plurality of external object types.
[0114] The blocking monitor 4200 may identify an external object, track the external object, and present the position information, path information, or both, or a combination thereof, of the external object. For example, the blocking monitor 4200 may be based on operating environment information (e.g., the current position of the external object), information indicating the current trajectory and / or speed of the external object, information indicating the classification type of the external object (e.g., a pedestrian or a remote vehicle), vehicle traffic network information (e.g., a crosswalk close to the external object), previously identified or tracked information associated with the external object, or any combination thereof, to identify the external object and identify the predicted path of the external object. The predicted path may indicate a sequence of predicted spatial positions, predicted temporal positions, and corresponding probabilities.
[0115] The blocking monitor 4200 may communicate the probability of usefulness, or the corresponding blocking probability, to the autonomous driving vehicle operation management controller 4100. The autonomous driving vehicle operation management controller 4100 may communicate the probability of usefulness or the corresponding blocking probability to each instantiated instance of the scenario-specific operation control evaluation module 4400.
[0116] Although not explicitly shown in FIG. 4, the autonomous driving vehicle operation management system 4000 may include a prediction module that can generate prediction information and transmit it to the blocking monitor 4200, and the blocking monitor 4200 may output the probability of usefulness information to one or more of the operation environment monitors 4300.
[0117] Each SSOCEM4400 may model a respective individual vehicle operation scenario. The autonomous vehicle operation management system 4000 includes any number of SSOCEM4400s that each model a respective individual vehicle operation scenario. Modeling an individual vehicle operation scenario may include generating and / or maintaining state information representing aspects of the vehicle's operating environment corresponding to the individual vehicle operation scenario, identifying potential interactions between the respective modeled aspects of the corresponding states, and determining candidate actions to resolve the model. More simply stated, an SSOCEM4400 may include one or more models configured to determine one or more vehicle control actions for handling a scenario given a set of inputs. The models may include, but are not limited to, a partially observable Markov decision process (POMDP) model, a Markov decision process (MDP) model, a classical planning (CP) model, a partially observable stochastic game (POSG) model, a decentralized partially observable Markov decision process (Dec-POMDP) model, a reinforcement learning (RL) model, an artificial neural network, hard-coded expert logic, or any other suitable type of model. Examples of different types of models are provided below. Each SSOCEM4400 includes computer-executable instructions that define how the model operates and how the model is utilized.
[0118] SSOCEM4400 may implement a CP model that can be a single-agent model that models individual vehicle operation scenarios based on defined input states. The defined input states may indicate the respective non-probabilistic states of the elements of the operating environment of the autonomous vehicle for individual vehicle operation scenarios. In the CP model, one or more aspects (e.g., geospatial position) of the modeled elements (e.g., external objects) associated with a temporal position may differ non-probabilistically from the corresponding aspects associated with another temporal position, such as the immediately subsequent temporal position, by a defined or fixed amount. For example, at a first temporal position, a remote vehicle may have a first geospatial position, and at an immediately subsequent second temporal position, the remote vehicle may have a second geospatial position that is different from the first geospatial position by a defined geospatial distance, such as a defined number of meters, along the predicted path of the remote vehicle.
[0119] SSOCEM4400 may implement a discrete-time stochastic control process, such as an MDP model, which can be a single-agent model that models individual vehicle operation scenarios based on defined input states. Changes in the operating environment of the autonomous vehicle, such as changes in the position of external objects, may be modeled as stochastic changes. The MDP model may utilize more processing resources and may model individual vehicle operation scenarios more accurately than the CP model.
[0120] The MDP model may model individual vehicle operation scenarios using a set of states, a set of actions, a set of state transition probabilities, a reward function, or a combination thereof. In some embodiments, modeling an individual vehicle operation scenario may include using a discount factor that can adjust or discount the output of the reward function applied over a subsequent temporary period.
[0121] The set of states may include the current state of the MDP model, one or more possible subsequent states of the MDP model, or a combination thereof. A state represents an identified situation that can probabilistically affect the operation of the vehicle in the vehicle's operating environment at an individual temporal position, which can be a predicted situation of each defined aspect, such as external objects and traffic control devices, of the vehicle's operating environment. For example, a remote vehicle operating in the vicinity of a vehicle can potentially affect the operation of the vehicle and can be represented in the MDP model. The MDP model may include representing, for each temporal position, its geospatial position, its route, direction of travel, or both, its speed, its acceleration or deceleration rate, or a combination thereof, which is the following identified or predicted information of the remote vehicle. At instantiation, the current state of the MDP model may correspond to the concurrent state or situation of the operating environment.
[0122] Any number or density of states may be used, but the number or density of states included in the model may be limited to a defined maximum number of states. For example, the model may include the 300 most likely states for the corresponding scenario.
[0123] The set of actions may include vehicle control actions available to the MDP model at each state within the set of states. Each set of actions may be defined for each individual vehicle operation scenario.
[0124] The set of state transition probabilities may probabilistically represent potential or predicted changes to the vehicle's operating environment as represented by the state in response to an action. For example, the state transition probabilities may indicate the probability that the operating environment corresponds to each state at each temporal position immediately following the current temporal position corresponding to the current state in response to the vehicle traversing the vehicle traffic network according to each action.
[0125] The set of state transition probabilities may be identified based on the operating environment information. For example, the operating environment information may include traffic conditions such as area type (e.g., urban or rural), time of day, ambient light level, weather conditions, rush hour situation, traffic congestion related to events, or predicted traffic conditions such as driver behavior related to holidays, road conditions, jurisdictional conditions such as country, state, or local municipality situation, or any other situation or combination of situations that may affect the operation of the vehicle.
[0126] Examples of state transition probabilities associated with a pedestrian's vehicle operation scenario may include a defined probability that a pedestrian crosses the road in violation of traffic regulations (e.g., based on the geospatial distance between the pedestrian and each road segment), a defined probability that a pedestrian stops at an intersection, a defined probability of a pedestrian crossing a crosswalk, a defined probability that a pedestrian yields the right of way to an autonomous vehicle at a crosswalk, and any other probability associated with the pedestrian's vehicle operation scenario.
[0127] Examples of state transition probabilities associated with a vehicle operation scenario at an intersection may include a defined probability that a remote vehicle reaches the intersection, a defined probability that a remote vehicle obstructs the progress of an autonomous vehicle, a defined probability (such as in the case where there is no right of way) that a remote vehicle crosses the intersection extremely closely behind a second remote vehicle crossing the intersection (piggybacking), a defined probability that a remote vehicle stops adjacent to the intersection in accordance with a traffic control device, regulation, or other indication of right of way before crossing the intersection, a defined probability that a remote vehicle crosses the intersection, a defined probability that a remote vehicle deviates from a predicted route proximal to the intersection, a defined probability that a remote vehicle deviates from a predicted right of way priority, and any other probability associated with the vehicle operation scenario at the intersection.
[0128] Examples of state transition probabilities associated with a vehicle operation scenario for lane change may include defined probabilities for a remote vehicle behind the vehicle to increase speed or a remote vehicle in front of the vehicle to decrease speed, i.e., defined probabilities for a remote vehicle to change speed, a defined probability for a remote vehicle in front of the vehicle to change lanes, a defined probability for a remote vehicle in proximity to the speed change of the vehicle to enable the vehicle to merge into a lane, or any other probability associated with the vehicle operation scenario for lane change.
[0129] The reward function may determine each positive or negative (cost) value that occurs for each combination of state and action. This occurrence represents the expected value of the vehicle traversing the vehicle traffic network from the corresponding state to the subsequent state according to the corresponding vehicle control action.
[0130] For example, the POMDP model may include an autonomous vehicle at a first geospatial location and a first temporal location corresponding to a first state. The model may indicate that the vehicle identifies and executes or attempts to execute a vehicle control action to traverse the vehicle traffic network from the first geospatial location to a second geospatial location at a second temporal location immediately after the first temporal location. The set of observations corresponding to the second temporal location may include operation environment information identified corresponding to the second temporal location, such as the geospatial location information of the vehicle, the geospatial location information of one or more external objects, the probability of usefulness, or the predicted route information.
[0131] The set of conditional observation probabilities may include the probabilities of making respective observations based on the operating environment of the autonomous vehicle. For example, the autonomous vehicle may approach an intersection by crossing a first road, and at the same time, a remote vehicle may approach the intersection by crossing a second road. The autonomous vehicle may identify and evaluate the operating environment information, such as sensor data corresponding to the intersection, which may include the operating environment information corresponding to the remote vehicle. The operating environment information may be inaccurate, incomplete, or incorrect. In the MDP model, the autonomous vehicle may identify the remote vehicle non-probabilistically, which may include identifying its position, or predicted route, etc. The identified information, such as the position identified based on inaccurate operating environment information, may be inaccurate or incorrect. In the POMDP model, the autonomous vehicle may identify information for probabilistically identifying the remote vehicle, such as probabilistically identifying the position information of the remote vehicle. The conditional observation probability corresponding to observing or probabilistically identifying the position of the remote vehicle represents the probability that the identified operating environment information accurately represents the position of the remote vehicle.
[0132] The set of conditional observation probabilities may be identified based on the operating environment information, such as the operating environment information described with respect to the reward function.
[0133] SSOCEM4400 may implement a Dec-POMDP model that can be a multi-agent model for modeling individual vehicle operation scenarios. The Dec-POMDP model may be similar to the POMDP model, except that the POMDP model models an appropriate subset, such as one of the vehicle and external objects, and the Dec-POMDP models the set of autonomous vehicles and external objects.
[0134] SSOCEM4400 may implement a POSG model that can be a multi-agent model for modeling individual vehicle operation scenarios. The POSG model may be similar to the Dec-POMDP model, except that the Dec-POMDP model includes a reward function for the vehicle, and the POSG model includes a reward function for the vehicle and respective reward functions for each external object.
[0135] SSOCEM4400 may implement an RL model that can be a learning model for modeling individual vehicle operation scenarios. The RL model may be similar to the MDP model or the POMDP model, except that defined state transition probabilities, observation probabilities, reward functions, or any combination thereof may be omitted from the model. Instead, for example, the RL model may be a model-based RL model that generates state transition probabilities, observation probabilities, reward functions, or any combination thereof based on one or more modeled or observed events.
[0136] In the RL model, the model may evaluate one or more events or interactions that can include simulated events, and in response to each event, generate or modify the corresponding model or its solution. Simulated events may include, for example, crossing an intersection, crossing a vehicle traffic network near a pedestrian, or changing lanes. As an example of using the RL model to cross an intersection, an RL model that shows candidate actions for crossing the intersection is included. Then, the autonomous vehicle uses the candidate actions as vehicle control actions with respect to the temporal position to cross the intersection. The result of crossing the intersection using the candidate actions may be determined, and the RL model may be updated based on the result.
[0137] The autonomous driving vehicle operation management system 4000 may include models of any number or type of combination. For example, the pedestrian SSOCEM 4410, the intersection SSOCEM 4420, and the lane change SSOCEM 4430 may implement POMDP models. In another example, the pedestrian SSOCEM 4410 may implement an MDP model, and the intersection SSOCEM 4420 and the lane change SSOCEM 4430 may implement POMDP models. Further, the autonomous driving vehicle operation management controller 4100 may instantiate any number of instances of the SSOCEM 4400 based on the operation environment information. The module 4440 is shown using a dashed line to indicate that the autonomous driving vehicle operation management system 4000 may include any number or additional types of the SSOCEM 4400.
[0138] One or more of the autonomous driving vehicle operation management controller 4100, the blocking monitor 4200, the operation environment monitor 4300, or the SSOCEM 4400 may operate continuously or periodically, such as at a frequency of 10 Hertz (10 Hz). For example, the autonomous driving vehicle operation management controller 4100 may identify vehicle control actions many times, such as 10 times per second. The operation frequencies of each component of the autonomous driving vehicle operation management system 4000 may or may not be synchronized, and the operation speed of one or more of the autonomous driving vehicle operation management controller 4100, the blocking monitor 4200, the operation environment monitor 4300, or the SSOCEM 4400 may be independent of other operation speeds.
[0139] FIG. 5 is a diagram of an example of a competence awareness system (CAS) 5000 according to an embodiment of the present disclosure. The CAS 5000 includes an autonomy cognitive agent (ACA) 5002. Any SSOCEM such as one of the SSOCEMs 4410, 4420, 4430, 4440 in FIG. 4 can be an ACA as described below with respect to the ACA 5002. The ACA 5002 can have multiple levels of autonomy. The CAS 5000 can be the autonomous driving vehicle operation management system 4000 in FIG. 4.
[0140] ACA5002 operates at multiple levels of autonomous operation (i.e., autonomy levels) and can plan using multiple levels of autonomous operation. When the AV executes vehicle control actions at an autonomy level, the AV (more specifically, the ACA) can receive a feedback signal (or simply, feedback) in response to the execution of the vehicle control actions by the AV. For example, the AV may execute vehicle control actions while being monitored by a human. The vehicle can be remotely monitored by a remote operator. The vehicle can be monitored by a human inside the vehicle. The human can cause a signal indicating feedback to be sent to or received by the AV (such as a module, component, or circuit within the AV like ACA5002).
[0141] ACA5002 can determine in advance which level of autonomy to engage. Thus, ACA5002 can have a gradient of autonomy levels. Thus, ACA5002 can maintain robustness without requiring the same effort as all - or - nothing (i.e., binary) autonomy. Thus, ACA5002 can anticipate and avoid situations where it cannot autonomously navigate and can anticipate and issue a request for help before the AV gets stuck. By anticipating and requesting help, wasted operation time and resources can be reduced.
[0142] As will be further described below, ACA5002 can learn the autonomy level to engage when crossing the same or a similar part of the vehicle traffic network from its experience of crossing a part of the vehicle traffic network. That is, ACA5002 can learn from its experience and predict the feedback received from a human operator such as a remote operator or a human inside the vehicle.
[0143] ACA5002 can modify the selection of the autonomy level based on new knowledge obtained while crossing a part of the vehicle traffic network. The learned prediction can be used to modify how ACA5002 crosses a part of the vehicle traffic network. For example, if ACA5002 has a high confidence in the response received from a human, a request for human assistance should not be issued. Thus, ACA5002 requests assistance only in situations where ACA5002 is convinced that it cannot autonomously handle the situation. Further, ACA5002 can anticipate and avoid situations where ACA5002 (based on its competence level) determines that it needs human assistance. Therefore, ACA5002 can prioritize situations where it can navigate autonomously. For example, ACA5002 can replan (e.g., determine a new route) to avoid a situation.
[0144] ACA5002 uses the domain model (DM) 5004, the human model (HM) 5006, and the autonomy model (AM) 5008 as inputs. HM5006 can also be referred to as a feedback model. ACA5002 updates HM5006 as indicated by arrow 5012 and further described below. ACA5002 updates AM5008 as indicated by arrow 5010 and further described below.
[0145] As described above, DM5004 can model the environment in which ACA5002 operates (i.e., the operating environment).
[0146] AM5008 can model the level of autonomy at which ACA5002 can operate, when (e.g., under what situations or constraints) ACA5002 can become operational at each of the autonomy levels, and what the utility of each of the autonomy levels is. The utility can indicate the expected value of executing vehicle control actions at each autonomy level. The utility values can be used to represent preferences between autonomy levels.
[0147] HM5006 can describe a feedback model that models the feedback that ACA5002 can receive from a human (e.g., a remote operator), how costly each type of feedback is, and how likely it is that ACA5002 will receive each type of feedback.
[0148] As described above, DM5004 can model the environment in which ACA5002 operates (i.e., the operating environment). For example, DM5004 can describe (e.g., include) the environmental transition dynamics and / or cost dynamics for ACA5002. In one example, DM5004 can be modeled as a Stochastic Shortest Path (SSP) problem. As is known, SSP is a formal decision-making model for inference in a fully observable stochastic environment aimed at finding the minimum-cost path from a start state to a goal state. For example, the goal of ACA5002 can be to successfully cross an intersection. Thus, the start state can be the first temporal position before the intersection (e.g., 50 meters before the intersection), and the goal state can be the second temporal position after passing through the intersection (e.g., 50 meters after the intersection). When ACA5002 approaches the intersection, ACA5002 plans a trajectory that includes a set of actions to achieve the goal. As a result of the plan, ACA5002 selects the next action to execute at the level of autonomy according to HM5006 and AM5008.
[0149] DM5004 can be formally modeled as a tuple 〈S, A, T, C, s o , G〉. S can be a finite set of states (i.e., a set among the set of operating environment information). A can be a finite set of actions (i.e., a set of vehicle control actions). T (i.e., T: S × A → Δ |S| ) can represent a transition function that describes the probability distribution over the successor states when taking action a ∈ A in state s ∈ S. C (i.e.,
Number
[0150] A solution to the SSP of DM5004 can be a policy π: S → A. That is, under policy π, action a (i.e., π(s)) is selected for state s. That is, policy π can indicate that action π(s) ∈ A should be taken in state s. Policy π is the value function V π (s) that can represent the predicted cumulative cost of reaching the target state G from state s following policy π. π : S → C can be included. That is, the value function provides the cost (i.e., value) of each intermediate state s i from the start state to reaching the target state. The optimal policy π * minimizes the predicted cumulative cost.
[0151] Therefore, for all state configurations of interest, the policy can be used to determine the action that the AV takes in that state. Examples of state descriptions can be "There is a pedestrian in front of the AV" and "The AV is at a stop sign". Such states can be associated with the action "stop" to avoid hitting the pedestrian. Another example of a state description can be "The AV is within an intersection" and "There is another vehicle at the stop sign". Such states can be associated with the action "go" so that the AV can complete crossing the intersection.
[0152] Intuitively speaking, DM5004 can include descriptors for how the world (i.e., the operating environment) changes in one time step for all combinations of states. DM5004 can include the concepts of good states and bad states. That is, DM5004 can include descriptors for which configurations of the world are good. For example, in a scenario of crossing an intersection, a good state is a state where the AV has completed crossing the intersection, and a bad state can correspond to the AV colliding with another vehicle or the AV violating the law. Negative rewards can be associated with bad states, and positive rewards can be associated with good states. Given such descriptors and rewards, the optimal policy π * can be calculated as a function of how the operating environment evolves over time.
[0153] Other fully observable probabilistic models such as MDPs may be used with minimal changes, but partially observable MDPs (POMDPs) can introduce additional sources of uncertainty, especially with respect to human interactions, making their use difficult. AM describes the levels of autonomy that an agent can operate at, the restrictions on the situations in which each level is permitted, the utility of each level, the set of system sub-competencies, or any combination thereof.
[0154] As described above, AM5010 can model the range of autonomous operations (i.e., levels of autonomy) within which ACA5002 can operate. The levels of autonomy can indicate both the actual different forms or ranges of autonomous operations (as will be described below with respect to the set L of levels of autonomy) and (as will be described below with respect to the autonomy profile κ) when each of the levels of autonomy can be permitted under some external constraints.
[0155] AM5008 can be formally modeled as a tuple 〈L, κ, μ, T〉, where L represents a finite set of autonomy levels, and each level l ∈ L corresponds to some set of constraints on the autonomous behavior of ACA5002. κ (i.e., κ: S × A → P(L)) is the autonomy profile, which describes which autonomy levels l ∈ L are permitted when performing action a ∈ A in state s ∈ S. μ (i.e., [Number] ) is the cost of autonomy, which describes the cost of taking action a ∈ A at level l ∈ L in state s ∈ S, assuming the agent has just acted at level l’ ∈ L. T is a set of sub-competencies, and each sub-competency τ i ∈ T is a mapping τ i : S × A → Δ |S| .
[0156] L = {l0,..., l n} can be a set of autonomy levels. Each autonomy level l i can correspond to a set of constraints on the autonomous behavior of ACA5002. In one example, the set of actions L can be a partially ordered set (i.e., a poset). That is, the actions in the set L can have an order or sequence, for example, indicating an improvement in the level of autonomy.
[0157] In one example, the autonomy levels of the set L can include four autonomy levels, namely, the "no autonomy" level (i.e., l0), the "verified autonomy" level (i.e., l1), the "supervised autonomy" level (i.e., l2), and the "unsupervised autonomy" level (i.e., l4). The disclosure herein is not limited to such a set of autonomy levels L. That is, other autonomy levels with different semantics are possible.
[0158] The "no autonomy" level l0 can indicate that ACA5002 requires a human to perform an action instead of ACA5002 itself. The no-autonomy level can be summarized as ACA5002 requiring a human to fully control the AV so that the human can escape from a certain situation (e.g., an obstacle scenario).
[0159] The "verified autonomy" level l1 can indicate that ACA5002 must query and receive explicit approval from a human operator even before attempting a selected (e.g., identified, determined, etc.) action. For example, in a series of actions (i.e., a plan) determined by ACA5002, ACA must seek explicit approval for each action before the action is executed.
[0160] The "supervised autonomy" level l2 can indicate that ACA5002 can operate autonomously as long as there is a human monitoring (e.g., remotely, or otherwise) ACA5002. In the "supervised autonomy" level l2, a human can intervene if something goes wrong while an action is being executed autonomously. For example, as long as a human is monitoring the AV, a series of operations can be executed. If an obstacle is detected before or after an operation in the series of operations, ACA can request assistance from a human (e.g., a remote operator).
[0161] To clarify the demarcation between the "verified autonomy" level l1 and the "supervised autonomy" level l2, an example is given here. At the "supervised autonomy" level l2, the monitoring does not need to be remote. For example, the test procedure for an AV can be considered "supervised autonomy" only when there is a human supervisor inside the AV who can override and control the AV in dangerous situations although the AV can drive autonomously. As a further demarcation, the "verified autonomy" level l1 can require that ACA5002 receive explicit permission from a human (either inside or remote from the AV) before performing the desired action. In particular, receiving explicit permission can mean that ACA5002 should stop until it receives the permission. On the other hand, at the "supervised autonomy" level l2, such a requirement does not exist as long as there is a human supervisor. That is, ACA5002 can continuously perform its desired operation without the need to stop and relying on the authority of the human supervisor who overrides in case of potential danger.
[0162] The "unsupervised autonomy" level l3 can indicate that ACA5002 can be in full autonomous operation without the need for human approval, supervision, or monitoring.
[0163] The autonomy profile κ (i.e., κ: S × A → P(L)) can map a state s ∈ S and an action a ∈ A to a subset of the set L of autonomy levels. P(L) represents the power set of the set L of autonomy levels. The autonomy profile κ can define the constraints on the permitted autonomy levels for any situation (i.e., the state of DM5004). Given the current state of the environment and the next action to be performed, the autonomy profile κ defines the set of acceptable autonomy levels.
[0164] Constraints can be hard constraints or can include hard constraints. For example, constraints can be technical, legal, or ethical constraints. By way of illustration, a non-limiting example of a legal constraint could be that an autonomous vehicle cannot operate autonomously (i.e., at unsupervised autonomy level l4) within a school zone. A non-limiting example of a social constraint could be the road rule that when the traffic signal turns green, oncoming vehicles yield the right of way to the vehicle that turns left first. Thus, the constraint is that the vehicle that turns left first must proceed without waiting for the traffic to clear.
[0165] Constraints can be, can include, or can be used as temporary conservative constraints that can be updated over time as ACA5002 improves. The autonomy profile κ can constrain the space of all policies (π), such that ACA5002 can only follow policies that never violate the autonomy profile κ.
[0166] The utility μ represents the following: Assuming that the operation at time step t was executed at autonomy level l, what is the utility of executing another operation at time step t + 1 at another autonomy level l'? The operation at time step t + 1 does not have to be the same operation as the one taken at time step t, but it can be the same operation. In some situations, there can be a negative utility associated with switching autonomy levels. For example, in dynamic situations (e.g., complex intersections), ACA5002 may switch between "supervised autonomy" and "unsupervised autonomy" every time step if no utility μ is given. Constantly switching autonomy levels can actually be more unpleasant for a human who has to constantly switch attention than simply remaining in the "supervised" mode throughout the time.
[0167] As will be further described below, AM5008 can evolve. That is, AM5008 can be trained based on the experience of ACA5002. By way of example, assume that a first AV is deployed in a first market (e.g., Japan) and a second AV is deployed in a second market (e.g., France). The first AV and the second AV may initially include the same autonomy model operating in a binary autonomy mode. That is, the AV can secretly handle situations programmed to be recognized and traversed in advance, or whether the AV requires assistance from a human (e.g., a remote operator). Since each of the first market and the second market may have different (e.g., social) road rules, the autonomy model of the first AV evolves to be different from the autonomy model of the second AV based on the feedback each receives from humans. When the autonomy models learn the situations they are capable of handling in their respective markets, ACA no longer needs to require assistance from a human (e.g., a remote operator) for the learned situations (i.e., scenarios).
[0168] Without loss of generality, L can be assumed to be a completely ordered set, and if the level of autonomy can change from one level to another, it can be extended to any graph where two levels are connected. The constraints corresponding to each level of autonomy can be either essentially technical, i.e., internally imposed constraints such as requiring human supervision during adverse weather conditions that can be known in advance to cause errors, and externally imposed constraints such as those that are essentially ethical or legal. Furthermore, κ can be defined not only to reflect fixed constraints but also to be updated over time and to reflect temporary constraints that help enable more conservative autonomous behavior while the system is still learning. Each constraint may be associated with a corresponding form of human assistance or involvement. For example, at the level of "supervised autonomy", the agent can operate completely autonomously, conditional on the presence of a human who actively monitors the agent's execution and can override any action that is considered unsafe or undesirable. Thus, the higher the level of autonomy, the lower the cost of human involvement, but this is not a requirement of the model.
[0169] The competence of the system may depend on the behavior of the subsystem that notifies the main process that the system is planning. For example, a perception system that is likely to generate incorrect state updates in some situations may contribute to lower competence in those situations. Sub-competence can be a mathematical representation of the behavior of the subsystem that notifies the main process at an abstraction level of what the planning model is inferring. Perceptual sub-competence can be defined as the likelihood of perception failure in a given state. In this case, perception can be a subsystem that notifies the main process by providing state update information. Sub-competence that is completely known in advance or assumed to be completely known in advance may be included in the domain model as part of the transition function and thus may be omitted from the autonomy model. Sub-competence that is not known in advance is a function based on observations and data collected while operating online
Number
[0170] As described above, HM5006 (i.e., the feedback model) can model the reliability of ACA5002 regarding the interaction between ACA5002 and a human operator. HM5006 can be formally represented as a tuple 〈Σ,λ,ρ,τ H 〉, where Σ represents a finite set of feedback signals that the agent can receive from a human, λ represents a feedback profile, and represents the probability distribution of the feedback signal that the agent will receive when executing action a ∈ A at level l ∈ L in state s ∈ S, assuming that the agent has just acted at level l’ ∈ L. ρ represents a human cost function, representing the cost to the human when executing action a ∈ A at level l ∈ L in state s ∈ S, assuming that the agent has just acted at level l’ ∈ L. τ H represents the human state transition function, representing the probability distribution of the subsequent state s’ ∈ S when the human controls the system when the agent attempts to execute action a ∈ A in state s ∈ S.
[0171] Σ = {σ0,...,σ nLet \(\{\}\) be the set of possible feedback signals that ACA5002 can receive from a human operator. Non-limiting examples of feedback signals are described below with respect to Figure 6. The feedback profile \(\lambda\) can represent the probability that ACA5002 receives a signal \(\sigma\in\Sigma\) when it executes an action \(a\in A\) at an autonomy level \(l'\in L\), assuming that ACA5002 is in state \(s\in S\) and has just operated at an autonomy level \(l\in L\). Thus, the feedback profile \(\lambda\) can be symbolically represented as \(\lambda:S\times L\times A\times L\rightarrow\Delta\) |Σ| It can be symbolically represented as.
[0172] "Has just operated at ~" can mean that, at time step \(t\), "has just operated at ~" can mean that it can mean the autonomy level at which the action taken by the ACA at time step \(t - 1\) was executed. As an example, assume that at time step \(t\), the ACA executes an action \(a\) at an autonomy level \(l2\) (i.e., "supervised autonomy"). Thus, the human is already participating and observing the behavior of the ACA. At time step \(t + 1\), if the ACA executes action \(a'\) again at autonomy level \(l2\), the probability that the human will override action \(a'\) can be lower than when the ACA executes action \(a\) at autonomy level \(l3\) (i.e., "unsupervised autonomy"), where the human may be more surprised and thus more likely to override the action.
[0173] The human cost function \(\rho\) can return a positive cost to a human who executes an action \(a\in A\) at an autonomy level \(l'\in L\), assuming that ACA5002 is in state \(s\in S\) and has just operated at an autonomy level \(l\in L\). The human cost function \(\rho\) can be symbolically
Number
[0174] The human state transition function τ can represent the probability that ACA5002 selects to execute an action a ∈ A in state s ∈ S and that when the human controls the AV, the human (e.g., a remote operator) transitions ACA5002 to state s' ∈ S. "Transition ACA5002 to state s" means that the human operates the AV so that state s is realized. The human state transition function τ H is such that τ H : S × A → Δ |S| can be symbolically represented as. For example, assume that the state is s (e.g., s = "at an intersection") and ACA is about to perform action a (e.g., a = "turn left"), but the human overrides ACA and takes over control. In this case, the human state transition function τ H represents the probability that the human transitions ACA to some state (e.g., completes a left turn or instead goes straight) given the state ACA was in (i.e., state s) and the action the agent was about to perform (i.e., turn left).
[0175] In practice, note that the feedback profile λ and the human state transition function τ H are not known in advance. Thus, ACA5002 can maintain respective estimated values of the feedback profile λ and the human state transition function τ H based on previous data collected by ACA5002 in the same or similar situations. The update of HM5006 is indicated by arrow 5012. Thus, as further explained with respect to FIG. 6, after ACA5002 executes an action (in the action execution phase), system 5000 can record the feedback, if any, that ACA5002 received from the human operator and use the feedback to update at least one of the feedback profile λ or the human state transition function τ H
[0176] System 5000 (i.e., Competence Awareness System (CAS)), and more specifically ACA5002, can be considered as a solution (e.g., definition, decision, etc.) to problems that combine DM5004, HM5006, and AM5008 in the context of automated planning and decision-making.
[0177] DM5004 can represent the SSP that serves as the fundamental basis for ACA5002 to find solutions. However, ACA5002 can proactively generate plans (e.g., for different autonomy levels and using different autonomy levels) that operate across multiple levels of autonomy using AM5008. This is in contrast to autonomous agents that can adjust the plan during plan execution. The proactively generated plan can receive a set of constraints κ. ACA5002 can use HM5006 to predict in advance the likelihood of each feedback signal, and as a result, ACA5002 can avoid situations where it seems unlikely to operate autonomously.
[0178] System 5000 can combine all three of DM5004, HM5006, and AM5008 into one decision-making framework. System 5000 (more specifically, ACA5002) is used to solve the problem of generating a policy to achieve its task (e.g., successfully cross an intersection).
[0179] The problem can be formally defined as an extended SSP, the details of which are presented here. The Competence Awareness System (CAS) can be represented as a tuple
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0180] CAS state
Number
Number
Number
Number
[0181] A solution for a given CAS can be a policy π that maps states and levels
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0182] The policy is restricted to be selected from Π κ Therefore, when the autonomy profile κ is changed, the space of permitted policies also changes because the policy is restricted to be selected from Π. This, in turn, may mean that the optimal policy π * is only as good as the function κ intuitively. Thus, there is a trade-off when determining the initial constraint (κ) on the permitted autonomy.
[0183] In an implementation, a conservative approach can be chosen that significantly constrains the system, such as setting |κ(s,a)| = 1 for all (s,a) ∈ S×A, thereby reducing the complexity of the problem for solving the underlying domain model with a deterministic level. However, doing so risks having a policy that is not globally optimal with respect to the level of autonomy L and may make it impossible to reach a globally optimal policy depending on the initial autonomy profile κ.
[0184] In another implementation, a risky approach is chosen by not constraining the system at all beforehand, thereby allowing the system to completely decide on the choice of the level of autonomy. This approach includes a policy that is necessarily optimal (according to the model of ACA), but it is naturally slower because the policy space is wider, and it is inherently less safe because ACA may act at an undesirable level, rendering the overall purpose of the model ineffective in a sense.
[0185] In yet another implementation, in most domains, the ideal initialization can be closer to the midpoint of the above extreme values. The autonomy profile κ can have fewer constraints in situations where the predicted cost of failure is relatively low and more constraints in situations where the cost of failure is high. For example, in AVs, the autonomy profile κ can be more initially constrained in situations that include a chaotic environment such as pedestrians, poor visibility, or a large intersection with multiple vehicles, but driving along a highway is generally less risky and can result in far fewer benefits from a constrained autonomy profile.
[0186] The components of the CAS model are the system's ability to adjust the autonomy profile over time using what the system has learned to optimize autonomy by reducing unnecessary dependence on human assistance, regardless of how the autonomy profile is initialized. However, before operating at a new level of autonomy, the system may not know anything about how humans will interact with the system at that level, i.e., the feedback profile at that new level may be initialized by default to some baseline distribution. As a result, the system may need to explore levels of autonomy that may be considered more cost-effective than its current level, and as a result, the system may generate the data necessary to improve the accuracy and reliability of its feedback profile at those levels.
[0187] However, if the system is allowed to change its own autonomy profile, without careful consideration, it may lead to serious consequences in the real world. Therefore, the embodiments disclosed herein perform gated exploration, and the system is configured to obtain permission from a human before exploring a new (i.e., unpermitted) level of autonomy. Thus, the system must first query the human to update the autonomy profile to enable such exploration, and gates the exploration of unpermitted levels by human authorities to prevent the agent from randomly performing dangerous actions.
[0188] The embodiments disclosed herein may use a variant of the ε-greedy exploration-exploitation strategy where ε is not fixed and instead is proportional to the relatively predictable cost of performing a given action at each level of autonomy. The probability of exploring a level l’ adjacent to the current level l in L is proportional to the softmax of the negative q-values of operating at level l’ over all levels adjacent to l. [Number] where adj(l, l'') is 1 if l and l' are adjacent in L or l = l', and 0 otherwise. To ensure that all permitted levels are explored efficiently, for each l ∈ L, a potential γ l ∈ [0, 1] can be implemented with a potential-based mechanism that is maintained and updated as follows at each level exploration step. [Number]
[0189] Using the various properties of the CAS, important results of the competence recognition system can be proven. In the example provided below, it can be assumed that there is a single human authority with whom the semi-autonomous systems within the CAS interact. The single human authority can be represented as H. H is a tuple <F H , λH , κ H can be represented by >, where -F H is the set of features used by H when providing feedback, -
Number
Number
[0190] Feedback consistency is the property of how consistently a human authority provides inappropriate feedback when the same query is given by the acting agent. In one example, let F H ⊂ F be the set of features H used by the human authority,
Number
Number
Number
Number
[0191] λ H Let be the stationary distribution of the feedback signal followed by the human authority. The competence of CASS denoted by X S is a mapping from H and each τ i ∈ T, given complete knowledge, to the optimal (minimum-cost) level of autonomy. Formally,
Number
Number
Number
Number
[0192] is the most beneficial (e.g., cost-effective) level of autonomy if it can know the true human feedback distribution and its own sub-competence. When L is a partially ordered set, this is generally max(κ
Number
Number
[0193] This definition of competence depends on λ H and is thus a definition of the competence of the entire human-agent system and is clearly not just a measure of the underlying agent's technical capabilities (i.e., D and T). As a consequence of this fact, a CAS that has only as much ability as a human authority figure thinks it has, and that is not fully understood by a human authority figure, may result in a system with lower competence than a human authority figure who knows the system's limits and capabilities. One reason to model competence in this way may be to avoid relying on any thresholding based on an evaluation metric to determine whether a system is capable or not.
[0194] CAS S is such that any new feedback drawn from the true distribution λ H is not expected to change the optimal level of autonomy for any
Number
Number
[0195] and for all actions a ∈ A,
Number
Mathematics
[0196] In the first proposition,
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
[0197] In the second proposition,
Mathematics
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0198] Let S be CAS. S is in a state in the following cases
Number
Number
Number
[0199] Under certain conditions, the competence recognition system S can be guaranteed to reach level - optimality. Thus, the system can be guaranteed to reach the point of operating at its true competence in all situations. If the set T of sub - competences in A is not empty (i.e., there exist sub - competences that are assumed to be completely unknown beforehand), it can be observed that the existence of a mechanism for online learning of those sub - competences may be necessary. If there are some sub - competences that are not known completely beforehand and the system has no means of learning, the true competence of the system cannot be learned online and level - optimality cannot be reached. Thus, if there is sufficient data, it can be assumed that all sub - competences τ i ∈ T can be learned online. In other words, each estimator
Number
[0200] To prove that the competence recognition system will reach level - optimality, the concept of gate search can be relied upon. However, for the following utilization, for any given
Number
Number
Number
[0201] Let S be a CAS and κ t represent the autonomy profile κ at time t.
Number
Number
Number
Number
[0202] In the first theorem, let S be a CAS that follows a gate-controlled exploration strategy and performs utilization under stationarity, where
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0203] Many problems in the open world are too complex, and even with domain expertise, it is not possible to fully specify in advance all the features that will be relevant during the system deployment process. Therefore, the CAS model, while enabling a semi-autonomous system to optimize its autonomy over time, may appear to be limited by the features within its fixed model that are insufficient to fully model human feedback.
[0204] For example, consider a robot deployed on a campus with the task of delivering packages to various offices in different buildings. Initially, the robot may be able to detect doors and know that it must seek approval or supervision before opening them, but since the nature of the doors can vary from building to building and campus to campus, its initial model does not use additional information about the doors. Human supervisors may use additional features of the doors, such as the size, shape, and whether the door can be pushed open. Since the CAS model does not represent these features, human feedback may appear inconsistent in the presence of doors because feedback is effectively normalized across all these additional features. This, in turn, can lead to a potential increase in human dependence due to reduced competence, inadequate performance, and lack of information. To address this shortcoming, methods such as the method shown in Figure 6 can be implemented.
[0205] FIG. 6 is a flow diagram of an example of a method 6000 for providing the CAS with the ability to improve its competence over time. At 6002, the method 6000 includes increasing the granularity of its state representation by online model updates. At 6004, the method 6000 includes identifying states that are considered indistinguishable under the current feedback profile of the system (i.e., states for which human feedback cannot be predicted with a high degree of confidence). At 6006, the method 6000 includes determining the features or set of features that best discriminate human feedback that is available to the system but not currently being used. This can result in a finer description of the boundaries between regions of the state space having different levels of competence.
[0206] FIG. 7 is a diagram of an exemplary example of the method of FIG. 6. This example is an illustration of one embodiment in a navigation task for a robotic system (e.g., a robot) that can utilize two important characteristics of the CAS model. The first characteristic that can be utilized can be the existing information available in a standard CAS model in the form of human feedback for identifying where new features should be added, without additional work being added by humans. The second characteristic that can be utilized can include the characteristics of the interaction between humans and agents for avoiding the need to directly modify the transition function or reward function and only directly modifying the state space. As a result, the entire process can be executed completely autonomously online.
[0207] FIG. 7 includes a navigation environment 7000 that includes several corridors and doors. The doors are shown as red (R) doors, blue (B) doors, and green (G) doors, with each color representing a different type of door. The navigation environment 7000 includes several paths that correspond to optimal paths under different granularities of state space representation. For example, path 7002 (shown as a solid line) may correspond to an optimal path for a robot to cross using an R door. Path 7004 (shown as a thick dashed line) may correspond to an optimal path for a robot to cross using a B door. Path 7006 (shown as a thin dashed line) may correspond to an optimal path for a robot to cross using a G door. When features are identified and added to the state representation, the system can better learn to utilize human assistance and crossing paths that better fit the system's competence.
[0208] The embodiments described herein may determine a state that cannot be discriminated. In one example, let S be a competence recognition system. When a robot system is deployed in an open world, both the exact environment in which the system will operate and the human authorities with whom the system will interact may not be known in advance. Simply including all possible features available to the system from perception or external sources in its planning model may add no useful information and, in the case of many features that only increase the number of states, may make the planning cumbersome without benefit. Thus, assume that S has available a complete feature space that can be partitioned into an active feature space used by S and an inactive feature space not yet used by S within its planning model. However, when S receives additional feedback over time, S learns to utilize some of the inactive features, adds them to its state representation, and more effectively aligns those features with features used by human authorities.
[0209] For example, the complete feature space available to S from that sensor or other external source can be divided into an active feature space used by S and a non-active feature space not yet used by S. As S receives additional feedback over time, S will learn to utilize non-active features in order to more effectively align with the features used by human authorities.
[0210] If the complete feature space F = {F1, F2,..., F n} available to S is given, the active feature space is
Number
Number
Number
Number
[0211] In one example, the human authority H has
Number
Number
Number
Number
[0212] Here, δ is called the discrimination slack and is used to determine the required prediction confidence for states declared as non-discriminable. For example, the lower the discrimination slack is set, the higher the required confidence becomes. The discrimination slack helps provide a formal trade-off mechanism between increasing the complexity of the underlying planning model and the completeness of the competence recognition model. The decision of how to set δ can be made by expertise, offline evaluation, and dynamic online adjustment.
[0213] Given the concept of an indistinguishable state, the central idea of this approach can be defined. A discriminator is any subset of the inactive feature space that can help an agent better discriminate feedback from H about states the agent cannot distinguish. As an example, initially consider only the presence of a door within its active feature space, but consider an agent that has additional features representing the size of the door within its inactive feature space. Without these additional features, the agent may perceive equal approvals and denials from human authorities, leading to a feedback profile with a low probability of any feedback signal for all doors, while H has consistently not permitted the robot to open doors of a particular size to prevent self - damage. By including these features representing the size of the door within the active feature space, the agent's new feedback profile may be able to predict the correct feedback signal for both small and large doors with a high probability.
[0214] The embodiments disclosed herein may perform iterative state - space refinement. In one example, a discriminator is any subset
Number
Number
[0215] FIG. 8 is a diagram showing an example of a single-step state space refinement algorithm 8000. The single-step state space refinement algorithm 8000 represents pseudocode for improving the competence of the CAS by iterative partitioning of the state space by adding new features to the state representation over time. As shown in FIG. 8, the single-step state space refinement algorithm first identifies 8002 the set of currently indistinguishable states. To avoid accidentally and indiscriminately labeling sparsely sampled state-action pairs,
Number
[0216] Next, the single-step state space refinement algorithm 8000 samples 8006 indistinguishable states and identifies 8008 the most likely discriminator for that state using any standard feature selection technique such as minimum redundancy maximum relevance (mRMR). For each potential discriminator, a new feedback profile is trained 8010 using a portion of the entire dataset in the state where the discriminator is temporarily added to the set of active features. The discriminator that results in the best performing feedback profile is selected 8012. In one example, the discriminator that results in the best performing feedback profile may be the highest Matthews coefficient. If the verification is successful, the discriminator is added to the set of active features and the system is updated 8014.
[0217] Two assumptions can be made in the design and use of the single-step state space refinement algorithm 8000. First, the initial transition function provided by the domain model is such that the agent is κ HIt can be assumed to be sufficiently correct for any scenario that enables autonomous action. In some examples, competence can be improved by iteratively refining the state space. Also, as the human authority gains a deeper understanding of the agent's capabilities, it may be possible to enhance competence by directly updating the transition function and replanning.
[0218] Second, it can be assumed that the human authority has a sufficient understanding of the agent's capabilities both to prevent the execution of actions that the agent cannot successfully perform and to provide consistent feedback. This assumption can be made for two reasons. First, there may be various ways to deepen the authority's understanding of the system's capabilities so that the authority can have appropriate trust or dependence on the system. These can include pre-deployment training, standardized feedback criteria, and system expertise. Second, the recognition of potential obstacles and the handling of obstacle recovery are separate areas of active investigation orthogonal to the examples described herein.
[0219] Under these assumptions, directly updating the transition function or reward function of the domain model may not be required at any point in time. It is sufficient if the agent can distinguish between actions that have the competence to execute autonomously and actions that require human involvement.
[0220] Adding discriminators does not prevent discriminable states from being discriminable. Any given discriminable state will either be affected by the discriminator or not. If the state is not affected, the feedback profile for the state does not change. If the state is affected, the initial state within the problem no longer exists by definition. More importantly, if a sufficient set of features is provided, it can be guaranteed that all states will ultimately be appropriately discriminated.
[0221] The following theorem states that if all the features that humans use to determine their feedback are available to the robot, there must exist a point in time when the robot has completely discriminated all states, and there are no states that cannot be discriminated beyond that point. I t Let be the number of states that cannot be discriminated at time t,
Number
Number
Number
Number
Number
[0222] To prove the above theorem, first, assume F H ⊆ F, and if there exists a point such that
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0223] Figure 9 is an example of a flowchart of a technique 9000 for competence recognition decision-making according to an embodiment of the present disclosure. The technique 9000 can be implemented by a competence recognition system (CAS) such as the competence recognition system (CAS) 5000 of FIG. 5. The technique 9000 can be implemented by an autonomous cognitive agent such as the ACA 5002 of FIG. 5. Some or all of the operations of the technique 9000 can be implemented by the SSOCEM 4400 or by another component of the autonomous vehicle operation management system 4000 of FIG. 4.
[0224] The flowchart 9000 shows how a CAS, or an ACA of a CAS, can select an action, receive feedback, and update the model(s) that are the subject(s) of the experience(s). The flowchart 9000 will be described with reference to FIG. 7.
[0225] Figure 10 is a diagram of an autonomous driving scenario 10000 used to explain the technique 9000 of FIG. 9. The scenario 10000 includes an intersection 10002. A vehicle 10004 is approaching the intersection 10000. A stop line 10006 (i.e., a stop sign) is a line at which the vehicle 10004 is to stop in order to ensure safe progress along a track 10008 (i.e., a route). The vehicle 10004 can be the vehicle 1000 of FIG. 1. The vehicle 10004 can be one of the vehicles 2100 / 2110 of FIG. 2. The vehicle 10004 can include an autonomous vehicle operation management system such as the autonomous vehicle operation management system 4000 of FIG. 4. The vehicle 10004 can include a competence recognition system (CAS) such as the system 5000 of FIG. 5. Thus, the vehicle 10004 can be an autonomous vehicle or a semi-autonomous vehicle. A human operator (e.g., a remote operator or an in-vehicle operator) can play a role of monitoring and assisting the vehicle 10004 (remotely), such as in response to issuing a remote request for assistance from the remote operator by the vehicle 10004.
[0226] Flowchart 9000 is described with respect to a domain model (such as DM5004 in FIG. 5) related to an intersection scenario such as intersection 10002 in FIG. 10.
[0227] The domain model of the intersection scenario in FIG. 10 can include abstracted (e.g., symbolic) information extracted based on sensor (lidar, radar, camera, etc.) information. For example, the domain model can include information such as the position of vehicle 10004 (e.g., "approaching the stop line of the intersection"), other related world objects (e.g., vehicle 10010 which is the "vehicle on the left"), road configuration (e.g., "the east-west road is through traffic" means there are no traffic signals, stop signs, yield signs, etc.), and other related information for the scenario (e.g., "the object on the left has not moved for a long time", "the left sensor is blocked", etc.).
[0228] As can be understood, an identified obstacle or shield (such as vehicle 2010) can be so identified when vehicle 10004 is at a certain distance from the obstacle. However, as vehicle 10004 approaches the shield or obstacle, vehicle 10004 can be determined that the shield or obstacle is not so. This may be due to noise in the sensor data. Thus, the state associated with the scenario is corrected at 9002 in FIG. 9 and a new plan is calculated.
[0229] More generally, a scenario can be described in an abstract state space and / or as a combination of objects related to the detected scenario. The state can include all the objects necessary for autonomous or at least semi-autonomous success. For each of these objects, the state space of the AV (e.g., vehicle 10004) with respect to that object (e.g., vehicle 10010) can be maintained. The combination of all the state spaces together can form the domain space in which competence modeling is performed (e.g., can constitute the domain space, can be the domain space, etc.). Thus, competence is thus modeled with respect to the entity-entity pairing for each pair of AV-other entity related to the scenario. For example, if scenario 10000 includes a pedestrian, for the pairing of vehicle 10004 and the pedestrian, competence modeling (i.e., determining the level of autonomy) is determined separately.
[0230] The states used in the domain model, or the states used to evolve the system, can be specified at different levels of granularity. By way of illustration and without limitation, the domain model can be for a particular intersection 10002, the domain model can be for all similarly configured intersections within any geographical area of interest, the domain model can be for crossing one or more intersections at a certain time of day, and the domain model can include additional metadata information such as an event venue proximal to the location of the intersection and that a concert has just ended, etc.
[0231] Flowchart 9000 is described with respect to a set A of actions. The set of actions is domain - related. With respect to the domain of intersections, the actions can include the actions "go", "stop", and "creep forward". However, the set of actions can include more, fewer, other actions, or combinations thereof. In another scenario, the set of actions can be a different set of actions. For example, with respect to a passing scenario, the actions can include the actions "follow", "stop", "overtake on the left", and "overtake on the right".
[0232] The action "go" can mean proceeding along track 10008. The action "stop" means that the vehicle should remain stationary in the next step / action step or should stop if it was moving. The action "creep forward" can mean that vehicle 10004 should creep slightly forward from a stop position such as stop line 10006. The "creep forward" action can be useful, for example, when the sensors of vehicle 10004 are blocked at the current position of vehicle 10004. For example, in scenario 10000, the field of view of the left - hand sensor of AV10004 is blocked by vehicle 10010. Thus, similar to what a human driver could do, the "creep forward" action moves vehicle 10004 slightly forward into the intersection in order to try to see past vehicle 10010.
[0233] Flowchart 9000 is described with respect to a set L of autonomy levels that includes four autonomy levels L = {l0, l1, l2, l3}.
[0234] Autonomy level l0 can indicate "no autonomy", which means that direct human assistance is required and that the ACA of vehicle 10004 has determined, if at all, that it does not have the ability to execute the determined action.
[0235] The autonomy level l1 can indicate "verified autonomy", which can mean that the ACA of vehicle 10004 queries and receives approval from a human (e.g., a remote operator) before executing the selected action. Since the ACA of vehicle 10004 has determined that it does not have the ability to fully autonomously execute the determined (e.g., selected) action, assistance is required.
[0236] The autonomy level l2 can indicate "supervised autonomy", which requires that a human be present (e.g., remotely monitoring) and be able to intervene (e.g., override the system) if an obstacle occurs while the selected action is being executed.
[0237] The autonomy level l3 can indicate "unsupervised autonomy", which can mean that the selected action can be executed (i.e., performed) without any human intervention, supervision, monitoring, etc. That is, the autonomy level l3 indicates that the ACA has sufficient ability to execute the action.
[0238] The flowchart 9000 is described with respect to a set Σ of feedback signals that includes four feedback signals, i.e., the set Σ = {no feedback, approval, denial, override}. For ease of reference, the "no feedback", "approval", "denial", and "override" signals are respectively
Number
Number
Number
Number
[0239] Furthermore, flowchart 9000 assumes that feedback "approval" and "denial" can only be received at autonomy level l1 (i.e., "verified autonomy"), and feedback "override" can only be received at autonomy level l2 (i.e., "supervised autonomy").
[0240] When L and Σ are given, the state transition function T of this CAS can be specified. [Number] and [Number] Then,
Number
Number
[0241] In Equation (7), [·] represents the Iverson bracket. Equation (7) is as follows: when ACA operates at autonomy level l0 (i.e., "no autonomy"), ACA can follow the transition dynamics of the human (e.g., remote operator) who performs the control; when ACA operates at autonomy level l1 (i.e., "verified autonomy"), the probability that vehicle 10004 reaches state s' is the product of the probability that ACA is approved to take an action and the probability that ACA successfully follows T, plus the probability that the action is denied and the state remains the same; when the agent operates at autonomy level l2 (i.e., "supervised autonomy"), the probability that vehicle 5004 reaches state s' is the probability that ACA successfully follows T without any human intervention, plus the probability that the human overrides the action selected by ACA and the human puts the vehicle in that state (i.e., τ H ); and when ACA operates at autonomy level l3 (i.e., "unsupervised autonomy"), ACA follows the transition dynamics of its own model (e.g., domain model). It can be summarized as such.
[0242] State
Number
Number
Number
Number
[0243] Equation (8) represents the state in the extended SSP
Number
Number
Number
Number
Number
Number
[0244] Although the model of the competence recognition system or the self-awareness agent has been described therein, Technique 9000 can be summarized as including three stages: an action selection stage, an action execution stage, and a model update stage. In the action selection stage, actions and levels of self-awareness are selected. In the action execution stage, operations are executed according to the level of self-awareness, and feedback from a human operator or the like is received. In the model update stage, Technique 9000 updates the parameters of the feedback model (e.g., HM5006 in FIG. 5) according to new experiences. More specifically, Technique 9000 updates the feedback profile λ and / or the human state transition function τ H thereof.
[0245] In 9002, Technique 9000 can detect the state s of the world around a vehicle such as vehicle 10004 in FIG. 10. The state s is the state of the domain model (e.g., DM5004 in FIG. 5). The state s can be detected using vehicle sensors such as one or more of the sensors 1360 in FIG. 1.
[0246] In 9004, the detected state s can be used to select a policy π. As described above, the policy π can be restricted by a self-awareness profile κ that can include one or more constraints. For example, if the level of self-awareness selected by the ACA is not within the self-awareness profile κ, the ACA can replan to select new actions and a new level of self-awareness permitted by the self-awareness profile.
[0247] The corresponding competence probability can be associated with each level of self-awareness. That is, the competence probability can be associated with taking an action in a particular state.
[0248] For example, a 10% competence probability associated with choosing to execute an action (e.g., "go") autonomously (i.e., with "unsupervised autonomy") in a particular state can mean that there is a 10% chance of autonomously executing the action in that particular state at the correct level of autonomy. Thus, if the action is executed autonomously but the resulting state is a bad state, there can be a 0.1 probability of a large penalty being imposed. Conversely, the probability of correctly choosing the level of competence can be 0.9. Thus, competence is incorporated into the probability of a state transition, i.e., the probability of transitioning from s at time T to s' at time T+1, and therefore, when rewards are evaluated over time, the system will autonomously choose whether to perform all the calculations in the same way that insurance mathematics calculates similar risks. It basically performs a large-scale risk assessment using all of these probabilities over time and determines at which points it should be autonomous and when it should ask for help and when it should ask for approval.
[0249] The competence probability can be considered (e.g., used) when calculating the policy π. The competence probability can be a factor in the state transition from the first state s1 at time T to the second state s2 at time T+1.
[0250] The corresponding cost can be associated with the selected level of autonomy. In one example, the cost (e.g., remuneration) can be in units of time. However, other units of measurement can also be used. For example, since a remote operator first needs to obtain a situation awareness of the situation, analyze the situation, control the AV, and actually control the AV to execute an action, it takes a long time to seek help from a human. Such an interaction may take, for example, 60 seconds and can be considered to be extremely costly. On the other hand, in the case of the autonomy level l3 (i.e., "unsupervised autonomy"), no delay occurs. Therefore, when calculating the policy π, the cost in terms of time or some other unit, which selects different levels of autonomy that involve humans to different extents in different amounts, can be considered (e.g., used).
[0251] In 9006, the technology 9000 can output an action a and a level l of autonomy. Therefore, in 9004, a tuple (a, l) can be output. The level l of autonomy determines how the ACA of the vehicle 10004 executes the action a. Examples of interpreting the tuple (a, l) include "The ACA selects a stop action, but the ACA will seek help as to whether stopping is the correct action" or "The ACA selects a go action because the ACA is fully capable and will go."
[0252] When the level of autonomy is l0 (i.e., "no autonomy"), the technology 9000 proceeds to 9014. For example, a human operator can remotely control the vehicle 10004 or send a command to the vehicle 10004 to execute an action. After the manual operation in 9014 is completed, the technology 9000 proceeds to 9016.
[0253] When the level of autonomy is l1 (i.e., "verified autonomy"), Technique 9000 proceeds to 9008. In a simple example, the ACA of vehicle 10004 may determine that, since the view is blocked by vehicle 10010, the action to be taken is to "slowly move forward" towards the center of intersection 10002. However, before completing the action, vehicle 10004 can transfer the action to a human (e.g., a remote operator) for approval. The remote operator can send a "deny" feedback signal (and the ACA can receive the feedback signal), whereby the level of autonomy is switched to "no autonomy" at 9014. The ACA can receive an "approve" feedback signal. In the case of an "approve" feedback signal, Technique 9000 proceeds to 9010 and executes the selected action.
[0254] When the level of autonomy is l2 (i.e., "supervised autonomy"), Technique 9000 proceeds to 9010. At 9010, before attempting the selected action, the ACA of vehicle 10004 can ensure that the remote operator is monitoring vehicle 10004 while the action is being executed. While the ACA is executing the selected action, the ACA can complete the action without receiving feedback from a human at 9012. Thus, after the action is completed, Technique 9000 proceeds to 9016. On the other hand, while the action is being executed, at 9014, the ACA can receive an override feedback signal, whereby the level of autonomy transitions to the "no autonomy" level.
[0255] When the level of autonomy is l3 (i.e., "unsupervised autonomy"), the action is autonomously executed by vehicle 5004, and Technique 9000 proceeds to 9016.
[0256] At 9016, Technique 9000 updates the feedback profile λ and the human state transition function τ based on the received feedback or feedback that may be feedback-free. H to update.
[0257] How the feedback profile λ is updated may depend on the classifier being used. However, generally, the feedback profile λ can be updated by augmenting the data set of the received human feedback signal and retraining the classifier (decision maker) on the new data.
[0258] Human state transition function τ H can be updated by observing to which new state the ACA goes when the human takes over control and generally taking maximum a posteriori estimation. In practice, this can be done simply by counting the frequencies of all the states that the ACA finally reaches when there was an intention or attempt to take an action
Number
Number
[0259] The model can be updated using model-free reinforcement learning, model-based reinforcement learning, or some other learning technique. In model-free reinforcement learning, the probability values can be adjusted up and down. In model-based reinforcement learning, the probability values can be calculated based on the number of times the target state within the scenario was successfully reached compared to the total number of times the scenario was encountered.
[0260] In an example of reinforcement learning, an estimated value of how much time (e.g., cost) it takes for the ACA to reach its target state (e.g., passing through an intersection) is maintained. Such time can be maintained at different levels of autonomy. If the target state has not been reached (e.g., not crossed the intersection), if human assistance is requested, or if the human assists or rejects the ACA, etc., the probability can be adjusted downward. In the case of "unsupervised autonomy" or "without feedback", the probability can be adjusted upward. That is, for all different combinations of states, the probability can always be adjusted up and down.
[0261] In some implementations, from 9016, technology 9000 can proceed to 9018, and in other implementations, technology 9000 can return to 9002. That is, block 9018 can be an optional block.
[0262] At 9018, technology 9000 can engage in gate-controlled exploration (GE).
[0263] As described above, the basic component of the CAS is that the system can optimize its autonomy by reducing unnecessary dependence on humans by using what it has learned to adjust the autonomy profile κ over time. However, before operating at a new level of autonomy, the ACA may not have knowledge of how humans will interact with the ACA at that level of autonomy. Therefore, the feedback profile at the new level of autonomy can be uniformly random because no data is being received. As a result, the CAS needs to "explore" levels of autonomy for which it has reason to believe the CAS may be more cost-effective than its current level. Thus, the CAS can generate the data necessary to improve the accuracy and reliability of its feedback profile at those levels.
[0264] However, any kind of random or pseudo-random search (e.g., random search, ε-greedy search) can lead to frequent disruptions and, in the real world, can result in serious consequences (e.g., causing collisions, traffic jams, etc.). For this purpose, a simple extension of a conventional search method called gate-controlled search can be used herein.
[0265] In gate-controlled search, the CAS can still follow a random or pseudo-random search policy to attempt to act at an autonomy level not permitted by the autonomy profile κ. When this occurs, instead of simply executing the action, the ACA can first request that the human permit the ACA to change the autonomy profile κ, whereby the autonomy level the ACA attempts to reach is then permitted.
[0266] In the literature of reinforcement learning with an unknown domain, the agent has to trade off between exploiting the information it has and simply taking the actions that had the best execution in the past, or exploring new actions and new states (or what was merely sub-optimal in the past) that may lead to better results.
[0267] This concept is utilized herein. In the model described herein, it is assumed that the optimal autonomy level is not known in advance; otherwise, the ACA could simply be designed to execute at that level. Therefore, the ACA has to learn over time what its optimal level is. However, since the ACA operates in the real world while learning, the ACA has to make this same trade-off. That is, the ACA has to consider the history of the received human feedback signals and decide whether to operate at the current optimal autonomy level or to try a new, possibly higher, autonomy level where the feedback signal is absent or limited. This is the "exploration" part of gate-controlled search (GE).
[0268] However, simply enabling the ACA to operate at a level of autonomy where it is not permitted in order to try whether it is better would defeat the purpose of the model described herein. Therefore, when wanting to change the highest level of autonomy at which the ACA can operate, one must first consult a human and change that permitted level of autonomy. This is the “gate-controlled” part of the gate-controlled search (GE).
[0269] In general, the ACA is expected to explore higher levels of autonomy, but assuming a conservative initial model, this is not a requirement, and it is noteworthy that in fact the agent can explore downward. That is, the ACA can actually lower the highest level of autonomy at which the ACA can act by consulting a human, depending on the situation. By doing so, the ACA is forced to act at a level of autonomy that it would not generally do, and it is possible to obtain more data at that level and improve the quality of the model.
[0270] From 9018, the technique 9000 returns to 9002 and repeats the above-described operations, whereby the technique 9000 detects the current state of the world, selects the actions and autonomy levels to execute, and so on.
[0271] To test the competence recognition system, the CAS model can be implemented in two simulated autonomous vehicle domains at different levels of abstraction. The first domain can be a high-level navigation problem where the autonomous vehicle has to plan and execute the optimal route between two locations, conditioned on its knowledge of different intersections and roads and its different operations that can be executed in the previous domain, i.e., its own competence in passing obstacles blocking the lane.
[0272] In the navigation domain of an autonomous vehicle, the autonomous vehicle operates within a known map represented by a directed graph G=(V,E), where each vertex v∈V represents an intersection and each edge e∈E represents a road. The autonomous vehicle is tasked with simply navigating the map from a start node to a goal node. The state of each vertex v∈S is represented by a tuple 〈ID,p,o,v,θ〉, where ID is the ID of the vertex, p is a boolean value indicating the presence of pedestrians, o is a boolean value indicating the presence of obstacles, v is the number of other vehicles if any, and θ is the direction of travel of the vehicle. The state of each edge e∈S is represented by a tuple 〈u,v,l,θ,o〉, where u and v are the IDs of the start and end states of the edge respectively, l is the number of lanes along the edge, θ is the direction of travel, and o is a boolean value representing the presence of an obstacle blocking the AV's lane. Further, each edge is associated with a travel length and speed. Model parameters such as the probability of encountering a pedestrian at a vertex or an obstacle on an edge are given as part of the model input.
[0273] In the state of a vertex, the agent can perform one of straight, right turn, left turn, U-turn, or wait. All maneuvers are assumed to succeed deterministically. In the state of an edge, the agent can perform one of continue, overtake an obstacle, or wait. Overtaking is assumed to succeed with a probability [0.2,0.5,0.8] depending on the number of lanes, continue deterministically fails in the presence of an obstacle, and otherwise transitions the agent to the end vertex of the edge with probability p∝speed / length. Each action has a unit cost in the domain model.
[0274] The self-discipline profile κ is initialized to L in the state of an obstacle-free edge and in the node state without pedestrians, shelters, other vehicles, or when in action waiting. In all other cases, κ is initialized to {l0, l1}. The feedback profile λ is initialized to be uniformly random across possible feedback signals. Since human control is required for manual control of the vehicle, the cost for a human to operate at l0 is 8.0, the cost for a human to operate at l1 is 3.0, the cost at level l2 is 1.0, and there is no additional cost for a human when operating at l3. The system incurs a cost of 1.0 when receiving a negative response at l1 and a cost of 3.0 when receiving an override at l2.
[0275] In the obstacle-passing domain of the autonomous vehicle, the autonomous vehicle must overtake an obstacle blocking the lane on a one-lane-per-side road, and importantly, for this purpose, it must enter the oncoming lane. The state s ∈ S is represented by the tuple 〈p, o, t, d, w〉, where p is the position of the vehicle (0 - 4), o is the position of the nearest oncoming vehicle (1 - 3, 0 if none, -1 if unknown), t represents the presence of a following vehicle, d represents whether the obstacle is dynamic (e.g., a slow-moving tractor) or static (i.e., debris or a stalled vehicle), and w represents whether the nearest oncoming vehicle has stopped.
[0276] The autonomous vehicle can execute the following actions: wait, creep forward, and go. Creep forward advances the position of the AV by 1 with a probability of 0.5 unless at position 0, in which case it deterministically advances to the edging position 1. Go deterministically advances the position of the AV by 1 at all positions except 0, and at position 0, it advances the AV to position 2 (i.e., skips the edging position). All actions have a unit cost as long as the AV and oncoming vehicles do not share or cross positions where a high-cost crash would occur.
[0277] The self-discipline profile κ is initialized to {l2} in all cases, i.e., in domains where such safety is important, initially, it is expected that humans are always ready to recognize and override the system. As described above, the feedback profile λ is initialized uniformly at random. When the CAS operates at l0 and it is assumed that the operation is successfully completed (i.e., the human does not return control while passing an obstacle), a large cost of 10.0 is incurred for the human, while the cost is 1.0 when supervising at l2 and no cost is incurred at l3. The system incurs a large penalty of 12.0 when overridden by a human.
[0278] To evaluate the iterative state-space refinement approach, a simulated domain can be implemented that includes a mobile robot tasked with delivering packages to various rooms within various buildings across a small campus. To achieve its goal, the robot has to handle two major obstacles: doors and crosswalks, and inappropriate handling of them by the agent can lead to costly obstacles. In this case, the initial domain model of the CAS may lack specific functions used by humans when determining such feedback.
[0279] In the delivery robot domain, the robot operates within a known map and is tasked with delivering packages from one office to another within a campus environment. The robot must safely navigate an environment that includes closed doors and crosswalks across major roads. A state s ∈ S is represented by the tuple 〈x, y, θ, o〉, where x, y, and θ are the robot's pose, and o represents the presence of a door, traffic conditions (when on a crosswalk), or no obstacles at the robot's current position. Additional information is available from sensor information about each obstacle but is not pre - used in the domain model. In the case of a door, this information includes the door's color, height, width, and opening type (i.e., pull or push). In the case of a crosswalk, this includes whether there are obstacles blocking the view and whether the road is one - way or two - way. Additionally, the time is also always known to the robot.
[0280] The robot can execute the following actions: move, open, wait, cross, and hand - over. Move advances the agent in their direction of travel. Open is not performed if the robot is not at the door, but otherwise, if the robot can open the door, it deterministically opens the door, and if not, a penalty is imposed on the robot (note: the initial model of the robot does not distinguish and assumes that all doors can be opened). Wait is not performed unless the robot is on a crosswalk where traffic conditions can change. Cross is not performed unless the robot is on a crosswalk. When on a crosswalk, the robot deterministically crosses if the traffic volume is nil (empty), crosses successfully with a 50% probability if the traffic volume is low (otherwise stays in the same place), and when the traffic volume is high, the probability of crossing is 10%, the probability of crashing and getting stuck is 10%, and the probability of staying in the same place is 80%.
[0281] When the robot is operating before delivering the package, a negative unit reward is generated at each time step. If the robot attempts to open a specific type of door that it does not have the ability to open, the robot itself may be damaged and incur a small penalty (-10). If the robot collides with a vehicle while crossing the road, it incurs a very large penalty (-100).
[0282] The autonomy profile κ is initialized to L when there are no obstacles or when performing action waiting, and is initialized to {l0, l1} otherwise. As described above, the feedback profile λ is initialized uniformly at random. The cost for operating at level l0 is 10.0, the cost for operating at level l1 is 2.0, the cost for operating at level l2 is 1.0, and there is no additional cost at level l3. In this domain, no additional cost is incurred for the received feedback.
[0283] FIG. 11 is a flowchart diagram of an example of a technique 11000 for autonomous driving by an autonomous vehicle (AV) according to an embodiment of the present disclosure. The technique 11000 of FIG. 11 can be implemented by the competence recognition system or the autonomy recognition agent of FIG. 5. The technique 11000 can be implemented in the vehicle 1000 shown in FIG. 1, one of the vehicles 2100 / 2110 shown in FIG. 2, a semi-autonomous vehicle, or any other AV that performs autonomous driving.
[0284] At 11110, the technique 11000 detects the environmental state based on the sensor data. In one example, the environmental state can be a contributing state as described above with respect to the extended SSP problem. In one example, the state can be a state s of the set of states S. The state can be detected as described with respect to 9002 of FIG. 9.
[0285] At 11120, the technique 11000 selects an action based on the environmental state. The action can be selected as described with respect to 9006 of FIG. 9.
[0286] At 11130, the technique 11000 identifies a set of currently indistinguishable states. Conditional on the assumption that there exists a true correct feedback signal that is returned by a human with at least a probability of ε for all state-action pairs, to avoid accidentally and indiscriminately labeling sparsely sampled state-action pairs, the probability of observing all labeled instances of its elements within the existing dataset D is at least a certain threshold p ε Only when this is the case can the process be restricted to consider state-action pairs.
[0287] At 11140, the technique 11000 samples states that are indistinguishable from the set of indistinguishable states and identifies one or more potential discriminators for those states using a feature selection technique such as mRMR.
[0288] At 11150, for each potential discriminator, a new feedback profile can be trained using a portion of the entire dataset in a state where the discriminator is temporarily added to the set of active features. If the verification is successful, the discriminator is added to the set of active features and the system is updated.
[0289] At 11160, technique 11000 determines an autonomy level associated with an environmental state and an action. The autonomy level can be selected based on at least an autonomy model and a feedback model, as described with respect to 9004 in FIG. 9. The autonomy model can be as described with respect to AM5008 in FIG. 5. Thus, in one example, the autonomy model can include a utility model and an autonomy profile. The utility model can describe the utility of performing a first action at a first autonomy level for a first environmental state when the AV has just operated at a second autonomy level. The autonomy profile can map each environmental state to each action and can define constraints on the level of autonomy permitted for a particular environmental state. The feedback model can be as described with respect to HM5006 in FIG. 5.
[0290] In one example, the autonomy level can be selected from a set of autonomy levels that includes a first autonomy level indicating "no autonomy", a second autonomy level indicating "autonomy with verification", a third autonomy level indicating "autonomy with supervision", and a fourth autonomy level indicating "autonomy without supervision".
[0291] At 11170, technique 11000 performs an action according to the autonomy level.
[0292] In the example of 11170, the autonomy level can be the second autonomy level indicating "autonomy with verification", and performing an action according to the autonomy level can include receiving an approval feedback signal or a denial feedback signal for the action. The feedback signal can be received from a human. In one example, the human can be a human in the vehicle. In one example, the human can be a remote operator. In one example, performing an action according to the autonomy level can include querying for an approval feedback signal before receiving the approval feedback signal.
[0293] In the example of 11170, the autonomy level can be a third autonomy level indicating "autonomy with supervision", and executing an action according to the autonomy level can include determining that the AV is being monitored by a human before executing the action. For example, technology 11000 can receive a signal from a human indicating that the human is monitoring the AV.
[0294] In the example of 11170, as described with respect to FIG. 9, an approval feedback signal can be received, and executing an action according to the autonomy level can include determining that the AV is being monitored, executing the action in response to determining that the AV is being monitored, receiving an override signal, and in response to receiving the override signal, stopping the action and switching the AV to a manual operation mode.
[0295] In the example of 11170, executing an action according to the autonomy level can include determining whether to request human approval for the action before executing the action based on the autonomy level.
[0296] In one example, as described with respect to 9014 of FIG. 9, the autonomy level can be "no autonomy" such that the AV is not permitted to execute autonomous actions, and executing an action according to the autonomy level can include enabling the AV to be manually controlled by a human.
[0297] In one example, the technique 11000 can include updating at least one of an autonomy profile, a feedback profile, or a human transition function in response to performing an action. Updating each of the autonomy profile, the feedback profile, or the human transition function for an AV can be based on a feedback signal received by the AV itself. However, the updating can be based on feedback signals received by many AVs (such as AVs within a fleet). For example, all signals sent to an AV can be aggregated at a central location, the model(s) can be updated, and the updated model can be redistributed to the AV. However, other ways of updating the model based on feedback signals received by multiple AVs (i.e., sent to multiple AVs) are possible.
[0298] As described above, one aspect of the disclosed implementations can include a system including a memory, such as the memory 1340 of FIG. 1, and a processor, such as the processor 1330 of FIG. 1. The memory can include instructions executable by the processor to calculate a policy for solving a task, as described above with respect to the extended stochastic shortest path (SSP) problem. The policy can map environmental states and autonomy levels to actions and autonomy levels. Calculating the policy can include generating a plan that operates across multiple levels of autonomy.
[0299] In one example, generating a plan that operates across multiple levels of autonomy can include generating the plan subject to constraints on the permitted level of autonomy in each state, as described above with respect to the autonomy profile κ.
[0300] In one example, the instructions can include instructions for updating a feedback profile. The feedback profile can represent a first probability of receiving a first signal when the system is in a first state and has just operated at a first level of autonomy, and the system executes a first action at a second level of autonomy.
[0301] In one example, the instructions can include instructions for updating a human state transition function. The human state transition function can represent a second probability that when the system is selected to execute a second action in a first state and the human performs manual control, the human transitions to a second state of the environmental model.
[0302] In one example, the instructions can include instructions for updating an autonomy profile. The autonomy profile can define a set of acceptable levels of autonomy given a current state and a next action to be executed.
[0303] As used herein, the term "instructions" can include an instruction or expression for executing any method disclosed herein, or any part or plurality of parts thereof, and can be implemented in hardware, software, or any combination thereof. For example, the instructions can be implemented as information such as a computer program stored in a memory that can be executed by a processor to execute any of the respective methods, algorithms, aspects, or combinations thereof as described herein. An instruction or a part thereof can be implemented as a dedicated processor or circuit that can include dedicated hardware for executing any of the methods, algorithms, aspects, or combinations thereof as described herein. In some implementations, parts of the instructions can be distributed across multiple processors on a single device, multiple devices, which can communicate via a network such as a local area network, a wide area network, the Internet, or a combination thereof.
[0304] As used herein, the terms "example", "embodiment", "implementation", "aspect", "feature", or "element" serve as examples, instances, or illustrations. Unless explicitly indicated, any example, embodiment, implementation, aspect, feature, or element is independent of any other example, embodiment, implementation, aspect, feature, or element, and can be used in combination with any other example, embodiment, implementation, aspect, feature, or element.
[0305] As used herein, the terms "determine", "identify", or any variation thereof include selecting, ascertaining, calculating, retrieving, receiving, determining, establishing, obtaining, or otherwise identifying or determining in any way using one or more of the devices shown and described herein.
[0306] As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or" unless otherwise specified or clear from the context. Further, the articles "a" and "an" used in this application and the appended claims are generally to be construed to mean "one or more" unless otherwise specified or clear from the context that the singular form is intended.
[0307] Furthermore, for simplicity of explanation, the figures and descriptions herein may include a series or sequence of steps or stages, but the elements of the methods disclosed herein may be performed in various orders or simultaneously. Further, the elements of the methods disclosed herein may occur with other elements not explicitly presented and described herein. Further, not all elements of the methods described herein are required to practice the methods according to the disclosure. Although aspects, features, and elements are described herein in specific combinations, each aspect, feature, or element can be used independently or in various combinations regardless of the presence or absence of other aspects, features, and elements.
[0308] The above aspects, examples, and implementation forms are described to enable easy understanding of the present disclosure and are not limiting. On the contrary, the present disclosure includes various modifications and equivalent configurations included within the scope of the appended claims, and the scope should be given the broadest interpretation to encompass all such modifications and equivalent structures as permitted under the law.
Claims
1. A method of autonomous driving by an AV (Autonomous Vehicle), comprising: detecting an environmental state based on sensor data; selecting an action based on the environmental state; identifying a set of current indistinguishable states; identifying discriminators from the set of current indistinguishable states; training a feedback model for the discriminators; determining an autonomy level associated with the environmental state and the action, the autonomy level being selected based at least on an autonomy model and the feedback model; executing the action according to the autonomy level. A method as described above.
2. The method according to claim 1, wherein the autonomy level is selected from a set including a first autonomy level indicating "no autonomy", a second autonomy level indicating "autonomy with verification", a third autonomy level indicating "autonomy with supervision", and a fourth autonomy level indicating "autonomy without supervision".
3. The method according to claim 2, wherein the autonomy level is the second autonomy level indicating "autonomy with verification", and executing the action according to the autonomy level includes: receiving an approval feedback signal or a denial feedback signal for the action. Including the method according to claim 2.
4. The method according to claim 3, wherein executing the action according to the autonomy level includes: inquiring about the approval feedback signal before receiving the approval feedback signal. Including the method according to claim 3.
5. The method according to claim 3, wherein the autonomy level is the third autonomy level indicating "autonomy with supervision", and executing the action according to the autonomy level includes: determining that the AV is being monitored by a human before executing the action. Including the method according to claim 3.
6. The method according to claim 3, wherein executing the action according to the autonomy level and the approval feedback signal includes: determining that the AV is being monitored; responding to the determination that the AV is being monitored by executing the action; receiving an override signal; responding to receiving the override signal by stopping the action and switching the AV to a manual operation mode. Including the method according to claim 3.
7. Executing the action according to the autonomy level includes: Based on the autonomy level, determining whether to request approval from a human for the action before executing the action The method according to claim 1, comprising:
8. The autonomy level is "no autonomy" so that the AV is not permitted to execute autonomous actions, Executing the action according to the autonomy level, Enabling the AV to be manually controlled by a human including The method according to claim 1.
9. The autonomy model includes a utility model and an autonomy profile, The utility model describes the utility of executing a first action at a first autonomy level for a first environmental state when the AV has just operated at a second autonomy level, The autonomy profile maps each environmental state to each action and defines constraints on the permitted autonomy levels for specific environmental states, The method according to claim 1.
10. Updating at least one of an autonomy profile, a feedback profile, or a human transition function in response to executing the action The method according to claim 1, further comprising:
11. A system for autonomy, comprising: a memory; a processor, wherein the processor executes instructions stored in the memory to calculate a policy for solving a task by solving an extended probabilistic shortest path (SSP) problem, identify discriminators from a set of currently indistinguishable states, and train a feedback model for the discriminators configured to The policy maps environmental states and autonomy levels to actions and autonomy levels, Calculating the policy, Generating a plan that operates across multiple autonomy levels including a system
12. Generating a plan that operates across the multiple autonomy levels, Generating a plan subject to constraints on the permitted autonomy levels in each state The system according to claim 11, comprising:
13. The system according to claim 12, wherein the constraints map states and actions to a subset of autonomy levels.
14. The instructions are When the system is in the first state and has just operated at the first autonomy level, update the feedback model representing the first probability of receiving a first signal when the system executes a first action at the second autonomy level. The system according to claim 11, further comprising an instruction.
15. The instruction is Select the system to execute a second action in the first state, and update the human state transition function representing the second probability that the human transitions to the second state of the environmental model when the human performs manual control. The system according to claim 11, further comprising an instruction.
16. The instruction is The system according to claim 11, further comprising an instruction to update an autonomy profile, wherein the autonomy profile defines a set of acceptable autonomy levels given a current state and a next action to be executed.
17. A method for autonomous driving, comprising: Calculating a policy for solving a task by solving an extended probabilistic shortest path (SSP) problem; Identifying discriminators from the current set of indistinguishable states; Training a feedback model for the discriminators; wherein the policy maps environmental states and autonomy levels to actions and autonomy levels; calculating the policy includes generating a plan that operates across multiple levels of autonomy; including a method.
18. Generating a plan that operates across the multiple levels of autonomy includes generating a plan subject to constraints on the permitted levels of autonomy in each state. The method according to claim 17.
19. The method according to claim 18, wherein the constraints map states and actions to a subset of levels of autonomy.
20. When the agent is in the first state and has just operated at the first autonomy level, update the feedback model representing the first probability of receiving a first signal when the agent executes a first action at the second autonomy level; Select the agent to execute a second action in the third state, and update the human state transition function representing the second probability that the human transitions to the second state of the environmental model when the human performs manual control; further comprising updating the autonomy profile, wherein the autonomy profile defines a set of acceptable autonomy levels given the current state and the next action to be performed The method according to claim 17
Citation Information
Patent Citations
Automatic operation control device and automatic operation control method
JP2018030425A
Driving control device and HMI control device
JP2021113043A