Yield Scenario Encoding for Autonomous Systems
By generating waiting element data structures and using waiting element engines to determine the give way behavior, the problem that traditional autonomous vehicles cannot safely negotiate the give way protocol is solved, safe and predictable give way negotiation is achieved, and the safety and driving experience of autonomous vehicles are improved.
Patent Information
- Application Number
- CN202210592215.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-27
- Filing Date
- 2022-05-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Traditional autonomous vehicles cannot safely and predictably encode and execute the transfer agreement, resulting in the inability to safely and effectively negotiate the transfer in scenarios such as intersections and merged lanes, affecting the widespread deployment of autonomous vehicles and trucks.
By generating the waiting element data structure, the autonomous system's give way scenarios are encoded, including the self-path and the geometry and competition state of the competitor path, and the waiting element engine and the give way planner determine appropriate give way behavior to ensure that the autonomous system safely negotiates the give way scenarios.
The autonomous system is implemented to negotiate in a safe and predictable way in the given-travel scenario, and the allowable behavior is in line with traffic rules, improving the safety and predictability of autonomous vehicles, and reducing drivers' anxiety and tension.
Smart Images

Figure CN116030652B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application is related to U.S. Patent Application No. 17 / 395,318, filed on August 5, 2021, entitled "Behavior Planning for Autonomous Vehicles in Yield Scenarios", which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION
[0003] Advances in machine vision methods, neural network architectures, and computing substrates have begun to support autonomous vehicles - such as, but not limited to, land-based autonomous vehicles (e.g., autonomous cars and trucks) and robots. For public and government regulators to accept the widespread deployment of autonomous cars and trucks on the road, autonomous cars and trucks must achieve a safety level that exceeds that of current average human drivers. Safe and effective driving requires that all drivers trust that other vehicles in the area will yield appropriately when obligated. If a vehicle fails to yield, due to the "unpredictability" of other drivers, the drivers of other nearby vehicles may not be able to continue driving in a safe and effective manner, e.g., a driver who has been presented with a behavioral cue that a vehicle may fail to yield when obligated. Thus, a necessary condition for the deployment of autonomous cars and trucks includes the success of autonomous cars and trucks in "safe and polite" negotiation of yield scenarios (e.g., intersections and merge lanes).
[0004] Typically, local traffic laws and driving protocols in the area specify which vehicle operators (and under what conditions) are responsible or obligated to yield to others. Such regulations include traffic laws (e.g., a vehicle must yield to pedestrians at a crosswalk), situation-specific signs (e.g., street signs indicating which approaching roads at an intersection are responsible for yielding to other approaching roads), and other real-time cues (e.g., multiple cars arriving at a traffic circle almost simultaneously). However, traditional autonomous vehicles are unable to encode and deploy such protocols. Instead, traditional systems may be designed to avoid collisions while failing to consider yield protocols and thus unable to navigate yield scenarios safely and predictably. SUMMARY OF THE INVENTION
[0005] Embodiments of the present disclosure relate to encoding yield scenarios for autonomous systems (e.g., manned or unmanned vehicles or robots). Systems and methods are disclosed for providing real-time control of an autonomous system when the system encounters a yield scenario.
[0006] Compared with traditional systems such as those described above, the disclosed embodiments enable an autonomous system to negotiate a yield scenario in a safe and predictable manner. In at least one embodiment, in response to detecting a yield scenario, a data structure is generated that encodes the geometry of the ego path, the geometry of competitor paths including at least one conflict point with the ego path, and the conflict state associated with at least one conflict point. The geometry of the yield scenario context can also be encoded, such as the geometry defining the internal ground area of an intersection (e.g., as a polygon), entry or exit lines, etc. The data structure is passed to the yield planner of the autonomous system. The yield planner determines the yield behavior of the autonomous system based at least on the data structure. The control system of the autonomous system can operate the autonomous system according to the yield behavior such that the autonomous system safely negotiates the yield scenario.
[0007] In at least one embodiment, a yield scenario (e.g., an intersection or merge yield scenario) can be detected based at least on analyzing sensor data generated by at least one sensor of an autonomous vehicle. Map localization and / or perception can be used to determine various information associated with the yield scenario. For example, a first path of the autonomous vehicle and a second path of a competitor (e.g., another vehicle or other object) can be determined through the yield scenario. There may be at least one conflict point between the paths, which may indicate a potential collision if the paths are traversed. To determine the conflict state of at least one conflict point (defining how the vehicle should behave), the system can determine one or more traffic rules applicable to the yield scenario. A wait element data structure (also referred to as a wait element) can then encode the information used by the vehicle to navigate the yield scenario, such as the geometry of the paths, the conflict state, and other information. For example, the wait element can be provided to a control agent of the vehicle. The control agent can be enabled to use the wait element to determine the yield behavior of the first vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following describes in detail the system and method for encoding a yield scenario for an autonomous system of the present invention with reference to the accompanying drawings, where:
[0009] Figure 1 is an example of a yield scenario according to some embodiments of the present disclosure;
[0010] Figure 2 illustrates a non-limiting example of a wait element data structure and a non-limiting example of a conflict state data structure according to some embodiments of the present disclosure;
[0011] Figure 3 illustrates a non-limiting example of a wait element engine according to various embodiments of the present disclosure;
[0012] Figure 4is a flowchart showing a method for encoding a yield scenario for an autonomous vehicle (e.g., a self-vehicle) according to some embodiments of the present disclosure;
[0013] Figure 5 is a flowchart showing a method for encoding a yield scenario for an autonomous vehicle (e.g., a self-vehicle) according to some embodiments of the present disclosure;
[0014] Figure 6 is a flowchart showing a method 600 for resolving a competing state between vehicle paths according to some embodiments of the present disclosure;
[0015] Figure 7A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;
[0016] Figure 7B is according to some embodiments of the present disclosure Figure 7A an example of the camera positions and fields of view of an exemplary autonomous vehicle;
[0017] Figure 7C is according to some embodiments of the present disclosure Figure 7A a block diagram of an example system architecture of an example autonomous vehicle;
[0018] Figure 7D is a system diagram of the communication between a cloud-based server and Figure 7A an example autonomous vehicle according to some embodiments of the present disclosure;
[0019] Figure 8 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0020] Figure 9 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0021] Systems and methods related to encoding yield scenarios for autonomous vehicles are disclosed. Although the present disclosure may be directed to an example autonomous vehicle 700 (also referred to herein as "vehicle 700" or "self-vehicle 700"), examples of which are referenced Figures 7A - 7DFor example, the systems and methods described herein can be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more Adaptive Driver Assistance Systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles connected to one or more trailers, aircraft, vessels, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other vehicle types, but are not limited thereto. Additionally, while the present disclosure may be described with respect to controlling an autonomous vehicle to negotiate a yield scenario, this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technological space where an autonomous control system can be used.
[0022] During the normal operation of an autonomous vehicle, the control agent must avoid moving and non-moving obstacles (e.g., other vehicles, pedestrians, bicycles, lane obstacles, etc.). In addition to avoiding collisions, the agent also has the fundamental responsibility to yield to other road users in certain scenarios (e.g., “yield conditions”). Such yield conditions can exist at (controlled and uncontrolled) intersections, crosswalks, merge lanes, highway (or interstate) on-ramps / off-ramps, roundabouts, etc., such as navigating a parking structure and / or a parking lot. To allow other users to safely, confidently, and effectively “clear” the yield condition, the yield behavior may involve slowing down or even bringing the vehicle to a complete stop. For example, at an unmarked intersection where another vehicle has previously arrived, the subsequently arriving vehicle may deploy an appropriate yield behavior by slowing down to allow the first vehicle to safely pass through the intersection. This yield behavior ensures that the subsequent vehicle does not enter the intersection until the first vehicle has safely passed through the intersection. In such yield conditions, one or more users may have a clearly defined obligation (or responsibility) to yield to other users.
[0023] Yielding behavior not only provides the utility of avoiding collisions. Appropriate yielding behavior can ensure "polite and expected" driving dynamics, which are necessary for safe and efficient transportation. For example, even if an agent takes action to avoid a potential collision under yielding conditions (e.g., by accelerating through an intersection), failure to yield when obligated to do so can create tense and anxiety-provoking driving conditions for all users in the area. Even if a person accelerates to avoid a collision, a non-yielding vehicle can create anxiety, a sense of danger, and anger (e.g., road rage) among other drivers, bicyclists, and pedestrians. That is, even if a collision is avoided by taking aggressive action; the collision is not avoided in the "safe and polite manner" expected by other users. Thus, when operating an autonomous vehicle, the agent of the autonomous vehicle may be obligated (e.g., legally or normatively) to employ one or more behavioral yielding strategies when approaching a yielding scenario.
[0024] The present disclosure provides, in part, a "waiting element engine" for an autonomous vehicle ("ego vehicle") that can actively monitor the approach of one or more yielding scenarios (e.g., a vehicle is approaching an intersection, a vehicle is passing through an on / off ramp, or a vehicle is preparing to change lanes). The waiting element engine can generate one or more "waiting element" data structures that encode a representation of the yielding scenario. The waiting element can be provided as an input to the "yield planner" of the ego vehicle. Examples of yielding scenarios include intersections (e.g., crossroads) and lane merges (e.g., lane merging on an entrance / on-ramp). Each yielding scenario may be associated with at least two participants: the ego vehicle and at least one competitor (e.g., another vehicle, a pedestrian, a bicyclist, etc.). Competitors can include another vehicle (e.g., autonomous, semi-autonomous, and / or traditionally manually operated), as well as individuals (pedestrians and bicyclists).
[0025] In various embodiments, a yielding scenario can be associated with more than one competitor (e.g., the ego vehicle is approaching an intersection where there are multiple other vehicles, pedestrians, and / or bicyclists, the ego vehicle is merging into a lane with multiple other vehicles, etc.). Each participant associated with the yielding scenario may be associated with one or more "potential paths" or lanes. For a participant in a yielding scenario, a potential path or lane may include a given set of current or possible spatial paths, e.g., the current coordinates of the participant in a spatial velocity phase space. Thus, the potential path of a participant may depend not only on their current spatial and velocity coordinates, but also on the limitations of the vehicle (or individual or entity) with respect to acceleration, deceleration (e.g., braking force), and maneuverability (e.g., turning radius, traction control, etc.).
[0026] When a yielding condition is detected, a waiting element engine can receive and / or generate environmental data originating from various sources related to a yielding scenario (e.g., in-vehicle and / or out-of-vehicle sensors and / or detectors, perception-based data, map-based data, geographical location data, etc.). The waiting element engine can analyze and fuse the various data, as well as check, parse, and match the data against various yielding-related traffic rules based on the fused and analyzed data to generate various "waiting geometries" and "competing states", which can characterize the yielding scenario of the ego vehicle. For example, data encoding various aspects of the potential paths and scenario geometries of the vehicle can be parsed and matched against one or more yielding or traffic rules to determine the competing state of the yielding scenario (e.g., take way, stop at the entrance, yield from the entrance, etc.). The waiting geometries can be grouped into one or more "waiting groups", where a waiting group can refer to all the waiting elements of the yielding scenario. The waiting geometries and competing states can be encoded in a "waiting element" data structure. One or more waiting element data structures can be provided to the "yielding planner" of the autonomous vehicle for controlling the vehicle.
[0027] The yielding planner can receive the waiting element data structure and determine an appropriate yielding behavior. When the control agent of the ego vehicle adopts the determined yielding behavior (e.g., as defined by the competing state), the ego vehicle can safely meet its required and expected yielding obligations while avoiding collisions.
[0028] In at least one embodiment, the waiting element engine can receive and / or obtain various input data, which can include geometry-related data, signal-related data, and map-related data. Obtaining (or receiving) the data can be done using sensing, perception, and / or detection techniques, which can utilize geometric or visual perception, map perception (possibly including localization), and signal perception. In at least one embodiment, the perception data can include lane map data. The lane map data can include one or more paths that can be assigned as the potential paths of the ego vehicle (e.g., ego paths) and one or more paths that can be assigned as the potential paths of one or more competitors (e.g., competitor paths). Other input data may include various raw sensor data from the ego vehicle or competitors. The geometric input data can include various information about the environmental geometry, such as applied to the potential paths associated with the yielding scenario and / or background context. The signal input data can include and / or encode traffic signals, such as traffic lights, traffic signs, stop signs, yield signs, right-of-way signs, such as major road signs, speed signs, as well as gestures or other body postures used for ground traffic signaling.
[0029] In various embodiments, a waiting element data structure can be generated for each possible pairing of a self-vehicle potential path and a competitor potential path. In a non-limiting example of a yield scenario, the yield scenario is associated with one self-vehicle and j competitors, where j is a positive integer. The self-vehicle may be associated with i potential paths, and each of the j competitors is associated with k potential paths, where i and k are also positive integers. In such an example, the waiting element engine can generate i × j × k individual waiting elements. Thus, each waiting element can be associated with a self-vehicle potential path and a competitor potential path. The waiting element can encode the "waiting geometry" of the self-path, the waiting geometry of the competitor path, and the waiting geometry of the context of the two paths. The waiting element can further encode the "competition state" between the two paths.
[0030] In short, the waiting geometry of a declared path (e.g., a self-path or a competitor path) can include a set of field-value pairs or other data types or elements of the path for encoding various aspects of the declared path. Such fields for the waiting geometry of a path can include, but are not limited to, entry line, exit line, entry and exit competitor areas, intersection entry line and interior ground, competition points between the self-path and the competitor path (optional explicit encoding of one of the intersections or merge points between the paths), etc. The competition state of the waiting element (e.g., the competition state) can be or define an instruction to the yield planner on how the self-vehicle should yield or not yield with respect to the waiting element. Such states include, but are not limited to: do not yield, stop at entry, yield from entry, etc.
[0031] In at least one embodiment, in order to generate waiting elements, geometric input data, which can be referred to as waiting geometry data, can be "fused" with lane map data and map data. The "fused geometry" data can then be classified (e.g., as a left turn, right turn, U-turn, etc.) and associated with one or more paths. Signal data can be fused with map data. The signal state (e.g., green light, red light, inactive, etc.) can be determined from the fused signal data. The fused, classified, and associated geometric data, together with map data, signal state, and other data, can be fed as input to the "competition state resolver" of the waiting element engine. The condition state resolver can use geometric, signal, map, and other sensor data, as well as traffic rules, to determine the competition state and parse the data into waiting elements.
[0032] Reference Figure 1 , Figure 1 FIG. 13 shows an example of a yield scenario 100 in accordance with some embodiments of the present disclosure. Figure 1The non - restrictive yield scenario 100 is an example of a yield scenario at an intersection (or junction). Other types of yield scenarios include at least merge yield scenarios (e.g., at a highway on - ramp). In this non - restrictive example of the intersection yield scenario 100, three vehicles are approaching a four - way intersection. These three vehicles include a first vehicle 102 (e.g., the ego vehicle), a second vehicle 104 (e.g., a first competitor), and a third vehicle 106 (e.g., a second competitor). The wait element engine 130 is used to generate one or more wait element data structures (e.g., wait element_1 110 and wait element_2 120) for the yield scenario 100. In some embodiments, the wait element engine 130 can be on - board the ego vehicle 102. In other embodiments, the wait element engine 130 can be at least partially remote from the ego vehicle 102. In such embodiments, the ego vehicle 102 can access the wait element engine 130 via one or more communication networks.
[0033] At least in combination with Figure 3 The various embodiments of the wait element engine are discussed at least in combination with the wait element engine 300. 3. It should be understood that such and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, commands, function groupings, etc.) can be used in addition to or instead of those shown, and some elements can be omitted entirely. Moreover, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components and implemented in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, the various functions can be performed by a processor executing instructions stored in a memory.
[0034] Wait element data structures, such as but not limited to wait element_1 110 and wait element_2 120, can be associated with a pair of declared paths, where one of the declared paths is the declared path of ego vehicle 102 and the other of the pair of declared paths is the declared path of a competitor (e.g., first competitor 104 or second competitor 106). Each wait element can encode a wait geometry for the declared path of ego vehicle 102 (e.g., ego wait geometry 112 of wait element_1 110 or ego wait geometry 122 of wait element_2 120) and a wait geometry for the declared path of the competitor (e.g., competitor_1 wait geometry 114 of wait element_1 110 or competitor_1 wait geometry 124 of wait element_2 120). Thus, wait geometry_1 110 can be associated with a single declared path of ego vehicle 102 and a single declared path of first competitor 104 (in other examples, more paths can be associated with the wait element). Similarly, wait geometry_2 120 is associated with a single declared path of ego vehicle 102 and a single declared path of second competitor 106. Note that the declared path of ego vehicle 102 associated with wait element_1 110 can be (but does not have to be) the same declared path of ego vehicle 102 associated with wait element_2 120. Various embodiments of the wait element geometry are discussed at least in conjunction with Figure 2 are discussed.
[0035] In addition to the wait geometry for the pair of declared paths, each wait element can encode a wait geometry for context (e.g., wait geometry context 116 of wait element_1 110 and wait geometry context 126 of wait element_2 120)). Further, each wait element can encode the competition state for the pair of declared paths (e.g., competition state_1 118 of wait element_1 110 and competition state_2 of wait element_2 120). Various embodiments of the wait geometry context and competition state are discussed at least in conjunction with Figure 2 are discussed.
[0036] More generally, a waiting element may include (or encode) a waiting geometry of a self-path (e.g., a declared path of the self-vehicle) and a waiting geometry of a competitor path (e.g., a declared path of a competitor in a yield scenario 100), some subset of waiting geometry context and competition state. The self-waiting geometry 112, competitor_1 waiting geometry 114, waiting geometry context 116, and competition state_1 118 (of waiting element_1 110) may be data objects and / or data structures. Similarly, the self-waiting geometry 122, competitor_2 waiting geometry 124, waiting geometry context 126, and competition state_2 128 (of waiting element_2 120) may be data objects and / or data structures. In various embodiments, if the data values (or elements) of one or more of these data objects are not readily available (or not applicable to a given yield scenario), the data encoding of these missing elements may be set to "invalid" and / or "not applicable".
[0037] In one or more embodiments, waiting elements (e.g., waiting element_1 110 and waiting element_2 120) constitute the "atoms" of how information (or data) about a waiting condition (e.g., yield scenario 100) is encoded. The waiting elements may be provided as input to a yield planner ( Figure 1 not shown) of the self-vehicle 102. Based at least on the encoding of the waiting elements, the yield planner may determine an appropriate yield behavior to safely negotiate the yield scenario 100. As discussed herein, the waiting elements may be determined and / or generated by employing at least one of two methods (and / or a combination thereof). One method includes applying a set of yield-related traffic rules (e.g., yield heuristics) to the yield scenario 100. Another method includes applying mapping and real-time perception of geometry-related and / or signal state data to the yield scenario 100. As noted, in various embodiments, these two methods may be combined in various forms to employ map data, real-time perception of geometry / signal data, and yield heuristics.
[0038] Turning our attention to Figure 2 , Figure 2 FIG. shows a non-limiting example of a waiting element data structure 200 and a non-limiting example of a competition state data structure 210 (also referred to as competition state) according to some embodiments of the present disclosure. In various embodiments, the waiting geometry 200 and / or the competition state 210 may be data objects, or any other such structured data. Generally, the waiting geometry 200 may represent when additional information (e.g., related to a yield scenario (e.g., Figure 1The geometric and associated metadata resulting from applying information and / or data related to the yielding scenario 100) to a lane map (e.g., the ego vehicle declared path and / or the competitor declared path). That is, the wait geometry 200 can be applied to the ego path (e.g., the declared path of the ego vehicle), the competitor path (e.g., the declared path of the competitor), or the background context. For example, if the wait geometry 200 is applied to the ego path, the wait geometry 200 can be similar to Figure 1 the ego wait geometry 112 of the wait element_1 110 and / or the ego wait geometry 122 of the wait element_2 120. If the wait geometry 200 applies to the competitor path, the wait geometry 200 can be similar to the competitor_1 wait geometry 114 of the wait element_1 110 and / or the competitor_2 wait geometry 124 of the wait element_2 120. If the wait geometry 200 is applied to the background context, the wait geometry 200 can be similar to the wait geometry context 116 of the wait element_1 110 and / or the wait geometry context 126 of the wait element_2 120. Such context wait geometries can be related to the boundaries and / or the interior ground area of the intersection, or the presence of the intersection entry line.
[0039] In various non-limiting embodiments, the wait geometry can encode at least some of its data into field-value pairs. Thus, a set of fields (e.g., the wait geometry fields 202) can be associated with (or encoded in) the wait geometry 200. One or more values can be associated with each field of the wait geometry fields 202 to encode a set of field-value pairs. Note that the values can be data structures or data. In some embodiments, the value of a particular field can be another field, such that the wait geometry 200 can encode one or more data trees. As Figure 2As shown, such a field 202 may include, but is not limited to: an entry line (e.g., of a corresponding self-path or competitor path), an exit line (e.g., of a corresponding self-path or competitor path), an entry into a competitor area (e.g., of a corresponding self-path or competitor path), an exit from a competitor area (e.g., of a corresponding self-path or competitor path), and an intersection entry line (e.g., of a corresponding self-path or competitor path). In some embodiments, field 202 may include coordinates or other information defining the location of a boundary and / or an internal ground area (as part of the general context of a waiting group), as well as one or more competition points between the self-path and the competitor path (optional explicit coding of one of the intersections or merge points between the paths). Field 202 may additionally include a speed limit applied in the general context (which will be considered to apply between the entry line and the exit line). The value of each of these fields may be encoded as invalid to accommodate coding of waiting conditions where these fields do not apply to a particular yield scenario (e.g., a traffic light at an entrance ramp has only one self-path and one entry line, but no exit line, competitor path, or internal ground). Another example is the coding of a waiting group for a new speed limit that contains only one entry line and one speed limit in the entire context, with everything else set to invalid. The exit line will be interpreted as infinite or until further notice, and the same applies to other attributes.
[0040] The entry line of the self-path may encode the stopping points of several yield behaviors. The entry line may also represent the start of a general competition area, ended by the exit line, which may indicate which section of the self-path needs to be cleared to clear the waiting conditions of the waiting group. The internal ground area may represent the internal ground of the intersection as a polygonal area that can be referenced in a coordinate space. The internal ground area may cover the portion between the entry line and the corresponding exit line (sometimes the exit line may extend out, e.g., beyond a crosswalk, even if the internal ground does not). Entering the competitor area and the internal ground provide context for analyzing other participants. This may be performed by assigning participants to paths and areas (in a non-exclusive manner). The yield planner may use the geometry of the self-path and the competitor path, as well as the competition points, to implement yield behaviors as needed. The geometry may also be used to determine which rules to apply.
[0041] In some examples, a competition point may indicate or represent an explicit geometric point. In other examples, a competition point may refer to an abstract concept that is a particular competition for which a waiting element refers and / or is encoding its state. In such examples, the competition state at the competition point may be the payload of the entire competition state resolution process. For each competition point, it may provide a determination of how the ego vehicle should yield or not yield with respect to that competition point. In this sense, the competition point may also indicate the choice of the ego path, access to competitor paths, and via those competitor paths, the actual competitors, and the manner in which the ego vehicle behaves with respect to them.
[0042] The waiting geometries can be collected (e.g., logically organized) into groups, where the semantic meaning of a waiting group can be that all the waiting conditions in the group can be considered together and in particular cleared together, such that the ego vehicle is not in the middle (e.g., when the ego vehicle is still on the oncoming traffic path, the ego vehicle should not get stuck waiting for a pedestrian at the end of a left turn, so the oncoming traffic competition can be considered together with the crosswalk competition in the same waiting group).
[0043] Referring to the competition state 210, a waiting element engine, such as but not limited to Figure 3 the waiting element engine 300, may include a competition state resolver (e.g., the competition state resolver 340 of the waiting element engine 300). Such a competition state resolver may perform a competition state resolution process. The goal of the competition state resolution process may be to provide a competition state (e.g., the competition state 210) for each waiting element (e.g., Figure 1 the waiting element_1 110 and the waiting element_2 120). The competition state 210 of a waiting element may be an instruction to the yield planner regarding how the ego vehicle should yield or not yield with respect to this waiting element, as a matter of rules, expectations, formal or informal conventions or norms.
[0044] In some non - limiting embodiments, the competing state 210 may not indicate what actually happens in a yield scenario, what is physically possible in a yield scenario, or whether the ego - vehicle might be forced to yield even though it has the right - of - way in a yield scenario (e.g., an intersection yield scenario or a merge yield scenario). Instead, the competing state 210 can indicate what should happen according to convention. Then, it may be the responsibility of the yield planner to actually enforce yielding in the sense of considering what should happen (e.g., as encoded in the competing state 210), whether the ego - vehicle is actually in a position to stop and follow that instruction, and whether other participants (e.g., competitors in the yield scenario) appear to be fulfilling their expected yielding obligations and taking appropriate actions. For example, the yield planner may determine that even though the competing state is "TakeWay" (no - yield), the competitor has not yielded (essentially detecting "honk appropriately") and decide to yield even though that is not what should happen. The yield planner can implement yielding behavior by analyzing all competing in a wait - group until all waiting elements in the group can be jointly cleared. All competing in the wait - group can be jointly adhered to, meaning that the most restrictive competing can define the expected yielding behavior of the ego - vehicle. For example, if one competing state of the wait - group is "no - yield" and another is "Stop at Entry", the ego - vehicle may stay at the entry line.
[0045] As Figure 2As shown, the competition state 210 (e.g., encoded in the waiting element) can include one or more of the seventeen states listed in the competition state 210. Note that this list of possible competition states is non-exhaustive, and in other embodiments, the competition state 210 can include additional and / or alternative competition states. The do not yield state may indicate that it is expected that the competitor yields. Thus, the do not yield state can indicate that there is no formal constraint on the corresponding waiting element (except for the yield planner to observe the competitors associated with the waiting element and ensure that they yield as expected). For various competition states, the keyword Transient can be used to indicate that the corresponding competition state can be updated and / or evolved in the near future. Thus, the do not yield transient state can indicate that the do not yield state currently applies, but may soon change to a more restrictive state. A typical example is the "yellow" state of a traffic light, which may be encoded by the Take Way Transient. Stop at Entry may indicate that the instruction is to stop at the entry line, wait for further instructions, and not proceed until the competition state changes. Yield from Entry state can indicate that the ego vehicle should remain at the entry line until the time when it is expected that the competition is cleared. For such states of Stop at Entry, the rule may not enforce a pre-stop, but the control agent of the ego vehicle should ensure that the competition is cleared before the ego vehicle passes the entry line, which typically results in a pre-stop. Also note that since the waiting conditions in the waiting group can be considered jointly, this generally means that in practice, the control agent should ensure that all competitions in the waiting group are cleared before the ego vehicle passes the entry line. In other words, if one competition in the waiting group has Yield from Entry, all other competitions in the waiting group may inherit the same competition state when analyzed by the yield planner, and if one has Stop at Entry, all waiting elements in the waiting group can inherit the pre-stop. The Yield From Entry transient state can be a transient version of the Yield from Entry state.
[0046] The Yield Contention Point state can indicate that the rules may not enforce an early stop, or that the ego vehicle may not have to officially stop at the entry line while waiting for the contention to clear (although there is nothing wrong with doing so in principle). The control agent may have to ensure that the ego vehicle yields correctly to the competitors associated with this contention, that the ego vehicle does not block the contention, and that the ego vehicle behaves in such a way that the competitors are clear about the contention to which the ego vehicle is yielding. This may mean that the ego vehicle turns left forward at the intersection, but slowly enough and with enough margin so that oncoming traffic understands that the ego vehicle seems to intend to yield and clearly does not impede oncoming traffic. The Yield Contention Point Transient state can be a transient version of the Yield Contention Point state. The Stop at Entry then Yield from Entry state may be equivalent to (or at least similar to) the Yield from Entry state, but with the additional condition that an early stop is required at the entry line. The Stop at Entry then Yield Contention Point state may be equivalent to (or at least similar to) the Yield Contention Point state, but with the additional condition that an early stop is required at the entry line.
[0047] The Stop at Entry then Yield ContentionPoint Transient state can be a transient version of the Stop at Entry then YieldContention Point state. The Stopped First hasPrecedence state may be a typical "American multi-stop" situation. The right of way can be determined as a first-in-first-out queue, where "entry" is defined as approaching the intersection as the first participant from that contention path (possibly within the corresponding competitor area on the entry line pointing towards the inner ground) and stopping. In other words, this contention state may mean further processing of "who stopped first" to actually resolve into a non-yield or yield from entry state for each participant associated with that competitor path.
[0048] The "Negotiate" state can indicate that there is no known basis for determining the right of way, such as for a highway merge (equally sized highways merging with a similar straight shape) without hints from traffic rules, map statistics, geometry, or road size. The "Stop at Entry then Negotiate" state may be equivalent to (or at least similar to) the "Negotiate" state, but with the additional condition that a pre-stop is required at the entry line. This state can be adopted when there is no agreement but there is a clear entry line. The "Not Allowed" state can be used to enable the encoding of something not being allowed. For example, a waiting element may include a left turn path through traffic to enter a parking lot, and there may be a signal indicating a prohibited turn (e.g., crossing a double yellow solid line). In this case, the ego vehicle may signal (to a competitor) to proceed in such a way that not only is an immediate stop required, but it will never change and is simply not allowed. This state may be useful for a yield planner as it is considering multiple options for the ego path (e.g., multiple ego paths can be considered simultaneously). The "Stop and Request Takeover" state can indicate that the ego vehicle has encountered something determined to be outside the operational design domain (e.g., a signal or marking indicating road construction may have been detected and the ego vehicle has not yet implemented handling such a condition). In this state, the control agent (or yield planner) may request that the ego vehicle decelerate, stop, and request takeover behavior. The "Unknown" state can be used to encode predictions of future competitive states, and in this case, it may be useful to be able to encode situations where there is no knowledge or prediction.
[0049] Figure 3 FIG. 4 shows a non-limiting example of a waiting element engine 300 according to some embodiments of the present disclosure. As described throughout, the waiting element engine 300 may be onboard an autonomous vehicle (e.g., Figure 1 the ego vehicle 102). In other embodiments, the autonomous vehicle may access a remote waiting element engine via one or more communication networks. As discussed throughout, when the ego vehicle approaches an intersection or a merge yield scenario (e.g., Figure 1 the intersection yield scenario 100), the waiting element engine 300 may generate one or more waiting elements (e.g., Figure 1 the waiting element_1 110, Figure 1 the waiting element_2 120, and Figure 3 the waiting element 3 130) as output. The various inputs to the waiting element engine 300 are discussed below.
[0050] The waiting element engine 300 may include a waiting geometry sensor 302, a mapper 304, a signal sensor 306, a lane illustrator 350, and / or another sensor data receiver 308. The waiting geometry sensor 302, the mapper 204, the signal sensor 306, the lane illustrator 250, and the other sensor data receiver 308 receive various inputs as described below. The waiting element engine 300 may also include a waiting geometry fuser 322, a geometry classifier 324, and a geometry correlator 326. The waiting element engine 300 may also include a signal fuser 328, a signal state estimator 330, and a competition state resolver 340. The competition state resolver 340 may include a condition checker 342, a base rule parser 344, a mapping rule checker 346, and a waiting element fuser 348. The output waiting elements 310 may include self-waiting geometry 312, competitor waiting geometry 314, context waiting geometry 316, and a competition state 318.
[0051] The lane illustrator 350 is generally responsible for receiving one or more lane maps as inputs to the waiting element engine 300. The lane maps may be received in response to approaching and / or detecting a yield scenario. The lane maps may include a set of declared paths from the same lane bundle and include a set of self-paths 352 and a set of competitor paths 354. The self-paths 352 may be received or generated from the output of one or more neural networks or other machine learning models, maps, and / or vehicle trajectories. For example but not limited to, the machine learning models described herein may include any type of machine learning model, such as one or more machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid machines, etc.), and / or other types of machine learning models.
[0052] The ego path 352 can be fused in a set to form a new lane map. The lane map can be structured as an array of lane maps, allowing flexibility in easily using any combination of lane maps. One or more competitor paths 354 can overlap with one or more ego paths 352 at one or more competition points. A particular competition point can be classified as one of two main competition point types: an intersection competition point or a merge competition point. For intersection competition, by way of example and not limitation, the ego path can intersect the competitor path at a single point (or a small neighborhood of points). For a merge competition point, the ego path may encounter the competitor path and join the competitor path within at least a segment of the competitor path (or vice versa). The competitor paths 354 can be received or generated from the output of one or more neural networks or other machine learning models, maps, or intersection parsing. While the ego path 352 is the declared path of the ego vehicle, the competitor paths are the declared paths from vehicles, bicycles, pedestrians, trams, and trains, or any other participants in a yield scenario. Similar to the ego path 352, the competitor paths 354 can be merged from their sources (which can be similar or identical to the source of the ego path 352) to form a whole.
[0053] The wait geometry sensor 302 and the signal sensor 306 can generally perform "wait sensing". The wait geometry sensor 302 receives various geometry-related inputs (in response to identifying and / or detecting the ego vehicle approaching a yield scenario) and generates wait geometry for output (e.g., Figure 2 the wait geometry 200). Thus, the wait geometry sensor 302 can generate wait geometry (e.g., ego wait geometry 312) for each ego path in the ego path set 352. Similarly, the wait geometry sensor 302 can generate wait geometry (e.g., competitor wait geometry 314) for each competitor path in the competitor path set 354. Similarly, the signal sensor 306 receives various signal-related inputs (in response to identifying and / or detecting the ego vehicle approaching a yield scenario) and generates signals for output.
[0054] The wait geometry (e.g., ego wait geometry 312 or competitor wait geometry 314) can include the geometry and metadata generated when additional information about the wait condition is applied to the lane map. The wait geometry can be applied to the ego path (e.g., the entry line), the competitor path (e.g., the competitor region), or the background context (e.g., the interior ground of the intersection, or the presence of the intersection entry line). As at least in combination with Figure 2As discussed for the waiting geometry 200, specific waiting geometries may include the following field-value pairs: entry and exit lines (for self and competitor paths), entry and exit competitor regions (for self and competitor paths), intersection entry lines and internal ground (as part of the general context of a waiting group), and competition points between self and competitor paths (optional explicit encoding of one of the intersections or merge points between paths). The waiting geometry may include field-value pairs for speed limits applied in the general context (deemed applicable between the entry and exit lines). The value of each of these fields may be encoded as invalid to accommodate encoded waiting conditions for situations where the field does not apply to a yielding scenario (e.g., an on-ramp traffic light has only one self path and one entry line, but no exit line, competitor path, or internal ground). Another example may be a waiting group encoding a new speed limit that includes only one entry line and one speed limit in the overall context, and all else is set to invalid. The exit line may then be interpreted as infinite or until further notice, and the same applies to other attributes.
[0055] The value of the entry line field for the self path may encode the declared path stop point of the self vehicle for various possible yielding behaviors. Such an entry line may also encode the start of a general competition region, ending with the exit line, which may indicate which segment of the self path needs to be cleared for the waiting group to clear the waiting condition. The internal ground region may represent the internal ground of the intersection as a polygon region. The internal ground region may cover the path segment between the entry and exit lines (sometimes the exit line is moved out, e.g., beyond a crosswalk, even if the internal ground is not). The entry competitor region and the internal ground can provide context for analyzing other participants. This can be achieved through obstacles in the path analysis (OIPA) that assign participants to paths and regions (in a non-mutually exclusive manner). The yielding planner may use the geometry of the self and competitor paths and the competition points to implement yielding as needed. The waiting geometry can also be used to determine which rules apply.
[0056] The competition point may represent one or more explicit geometric points (e.g., the intersection of the self and competitor paths). In some embodiments, the competition point may be considered an abstract concept that is the specific competition to which the waiting element refers and for which its state is being encoded. In the latter sense, the competition state at the competition point may be the payload of a competition state resolution process (as performed by the competition state resolver 340). For each competition point, the competition state may provide a determination of "in what way the self vehicle should yield or not yield with respect to this competition point". In this sense, the competition point can also indicate how the actual competitors and the self vehicle should advance relative to them, given the self path choice, by accessing the competitor path and passing through the competitor path.
[0057] The waits can be grouped into one or more wait groups, where the semantic meaning of a wait group is that all wait conditions in the group can be considered jointly, and in particular can be cleared jointly, so that the ego vehicle does not remain in the middle (e.g., at the end of a left turn, the ego vehicle is not stuck waiting for a pedestrian while the ego vehicle is still in the path of oncoming traffic, so the crosswalk competition in the same wait group needs to be considered jointly with the oncoming traffic competition).
[0058] The mapper 304 can receive map data in response to identifying and / or detecting an ego vehicle approaching a yield scenario (e.g., at least based on map positioning). The map data can include one or more 2D or 3D maps of the environment of the upcoming yield scenario. The signal sensor 306 receives signal-related data in response to identifying and / or detecting an ego vehicle approaching a yield scenario. The signal-related data can include signal data encoding indications of things such as traffic lights and traffic signs, e.g., stop signs, yield signs, right-of-way signs such as major road signs, and speed signs. The signal data can also encode outputs generated from one or more neural networks or other machine learning models, e.g., whether the intersection is a traffic sign intersection, a stop sign intersection, unmarked, a roundabout, a highway on-ramp, a toll booth, or other types. The signal data can encode indications of police, flagmen or road workers directing traffic, barriers blocking the road, and street lights at crosswalks. The signal data can encode indications of traffic cones, merge arrows, and all temporary items placed on the road to redirect traffic. In various embodiments, the signal sensor 306 can separate the presence of a signal from its state. The signal sensor 306 can perform real-time signal sensing, which provides the presence and state of the signal (e.g., there is a traffic light at these 2D or 3D coordinates and its current state is yellow). In some embodiments, the map (e.g., the map received by the mapper 304) can include information about the possible presence and / or location of signals, e.g., the presence of a traffic light or a stop sign, while the state of the signal (such as the state of a traffic signal) can be provided by on-site sensing (or infrastructure-to-vehicle communication) by the signal sensor 306. Signal detection (by the signal sensor 306) can provide the presence, type, and spatial attributes of the signal, e.g., the 3D position and bounding box of a traffic signal or sign or a police officer detecting traffic.
[0059] The Wait Geometry Fuser 322 can "fuse" or combine the map data from the Mapper 304 with the geometry data (e.g., wait elements) from the Wait Geometry Sensor 302 through a process called wait geometry fusion. Wait geometry fusion can optionally be performed to obtain improved geometry information from the combination of real-time perception (e.g., real-time geometry perception performed by the Wait Geometry Sensor 302) and map information (received and / or provided by the Mapper 304). For example, the Wait Geometry Sensor 302 can detect the presence of an intersection through live (e.g., real-time) perception. The map can also annotate intersections in the map data. In some embodiments, if an intersection exists in either source, the map and geometry data can be fused to instantiate the intersection, and if they are from both sources, the link detection. Similarly, the entry line can be detected in real time (by the Wait Geometry Sensor 302), and / or provided in the map (by the Mapper 304) based at least on previous map streams that contain live detections or actual stop points. Competitive areas, internal ground, path geometry, and competitive points are all entities that can be detected live and can also be injected into the map stream to facilitate future driving. The actual driving path ("de facto lane map") can also be mined from multiple drives. Thus, wait geometry fusion can coordinate and correlate multiple sources to provide a clear wait geometry for further processing stages.
[0060] The Signal Fuser 328 can "fuse" or combine the map data from the Mapper 304 with the signal data from the Signal Sensor 306 through a process called signal fusion. Signal fusion can provide an option to coordinate the information from live perception (real-time perception by the Signal Sensor 306) with the information from the map (received and / or provided by the Mapper 304), such as using the confirmed presence of traffic lights, signs, or intersection types to assist live perception. Similar to geometry fusion, signal fusion can be optional in some embodiments. This signal fusion provides the possibility of performing state estimation on traffic lights that are difficult to detect, and / or also provides the results from real-time signal perception.
[0061] The Signal State Estimator 330 is generally responsible for the process called signal state estimation. Signal state estimation can determine and / or provide the state of traffic lights, the gestures of police officers directing traffic, the "stop" or "slow" signs of flagmen, the state of streetlights at crosswalks, or the position of barriers blocking traffic. The results are typically selected from an enumeration class (e.g., green / yellow / red) or several combinations.
[0062] The geometric classifier 324 is generally responsible for the geometric classification process. The geometric classification will wait for the geometry to be classified into discrete classes. For example, the geometric classifier 324 can classify the ego path (e.g., encoded in the ego wait geometry 312) and the competitor path (e.g., encoded in the competitor wait geometry 314) into classes such as but not limited to: "left turn", "go straight", "right turn", "U-turn", etc. This classification can be performed to standardize the paths so that general language rules such as "turning right on a red light is not allowed in Manhattan" can be applied to the paths during the competitive state resolution process performed by the competitive state resolver 340. To apply such rules, the paths can be classified into classes that include general language (e.g., "right turn"). Note that this classification can benefit from the context of other paths (e.g., a relatively straight shape may be a right turn if it is the rightmost path, while it may not be if there is also a very sharp right turn). The geometric classification can also determine the waiting elements if the ego or the competitor is from the right.
[0063] The geometric classification can also be applied to path pairs. For example, the geometric classifier 324 can determine whether two paths (e.g., the ego path and the competitor path) are crossing or merging (if not explicitly given in the lane map), where the competitive point is, and which paths are from the right (to support the right-of-way rule that is generally applicable in Europe and in some cases in the United States). This can be performed by checking whether the direction vectors of the two paths at the competitive point are significantly different from parallel lines, and if so, the sign of the 2D vector cross product between them (applied to the sign of the determinant of the 2x2 matrix formed by stacking the ego path direction vector as the top row and the competitor path direction vector as the bottom row). If the sign is positive, the competitor is from the right. Note that this definition may mean that for an ego left turn through oncoming traffic, the oncoming traffic is considered to be from the right (since this is the case at the competitive point). If the direction vectors are almost parallel at the competitive point (usually because it is a merge), the vectors at the corresponding entry lines can be used).
[0064] The geometric classifier 324 can also determine or merge from a live perception source (e.g., the wait geometry sensor 302) whether the competitor path is from a "short road" (e.g., a lane, a gas station, or a parking lot), or is "significantly larger" or "significantly smaller" than a "short road". That is, the roads can be classified through the geometric classification process (e.g., classified as a "short road"). The geometric classifier 324 can provide class predicate classifications that allow logical rules to be applied to a coherent set of input variables. Some of this information may come directly from the output of one or more neural networks or other machine learning models rather than through geometric determination of the wait geometry data structure.
[0065] The geometric classifier 324 can also determine whether a path has attributes such as, but not limited to: "crossing a line", which can be a definitive determination of whether the path crosses a line (e.g., the path can be classified as "crossing a line"). This classification can be used as a hint for certain rules when the priority is unclear how to arbitrate. For example, if two paths are in competition and are otherwise equivalent, but one crosses a line and the other does not, the path that does not cross the line may have priority. A path that turns left through oncoming traffic can be classified with an attribute of the type "crossing a dashed line", "crossing a solid line", or "crossing a double solid yellow line" so that country / region-specific rules can be applied to determine whether this is allowed. The geometric classifier 324 can also set variables such as, but not limited to, "crossing" to true or false for a waiting group (and thus for each waiting element). This can also be used to hint at some rules (e.g., to distinguish whether a crosswalk is adjacent to an intersection and how to handle it).
[0066] The geometric correlator 326 can associate signals with paths through a process called geometric correlation. Geometric correlation determines which signals apply to a path. This may answer questions such as "Is this light close enough to this path to apply to it?", "Is this light the closest / most relevant of this type to this path?", "Is this sign intended to apply to this path?". For the association of a light to a path, it may not be easy to separate from the rules because, for example, it is difficult to know if a light applies to a left turn and then it lights up a green arrow to resolve the ambiguity. Also note that this analysis may generally benefit from all signals and paths being considered together. For example, in the absence of other paths, a light offset to the right may apply to the ego path, but not in another scenario where there is a path further to the right. Similarly, in the absence of other lights, a light offset to the right may apply to the ego path, but not in another scenario where there is a light directly above the path. Thus, geometric correlation can consider the entire scene as well as the signal states (and even the raw sensor data) as needed. For the same reason, the architecture can allow the geometric classification performed by the geometric classifier 324 and the geometric correlation performed by the geometric correlator 326 to run jointly and have access to the waiting geometry, signals, and even the raw sensor data. In this sense, the process may assign links between signals such as traffic lights, signs, and paths. In at least one embodiment, geometric correlation can first check if a light or sign is at a distance that allows it to be linked to a path, and then whether it has indeed found the closest (in a sense) applicable light or sign of each type, determining the priority order (e.g., a left turn light has the highest priority for a left turn, but the closest regular light also applies, although it has a second priority). Note that if the light changes state and resolves some form of ambiguity, the association may change immediately.
[0067] The outputs of the waiting geometry fuser 322, mapper 304, signal state estimator 330, geometry classifier 324, geometry correlator 326, and other sensor data receivers 308 can be fed (as inputs) into the competing state resolver 340. The competing state resolver 340 can perform a process referred to as the competing state resolution process. The goal of the competing state resolution process can be to provide a competing state (e.g., competing state 318) to the waiting element 310 (and / or other waiting elements). As a matter of rule, expectation, formal or informal convention or specification, the competing state 318 of the waiting element 310 can be an instruction to the yield planner regarding how the ego vehicle should yield or not yield relative to that waiting element. In some non-limiting embodiments, the competing state 318 may not indicate what actually happens in a yield scenario, what is physically possible in a yield scenario, or whether the ego vehicle may be forced to yield even though it has the right of way in a yield scenario (e.g., a crosswalk yield scenario or a merge yield scenario). Instead, the competing state 318 can indicate what should happen according to convention. The actual implementation of yielding may be the responsibility of the yield planner, in the sense that it will consider what should happen (e.g., as encoded in the competing state 318), whether the ego vehicle is actually in a position to stop and follow that instruction, and whether other participants (e.g., competitors in a yield scenario) appear to be fulfilling their expected yielding obligations and take appropriate action. In other words, the yield planner can determine that even though the competing state is not to yield, the competitor is not yielding (essentially detecting a "fit to honk") and decide to yield even though that is not what should happen. The yield planner can implement the yielding behavior, analyzing all the competing in the waiting group until all the waiting elements in the group can be jointly cleared. All the competing in the waiting group can be jointly adhered to, meaning that the most restrictive competing can define the expected yielding behavior of the ego vehicle. For example, if one competing state of the waiting group is not to yield and another is to stop at the entrance, the ego vehicle may stay at the entrance line. The competing state 318 can be similar to Figure 2 the competing state 210. The various possible state values of the competing state 318 are discussed in conjunction with the competing state 210.
[0068] The contention state parsing process of the contention state parser 340 can be based at least on different basic "road rules" (or "ground rules") that can vary from country to country, state to state, region to region, etc. The ground rules can be basic logical rules applied to the waiting geometry (e.g., self-waiting geometry 312 and competitor-waiting geometry 314) and signals after they have been reduced to basic enumerated variable states through a geometric classification process (e.g., by geometric classifier 324) and a geometric association process (e.g., by geometric associator 326). After these reductions, each waiting element may have a well-defined set of geometric classes and signal states applicable to the waiting geometry. In addition to the ground rules, the contention state parsing process can also use "mapping rules". The mapping rules can include an array of waiting elements (e.g., waiting element 310) and proposition pairs linked to contention states.
[0069] Each proposition can be conditioned on any number of signal states (and is also allowed to be conditioned on no signal states). If a proposition evaluates to true, the contention state and / or its proposition can be paired with the indicated contention state. The semantics of the mapping rules may be that each proposition is evaluated in order, and the first proposition that evaluates to true defines the contention state. The state already included in the waiting element can be considered the default contention state, which is selected if none of the propositions evaluate to true. The condition checker 342 of the contention state parser 340 is typically responsible for performing such evaluations and linking the propositions.
[0070] In at least one embodiment, the condition checker 342 can typically determine which signals are valid or invalid (or active or inactive) based at least on one or more sensed and / or determined environmental conditions. For example, depending on weather conditions (e.g., rain, snow, fog, wind), time of day, day of the week, the presence of other signs (e.g., a road construction sign may supersede other signs), the vehicle type of the autonomous vehicle (e.g., car vs. truck), etc., certain signs may or may not apply. For example, if there is a mapping rule with conditions that only apply from time X to time Y, the condition checker 342 can be used to determine and mark whether the mapping rule currently applies.
[0071] The mapping rule matcher 346 is generally responsible for performing the mapping rule matching process. The mapping rule matching process can take mapping rules and match their components (or data components) to items determined to actually exist, and generate waiting elements in the process. For example, the mapping rule matcher 346 can take the signal and other inputs determined to be applied by the condition checker 342 and parse them into one or more mapping rules. For example, the mapping rule matcher 346 can determine the application of one or more mapping rules based on the associated traffic light signal being green, while the mapping rule matcher 346 may not determine whether the mapping rule applies when the traffic light signal is red (but may determine that a different rule applies). These determinations can be based at least on the conditions (such as traffic light status) that the mapping rule matcher 346 knows are actually applied according to the determinations of the condition checker 342.
[0072] To increase the ability to benefit from mutual exclusion constraints and generally make overall decisions, the mapping rule matching process performed by the mapping rule matcher 346 can start with the process of establishing correspondences between paths in the map and paths, waiting geometries, and signals with the waiting geometries and signals determined to actually exist by the condition checker 342. Note that in some cases, these entities may first come from the map (e.g., the waiting element engine 300 can be configured to receive a lane map from the map and consider mapping rules with one of the same paths and whether it matches), so it can perform 'by-id' matching that has been established during lane map fusion, waiting geometry fusion, or signal fusion. However, to increase flexibility and generality, the mapping rule matcher 346 can perform the matching without using information from the map. For example, the mapping rule matcher 346 can use a live perception lane map (e.g., via the waiting geometry sensor 302) and apply mapping rules from the map that associate traffic lights with the ego path, and avoid architectural complexities caused by having to propagate map identification symbols through lane maps, waiting geometries, and signals all the way. If the path, waiting geometry, or signal actually comes from the map, its geometry should be the same (almost the same if successfully fused / blended), so the matching should be correctly restored. The matching can also match entities that are not exactly the same. For example, it may be expected that roughly similar left-turn shapes and positions will match (again note that if there are two parallel left-turns, considering them jointly will help with the matching). Thus, the process can essentially perform a matching that corresponds the'map scene' with the 'actual scene', which has been determined by any combination of live perception and map localization. The result may be a one-to-one correspondence between a subset of the actual entities and a subset of the map entities.
[0073] In various embodiments, each mapping rule may produce an output wait element (e.g., wait element 310 or its precursor) by matching all of its entities and resolving all of its signals. Many mapping rules may contain a valid self-path (since many competing states are conditionally implemented thereon and would not make sense without it). If the self-path does not find a match, there may be several reasons. If the localization fails in a known way, the mapping rule may not be used and the localization failure may be handled differently. However, if the localization is inaccurate, it may result in the self-path not matching. Another possibility is that the self-path is inaccurate in the map or the actual scenario. Another possibility is that the path is too far from the live perception or is occluded. In such cases, a conservative approach may force the self-path into the scenario. For this reason, the self-paths from the mapping rules for which no match was found, along with their wait geometries, may be added. The same process may be applied to competitor paths. In fact, in addition to signals, the entire set of wait elements from the mapping rules may be considered individually, although the correspondence may be important when considering the fusion of the mapping rules with the base rules performed by the wait element fuser 348. On the other hand, the signals may have to match in order to resolve their states. Any signal state that does not match may be set to Unknown, and the propositions in the mapping rule may account for this possibility and assign an appropriate competing state. For example, this typically requires setting the competing state to Stop at Entry when the state of a single traffic light is unknown. This may also require using one of several synchronized traffic lights to resolve the same wait element, defaulting to Stop at Entry only if they are all unknown. In other cases, when the green turn arrow traffic light is not visible but the green circle is visible and it is known that turning is always allowed in such cases, this may require returning a Yield Contention Point to turn left through oncoming traffic (although it is not known whether it is protected). This example is aggressive, but the design provides a high degree of flexibility rather than a high degree of complexity (another less flexible option is to list the possible states and enumerate all possible combinations and find the most restrictive competing state among all possibilities).
[0074] The base rule parser 344 is generally responsible for performing a base rule parsing process for propositional rules that can be derived in two ways. First, propositional rules can be derived from base rules through basic state variable estimation, geometric classification, and geometric association. Second, propositional rules can be derived from mapping rules. For example, the base rule parser 344 can use the condition checker 342 to determine the signals and other inputs to be applied and parse them into one or more base rules. In at least one embodiment, the base rule parser 344 can operate similarly to the mapping rule matcher 346, but applies general or common rules and conditions during driving that are independent of the history or observed behavior of the vehicle at the yield scenario location.
[0075] Although the base rules may vary from country to country, state to state, region to region, etc., they can be applied consistently between yield scenarios as long as the corresponding conditions are met. In contrast, mapping rules can apply driving rules and conditions based at least on the history or observed behavior of the vehicle at the yield scenario location or a similar yield scenario location. In at least one embodiment, the mapping rules can be encoded into the map data and applied based at least on positioning the autonomous vehicle on the map. However, the base rules can be applied regardless of the positioning and the yield scenario location. By providing the base rule parser 344, wait elements can be generated even when the map data is unavailable or cannot be applied or positioned to the current yield scenario. For example, in the case where the mapping rule matcher 346 cannot determine one or more components and / or elements of the wait element, the base rule parser 344 can fill in any gaps, and vice versa. Thus, the wait element can be generated entirely from the map data, entirely from the perception data, or from a combination of both types of data.
[0076] The Wait Element Fuser 348 is generally responsible for fusing or combining data corresponding to the parsed base rules and the matched mapping rules to resolve the competing state 318. For example, in at least one embodiment, the Base Rule Parser 344 and the Mapping Rule Matcher 346 can each generate corresponding wait elements and / or elements and / or their components. The Wait Element Fuser 348 can fuse any of these different aspects to form the wait element 310. In at least one embodiment, one or more aspects from the Base Rule Parser 344 and the Mapping Rule Matcher 346 may conflict. For example, the same field or data element may have different values. The Wait Element Fuser 348 can identify and / or detect such conflicts to determine one or more parsed values of the wait element 310. In at least one embodiment, the determinations of the Mapping Rule Matcher 346 can generally take precedence because they are location-based and the sensed data may not always be reliable. For example, this may be useful in cases where the Base Rule Parser 344 cannot derive the relevant rules that should be applied. For example, if there is no sign or other visual indicator that a left turn is not allowed at an intersection, the Base Rule Parser 344 may not be able to apply the corresponding base rule, even if the rule is not conventionally followed. However, the Mapping Rule Matcher 346 can apply the rule based at least on the historical driving of the autonomous vehicle through the intersection where the rule has been observed to apply. However, there may be temporary or new signals that may not be present or have a high enough confidence (e.g., based on inconsistent or too few observations, stale observations, etc.) to be included in the map data. The Base Rule Parser 344 can be used to consider such scenarios when parsing data into the wait element 310 (e.g., for construction signs, electronic signs, or other temporary or transient signals, based at least on the corresponding determinations of the Base Rule Parser 344 being assigned to or associated with these types of signals, the Base Rule Parser 344 can take precedence).
[0077] Although Figure 3 Not shown, the Wait Element Engine can optionally include a Path Obstacle Analyzer that performs Path Obstacle Analysis (OIPA). The OIPA can use the lane map, wait geometry, and Semantic Motion Segmentation (SMS) obstacle perception outputs to link participants to the path and wait geometry. The OIPA can be performed by rendering the path and region as an index image and then projecting the polygon shape of the participant into the image and integrating the overlap amount.
[0078] Additionally, the waiting element engine may include an occlusion analyzer. Occlusion analysis can provide occlusion understanding for OIPA results by obtaining a lane map and obstacle perception outputs (such as SMS and depth maps) and detecting segments in the lane map that may hide unseen participants. This may allow the yield planner to consider unseen participants as well as visible ones. The competitor path may come with an expected speed limit, and then the occluded portion of the path can be used to insert an unseen participant combination with the maximum speed at the nearest occluded position where it may currently be, with appropriate warnings for the expected convention (e.g., if the competition state is Stopped First has Precedence, then an unseen vehicle far behind its competition area and the entry line should not reasonably be expected to enter at maximum speed if it is not allowed to stop at its entry line, while if it is a Yield Contention Point, it can be assumed that it may enter at maximum speed). With this information, the yield planner can consider the unseen vehicle and correctly generate behaviors such as decelerating, waiting, or moving slowly forward until it can see the occluded area, or until the traffic light turns green in the case of a right-turn red light.
[0079] The waiting element engine 300 may additionally perform a Who-Stopped-First analysis. The Who-Stopped-First analysis can use motion analysis and OIPA results to determine which participant stopped first to support the competition state Stopped First has Precedence. The OIPA results can be used to determine whether a participant has entered the site, whether it is in its competition area or at the entry line. Motion can also be used to determine whether a participant is in motion, stopped but recently moved, or whether it may be a parked vehicle. Another analysis of the waiting element engine 300 can include a Who-Goes-First analysis. The Who-Goes-First analysis can be a machine learning analysis that estimates the competition state corresponding to each competitor. This analysis can be trained with many examples of future progress in which it can be determined whether a participant has passed a competition point before the ego vehicle and vice versa. Knowing the likelihood of who goes first may indicate yield expectations.
[0080] As described throughout, the output of the wait element 310 of the wait element engine 300 can be passed to the yield planner of the ego vehicle. Given a wait element with a parsed competitive state, an OIPA result with occlusion, who stopped first, who goes first, and obstacle perception output, the yield planner can implement a yield behavior for the ego vehicle. If needed and possible, the yield planner can cause a yield behavior in the ego vehicle and monitor the yielding of other participants when not yielding. A yield planner analysis can be performed to predict in advance what would happen if the ego vehicle proceeds on a declared path. If the ego vehicle continues to proceed and should yield, the ego vehicle may clear the competition before a competitor with the right of way is affected (e.g., forced to change their behavior). Yielding may include ensuring that the ego vehicle does not influence the competitors to change their behavior such that it deviates from their preferred or expected behavior.
[0081] Now referring to Figures 4 - 6 , each block of methods 400 - 600 and other methods described herein include computational processes that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The methods can also be embodied as computer - usable instructions stored on a computer storage medium. These methods can be provided by a stand - alone application, a service, or a hosted service (stand - alone or in combination with another hosted service) or a plug - in of another product, to name a few. Additionally, by way of example, methods 400 - 600 are described with respect to Figure 3 the wait element engine 300. However, these methods can be additionally or alternatively executed by any one system or any combination of systems, including but not limited to those described herein.
[0082] Figure 4 is a flowchart showing a method 400 for encoding a yield scenario for an autonomous vehicle (e.g., an ego vehicle) according to some embodiments of the present disclosure. The method can be executed by a wait element engine, such as but not limited to Figure 3 the wait element engine 300. In block B402, method 400 includes detecting and / or identifying an upcoming yield scenario. The yield scenario can correspond to an intersection (or junction) yield scenario or a merge yield scenario. The yield scenario can be associated with an autonomous vehicle (e.g., an ego vehicle) and one or more competitors.
[0083] In block B404, geometric data can be received. In some embodiments, the geometric data can be received in response to detecting the yield scenario. The geometric data can include geometric perception data. For example, the wait geometry perceptron 302 of the wait element engine 300 can receive real - time geometric data generated from one or more sensors of the autonomous vehicle.
[0084] At block B406, signal data can be received. In some embodiments, signal data can be received in response to detecting a yield scenario. The signal data can include signal perception data. For example, the signal sensor 306 of the wait element engine 300 can receive real-time signal data generated from one or more sensors of an autonomous vehicle.
[0085] At block B408, map data can be received. In some embodiments, map data can be received in response to detecting a yield scenario. For example, the mapper 304 of the wait element engine 300 can receive map data.
[0086] At block B410, lane map data can be received. In some embodiments, lane map data can be received in response to detecting a yield situation. The lane map data can include one or more self-paths and one or more competitor paths for one or more competitors in a yield scenario. For example, the lane mapper 350 of the wait element engine 300 can receive the self-path 352 of the ego vehicle (e.g., an autonomous vehicle) and the competitor paths 354 of one or more competitors associated with the yield scenario.
[0087] At block B412, the geometry of the self-path among one or more self-paths can be determined. The geometry of the self-path can be determined by geometric data, map data, signal data, and / or lane map data. Thus, the wait geometry sensor 302, mapper 304, signal sensor 306, lane mapper 350, or any combination thereof of the wait element engine 300 can generally be responsible for determining the geometry of the self-path. In some embodiments, the wait geometry fuser 322, geometry classifier 324, geometry correlator 326, signal fuser 328, signal state estimator 330 of the wait element engine 300, or any combination thereof can contribute to determining the geometry of the self-path.
[0088] Also in block B412, the geometry of the competitor paths of one or more competitors can be determined. Similar to the ego path, the geometry of the competitor paths can be determined by geometry data, map data, signal data, and / or lane diagram data. Thus, the waiting geometry sensor 302, mapper 304, signal sensor 306, lane diagrammer 350, or any combination thereof can generally be responsible for determining the geometry of the competitor paths. In some embodiments, the waiting geometry fuser 322 of the waiting element engine 300, the geometry classifier 324 of the waiting element engine 300, the geometry correlator 326 of the waiting element engine 300, the signal fuser 328 of the waiting element engine 300, the signal state estimator 330 of the waiting element engine 300, or any combination thereof can contribute to determining the geometry of the competitor paths. In some embodiments, in block 414, the geometry of the context of the path and / or the yield scenario is determined.
[0089] In block B414, the geometries of the ego path and the competitor paths can be encoded. The geometry of the ego path can be encoded in the ego waiting geometry (e.g., Figure 3 ego waiting geometry 312). The geometry of the competitor paths can be encoded in the competitor waiting geometry (e.g., Figure 3 competitor waiting geometry 314). In at least one embodiment, the geometry of the context for the path and / or yield geometry is encoded in the context waiting geometry, e.g., Figure 3 context waiting geometry 316).
[0090] In block B416, the competition state between the ego path and the competitor paths can be determined based at least on the determined geometries. Various embodiments for determining the competition state are discussed in connection with at least the waiting element engine 300, Figure 5 method 500, and / or Figure 6 method 600. However, briefly stated here, the competition state resolver 340 of the waiting element engine 300 can generally be responsible for determining and / or resolving the competition state between the ego path and the competitor paths. The competition state (e.g., Figure 3 competition state 318) can be encoded.
[0091] In block B418, a waiting element data structure (or data object) can be generated, e.g., Figure 3 waiting element 310. The waiting element data structure can include at least one of the geometry of the ego path, the geometry of the competitor paths, and the competition state. The waiting element can also include the context waiting geometry.
[0092] In B420, the waiting element can be provided to the yield planner of the autonomous vehicle.
[0093] Figure 5 is a flowchart showing a method 500 for encoding a yield scenario for an autonomous vehicle (e.g., a self-vehicle) according to some embodiments of the present disclosure. The method may be executed by a waiting element engine, such as but not limited to Figure 3 the waiting element engine 300. At block B502, method 500 includes sensing the waiting geometry of the self-path and the competitor path. A geometry sensor (e.g., the geometry sensor 302 of the waiting element engine 300) may sense the waiting geometry. In response to detecting and / or identifying a yield scenario of the self-vehicle (autonomous vehicle), the waiting geometry may be sensed. The self-path and the competitor path may be sensed by a lane illustrator (e.g., the lane illustrator 350 of the waiting element engine 300).
[0094] At block B504, one or more signals of the self-path / competitor path may be sensed. A signal sensor (e.g., the signal sensor 306) may sense the signals of the path.
[0095] At block B506, the waiting geometry may be fused with map data. A waiting geometry fuser (e.g., the waiting geometry fuser 322 of the waiting element engine 300) may fuse the waiting geometry with map data.
[0096] At block B508, the fused waiting geometry may be classified. A geometry classifier (e.g., the geometry classifier 324 of the waiting element engine 300) may classify the waiting geometry.
[0097] At block B510, the fused waiting geometry may be associated. A geometry associator (e.g., the geometry associator 326 of the waiting element engine 300) may associate the waiting geometry.
[0098] At block B512, the signal may be fused with map data. A signal fuser (e.g., the signal fuser 328 of the waiting element engine 300) may fuse the signal with map data.
[0099] At block B514, the state of the fused signal may be estimated. A signal state estimator (e.g., the signal state estimator 330 of the waiting element engine 300) may estimate the state of the fused signal.
[0100] At block B516, the competition state between the self-path and the competitor path may be resolved. Various embodiments of resolving the competition state are discussed in method 600 in conjunction with at least Figure 6 . However, briefly speaking here, a competition state resolver (e.g., the competition state resolver 340 of the waiting element engine 300) may resolve the competition state between the self-path and the competitor path.
[0101] In block B518, a waiting element can be generated. For example, a waiting element engine 300 can generate a waiting element 310.
[0102] In block B520, the waiting element can be provided to a system (such as a yield planner) that provides a guidance service to an autonomous vehicle (e.g., a self-vehicle).
[0103] Figure 6 is a flowchart showing a method 600 for parsing a competition state between vehicle paths according to some embodiments of the present disclosure. The method can be executed by a competition state parser, such as but not limited to Figure 3 the competition state parser 340. Various inputs can be provided to the competition state parser to parse the competition state between two paths (e.g., a self-path and a competitor path). For example, as Figure 3 shown, the map data of the two paths and the fused, classified, and / or associated geometries (e.g., waiting geometries) can be provided to the competition state parser 340. In addition, the fused signals (including the estimated signal states) can be provided to the competition state parser. Various other sensor data (from sensors on the autonomous vehicle) can be provided to the competition state parser. Each block of method 600 can adopt any such input data.
[0104] Method 600 includes checking the conditions of the competition state at block B602. A condition checker (e.g., the condition checker 342 of the competition state parser 340) can be used to check the conditions of the competition state.
[0105] In block B604, one or more base rules can be parsed. A base rule parser (e.g., the base rule parser 344 of the competition state parser 340) can parse the base rules.
[0106] In block B606, one or more mapping rules can be matched to the competition state. A mapping rule matcher (e.g., the mapping rule matcher 346 of the competition state parser 340) can match the competition state with one or more mapping rules.
[0107] In block B608, various data structures (e.g., waiting geometries and competition states) can be fused into a waiting element. A waiting element fuser (e.g., the waiting element fuser 348 of the competition state parser 340) can fuse the data structures.
[0108] In block B610, the fused data structures (e.g., waiting geometries and the parsed competition state) can be encapsulated into a waiting element (e.g., Figure 3 the waiting element 310).
[0109] Example autonomous vehicle
[0110] Figure 7A FIG. 700 illustrates an example autonomous vehicle 700 in accordance with some embodiments of the present disclosure. The autonomous vehicle 700 (alternatively, referred to herein as "vehicle 700") can include, but is not limited to, passenger vehicles such as cars, trucks, buses, first responder vehicles, shuttle vehicles, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, underwater vessels, drones, vehicles connected to trailers, and / or another type of vehicle (e.g., a driverless and / or one or more passenger-carrying vehicle). Autonomous vehicles are generally described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016 - 201806, issued June 15, 2018, Standard No. J3016 - 201609, issued September 30, 2016, and previous and future versions of the standard). The vehicle 700 may be capable of implementing functions corresponding to one or more of Levels 3 - 5 of the autonomous driving levels. The vehicle 700 may be capable of operating according to one or more of Levels 1 - 5 of the autonomous driving levels. For example, depending on the embodiment, the vehicle 700 may be capable of providing driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). As used herein, the term "autonomous" may include any and / or all types of autonomy of the vehicle 700 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assisted autonomy, being semi-autonomous, being mostly autonomous, or other designations.
[0111] The vehicle 700 can include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 700 can include a propulsion system 750, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. The propulsion system 750 can be connected to a drivetrain of the vehicle 700 that can include a transmission to effect propulsion of the vehicle 700. The propulsion system 750 can be controlled in response to receiving a signal from the throttle / accelerator 752.
[0112] A steering system 754 that may include a steering wheel can be used to steer vehicle 700 (e.g., along a desired path or route) while the propulsion system 750 is operating (e.g., while the vehicle is in motion). The steering system 754 can receive signals from a steering actuator 756. For fully autonomous (Level 5) functionality, the steering wheel can be optional.
[0113] A brake sensor system 746 can be used to operate vehicle brakes in response to receiving signals from a brake actuator 748 and / or a brake sensor.
[0114] One or more controllers 736 that may include one or more system-on-chips (SoCs) 704 ( Figure 7C ) and / or one or more GPUs can provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 700. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 748, operate the steering system 754 via one or more steering actuators 756, and operate the propulsion system 750 via one or more throttles / accelerators 752. One or more controllers 736 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 700. One or more controllers 736 can include a first controller 736 for autonomous driving functions, a second controller 736 for functional safety functions, a third controller 736 for artificial intelligence functions (e.g., computer vision), a fourth controller 736 for infotainment functions, a fifth controller 736 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 736 can handle two or more of the above functions, two or more controllers 736 can handle a single function, and / or any combination thereof.
[0115] One or more controllers 736 may provide signals for controlling one or more components and / or systems of vehicle 700 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data may be received from, for example and without limitation, a global navigation satellite system sensor 758 (e.g., a global positioning system sensor), a RADAR sensor 760, an ultrasonic sensor 762, a LIDAR sensor 764, an inertial measurement unit (IMU) sensor 766 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 796, a stereo camera 768, a wide-angle camera 770 (e.g., a fish-eye camera), an infrared camera 772, a surround camera 774 (e.g., a 360-degree camera), a remote and / or mid-range camera 798, a speed sensor 744 (e.g., for measuring the speed of vehicle 700), a vibration sensor 742, a steering sensor 740, a brake sensor (e.g., as part of a brake sensor system 746), and / or other sensor types.
[0116] One or more of the controllers 736 may receive inputs (e.g., represented by input data) from the instrument cluster 732 of vehicle 700 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 734, an audible annunciator, a speaker, and / or via other components of vehicle 700. These outputs may include information such as vehicle speed, rate, time, map data (e.g., Figure 7C the HD map 722), location data (e.g., the location of vehicle 700 on a map, for example), direction, the locations of other vehicles (e.g., occupancy grids), and information about objects and object states as perceived by the controller 736, etc. For example, the HMI display 734 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).
[0117] Vehicle 700 also includes a network interface 724, which may communicate via one or more networks using one or more wireless antennas 726 and / or a modem. For example, the network interface 724 may be capable of communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. One or more wireless antennas 726 may also enable communication between objects (e.g., vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth LE, Z-wave, ZigBee, etc. and / or one or more low-power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.
[0118] Figure 7BAn example of the camera positions and fields of view for an exemplary autonomous vehicle 700 in accordance with some embodiments of the present disclosure. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on the vehicle 700. Figure 7A The camera types for the cameras may include, but are not limited to, digital cameras that may be adapted to work with components and / or systems of the vehicle 700. The cameras may operate under an Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera types may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in an effort to increase light sensitivity.
[0119] In some examples, one or more of the cameras may be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera may be installed to provide functions including lane departure warning, traffic sign assistance, and smart headlight control. One or more of the cameras (e.g., all of the cameras) may record and provide image data (e.g., video) simultaneously.
[0120] One or more of the cameras may be mounted in a mounting assembly, such as a custom-designed (3-D printed) component, to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the cameras. Regarding the wing mirror mounting assembly, the wing mirror assembly may be custom 3-D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.
[0121]
[0122] A camera (e.g., a front camera) having a field of view that includes an environmental portion in front of vehicle 700 can be used for surround view to help identify forward paths and obstacles, and, with the help of one or more controllers 736 and / or a control SoC, assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.
[0123] A variety of cameras can be used in a front-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imager. Another example can be a wide-angle camera 770, which can be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 7B only one wide-angle camera is illustrated in, any number of wide-angle cameras 770 can be present on vehicle 700. Additionally, a tele camera 798 (e.g., a long-range stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The tele camera 798 can also be used for object detection and classification and basic object tracking.
[0124] One or more stereo cameras 768 can also be included in a front-facing configuration. The stereo camera 768 can include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (FPGA) with an integrated CAN or Ethernet interface on a single chip. Such a unit can be used to generate a 3-D map of the vehicle environment, including distance estimates for all points in the image. An alternative stereo camera 768 can include a compact stereo vision sensor that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 768 can be used in addition to or alternatively to those described herein.
[0125] A camera (e.g., a side-view camera) having a field of view that includes an environmental portion on the side of vehicle 700 can be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, a surround camera 774 (e.g., as Figure 7BThe four surround cameras 774 shown in the figure can be placed on the vehicle 700. The surround cameras 774 can include wide-angle cameras 770, fish-eye cameras, 360-degree cameras, and / or the like. By way of example, four fish-eye cameras can be placed at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 774 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward camera) can be utilized as the fourth surround camera.
[0126] A camera having a field of view that includes an environmental portion at the rear of the vehicle 700 (e.g., a rear-view camera) can be used to assist with parking, surround viewing, rear collision warning, and creating and updating an occupancy grid. A variety of cameras can be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 798, stereo cameras 768, infrared cameras 772, etc.).
[0127] Figure 7C For an example autonomous vehicle 700 in accordance with some embodiments of the present disclosure Figure 7A is a block diagram of an example system architecture. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) can be used in addition to or in place of those shown, and some elements can be omitted entirely. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0128] Figure 7C Each of the components, features, and systems in the vehicle 700 is illustrated as being connected via a bus 702. The bus 702 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as the "CAN bus"). The CAN can be a network within the vehicle 700 that is used to assist in controlling various features and functions of the vehicle 700, such as driving brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.
[0129] Although the bus 702 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or alternatively to a CAN bus, FlexRay and / or Ethernet can be used. Further, although the bus 702 is shown as a single line, this is not intended to be limiting. For example, there can be any number of buses 702, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 702 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 702 can be used for a collision avoidance function, and a second bus 702 can be used for drive control. In any example, each bus 702 can communicate with any component of the vehicle 700, and two or more buses 702 can communicate with the same component. In some examples, each SoC 704, each controller 736, and / or each computer within the vehicle can have access to the same input data (e.g., input from sensors of the vehicle 700), and can be connected to a common bus such as a CAN bus.
[0130] The vehicle 700 can include one or more controllers 736, such as those described herein with respect to Figure 7A The controllers 736 can be used for a variety of functions. The controllers 736 can be coupled to any other different components and systems of the vehicle 700, and can be used for control of the vehicle 700, artificial intelligence for the vehicle 700, infotainment for the vehicle 700, and / or the like.
[0131] The vehicle 700 can include one or more system-on-chips (SoCs) 704. The SoC 704 can include a CPU 706, a GPU 708, a processor 710, a cache 712, an accelerator 714, a data store 716, and / or other components and features not shown. In a variety of platforms and systems, the SoC 704 can be used to control the vehicle 700. For example, one or more SoCs 704 can be combined with an HD map 722 in a system (e.g., a system of the vehicle 700), and the HD map can obtain map refreshes and / or updates via a network interface 724 from one or more servers (e.g., Figure 7D one or more servers 778) of
[0132] The CPU 706 may include a CPU cluster or a CPU complex (alternatively referred to herein as a "CCPLEX"). The CPU 706 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 706 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 706 may include four dual-core clusters, each with a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 706 (e.g., the CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 706 can be active at any given time.
[0133] The CPU 706 may implement power management capabilities including one or more of the following: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 706 may further implement an enhanced algorithm for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this task is offloaded to the microcode.
[0134] The GPU 708 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 708 may be programmable and efficient for parallel workloads. In some examples, the GPU 708 may use an enhanced tensor instruction set. The GPU 708 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 708 may include at least eight streaming microprocessors. The GPU 708 may use a compute application programming interface (API). Additionally, the GPU 708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0135] In automotive and embedded use cases, the GPU 708 can be power optimized for best performance. For example, the GPU 708 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be limiting, and the GPU 708 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of compute and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to allow for more fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0136] The GPU 708 can include, in some examples, high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900 GB / s and / or a 16GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.
[0137] The GPU 708 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 708 to directly access the CPU 706 page tables. In such an example, when the GPU 708 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 706. In response, the CPU 706 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 708. In this way, the unified memory technology can allow a single unified virtual address space for the memory of both the CPU 706 and the GPU 708, thereby simplifying GPU 708 programming and porting applications to the GPU 708.
[0138] In addition, the GPU 708 may include an access counter that can track how often the GPU 708 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.
[0139] The SoC 704 may include any number of caches 712, including those described herein. For example, the cache 712 may include an L3 cache that is available to both the CPU 706 and the GPU 708 (e.g., which is connected to both the CPU 706 and the GPU 708). The cache 712 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.
[0140] The SoC 704 may include an arithmetic logic unit (ALU) that can be utilized in the processing of performing any of a variety of tasks or operations regarding the vehicle 700, such as processing a DNN. In addition, the SoC 704 may include a floating-point unit (FPU) (or other math co-processor or digital co-processor type) for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs integrated as execution units within the CPU 706 and / or the GPU 708.
[0141] The SoC 704 may include one or more accelerators 714 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 704 may include a hardware accelerator cluster that may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) may enable the hardware accelerator cluster to accelerate neural networks and other computations. The hardware accelerator cluster may be used to supplement the GPU 708 and offload some of the tasks of the GPU 708 (e.g., freeing up more cycles of the GPU 708 for performing other tasks). As an example, the accelerator 714 may be used for targeted workloads that are stable enough to be easily accelerated (e.g., perception, convolutional neural network (CNN), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNN) and fast RCNN (e.g., for object detection).
[0142] The accelerator 714 (e.g., a hardware accelerator cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that may be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating point operations, and inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU, and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0143] The DLA may quickly and efficiently execute neural networks, especially CNNs, for any of a variety of functions on processed or unprocessed data, such as, and not limited to: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner recognition using data from a camera sensor; and / or CNNs for security and / or safety-related events.
[0144] The DLA may perform any function of the GPU 708, and by using an inference accelerator, for example, a designer may target the DLA or the GPU 708 for any function. For example, a designer may focus the processing and floating point operations of a CNN on the DLA, and leave other functions to the GPU 708 and / or other accelerators 714.
[0145] The accelerator 714 (e.g., a hardware accelerator cluster) may include a Programmable Vision Accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and not limited to, any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA), and / or any number of vector processors.
[0146] The RISC cores can interact with an image sensor (e.g., the image sensor of any camera described herein), an image signal processor, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0147] DMA can enable components of the PVA to access system memory independently of the CPU 706. DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, DMA can support addressing up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0148] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.
[0149] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. Among other things, any number of PVAs may be included in a hardware accelerator cluster, and any number of vector processors may be included in each of these PVAs. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system security.
[0150] The accelerator 714 (e.g., a hardware accelerator cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 714. In some examples, the on-chip memory may include at least 4MB SRAM consisting of, for example and without limitation, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include (e.g., using APB) an on-chip computer vision network that interconnects the PVA and the DLA to the memory.
[0151] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transmission. This type of interface may conform to the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.
[0152] In some examples, SoC 704 can include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LIDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.
[0153] Accelerator 714 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving applications. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.
[0154] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3 - 5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0155] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.
[0156] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 766 related to the orientation and distance of the vehicle 700, a 3D position estimate of an object obtained from a neural network and / or other sensors (such as a LIDAR sensor 764 or a RADAR sensor 760), etc.
[0157] The SoC 704 can include one or more data stores 716 (e.g., memory). The data store 716 can be on-chip memory of the SoC 704, which can store a neural network to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 716 can be large enough in capacity to store multiple instances of the neural network. The data store 712 can include an L2 or L3 cache 712. References to the data store 716 can include references to memory associated with the PVA, DLA, and / or other accelerators 714 as described herein.
[0158] The SoC 704 may include one or more processors 710 (e.g., embedded processors). The processor 710 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 704 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low power state transitions, SoC 704 heat and temperature sensor management, and / or SoC 704 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 704 may use the ring oscillator to detect the temperature of the CPU 706, GPU 708, and / or accelerator 714. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 704 in a lower power state and / or place the vehicle 700 in a driver safety stop mode (e.g., safely stop the vehicle 700).
[0159] The processor 710 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0160] The processor 710 may also include an always-on processor engine, which may provide the necessary hardware features to support low power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0161] The processor 710 may also include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.
[0162] The processor 710 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0163] The processor 710 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0164] The processor 710 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 770, the surround camera 774, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.
[0165] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.
[0166] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 708 does not need to continuously render new surfaces, the video image compositor may be further used for user interface composition. Even when the GPU 708 is powered on and active for 3D rendering, the video image compositor may be used to relieve the burden on the GPU 708 to improve performance and responsiveness.
[0167] The SoC 704 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 704 may also include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.
[0168] The SoC 704 may also include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC 704 may be used to process data from cameras (via gigabit multimedia serial link and Ethernet connections), sensors (such as LIDAR sensor 764, RADAR sensor 760, etc. that may be connected via Ethernet), data from bus 702 (such as the speed of vehicle 700, steering wheel position, etc.), and data from GNSS sensor 758 (connected via Ethernet or CAN bus). The SoC 704 may also include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and may be used to free the CPU 706 from routine data management tasks.
[0169] The SoC 704 may be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thereby providing an integrated functional safety architecture for a platform that utilizes and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. The SoC 704 may be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with the CPU 706, GPU 708, and data storage 716, the accelerator 714 may provide a fast and efficient platform for level 3 - 5 autonomous vehicles.
[0170] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.
[0171] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware accelerator cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on the DLA or dGPU (such as GPU 720) may include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may also include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.
[0172] As another example, multiple neural networks can run simultaneously, as required for level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network deployed over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 708.
[0173] In some examples, the CNNs for face recognition and owner recognition can use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 700. A processing engine that is always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a security mode, to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 704 provides security against theft and / or carjacking.
[0174] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 796 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 704 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating, as recognized by the GNSS sensor 758. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 762, a control program can be used to execute an emergency vehicle safety routine, slowing down the vehicle, pulling over to the side of the road, stopping the vehicle, and / or idling the vehicle until the emergency vehicle passes.
[0175] The vehicle may include a CPU 718 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 704 via a high-speed interconnect (e.g., PCIe). The CPU 718 may include, for example, an X86 processor. The CPU 718 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 704, and / or monitoring the status and health of the controller 736 and / or the infotainment SoC 730.
[0176] The vehicle 700 may include a GPU 720 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 704 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 720 may provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and may be used to train and / or update neural networks at least in part based on inputs (e.g., sensor data) from sensors of the vehicle 700.
[0177] The vehicle 700 may further include a network interface 724, which may include one or more wireless antennas 726 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 724 may be used to enable wireless connections to the cloud (e.g., to the server 778 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the passenger's client device) via the Internet. To communicate with other vehicles, a direct link may be established between the two vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 700 with information about vehicles approaching the vehicle 700 (e.g., vehicles in front of, beside, and / or behind the vehicle 700). This function may be part of the cooperative adaptive cruise control function of the vehicle 700.
[0178] The network interface 724 may include an SoC that provides modulation and demodulation functions and enables the controller 736 to communicate via a wireless network. The network interface 724 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0179] Vehicle 700 may also include a data store 728 that may include off-chip (e.g., outside of SoC 704) storage devices. The data store 728 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.
[0180] Vehicle 700 may also include a GNSS sensor 758. The GNSS sensor 758 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 758 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.
[0181] Vehicle 700 may also include a RADAR sensor 760. The RADAR sensor 760 may be used by the vehicle 700 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 760 may use CAN and / or bus 702 (e.g., to transmit data generated by the RADAR sensor 760) for control and access to object tracking data and, in some examples, accesses Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 760 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.
[0182] The RADAR sensor 760 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250 m) achieved through two or more independent scans. The RADAR sensor 760 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of the vehicle 700 with minimal traffic interference from adjacent lanes at a higher rate. The other two antennas may expand the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 700.
[0183] As an example, a mid-range RADAR system can include a range of up to 760m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 750 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the blind spots behind and beside the vehicle.
[0184] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.
[0185] Vehicle 700 can also include ultrasonic sensors 762. Ultrasonic sensors 762 that can be placed in the front, rear, and / or sides of vehicle 700 can be used for parking assistance and / or creating and updating occupancy grids. A variety of ultrasonic sensors 762 can be used, and different ultrasonic sensors 762 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 762 can operate at ASIL B for functional safety levels.
[0186] Vehicle 700 can include a LIDAR sensor 764. The LIDAR sensor 764 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 764 can be at ASIL B for functional safety levels. In some examples, vehicle 700 can include multiple LIDAR sensors 764 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).
[0187] In some examples, the LIDAR sensor 764 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 764 can have, for example, an advertised range of approximately 700m, an accuracy of 2cm - 3cm, and support for a 700Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 764 can be used. In such examples, the LIDAR sensor 764 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 700. In such examples, the LIDAR sensor 764 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, with a range of 200m, even for low-reflectivity objects. The front-mounted LIDAR sensor 764 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0188] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses the flash of a laser as the emission source to illuminate the vehicle's surroundings up to approximately 200m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow for the generation of highly accurate and distortion-free images of the surroundings using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 700. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a fan. The flash LIDAR device can use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR and because flash LIDAR is a solid-state device without moving parts, the LIDAR sensor 764 can be less susceptible to motion blur, vibration, and / or shock.
[0189] The vehicle can also include an IMU sensor 766. In some examples, the IMU sensor 766 can be located at the center of the rear axle of the vehicle 700. The IMU sensor 766 can include, for example and without limitation, accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 766 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 766 can include an accelerometer, a gyroscope, and a magnetometer.
[0190] In some embodiments, the IMU sensor 766 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 766 can enable the vehicle 700 to estimate the heading without input from a magnetic sensor by directly observing the change in velocity from the GPS to the IMU sensor 766 and correlating it. In some examples, the IMU sensor 766 and the GNSS sensor 758 can be integrated into a single unit.
[0191] The vehicle can include a microphone 796 placed in and / or around the vehicle 700. Among other things, the microphone 796 can be used for emergency vehicle detection and identification.
[0192] The vehicle may also include any number of camera types, including a stereo camera 768, a wide-angle camera 770, an infrared camera 772, a surround camera 774, a long-range and / or mid-range camera 798, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 700. The camera types used depend on the embodiment and the requirements of the vehicle 700, and any combination of camera types can be used to provide the necessary coverage around the vehicle 700. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 7A and Figure 7B is described in more detail.
[0193] The vehicle 700 may also include a vibration sensor 742. The vibration sensor 742 can measure the vibration of components of the vehicle such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 742 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-rotating axle).
[0194] The vehicle 700 may include an ADAS system 738. In some examples, the ADAS system 738 may include a SoC. The ADAS system 738 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0195] The ACC system can use RADAR sensors 760, LIDAR sensors 764, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 700 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 700 to change lanes. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0196] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 724 and / or the wireless antenna 726 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., a vehicle immediately in front of vehicle 700 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given the information about the vehicle in front of vehicle 700, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.
[0197] The FCW system is designed to alert the driver to a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of, for example, sound, visual warnings, vibration, and / or rapid braking pulses.
[0198] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.
[0199] The LDW system provides visual, auditory, and / or tactile warnings such as steering wheel or seat vibration to alert the driver when vehicle 700 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0200] The LKA system is a variant of the LDW system. If the vehicle 700 starts to leave the lane, then the LKA system provides a steering input or braking to correct the vehicle 700.
[0201] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0202] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 700 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0203] Conventional ADAS systems can be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide if a safe condition truly exists and act accordingly. However, in an autonomous vehicle 700, in the case of conflicting results, the vehicle 700 itself must decide whether to heed the results from the primary computer or an auxiliary computer (e.g., the first controller 736 or the second controller 736). For example, in some embodiments, the ADAS system 738 can be a standby and / or auxiliary computer for providing perception information to a standby computer rationality module. The standby computer rationality monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 738 can be provided to the supervisory MCU. If the outputs from the primary computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0204] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the direction of the host computer regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine an appropriate result.
[0205] The supervisory MCU may be configured to run a neural network that is trained and configured to determine, at least in part based on outputs from the host computer and the secondary computer, conditions under which the secondary computer provides a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact dangerous, such as a drainage grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments including a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include components of the SoC 704 and / or be included as a component of the SoC 704.
[0206] In other examples, the ADAS system 738 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0207] In some examples, the output of the ADAS system 738 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 738 indicates a forward collision warning due to an object being immediately in front, the perception block can use that information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0208] Vehicle 700 can also include an infotainment SoC 730 (e.g., in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 730 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to vehicle 700. For example, the infotainment SoC 730 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, WiFi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 734, telematics device, control panel (e.g., for controlling various components, features, and / or systems, and / or interacting with them), and / or other components. The infotainment SoC 730 can further be used to provide information (e.g., visual and / or auditory) to the vehicle's user, such as information from the ADAS system 738, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0209] The infotainment SoC 730 can include GPU functionality. The infotainment SoC 730 can communicate with other devices, systems, and / or components of vehicle 700 via a bus 702 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 730 can be coupled to the supervisory MCU such that in the event of a failure of the main controller 736 (e.g., the main computer and / or the backup computer of vehicle 700), the GPU of the infotainment system can perform some autonomous driving functions. In such examples, the infotainment SoC 730 can place vehicle 700 in the driver safe parking mode as described herein.
[0210] Vehicle 700 may further include an instrument cluster 732 (such as a digital instrument panel, an electronic instrument cluster, a digital instrument faceplate, etc.). The instrument cluster 732 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 732 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. In other words, the instrument cluster 732 may be included as part of the infotainment SoC 730, or vice versa.
[0211] Figure 7D A system schematic diagram of communication between a cloud-based server and Figure 7A an exemplary autonomous vehicle 700 according to some embodiments of the present disclosure. The system 776 may include a server 778, a network 790, and vehicles including the vehicle 700. The server 778 may include multiple GPUs 784(A)-784(H) (collectively referred to herein as GPUs 784), PCIe switches 782(A)-782(H) (collectively referred to herein as PCIe switches 782), and / or CPUs 780(A)-780(B) (collectively referred to herein as CPUs 780). The GPUs 784, CPUs 780, and PCIe switches may be interconnected by high-speed interconnects such as, for example and without limitation, the NVLink interface 788 developed by NVIDIA and / or PCIe connections 786. In some examples, the GPUs 784 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 784 and the PCIe switches 782 are connected via a PCIe interconnect. Although eight GPUs 784, two CPUs 780, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 778 may include any number of GPUs 784, CPUs 780, and / or PCIe switches. For example, each of the servers 778 may include eight, sixteen, thirty-two, and / or more GPUs 784.
[0212] Server 778 can receive image data from the vehicle via network 790, the image data representing an image showing an unexpected or changed road condition such as a recently started road work. Server 778 can transmit neural network 792, updated neural network 792, and / or map information 794 via network 790 to the vehicle, including information about traffic and road conditions. Updates to the map information 794 can include updates to the HD map 722, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 792, updated neural network 792, and / or map information 794 can be represented and / or generated based on data received from new training and / or data from any number of vehicles in the environment and / or experience from training performed at the data center (e.g., using server 778 and / or other servers).
[0213] Server 778 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by the vehicle and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 790), and / or the machine learning model can be used by server 778 to remotely monitor the vehicle.
[0214] In some examples, server 778 can receive data from the vehicle and apply the data to the latest real-time neural network for real-time intelligent inference. Server 778 can include a deep learning supercomputer powered by GPU 784 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 778 can include a deep learning infrastructure of a data center powered only by a CPU.
[0215] The deep learning infrastructure of server 778 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 700. For example, the deep learning infrastructure can receive periodic updates from vehicle 700, such as an image sequence and / or objects located in the image sequence that vehicle 700 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 700. If the results do not match and the infrastructure concludes that the AI in vehicle 700 has failed, then server 778 can transmit a signal to vehicle 700, instructing the fail-safe computer in vehicle 700 to take control, notify the passengers, and complete a safe parking operation.
[0216] For inference, server 778 can include a GPU 784 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0217] Example computing device
[0218] Figure 8 is a block diagram of an example computing device 800 suitable for implementing some embodiments of the present disclosure. Computing device 800 can include an interconnect system 802 that directly or indirectly couples the following devices: a memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply 816, one or more presentation components 818 (e.g., (one or more) displays), and one or more logic units 820. In at least one embodiment, (one or more) computing devices 800 can include one or more virtual machines (VMs), and / or any of its components can include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of GPUs 808 can include one or more vGPUs, one or more of CPUs 806 can include one or more vCPUs, and / or one or more of logic units 820 can include one or more virtual logic units. Thus, (one or more) computing devices 800 can include discrete components (e.g., full GPUs dedicated to computing device 800), virtual components (e.g., a portion of a GPU dedicated to computing device 800), or a combination thereof.
[0219] AlthoughFigure 8 Each block is shown as being connected via circuitry through an interconnect system 802, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presentation component 818 (such as a display device) may be considered an I / O component 814 (e.g., if the display is a touch screen). As another example, the CPU 806 and / or the GPU 808 may include memory (e.g., the memory 804 may represent a storage device in addition to the memory of the GPU 808, the CPU 806, and / or other components). In other words, Figure 8 the computing device is illustrative only. No distinction is made between such categories as “workstation,” “server,” “laptop computer,” “desktop computer,” “tablet computer,” “client device,” “mobile device,” “handheld device,” “gaming console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types because all are considered within the scope of Figure 8 the computing device.
[0220] The interconnect system 802 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 802 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 806 may be directly connected to the memory 804. Further, the CPU 806 may be directly connected to the GPU 808. In cases where there are direct or point-to-point connections between components, the interconnect system 802 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 800.
[0221] The memory 804 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by the computing device 800. Computer-readable media may include volatile and nonvolatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.
[0222] A computer storage medium can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements such as an operating system). A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 800. As used herein, a computer storage medium does not include the signal itself.
[0223] A computer storage medium can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. The term "modulated data signal" can refer to a signal that sets or changes one or more of its characteristics in a manner that encodes information in the signal. By way of example, and not limitation, a computer storage medium can include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer-readable media.
[0224] CPU 806 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. Each of the CPUs 806 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. The CPUs 806 can include any type of processor and can include different types of processors depending on the type of computing device 800 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 800, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as a math coprocessor), computing device 800 can also include one or more CPUs 806.
[0225] In addition to or instead of one or more CPUs 806, one or more GPUs 808 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 800 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 808 may be an integrated GPU (e.g., one of the CPUs 806) and / or one or more of the GPUs 808 may be a discrete GPU. In an embodiment, one or more of the GPUs 808 may be a coprocessor of one or more of the CPUs 806. The GPUs 808 may be used by the computing device 800 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPUs 808 may be used for general-purpose computing on GPUs (GPGPU). The GPUs 808 may include hundreds or thousands of cores capable of concurrently handling hundreds or thousands of software threads. The GPUs 808 may generate pixel data of an output image in response to a rendering command (e.g., a rendering command received from the CPU 806 via a host interface). The GPUs 808 may include a graphics memory (e.g., a display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory may be included as part of the memory 804. The GPUs 808 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 808 may generate pixel data or GPGPU data for different parts of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0226] In addition to and / or instead of the CPU 806 and / or GPU 808, the logic unit 820 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 800 to perform one or more of the methods and / or processes described herein. In an embodiment, the (one or more) CPU 806, the (one or more) GPU 808, and / or the (one or more) logic units 820 may perform any combination of the methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 820 may be part of one or more of the CPU 806 and / or GPU 808 and / or integrated in one or more of the CPU 806 and / or GPU 808 and / or one or more of the logic units 820 may be discrete components or otherwise external to the CPU 806 and / or GPU 808. In an embodiment, one or more of the logic units 820 may be a coprocessor of one or more of the CPU 806 and / or one or more of the GPU 808.
[0227] Examples of the logic unit 820 include one or more processing cores and / or their components, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0228] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 800 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communication). The communication interface 810 may include components and functions that enable communication over any of a plurality of different networks, such as a wireless network (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), a wired network (e.g., via Ethernet or InfiniBand communication), a low power wide area network (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more of the logic units 820 and / or the communication interface 810 may include one or more data processing units (DPUs) to directly transfer data received over the network and / or via the interconnect system 802 to one or more GPUs 808 (e.g., the memory of one or more GPUs 808).
[0229] The I / O port 812 may enable the computing device 800 to be logically coupled to other devices including I / O components 814, one or more presentation components 818, and / or other components, some of which may be built into (e.g., integrated in) the computing device 800. Illustrative I / O components 814 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and the like. The I / O components 814 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some cases, the input may be transmitted to appropriate network elements for further processing. The NUI may implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of the computing device 800. The computing device 800 may include a depth camera for gesture detection and recognition, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof. Additionally, the computing device 800 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables the detection of motion. In some examples, the computing device 800 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0230] The power supply 816 may include hardwired power, battery power, or a combination thereof. The power supply 816 may provide power to the computing device 800 to enable the components of the computing device 800 to operate.
[0231] The presentation component 818 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components.
[0232] The presentation component 818 may receive data from other components (e.g., the GPU 808, the CPU 806, the DPU, etc.) and output the data (e.g., as an image, video, sound, etc.).
[0233] Example data center
[0234] Figure 9 An example data center 900 that may be used in at least one embodiment of the present disclosure is shown. The data center 900 may include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and / or an application layer 940.
[0235] As Figure 9As shown, the data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources ("node C.R.s") 916(1)-916(N), where "N" represents any whole positive integer. In at least one embodiment, the node C.R.s 916(1)-916(N) may include, but are not limited to, any number of central processing units ("CPU") or other processors (including DPU, accelerators, field programmable gate arrays (FPGA), graphics processors or graphics processing units (GPU), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VM"), power modules, and / or cooling modules, and so on. In some embodiments, one or more of the node C.R.s 916(1)-916(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R.s 916(1)-916(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the node C.R.s 916(1)-916(N) may correspond to a virtual machine (VM).
[0236] In at least one embodiment, the grouped computing resources 914 may include separate groupings of the node C.R.s 916 housed within one or more racks (not shown), or many racks within a data center located at different geographical locations (also not shown). Separate groupings of the node C.R.s 916 within the grouped computing resources 914 may include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several of the node C.R.s 916 including CPU, GPU, DPU, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any combination of any number of power modules, cooling modules, and / or network switches.
[0237] The resource coordinator 912 may configure or otherwise control one or more of the node C.R.s 916(1)-916(N) and / or the grouped computing resources 914. In at least one embodiment, the resource coordinator 912 may include a software design infrastructure ("SDI") management entity for the data center 900. The resource coordinator 912 may include hardware, software, or some combination thereof.
[0238] In at least one embodiment, as Figure 9As shown, the framework layer 920 may include a job scheduler 933, a configuration manager 934, a resource manager 936, and / or a distributed file system 938. The framework layer 920 may include a framework for software 932 that supports the software layer 930 and / or one or more applications 942 of the application layer 940. The software 932 or the application 942 may respectively include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 920 may be, but is not limited to, a free and open-source software web application framework (such as Apache Spark TM (hereinafter referred to as "Spark")) that can utilize the distributed file system 938 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 933 may include a Spark driver to facilitate scheduling of workloads supported by different layers of the data center 900. The configuration manager 934 may be able to configure different layers, such as the software layer 930 and the framework layer 920 (which includes Spark and the distributed file system 938 for supporting large-scale data processing). The resource manager 936 may be able to manage the clustered or grouped computing resources that are mapped to the distributed file system 938 and the job scheduler 933 or are allocated to support the distributed file system 938 and the job scheduler 933. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 914 in the data center infrastructure layer 910. The resource manager 936 may coordinate with the resource coordinator 912 to manage these mapped or allocated computing resources.
[0239] In at least one embodiment, the software 932 included in the software layer 930 may include software used by at least a portion of the node C.R.s 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0240] In at least one embodiment, the applications 942 included in the application layer 940 may include one or more types of applications used by at least a portion of the node C.R.s 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0241] In at least one embodiment, any one of the configuration manager 934, the resource manager 936, and the resource coordinator 912 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions can save the data center operator of the data center 900 from making potentially poor configuration decisions and potentially avoid underutilization and / or poorly performing parts of the data center.
[0242] According to one or more embodiments described herein, the data center 900 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information. For example, the (one or more) machine learning models can be trained by calculating weight parameters according to a neural network architecture by using the software and / or computing resources described above with respect to the data center 900. In at least one embodiment, the trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information by using the resources described above with respect to the data center 900 by using the weight parameters calculated by one or more training techniques (such as but not limited to those training techniques described herein).
[0243] In at least one embodiment, the data center 900 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or their corresponding virtual computing resources) to perform training and / or inference by using the above resources. In addition, one or more of the software and / or hardware resources described above can be configured to allow a user to train or execute a service for inferring information, such as image recognition, speech recognition, or other artificial intelligence services.
[0244] Example Network Environment
[0245] A network environment suitable for implementing the embodiments of the present disclosure can include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) can be implemented on one or more instances of the Figure 8 computing device 800 - for example, each device can include similar components, features, and / or functions of the (one or more) computing device 800. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device can be included as part of the data center 900, and an example of the data center 900 is described in more detail herein with respect to Figure 9 more detail.
[0246] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. The network can include multiple networks or one of multiple networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide a wireless connection.
[0247] A compatible network environment can include one or more peer-to-peer network environments (in which case, servers may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functions described herein for servers can be implemented on any number of client devices.
[0248] In at least one embodiment, the network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports software layers and / or one or more applications of an application layer. The software or application can respectively include network-based service software or applications. In an embodiment, one or more client devices can use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, but is not limited to, a free and open-source software web application framework that can perform large-scale data processing (e.g., "big data") using a distributed file system.
[0249] A cloud-based network environment can provide any combination of cloud computing and / or cloud storage that performs the computing and / or data storage functions (or one or more parts thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers in a state, region, country, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a part of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0250] (One or more) client devices can include what is described herein regardingFigure 8 At least some of the components, features, and functions of the example computing device(s) 800 described. By way of example and not limitation, a client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, spacecraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0251] The present disclosure may be described in the general context of machine - usable instructions or computer code, including computer - executable instructions such as program modules, executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general - purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0252] As used herein, the recitation of "and / or" with respect to two or more elements should be construed to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0253] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including different steps or combinations of steps similar to those described herein in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is expressly described.
Claims
1. A method, comprising: identifying a scenario of the first vehicle based at least on analyzing perception data generated by at least one sensor of the first vehicle in an environment; determining a first path of the first vehicle and a second path of a second vehicle in the scenario based at least on the perception data, wherein the second path includes at least one competing point with the first path; determining that one or more traffic rules in a set of traffic rules applicable to the scenario based at least on the perception data and the set of traffic rules corresponding to the environment; assigning a competition state to the at least one competing point based at least on the one or more traffic rules being applicable to the scenario; generating a waiting element associated with the scenario, the waiting element encoding a representation of a first geometry associated with the first path, a second geometry associated with the second path, and the competition state; and providing the waiting element to a control agent of the first vehicle, wherein the control agent is configured to use the waiting element to determine a yielding behavior of the first vehicle.
2. The method according to claim 1 further comprises: Detecting a current state of one or more traffic signals included in the environment using the perception data, wherein determining that the one or more traffic rules in the set of traffic rules applicable to the scenario based at least on the current state of the one or more traffic signals.
3. The method according to claim 1, wherein The perception data includes first geometry information associated with at least the first path, and the method further comprises: receiving map data representing second geometry information associated with at least the first path; and generating a representation of the first geometry of the first path based at least on fusing the first geometry information with the second geometry information.
4. The method according to claim 1, wherein, The perception data includes first signal information associated with one or more detected traffic signals in the environment, and the method further comprises: receiving map data from at least one map including second signal information associated with one or more located traffic signals; and generating fused signal information associated with at least one traffic signal based at least on fusing the first signal information with the second signal information, wherein determining that the one or more traffic rules in the set of traffic rules applicable to the scenario based at least on the fused signal information.
5. The method according to claim 1, wherein The perception data includes geometric perception data, and the method further comprises: classifying the first path as belonging to at least one category in a set of predetermined categories based at least on the geometric perception data, wherein determining that the one or more traffic rules in the set of traffic rules applicable to the scenario based at least on the first path belonging to the at least one category.
6. The method according to claim 1, wherein, The perception data includes signal perception data associated with one or more traffic signals and geometric data associated with the first path, and the method further comprises: assigning the one or more traffic signals to the first path based at least on evaluating a distance between the one or more traffic signals and the first path using the signal perception data and the geometric data.
7. The method according to claim 1, wherein The set of traffic rules includes at least one of one or more mapping rules associated with positioning the first vehicle to at least one map or one or more base rules associated with a geographical area of the first vehicle.
8. A processor, comprising: one or more circuits for identifying a scenario of the first vehicle based at least on analyzing sensor data generated by at least one sensor of the first vehicle in an environment; determining a first path of the first vehicle and a second path of a second vehicle in the scenario based at least on positioning the first vehicle to one or more maps, wherein the second path includes at least one competing point with the first path; determining that one or more traffic rules apply to the scenario based at least on the positioning of the first vehicle; assigning a competition state to the at least one competing point based at least on the one or more traffic rules applying to the scenario; generating a wait element associated with the scenario, the wait element encoding a representation of a first geometry associated with the first path, a second geometry associated with the second path, and the competition state; and providing the wait element to a control agent of the first vehicle, wherein the control agent is configured to use the wait element to determine a yielding behavior of the first vehicle.
9. The processor according to claim 8, wherein, The one or more circuits are further configured to detect a current state of one or more traffic signals included in the one or more maps using perception data generated by at least one sensor of the first vehicle in the environment, wherein it is determined that the one or more traffic rules apply to the scenario based at least on the current state of the one or more traffic signals.
10. The processor according to claim 8, wherein, The one or more circuits are further configured to: receive map data of the one or more maps based at least on the positioning, wherein the map data includes first geometric information associated with at least the first path; determine perception data generated by at least one sensor of the first vehicle in the environment, wherein the perception data includes second geometric information associated with at least the first path; and generate a representation of at least the first geometry of the first path based at least on fusing the first geometric information with the second geometric information.
11. The processor according to claim 8, wherein, The one or more circuits are further configured to: receive map data of the one or more maps based at least on the positioning, wherein the map data includes first signal information associated with one or more positioned traffic signals from the one or more maps; determine perception data generated by at least one sensor of the first vehicle in the environment, wherein the perception data includes second signal information associated with one or more detected traffic signals in the environment; and generate fused signal information associated with at least one traffic signal based at least on fusing the first signal information with the second signal information, wherein it is determined that the one or more traffic rules apply to the scenario based at least on the fused signal information.
12. The processor according to claim 8, wherein, The map data includes geometric data, and the one or more circuits are further configured to: classify the first path as belonging to at least one category of a set of predetermined categories based at least on the geometric data, wherein the one or more traffic rules are determined to be applicable to the scenario based at least on the first path belonging to the at least one category.
13. The processor according to claim 8, wherein, The map data includes signal data associated with one or more traffic signals, and the one or more circuits are further configured to assign the one or more traffic signals to the first path based at least on evaluating a distance between the one or more traffic signals and the first path using the signal data and geometric data associated with the first path.
14. The processor according to claim 8, wherein, The traffic rules include one or more mapping rules associated with positioning the first vehicle to at least one map or one or more basic rules associated with a geographic area of the first vehicle.
15. A system, comprising: one or more processing units, configured to determine that one or more traffic rules are applicable to a scenario, the scenario including a generated path of a vehicle in the environment, based at least on analyzing sensor data generated by at least one sensor of the vehicle in the usage environment; assign a competition state based at least on the one or more traffic rules being applicable to the scenario; generate a waiting element associated with the scenario, the waiting element encoding at least a first geometry associated with at least the generated path and the competition state; and provide the waiting element to a control agent of the vehicle, wherein the control agent is configured to use the waiting element to determine a yielding behavior of the vehicle.
16. The system according to claim 15, wherein, The one or more processing units are further configured to detect a current state of one or more traffic signals included in the environment, wherein the one or more traffic rules are determined to be applicable to the scenario based at least on the current state of the one or more traffic signals.
17. The system according to claim 15, wherein The one or more processing units are further configured to: receive perception data, the perception data including first geometric information associated with at least the path; receive map data, the map data representing second geometric information associated with at least the path; and generate at least the first geometry of the path based at least on fusing the first geometric information with the second geometric information.
18. The system according to claim 15, wherein the one or more processing units are further configured to: receive perception data, the perception data including first signal information associated with one or more detected traffic signals in the environment; receive map data from at least one map including second signal information associated with one or more positioned traffic signals; and generate fused signal information associated with at least one traffic signal based at least on fusing the first signal information with the second signal information, wherein the one or more traffic rules are determined to be applicable to the scenario based at least on the fused signal information.
19. The system according to claim 15, wherein the one or more processing units are further configured to classify the path as belonging to at least one category of a set of predetermined categories based at least on geometric perception data, and wherein the one or more traffic rules are determined to be applicable to the scenario based at least on the path belonging to the at least one category.
20. The system according to claim 15, wherein The one or more processing units are included in at least one of the following: A control system for an autonomous or semi-autonomous machine; A perception system for an autonomous or semi-autonomous machine; A system for performing simulation operations; A system for performing deep learning operations; A system implemented using edge devices; A system implemented using robots; A system comprising one or more virtual machines (VMs); A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2
Behavior planning for autonomous vehicles in yield scenarios
US11926346B2
Planning method of express lane and unit
CN108062863A
Automated driving systems and control logic for cloud-based scenario planning of autonomous vehicles
CN110271556A